Systematic reviews are the backbone of evidence-based medicine, but they are also notoriously slow, labor-intensive undertakings that can take teams of researchers months or even years to complete. A new study published in PLOS Digital Health suggests that unmodified pretrained large language models may be able to shoulder much of that burden, at least in one fast-moving corner of clinical research. The work, led by Taylor B. Harrison, Dian Hu, and colleagues, presents a retrieval-to-abstraction framework that uses a large language model to screen articles, preprocess text and table data, and extract full-text information, all without any task-specific fine-tuning of the underlying model.
The team targeted a domain where the need for agile evidence synthesis is particularly acute: randomized controlled trials (RCTs) that rely on digital health technologies. From wearable activity trackers and smartphone-based interventions to telehealth platforms and remote monitoring sensors, digital tools have multiplied the ways clinical trials collect and report outcomes. That proliferation has outpaced the reporting guidelines designed to standardize trial publication, most notably CONSORT, the widely adopted checklist that establishes common data elements for trial reporting. When researchers attempt to pool results across digital health trials, they routinely encounter substantial heterogeneity in how outcomes are described, measured, and tabulated.
This reporting heterogeneity is more than an inconvenience. Systematic reviews depend on the ability to locate and compare consistent data elements across dozens or hundreds of primary studies. When a trial reports its engagement metrics in one format and its equity-related participant characteristics in another, or omits key technology-adherence details altogether, reviewers must either exclude the study or spend additional time hunting through supplementary materials. The authors argue that machine-assisted evaluation of reporting adherence offers a pathway toward more nimble information management, one that can accommodate the diverse and disruptive nature of technology-enabled trials as they emerge.
The framework the researchers propose follows a two-stage logic that mirrors how human reviewers work, but compresses the timeline dramatically. In the retrieval stage, the large language model performs article screening, deciding which candidate publications meet inclusion criteria for a given review question. In the abstraction stage, the model handles text and table preprocessing and then conducts full-text review, extracting the specific data elements the reviewers need. Crucially, the approach relies on an unmodified pretrained model, meaning the researchers did not train or fine-tune the model on annotated medical literature. Instead, they designed prompts and workflows that steer the general-purpose model toward reliable performance on specialized tasks.
To test the framework, the team ran two case studies, each examining adherence to a different set of reporting guidelines in digital health technology-enabled RCTs. The first case study evaluated how well trials reported technology-related outcomes, such as measures of participant engagement with the digital intervention. The second focused on digital health equity reporting items, which capture whether trials describe how their technologies perform across diverse populations, including groups that have historically been underrepresented in digital health research. Both case studies were chosen because they represent areas where reporting practices remain inconsistent and where automated assessment could deliver immediate practical value.
The performance results are the study’s headline finding. Across both screening and information-extraction tasks, the large language model achieved F1 scores of at least 0.9, a level the authors describe as on par with inter-annotator agreement, the benchmark that measures how well two trained human reviewers agree with each other on the same tasks. In practical terms, this means the model’s judgments about which articles to include and its extractions of specific data points from full texts were as consistent with a reference standard as a second human expert would typically be. For a tool that requires no custom training, that level of agreement represents a meaningful threshold for real-world deployment.
The implications extend well beyond the two case studies. Systematic reviews of RCTs underpin clinical practice guidelines, regulatory decisions, and health technology assessments, yet the volume of published trials continues to grow faster than the workforce available to review them. A framework that can reliably automate screening and abstraction could shorten the lag between the publication of primary evidence and its incorporation into synthesized knowledge. In digital health specifically, where interventions and technologies evolve on timescales of months rather than years, that acceleration could determine whether reviews remain relevant by the time they are published.
The study also carries a pointed message for trial investigators and journal editors. The authors demonstrate the importance of common reporting elements across digital health technology-enabled RCTs, and their automated assessment of reporting adherence makes gaps visible at scale. If a machine can rapidly quantify how consistently trials report technology engagement or equity-related items, then funders, guideline developers, and editorial boards gain a new instrument for monitoring compliance and identifying where CONSORT extensions or digital-health-specific checklists need strengthening. Better reporting, in turn, feeds back into better synthesis, creating a virtuous cycle that the framework is positioned to support.
Methodologically, the retrieval-to-abstraction design reflects a broader shift in how researchers are approaching large language models in scientific workflows. Rather than building bespoke models for each task, the authors show that a single pretrained model, guided by carefully structured prompts and pipeline stages, can handle the heterogeneous formats found in real published literature, including the tables and supplementary text where trial outcomes often hide. The preprocessing steps that convert raw text and tables into model-readable inputs are a key part of the pipeline, addressing one of the persistent obstacles to automating literature review: the sheer variability of how information is presented across journals and study designs.
Cautious optimism seems the appropriate stance. The study reports case demonstrations in a specific domain rather than a comprehensive validation across all of medicine, and the authors frame their findings as supporting the scalability and rigor of accelerated evidence synthesis, not as proof that human reviewers are obsolete. Human expertise remains essential for defining review questions, adjudicating ambiguous cases, and interpreting clinical significance. But the demonstration that an unmodified large language model can match inter-annotator agreement on both screening and extraction in digital health trial reviews marks a tangible step toward a future in which the evidence base for clinical decisions is updated as quickly as the technologies it evaluates. As digital health continues to reshape how trials are run, the tools for reading those trials may finally be catching up.
Subject of Research: Large language model-assisted literature synthesis for evaluating reporting adherence in digital health technology-enabled randomized controlled trials
Article Title: Retrieval-to-abstraction LLM-assisted literature synthesis: Case demonstrations in digital health technology-enabled randomized controlled trial outcome reporting
Article References: Harrison, T. B., Hu, D., Jia, H., Tan, S. H., Chen, F., Zhang, Z., Lu, Q., Wang, J., Wang, L., Huang, M., Prokop, L. J., Hoy, M., St. Sauver, J., Fu, S., & Liu, H. (2026). Retrieval-to-abstraction LLM-assisted literature synthesis: Case demonstrations in digital health technology-enabled randomized controlled trial outcome reporting. PLOS Digital Health, 5(10), e0001768. https://doi.org/10.1371/journal.pdig.0001768
Image Credits: AI Generated
DOI: 10.1371/journal.pdig.0001768
Keywords: large language models, systematic reviews, randomized controlled trials, digital health, CONSORT reporting guidelines, evidence synthesis, natural language processing, clinical trial reporting, digital health equity, information extraction, article screening, PLOS Digital Health
News Source: Ophelia Keating. (October 10, 2026). AI Reads the Trials: LLM Framework Speeds Systematic Reviews of Digital Health RCTs. Scienmag.



