When people feel a persistent cough or chest tightness coming on, most do not call a doctor first. They reach for their phone. A new randomized trial from China suggests that what they find on that phone could soon change dramatically, and for the better. In a study published in Nature Health, researchers report that a GPT-4o-based chatbot called LungDiag, embedded directly inside the WeChat messaging platform, significantly improved the ability of ordinary people without medical training to work through respiratory illness cases compared with conventional mobile web search. The findings offer some of the strongest prospective evidence yet that purpose-built conversational artificial intelligence can outperform the search engines that billions of people already rely on for health answers.
The trial was ambitious in scale and design. Between April and June 2025, the research team recruited 2,400 adults with no formal medical education across 24 healthcare systems in China. Participants were randomly assigned in a one-to-one ratio to either the LungDiag chatbot or a browser-based search control, with dedicated AI and generative AI-assisted search functions deliberately disabled in the control arm to ensure a fair comparison. The study was single-blind, prospective, non-interventional and multicentre, and it was preregistered, a detail that matters enormously in a field where retrospective evaluations of large language models have often been criticized for flexibility in how results are reported. Each participant worked through simulated clinical vignettes describing respiratory conditions, and the primary endpoint was item-level accuracy across four domains: identifying risk factors and aetiology, reaching a preliminary diagnosis, and deciding on the appropriate triage level.
The statistical machinery behind the analysis was correspondingly rigorous. Rather than simply pooling correct and incorrect answers, the investigators used item-level generalized linear mixed-effects models with random intercepts for both participants and items, an approach that accounts for the fact that some people are naturally better at these tasks than others and that some vignettes are inherently harder. In the primary compliant-case analysis, 2,176 questionnaires were included, with 1,088 in each group. The result was clear: LungDiag achieved a model-based overall accuracy of 70.0 percent, with a 95 percent confidence interval of 65.2 to 74.8 percent, compared with 55.4 percent for the web search control, whose interval ran from 49.7 to 61.1 percent. That corresponds to a risk difference of 14.6 percentage points, a risk ratio of 1.26, and a P value below 0.001.
What is perhaps most striking is where the advantage came from. The largest gap between the two tools appeared on preliminary diagnosis, where the chatbot-assisted group reached 81.2 percent accuracy against just 54.5 percent for the search group, a difference of nearly 27 percentage points. In other words, when the task was to figure out what condition a described patient might have, the constrained AI assistant was dramatically more effective at guiding laypeople to the right answer. Gains on triage, the task of deciding how urgently care is needed, were real but more modest: 53.4 percent versus 47.1 percent. This asymmetry is scientifically interesting because it suggests that a curated respiratory knowledge layer layered on top of a general-purpose language model can sharpen diagnostic reasoning more readily than it can resolve the genuinely difficult judgment calls about urgency that even clinicians sometimes disagree on.
That triage limitation deserves close attention. For binary high-urgency triage, deciding whether a case requires immediate emergency attention, LungDiag showed high sensitivity at 85.6 percent but only moderate specificity at 62.1 percent. The control arm performed worse on both counts, at 80.9 percent and 56.0 percent respectively. High sensitivity means the chatbot rarely missed a true emergency, which is the safer failure mode, but the moderate specificity implies it flagged many non-emergencies as urgent. In a real-world deployment, that trade-off could funnel worried-well users toward emergency departments that are already strained, a consequence the authors themselves acknowledge the study cannot fully evaluate. The trial was a simulated-case comparison, and the researchers are explicit that the findings do not establish effects on actual care-seeking behavior, diagnostic delay, or clinical outcomes.
There is also a cost to the accuracy, and it is measured in seconds. Median task completion time was 527 seconds for the LungDiag group versus 461 seconds for the control, roughly a minute longer per case. In a triage context, where a user may be anxious or where minutes could matter, that overhead is not trivial. It likely reflects the conversational nature of the tool, which asks clarifying questions and walks users through structured reasoning rather than dumping a list of links. Whether users would tolerate that friction in real life, and whether the extra minute buys enough accuracy to be worth it, are questions the trial design cannot answer but that any deployment would need to weigh.
The architecture of LungDiag is what separates it from simply asking a chatbot for medical advice. The system is task-constrained, meaning it is deliberately limited to respiratory assessment rather than open-ended conversation, and it incorporates a respiratory knowledge layer derived from the team’s earlier LungDiag work on electronic health records across multiple centers. This retrieval-augmented approach grounds the model’s outputs in curated clinical resources rather than relying solely on patterns learned during pretraining, which is one plausible explanation for why laypeople guided by the tool made fewer diagnostic errors than those sifting through search results on their own. The exact internal prompt text, retrieval-engine settings, and knowledge layer resources were not publicly released because they contain proprietary implementation details, though the functional specification and architecture are described in the supplementary materials, and de-identified participant-level datasets and analysis code are openly available on GitHub and Zenodo.
The context of respiratory disease in China gives the study particular urgency. Chronic respiratory conditions impose an enormous and growing burden globally, and lower respiratory infections remain among the leading causes of death worldwide according to the Global Burden of Disease project. In a country where WeChat functions as a near-universal communication layer, embedding a competent triage assistant directly into an app that more than a billion people already open every day could reach populations that formal healthcare access struggles to serve, particularly in rural regions. The trial’s 24-site footprint across diverse Chinese healthcare systems, and subgroup analyses by sex, age, education and region, suggest the team was attentive to exactly this question of generalizability.
The study also sits within a rapidly accumulating evidence base on human-AI interaction in medicine. Recent randomized trials have shown that GPT-4 assistance can improve physician performance on patient care tasks, while other work has documented both the promise and the pitfalls of large language models in clinical decision-making, including failures on tasks requiring integration of longitudinal context. What distinguishes the new study is its focus on laypeople rather than clinicians, and its use of a nationwide messaging platform rather than a standalone application. Most people who seek health information online will never consult a medical AI through a dedicated clinical portal; they will use whatever is already on their phone. Testing the technology in that realistic environment, against the realistic alternative of web search, is what makes the 14.6 percentage-point improvement meaningful.
Cautions remain, and the authors do not shy away from them. Simulated vignettes are not sick patients; they compress clinical reality into text and cannot capture the physical examination findings, hesitation, and ambiguity of real presentations. Accuracy on vignettes does not guarantee that a chatbot will reduce dangerous delays in seeking care, and questions of medical liability, regulatory oversight, and the black-box nature of large language models loom over any deployment. Still, the trial sets a methodological benchmark: a preregistered, randomized, adequately powered, transparently analyzed comparison with public data and code. As health systems worldwide grapple with surging demand and clinician shortages, the message from this study is cautiously optimistic. A carefully constrained AI assistant, built on a general-purpose model but grounded in domain knowledge and delivered through the messaging apps people already use, can help non-experts reason about respiratory illness substantially better than the search bar they would otherwise turn to. The next test is whether that advantage survives contact with the messy reality of actual patients, actual anxiety, and actual emergency rooms.
Subject of Research: Accuracy of a GPT-4o-based respiratory triage chatbot versus web search for layperson diagnosis and triage
Article Title: Accuracy of a respiratory assessment chatbot in a nationwide messaging system for layperson diagnosis and triage: a randomized preregistered study
Article References: Liang, J., Wang, Y., Ni, J., Cai, Y., Chen, K., Wang, J., Shi, Y., Chen, B., Dong, C., Guo, T., Huang, J., Huang, Z., Guo, Z., Li, J., Liu, B., Pu, Z., Qi, S., Sun, R., Wang, R., … He, J. (2026). Accuracy of a respiratory assessment chatbot in a nationwide messaging system for layperson diagnosis and triage: a randomized preregistered study. Nature Health. https://doi.org/10.1038/s44360-026-00189-9
Image Credits: AI Generated
DOI: 10.1038/s44360-026-00189-9
Keywords: artificial intelligence, chatbot, GPT-4o, respiratory disease, triage, diagnosis, randomized controlled trial, WeChat, large language models, digital health, layperson health literacy, Nature Health
News Source: Ophelia Keating. (October 5, 2026). AI Chatbot Beats Web Search at Helping Laypeople Diagnose Respiratory Illness. Scienmag.



