Tens of thousands of artificial intelligence diagnostic tools have now been authorised for clinical use around the world, and the industry that produces them is estimated to be worth roughly 13.7 billion US dollars. Yet according to a new opinion piece published in PLOS Digital Health by Alex Mirugwe of Makerere University’s School of Public Health, the central assumption underpinning this entire market — that strong performance on a benchmark test translates into real benefit for patients — has never been adequately tested. The number that justifies most authorisations, benchmark accuracy, is simply a measure of how well an algorithm performs on a fixed, curated dataset. It says almost nothing about what the tool will do when confronted with a genuinely sick person in a busy ward, an under-resourced clinic, or a population whose biology the model has never seen.
The evidence gap is striking. A 2025 meta-analysis in npj Digital Medicine examined 83 studies of generative AI diagnostic models, specifically large language and multimodal models rather than medical AI as a whole, and found a pooled diagnostic accuracy of just 52.1 percent, with a 95 percent confidence interval running from 47.0 to 57.1 percent. That figure sits 15.8 percentage points below expert physicians, a statistically significant shortfall. The authors of the meta-analysis caution that the included studies were highly heterogeneous and most carried a high risk of bias, so the pooled number should be read as an indicator of how weak the evidence base in this subfield is, rather than as a precise estimate of how these systems actually perform. Even so, it is a sobering baseline for a technology being deployed at scale in hospitals.
Two real-world sepsis prediction systems illustrate how dramatically outcomes can diverge. The Epic Sepsis Model, deployed across health systems worldwide, was reported to identify high-risk patients in 87 percent of cases. But when researchers restricted the analysis to data available before clinicians themselves suspected sepsis, that figure fell to 63 percent, and to 53 percent before a blood culture was ordered. In other words, the model had largely learned to detect clinician suspicion rather than the underlying physiology of the disease. In practice it missed roughly two-thirds of cases and generated so many false alarms that staff at one major health system acknowledged only 13 percent of its alerts in 2023 — a figure that reflects alert fatigue and workflow burden as much as raw model failure, but which renders the system nearly useless at the bedside.
Contrast that with TREWS, the Targeted Real-time Early Warning System developed at Johns Hopkins with sustained input from nurses and intensivists. Because it was built around clinical workflow from the start, a prospective multisite trial was able to demonstrate genuine patient benefit: reduced sepsis mortality, shorter duration of organ failure, and more timely administration of antibiotics. Mirugwe is careful to note that the two systems differ on several dimensions, including training data, the definition of the prediction target, deployment context, and institutional support, so the contrast cannot be reduced to a single cause. The clearest methodological difference, however, is that TREWS treated clinical co-design as a technical requirement from the outset rather than as an optional consultation, and this appears to be a major contributing factor in its more favourable real-world performance.
Beneath these individual failures lies a structural one: the datasets on which diagnostic algorithms are trained frequently do not reflect the populations they serve. Representation bias has become a dominant failure mode in medical AI. In a decade-long review of authorised devices, fewer than 30 percent disclosed the demographic composition of their training data, although disclosure norms have been shifting over the period examined. For much of that cohort, clinicians had no way of knowing whose biology the algorithm had actually learned from. The consequences of this opacity are not hypothetical. MIT researchers showed in 2024 that popular diagnostic models analysing chest radiographs had acquired a superhuman ability to predict a patient’s race, sex, and age from images alone. The models were exploiting demographic shortcuts that correlated with pathology in the training cohort — correlations that reliably break down when the tool is applied elsewhere.
Geography compounds the problem. A diabetic retinopathy algorithm validated at 87 percent sensitivity delivered only 60 to 80 percent sensitivity in public-health screening settings in northern India, with specificity falling as low as 14 percent. Part of the gap is attributable to the lower image quality of the non-mydriatic cameras used in routine screening compared with the pristine images used during validation. Same disease, same imaging modality, different population — and effectively a different machine. Tools trained at tertiary academic centres carry case-mix assumptions that simply do not hold in district hospitals in sub-Saharan Africa or rural Southeast Asia, precisely the settings where AI is most aggressively promoted as a solution to chronic specialist shortages.
A second structural failure concerns who builds these systems. AI diagnostic tools are typically constructed by engineers, evaluated by statisticians, and regulated by agencies whose frameworks were originally designed for drugs and static devices. The clinician who must integrate the model’s output with a patient’s history, physical examination, and ward context is routinely absent from the development process. The predictable results are alert fatigue, automation bias, and the silent override, in which clinicians quietly learn to ignore recommendations they have found unreliable. Evidence suggests that teams combining clinicians and data scientists produce more generalisable models, because clinicians understand which features are genuinely predictive and which are confounded by patterns of care. Excluding that knowledge, Mirugwe argues, is not a resourcing choice but an epistemic error.
The third failure is regulatory. Frameworks vary across jurisdictions but share one weakness: they use pre-market accuracy as the gateway to deployment without requiring prospective validation in the populations who will actually use the tool. In the United States, 96.4 percent of AI medical devices were cleared through the 510(k) pathway without any clinical trial evidence. Of the 950 AI-enabled devices the Food and Drug Administration had authorised through November 2024, 60 were subject to recall, accounting for 182 recall events in total, and 43 percent of those recalls occurred within a year of authorisation, mostly for diagnostic errors. The European Union’s AI Act of 2024 imposes the most demanding pre-market regime but does not yet specify outcome-level trial requirements. China requires localised trials for Class III devices, India’s 2025 draft guidance classifies AI imaging as Class C requiring clinical validation, and Japan uses an adaptive framework. Across every jurisdiction, demographic disclosure remains the exception rather than the rule, and lower-income settings remain exposed to tools validated entirely elsewhere.
Mirugwe proposes three achievable reforms. First, prospective external validation on demographically and geographically diverse populations should be a non-negotiable prerequisite for authorisation, with the existing TRIPOD+AI and STARD-AI frameworks already providing the methodological scaffolding. Second, performance must be disaggregated by clinically relevant subgroup: a tool with 92 percent overall accuracy but 74 percent accuracy in one subgroup is selectively unreliable, and the patients it fails are predictable. Demographic disclosure should become universal. Third, clinicians must become structural partners in development from problem definition onwards, not consultees reviewing a finished product. Achieving this at scale would require international coordination, and the International Medical Device Regulators Forum offers a ready mechanism through which harmonised standards could be adopted across jurisdictions. The obstacles are real — much training data is proprietary or retrospective, and privacy-preserving techniques such as federated learning can make full disclosure and prospective validation harder to mandate — but these are arguments for phasing requirements in and developing validation methods suited to distributed data, not for treating clinical readiness as optional.
The gap between algorithmic performance and clinical benefit, the article concludes, is not an engineering problem awaiting a better model. It is a governance problem created by norms that reward accuracy on curated datasets, regulatory pathways that mistake on-paper accuracy for clinical readiness, and a publication culture that treats external validation as optional. For a patient in a Lagos clinic whose retinal image is read by an algorithm trained on a different continent, the difference between 87 percent sensitivity in a trial and 60 percent in the field is a substantially higher risk of a missed diagnosis. An AI diagnostic tool that passes a benchmark test has not passed a clinical one. TREWS demonstrates that the gap can be closed where co-design and rigorous validation are taken seriously, but until regulation, publication standards, and funding criteria embed that distinction, the gap will remain the predictable default of the current system rather than an inevitable one.
Subject of Research: The gap between benchmark accuracy and clinical benefit in authorised AI diagnostic tools
Article Title: When the algorithm passes the test but fails the patient
Article References: Mirugwe, A. (2026). When the algorithm passes the test but fails the patient. PLOS Digital Health, 5(10), e0001774. https://doi.org/10.1371/journal.pdig.0001774
Image Credits: AI Generated
DOI: 10.1371/journal.pdig.0001774
Keywords: medical AI, diagnostic algorithms, benchmark accuracy, clinical validation, regulation, representation bias, sepsis prediction, Epic Sepsis Model, TREWS, FDA, EU AI Act, health equity
News Source: Ophelia Keating. (October 10, 2026). High Test Scores, Poor Bedside Performance: The Hidden Flaw in Medical AI. Scienmag.



