Millions of people now type their symptoms into chatbots, tap through triage apps, and consult algorithms before they ever book an appointment with a family doctor. Yet a sweeping new analysis suggests that the science underpinning these tools is running far behind their real-world availability. A scoping review published in BMC Medicine by researchers at the University of British Columbia, Simon Fraser University, Harvard T.H. Chan School of Public Health, and Scripps Research Translational Institute has mapped the entire published literature on patient-facing artificial intelligence in primary health care through early 2025, and its findings reveal a striking imbalance: the tools are already in patients’ hands, but the evidence that they are safe and equitable remains thin.
The research team, led by Alireza Habibi and Kathleen McMahon of the University of British Columbia’s Department of Medicine, with co-lead authors Sian Hsiang-Te Tsuei and Jacqueline Kueper, set out to answer a deceptively simple question: what do we actually know about AI tools that patients themselves use in primary care settings? Unlike hospital-based AI, which typically assists clinicians behind the scenes, patient-facing AI interacts directly with the public, offering advice, triage, monitoring, and self-management support. That direct-to-consumer positioning makes the evidentiary stakes unusually high, because errors reach patients without a professional intermediary to catch them.
To build the map, the investigators followed the PRISMA Extension for Scoping Reviews, a rigorous reporting standard designed to make evidence syntheses transparent and reproducible. The search strategy was developed in collaboration with a medical research librarian and preregistered on the Open Science Framework before data collection began, a step that guards against selective reporting. The team combed nine databases and registries, including MEDLINE, Embase, CINAHL, the Cochrane Library, CENTRAL, Web of Science, Google Scholar, ClinicalTrials.gov, and medRxiv, capturing everything from peer-reviewed trials to preprints and registered studies through March 5, 2025. That breadth matters, because digital health research often lives in gray-literature corners that traditional reviews miss.
The numbers tell a story of rapid growth and uneven depth. From 3,469 records identified, 65 studies published between 2016 and 2025 met the inclusion criteria, which required original research on patient-facing AI tools in primary health care reporting health-system or technical outcomes. The field has clearly accelerated: nearly a decade of output fits within a single review, and the pace of publication has climbed steeply since the public release of large language models reignited interest in conversational health technology. Yet the volume of studies is modest relative to the scale of deployment, suggesting that many tools reach the market with little published evaluation at all.
What kinds of tools dominate the landscape? Chatbots lead the pack, appearing in 31 studies, or 48 percent of the included literature. Mobile health applications followed at 24 studies, or 37 percent, and symptom checker or triage tools appeared in 21 studies, or 32 percent. These categories overlap, since many apps embed conversational agents, but the pattern is clear: the technology patients encounter most often is built around dialogue, either typed or spoken, with an algorithm that interprets symptoms, answers questions, and nudges behavior. The dominant use cases were pre-visit triage, addressed in 43 studies or 66 percent, and self-management support, addressed in 42 studies or 65 percent. In other words, AI is positioning itself at the very front door of the health system, deciding who needs care and guiding those who do not.
The methodological profile of the research, however, reveals where the field’s priorities have rested. Validation studies were the most common design, accounting for 20 studies or 31 percent, followed by usability and user-centred design studies at 13 studies or 20 percent. These early-stage evaluations ask whether a tool computes accurately and whether users can navigate it, but they stop short of the questions that matter most for population health. Outcomes reported across the studies clustered around people-centredness, in 41 studies or 63 percent, and technical performance, in 29 studies or 45 percent. By contrast, patient safety was explicitly evaluated in only 5 studies, or 8 percent, and access in just 4 studies, or 6 percent.
That gap is the review’s central alarm. Safety evaluation in digital health typically asks whether a tool causes harm through misdiagnosis, delayed care, inappropriate reassurance, or harmful advice, and whether those risks are distributed fairly across age groups, languages, literacy levels, and socioeconomic strata. Access evaluation asks whether the tool widens or narrows existing inequities, a question of acute importance in primary care, which serves populations least likely to be represented in the training data of commercial AI systems. With only a handful of studies addressing either domain, the field essentially lacks the evidence base needed to certify that patient-facing AI does more good than harm at scale.
The deployment picture makes this evidentiary shortfall more urgent than it might first appear. The review found that 27 of the tools, or 42 percent of those studied, were publicly accessible at the time of the review. Patients are not waiting for randomized trials; they are downloading apps and chatting with symptom checkers today. The authors conclude that robust pre-deployment and post-deployment evidence is urgently needed, supported by institutions that uphold rigorous evaluation across the entire AI lifecycle, from design and training through validation, release, and ongoing monitoring. Post-deployment surveillance is particularly neglected in the literature, even though real-world performance of AI systems can drift as user populations, disease patterns, and underlying models change.
Technically, the review highlights a structural challenge in how these tools are built and tested. Validation studies of symptom checkers and chatbots often measure accuracy against reference standards, but accuracy in a curated test set does not translate directly into safety in a messy primary care population, where patients present with multiple conditions, vague symptoms, and limited health literacy. Usability studies confirm that people can and will use these tools, which is precisely why the absence of safety data is concerning: high adoption combined with unmeasured risk is the profile of a technology that could scale harm as easily as benefit. The authors’ call for lifecycle-spanning evaluation implies embedding safety and equity endpoints into trials and implementation studies, not treating them as afterthoughts once a product has shipped.
For clinicians, policymakers, and patients, the review offers a sobering baseline rather than a verdict of failure. Primary health care is arguably the most consequential arena for patient-facing AI, because it is where most of the world’s health encounters happen and where triage decisions determine whether people reach care in time. The 65 studies mapped here show genuine progress in human-centred design and technical refinement, and the field’s growth since 2016 suggests momentum that could, with the right incentives, produce the safety and access evidence now missing. But the core message is unambiguous: a substantial share of patient-facing AI tools is already live, the published literature has concentrated on early-stage questions of experience and performance, and closing the gap on safety and access is no longer optional. The technology has arrived at the front door of medicine; the evidence base must now catch up to it.
Subject of Research: Patient-facing artificial intelligence tools in primary health care
Article Title: Patient-facing artificial intelligence in primary health care: a scoping review of literature through early 2025
Article References: Habibi, A., McMahon, K., Hsiang-Te Tsuei, S., & Kueper, J. (2026). Patient-facing artificial intelligence in primary health care: a scoping review of literature through early 2025. BMC Medicine. https://doi.org/10.1186/s12916-026-05223-x
Image Credits: AI Generated
DOI: 10.1186/s12916-026-05223-x
Keywords: artificial intelligence, primary health care, scoping review, chatbots, symptom checkers, mHealth, patient safety, digital health, triage, self-management, health equity, BMC Medicine
News Source: Ophelia Keating. (October 6, 2026). AI Health Tools Reach Patients Before the Evidence Catches Up, Review Finds. Scienmag.



