Millions of people now ask artificial intelligence chatbots about their symptoms, their medications, and their diets, treating conversational AI systems as round-the-clock medical consultants that never require an appointment. A new perspective paper published in Nature Health argues that this quiet migration of health decision-making onto private AI platforms has created a structural blind spot in modern medicine: when a chatbot delivers harmful advice, the harm is almost impossible for clinicians, researchers, or regulators to detect, measure, or correct. The paper, co-authored by researchers at Binghamton University, Stanford University, Texas A&M University, and Indiana University, contends that this invisibility is not an accidental byproduct of the technology but a built-in feature of how commercial AI systems currently operate.
The central problem the authors identify is one of traceability. When a patient receives bad advice from a physician, there is a medical record, a licensing board, an insurance claim, and often a malpractice pathway that together leave a documented trail. When a patient receives bad advice from a chatbot, the conversation typically remains locked inside the AI company’s platform, the user has no formal mechanism for reporting the problem, and outside researchers have no way to audit what was said or to estimate how often things go wrong. Kaicheng Yang, an assistant professor in the School of Computing at Binghamton University’s Thomas J. Watson College of Engineering and Applied Science, put the situation bluntly: people are getting medical advice from chatbots, but nobody is looking at it, and if there is an error or an issue there is simply no way for anyone to know.
The paper illustrates the danger with a striking real-world case. A 60-year-old man asked ChatGPT how to reduce chloride in his diet, and based on the guidance he received, he replaced table salt with sodium bromide, a chemical cousin that was once used industrially but is toxic to the human nervous system. He was ultimately hospitalized with bromide toxicity after developing hallucinations and paranoia, symptoms that can mimic psychiatric illness and that historically required careful detective work to attribute to bromide exposure. The critical failure in this case, the authors argue, was not only the advice itself but the information vacuum that followed. Because the man could not retrieve the conversation he had with the chatbot, his physicians had no way of knowing what he had been told, making it nearly impossible to reconstruct the pathway that led to his hospitalization.
This information vacuum, the researchers contend, actively prevents oversight rather than merely failing to invite it. In their framing, the harms of AI-generated health advice are not only possible but structurally hidden from the very clinicians, researchers, and regulators who could otherwise detect and correct them. The distinction matters because it reframes the issue from one of occasional bad outputs to one of systemic accountability. A single erroneous answer from a chatbot is a quality-control problem; an ecosystem in which erroneous answers leave no verifiable trace is a governance problem. Without visibility into what chatbots actually tell people about their health, there can be no meaningful epidemiology of AI-related harm, no feedback loop for improving models, and no evidence base for regulation.
The paper also emphasizes that chatbots can be wrong in characteristic ways that make the problem worse. Large language models can generate erroneous answers with unwavering confidence, a property that makes mistakes harder for lay users to spot. They can oversimplify complex medical questions that require individualized clinical judgment. They can draw on outdated information or reproduce false claims that circulated widely in their training data. Each of these failure modes has been documented in evaluations of AI systems on medical question benchmarks, and each becomes more consequential when users treat the output as authoritative guidance rather than as a starting point for a conversation with a qualified professional.
Perhaps most provocatively, the authors argue that these risks no longer require users to seek out health advice at all. AI-generated summaries now appear at the top of search engine results pages, and AI assistants are being integrated directly into social media platforms such as X and Meta’s services. This means that a person scrolling through a feed can be exposed to health claims they never asked for, generated by systems with the same failure modes as standalone chatbots but delivered in a context where users are even less likely to scrutinize them. Yang noted that he is sometimes not even looking for health information, yet it simply shows up while browsing social media feeds, and he and his colleagues believe this could produce undesirable outcomes, particularly when medical misinformation or state actors attempting to manipulate online discussion enter the picture.
Whether a person actively solicits medical guidance from a chatbot or encounters AI-generated health content incidentally, the paper argues that the resulting harm rarely leaves traces that clinicians, regulators, or researchers can verify. This asymmetry between the scale of AI health information delivery and the near-total absence of monitoring infrastructure is what gives the paper its urgency. Traditional public health surveillance depends on detectable signals: emergency room visits, laboratory findings, adverse event reports, patterns in insurance data. The bromide toxicity case was detected only because the patient’s symptoms were severe enough to require hospitalization and because clinicians eventually identified the chemical cause. Countless subtler harms, such as delayed care, inappropriate self-treatment, or misplaced reassurance, would generate no such signal.
To close this gap, the authors propose a package of transparency and accountability measures aimed at different points in the ecosystem. AI companies, they argue, should give users access to their own health-related conversations so that patients can share them with clinicians, allowing physicians to investigate the pathway that led to a harmful event. Those companies should also develop disclosure systems through which health guidance can be flagged, reported, and formally investigated, creating the kind of adverse event reporting infrastructure that exists in pharmacovigilance and medical device regulation. Social media companies, in turn, should take greater responsibility by clearly labeling AI-generated health content and withholding such content until it has been vetted by medical governing bodies. Search engines should restrict the information used in AI-generated summaries to vetted sources. Policymakers, the authors suggest, should consider extending physician malpractice liability frameworks to AI chatbot companies, aligning the legal exposure of these systems with the medical role they have effectively assumed in many people’s lives.
The authors are skeptical that voluntary corporate action will suffice. Yang stated that this is not something society should rely on the companies to do, because their incentive is always to make more money, and building a robust transparency system goes against that incentive. Instead, he argued for some kind of third-party monitoring system, an independent evaluation mechanism that can examine how AI systems handle health queries without depending on the platforms themselves to disclose their own failures. This position echoes long-standing arguments in drug and device regulation, where independent oversight emerged precisely because manufacturers could not be trusted to police themselves when safety and profit conflicted.
The paper arrives at a moment when the volume of AI-mediated health information is growing faster than any mechanism for studying it. As chatbots become embedded in search engines, operating systems, and social networks, the boundary between seeking medical advice and simply consuming content continues to dissolve, and with it the last remaining points where harm could be observed and documented. The authors’ argument is ultimately a call to treat AI-generated health information as a public health exposure worthy of surveillance, reporting requirements, and independent audit. Without such measures, the damage caused by bad chatbot advice will remain what the paper’s title suggests it already is: invisible, not because it does not occur, but because no one has built the instruments to see it.
Subject of Research: The invisible risks and lack of oversight of AI-generated health information from chatbots
Article Title: When AI chatbots give bad advice, no one can see the damage
Article References: When AI chatbots give bad advice, no one can see the damage. (n.d.). Original publication
Image Credits: AI Generated
DOI: Not provided
Keywords: AI chatbots, health misinformation, Nature Health, medical advice, transparency, regulation, patient safety, large language models, social media, accountability, public health, Binghamton University
News Source: Denise Maddox. (October 9, 2026). Chatbot health advice fails invisibly, researchers warn. Scienmag.



