Large language models are rapidly moving into medicine, drafting answers to clinical questions, summarizing patient histories, and supporting diagnostic reasoning. But the yardsticks used to judge them have lagged behind. A new study published in PLOS Digital Health by Yi Liu and Vijaya B. Kolachalama argues that the standard practice of reporting benchmark accuracy tells us surprisingly little about whether a model’s answer is actually safe or clinically faithful. Their proposed solution, a metric called EntQA, shifts evaluation away from surface-level word matching and toward something closer to what a clinician would check: whether the medically important entities in a patient’s background and in the diagnostic question survive intact in the model’s response.
The problem the researchers set out to solve is subtle but consequential. When a language model answers a medical question, it may produce text that looks plausible and even scores well on automated tests, yet quietly drops a critical detail, a medication name, a comorbidity, a laboratory value, that changes the clinical meaning of the answer. Traditional evaluation metrics, most of them borrowed from general-purpose natural language processing, compare a generated response against a reference answer using token overlap or the similarity of embedding vectors. Those approaches can penalize a correct paraphrase and reward a fluent but clinically hollow response. Worse, many of them require a gold-standard reference answer or laborious manual annotation, which makes them expensive to apply at scale and impractical for the fast-moving cycle of model development that defines modern artificial intelligence.
EntQA takes a different philosophical stance: instead of asking how similar two texts are, it asks which clinically relevant biomedical concepts from the patient background and the question are preserved in the model’s answer. The metric is reference-free, meaning it does not need a pre-existing correct answer to compare against. It extracts entities from the inputs, the patient-specific context and the diagnostic question, and measures how well the generated response retains them. In doing so, it aims to capture two things that benchmark accuracy alone cannot: whether the model’s reasoning preserves patient-specific context, and whether it stays true to the diagnostic intent embedded in the question. Because it relies on entity retention rather than human raters or reference texts, it can be computed automatically across thousands of responses, making it scalable in a way that expert evaluation never will be.
The technical evaluation behind the metric was unusually broad. The authors tested EntQA across five established medical question-answering benchmarks and seven models from the Qwen 2.5 Instruct family, spanning a dramatic range of scale from 0.5 billion to 72 billion parameters. This design allowed them to ask two distinct questions. First, does the metric track actual answer quality, as measured by conventional accuracy? Second, does it track model capability as models grow larger, which is a proxy for the general expectation that bigger models reason better? A useful evaluation metric should correlate positively with both; a misleading one might reward verbosity or fluency regardless of correctness.
The results were striking. Across the benchmarks and model sizes, EntQA showed consistently positive associations with both accuracy and scaling. Group-level correlations with accuracy reached a Spearman rank correlation coefficient of 0.9286, an exceptionally strong relationship that suggests the metric is capturing something fundamental about answer quality rather than incidental stylistic features. Correlations with model scale reached 0.252, a more modest but still positive association, indicating that the metric broadly rises with model capability even if scale alone does not fully determine entity retention. In other words, larger models do tend to hold on to more clinically relevant information, but the metric also discriminates among models of similar size, which is exactly where evaluation is hardest.
Perhaps more revealing than EntQA’s own performance was the behavior of the conventional metrics it was compared against. Overlap-based measures and embedding-based similarity scores frequently exhibited weak or even negative correlations with accuracy across the same experiments. A negative correlation is the most damaging possible outcome for an evaluation metric: it means that in some settings, the metric would systematically prefer worse models over better ones. The authors’ findings suggest that the field’s reliance on these inherited metrics is not merely imprecise but potentially actively misleading when applied to clinical question answering, where the relationship between textual similarity and clinical correctness breaks down in ways it may not in general-domain tasks.
The interpretability of EntQA is one of its most practical advantages. Because the metric is built around named biomedical concepts, a failing score can be traced to specific entities that a model dropped or distorted. An evaluation team can see not just that a model scored poorly, but what it lost, whether it omitted a patient’s hypertension history, ignored a drug interaction, or drifted away from the question’s diagnostic focus. This kind of granular, entity-level feedback is far more actionable for developers than a single aggregate accuracy number, and it aligns naturally with how clinicians audit each other’s reasoning, by checking whether the salient facts of a case were carried through to the conclusion.
The implications extend beyond benchmarking. As healthcare systems begin deploying language models in real clinical workflows, the question of continuous monitoring becomes urgent. Models drift, prompts change, and patient populations differ from benchmark distributions. A reference-free metric that requires no gold-standard answers and no manual annotation can, in principle, run continuously over live model outputs, flagging degradation in clinical fidelity before it reaches patients. The authors position EntQA precisely as such a framework: a scalable and interpretable way to assess clinical fidelity and reasoning quality in healthcare language models without requiring external evidence or reference standards. That combination of scalability and interpretability has been the missing piece in most prior evaluation work.
There are, of course, natural limits to what any single metric can certify. Entity retention is a necessary condition for a clinically reliable answer but arguably not a sufficient one; a model could preserve all the right concepts and still assemble them into flawed reasoning. The moderate correlation with model scale also hints that entity preservation is influenced by factors beyond raw capability, possibly including how training data distribute clinical terminology. The study’s findings, grounded in the Qwen 2.5 model family and five benchmarks, will need extension to other architectures and to open-ended clinical generation tasks beyond question answering. Still, the core result stands: an entity-centric view of evaluation correlates with what the field actually cares about, while the metrics the field has been using often do not.
The study arrives at a moment when regulators, hospitals, and model developers are all searching for trustworthy ways to evaluate medical artificial intelligence. By demonstrating that a reference-free, entity-centric metric can achieve Spearman correlations with accuracy as high as 0.9286 while conventional metrics falter, Liu and Kolachalama have offered the field a concrete alternative to benchmark-accuracy theater. If adopted broadly, the approach could shift the conversation from how often a model picks the right multiple-choice letter to whether it genuinely carries a patient’s clinical picture through its reasoning, a shift that matters far more for the safety of the patients these systems will ultimately serve.
Subject of Research: Entity-centric, reference-free evaluation of large language model responses in medical question answering
Article Title: Entity-centric evaluation of large language model responses for medical question-answering tasks
Article References: Liu, Y., & Kolachalama, V. B. (2026). Entity-centric evaluation of large language model responses for medical question-answering tasks. PLOS Digital Health, 5(10), e0001752. https://doi.org/10.1371/journal.pdig.0001752
Image Credits: AI Generated
DOI: 10.1371/journal.pdig.0001752
Keywords: large language models, medical question answering, EntQA, clinical evaluation, natural language processing, biomedical entities, reference-free metrics, healthcare AI, Qwen 2.5, benchmark evaluation, clinical fidelity, model scaling
News Source: Ophelia Keating. (October 10, 2026). New Metric Judges Medical AI by the Entities It Keeps, Not the Words It Matches. Scienmag.



