• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Tuesday, September 22, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Health

AI Forces Medicine to Rethink What Disease, Labels, and Evidence Mean

Bioengineer by Bioengineer
September 21, 2026
in Health
Reading Time: 6 mins read
0
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Artificial intelligence is transforming medicine at breathtaking speed, diagnosing heart disease from electrocardiograms, flagging sepsis hours before clinicians suspect it, and drafting clinical notes in seconds. Yet a new commentary in the Journal of Medical Systems argues that the hardest problems facing clinical AI are not computational at all. They are philosophical. Antonis A. Armoundas of Massachusetts General Hospital and Harvard Medical School contends that the deepest obstacles to trustworthy medical AI lie in contested notions that clinicians use every day as though they were stable: what counts as a disease, where a diagnostic threshold should fall, and what kind of evidence justifies acting on a prediction. As health systems embed AI outputs into triage, documentation, risk prediction, and care pathways, these hidden conceptual choices increasingly determine whether algorithms help or harm patients.

The starting point of the argument is that the labels on which AI models are trained are not neutral descriptions of reality. Medical terminologies such as the International Classification of Diseases, used for reporting and billing, and SNOMED CT, used for clinical representation, were built as harmonization and interoperability tools. They standardize communication, reimbursement, and surveillance, but they also embed purpose-specific assumptions about what counts as the same disease. The same diagnostic label is routinely used for population reporting, bedside decision support, and research cohort definition, even though these tasks demand very different levels of granularity, validity, and explanatory scope. When supervised models are trained on diagnosis codes, the commentary warns, they inherit the structure and limitations of those codes and may end up learning disease as it is encoded rather than disease as it is clinically observed.

This mismatch is not merely theoretical. A clinician deciding whether a patient’s blood pressure crosses the threshold for hypertension is making a judgment shaped by time constraints, available resources, risk tolerance, and the downstream harms of labeling someone with a chronic condition. The threshold that turns a continuous variable into a diagnosis is rarely purely objective; it is a decision made under constraints. Classification matters because it determines the next clinical action, whether that is treatment, reassurance, or intensified surveillance. AI does not dissolve those choices, but it can make the trade-offs visible, especially when labels originally designed for coding are quietly reused for care. Models optimized for surveillance or billing may consequently misclassify clinical need even while achieving impressive accuracy statistics.

The commentary draws on a long-running debate in the philosophy of medicine about whether disease can be defined purely biologically or whether it necessarily includes evaluative components. The naturalist account associated with philosopher Christopher Boorse centers disease on dysfunction relative to species-typical functioning, whereas Jerome Wakefield’s influential harmful dysfunction analysis treats disorder as a hybrid of biological dysfunction and judged harm. These abstract debates surface in daily practice because many diagnoses that appear as categorical labels are actually derived from continuous quantitative measures: blood pressure, glucose, ejection fraction, bone density, symptom scales. The history of overweight and obesity illustrates how such boundaries can shift over time, with substantial clinical and social consequences for the people reclassified on either side of the line.

Because clinical action is discrete while physiology is continuous, thresholds are unavoidable. The practical question, the author argues, is whether a given threshold supports clinical decisions or masks clinically important heterogeneity. Three evaluation strategies can make threshold effects measurable. Threshold sensitivity analyses reveal how changing a cut-off alters diagnosis rates, treatment eligibility, false positives, false negatives, and subgroup performance. Multimodal and longitudinal models can identify heterogeneity within a labeled condition, distinguishing patients who share a diagnosis code but differ in underlying mechanism, disease trajectory, or likely treatment response. Decision-analytic methods can then test whether an alternative threshold or representation genuinely improves clinically relevant choices, such as when to monitor, treat, escalate therapy, or allocate preventive resources. In this framing, AI does not replace clinical judgment; it exposes the consequences of classification-based judgment to scrutiny.

Taking continuity seriously leads to a more radical proposal. Petersen and Ursin, cited in the commentary, argue that forcing complex states into diagnostic classes creates an information bottleneck that discards clinically meaningful variation and can propagate historical and social biases when treated as ground truth for model development. They advocate continuous disease assessment: modeling severity, progression, and treatment response as evolving estimates in physiological space, with a diagnosis treated as a revisable hypothesis rather than an endpoint. Longitudinal modeling, multimodal representation learning, and time-series prediction can estimate latent states and trajectories from heterogeneous data streams including laboratory results, imaging, wearable signals, and clinical text. The key philosophical move is to treat disease as state estimation under uncertainty rather than association with a class of codes, which is closer to how experienced clinicians actually update their beliefs over time. Network medicine reinforces this shift, treating disease phenotypes as products of interacting biological networks rather than isolated categories.

Evidence itself, the commentary insists, is a classification problem. Evidence-based medicine ranks studies through appraisal tools such as GRADE for rating certainty of evidence, RoB 2 for assessing risk of bias in randomized trials, and PROBAST for prediction model studies, with the recent PROBAST+AI extension addressing machine learning prediction tools. Yet empirical work shows that different appraisal tools applied to the same studies can yield divergent judgments, and even within a single tool, agreement between assessors can be limited. AI tools can support evidence appraisal by extracting study features, enabling sensitivity analyses, and making uncertainty explicit, but they can also embed appraisal rules in workflows that are difficult to scrutinize. When large language models are used for evidence screening, summarization, or quality appraisal, their outputs may vary with model version, prompt structure, evidence sources, settings, and repeated runs, demanding stability measurement and human adjudication before informing evidence judgments.

The stakes become vivid in the case of proxy endpoints. A landmark study by Obermeyer and colleagues showed that an algorithm using future healthcare costs as a proxy for health need produced substantial racial bias, because costs reflect differential access and spending rather than morbidity. A model can be accurate for its measured endpoint while still misclassifying clinical need if the endpoint, data source, and subgroup performance are never examined directly. The commentary argues that endpoint specification must be treated as a primary methodological step: developers should define the clinical concept of interest, justify any proxy, and evaluate whether the proxy behaves similarly across subgroups and settings. When differences reflect structural inequities such as unequal access to digital tools or care, statistical adjustment alone is insufficient; measurement and deployment must be redesigned so that model outputs reflect clinical need rather than access or utilization. Related systematic reviews in perioperative medicine, pediatric anesthesia, and cardiology have found high risk of bias, limited external validation, and uncertain generalizability as persistent barriers to clinical integration.

Diagnostic reasoning, on this account, is not label assignment but abductive hypothesis generation and refutation: clinicians propose explanations, seek discriminating evidence, revise as data accumulate, and treat diagnoses as provisional. Clinical AI should mirror this structure by reporting uncertainty, supporting updating as new data arrive, and aligning predictions with the patient’s course over time rather than offering one-time diagnostic accuracy. Precision medicine also demands causal understanding of which interventions will help which patients, a requirement that philosophers Francesco Russo and John Williamson argue depends on combining mechanistic evidence with difference-making probabilistic evidence, neither of which suffices alone. Causal inference frameworks formalize the distinction between association and intervention effects, clarifying the assumptions needed to translate observational data into causal claims.

The practical payoff is what the commentary calls decision-centered precision medicine. Developers should first define the clinical decision a model supports, then define the clinical concept of interest and disclose whether a direct measure or a proxy label is used, evaluate stability across different electronic health record systems and patient populations, and finally assess decision-relevant outcomes rather than label accuracy alone, including effects on calibration, missed diagnoses, safer treatment, and equity. Health systems selecting AI tools should ask what decision the tool supports, what endpoint was used, whether external validation covered a population like theirs, whether subgroup performance was reported, and whether uncertainty is displayed in a usable form. After deployment, calibration, errors, subgroup performance, and patient-important outcomes should be monitored continuously. In this vision, implementation is not the final step after development; it is part of the evidence process itself. By turning conceptual ambiguities into testable, governable design choices, AI may ultimately sharpen medicine’s understanding of its own most basic categories.

Subject of Research: Philosophical foundations of clinical artificial intelligence and precision medicine, including disease classification, diagnostic thresholds, and evidence appraisal

Article Title: Clinical AI and Precision Medicine: Philosophical Questions About Labels, Disease, and Evidence

Article References: Armoundas, A. A. (2026). Clinical AI and Precision Medicine: Philosophical Questions About Labels, Disease, and Evidence. Journal of Medical Systems, 50(1), Article 130. https://doi.org/10.1007/s10916-026-02457-3

Image Credits: AI Generated

DOI: 10.1007/s10916-026-02457-3

Keywords: clinical AI, precision medicine, disease classification, diagnostic thresholds, medical labels, proxy endpoints, evidence-based medicine, algorithmic bias, large language models, continuous disease assessment, electronic health records, health equity

Cite Scienmag News
APA MLA Chicago

Ophelia Keating. (September 21, 2026). AI Forces Medicine to Rethink What Disease, Labels, and Evidence Mean. Scienmag. https://scienmag.com/ai-forces-medicine-to-rethink-what-disease-labels-and-evidence-mean/

Ophelia Keating. “AI Forces Medicine to Rethink What Disease, Labels, and Evidence Mean.” Scienmag, 21 September 2026, https://scienmag.com/ai-forces-medicine-to-rethink-what-disease-labels-and-evidence-mean/. Accessed 21 September 2026.

Ophelia Keating. “AI Forces Medicine to Rethink What Disease, Labels, and Evidence Mean.” Scienmag. September 21, 2026. https://scienmag.com/ai-forces-medicine-to-rethink-what-disease-labels-and-evidence-mean/

Copy citation Download RIS

Tags: algorithmic biasclinical AIcontinuous disease assessmentdiagnostic thresholdsdisease classificationelectronic health recordsevidence-based medicinehealth equitylarge language modelsmedical labelsPrecision medicineproxy endpoints

Share12Tweet7Share2ShareShareShare1

Related Posts

Two Immune Drugs Combined Achieve Dramatic Recovery in Stubborn Psoriatic Arthritis Case

September 21, 2026

Engineered Cytisine Derivative Disrupts Redox Balance to Kill Lung Cancer Cells

September 21, 2026

Sports Foods Are Ultra-Processed: New Review Weighs Athlete Health Against Performance

September 21, 2026

Blocking Myostatin Rebuilds Wasted Muscle in Dynamin 2 Myopathy Mice, But Power Lags Behind

September 21, 2026

POPULAR NEWS

  • Free Open-Source Workflow Matches Costly Software in Screening Bauxite Residue’s Climate Footprint

    29 shares
    Share 12 Tweet 7
  • Two Immune Drugs Combined Achieve Dramatic Recovery in Stubborn Psoriatic Arthritis Case

    29 shares
    Share 12 Tweet 7
  • Waste Catalyst Turned Air Purifier Destroys Formaldehyde at Room Temperature

    29 shares
    Share 12 Tweet 7
  • Lung Bacteria Found to Worsen Influenza Through a Metabolic Trap in Immune Cells

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Free Open-Source Workflow Matches Costly Software in Screening Bauxite Residue’s Climate Footprint

Two Immune Drugs Combined Achieve Dramatic Recovery in Stubborn Psoriatic Arthritis Case

Waste Catalyst Turned Air Purifier Destroys Formaldehyde at Room Temperature

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.