• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Wednesday, October 7, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Health

AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy

by
October 7, 2026
in Health
Reading Time: 5 mins read
0
AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy

AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Every year, millions of patients undergo whole-body bone scintigraphy, a nuclear medicine scan that reveals hotspots of abnormal bone activity caused by cancer metastases, fractures, or benign disease. Deciding who truly needs the scan, and how urgently, has long depended on clinical intuition. Now a team at Dongzhimen Hospital of Beijing University of Traditional Chinese Medicine has built machine learning models that attempt to forecast the scan’s diagnostic outcome before a single image is taken, offering clinicians a statistical preview of what the camera will find. The study, published in BMC Medical Imaging, compares two very different modeling philosophies and delivers a sober lesson about how the passage of time can quietly undermine even well-built predictive tools.

The research team, led by Kuo Ma, Ziyang Du, and corresponding author Peilin Wu, assembled a retrospective cohort of 3,969 patients scheduled for whole-body bone scintigraphy. After excluding eleven cases with missing values, 3,958 patients remained. Each patient was assigned to one of three diagnostic categories: bone metastasis, fracture, or benign bone disease. The goal was ambitious in its simplicity: using only eleven clinical features available before the scan, could an algorithm reliably predict which of the three categories a patient would ultimately fall into? Such a tool would give physicians an adjunctive reference for risk stratification, potentially flagging high-risk patients for expedited imaging or additional workup.

The methodological design is where the study distinguishes itself from much of the medical machine learning literature. Rather than splitting patients randomly into training and test sets and calling it a day, the authors built a dual validation architecture. The entire dataset was first divided chronologically, by examination date, into a temporal training set and a temporal validation set at an 8:2 ratio. The temporal training set was then further partitioned randomly into a random training set and a random test set, again at 8:2. This two-tier structure allowed the team to measure not just how well the models performed on patients resembling their training data, but how they held up on patients examined later in time, a far more realistic simulation of clinical deployment.

Two algorithms were pitted against each other. The first, multinomial logistic regression, is a classical statistical workhorse that extends binary logistic regression to multiple classes, estimating the probability of each diagnostic category as a function of the input features through a set of linear equations. Its virtues are transparency and stability: every coefficient can be inspected, and the model’s behavior is fully determined by its parameters. The second, XGBoost, is a gradient-boosted decision tree ensemble that has dominated machine learning competitions for a decade. It builds hundreds of shallow trees sequentially, each one trained to correct the residual errors of its predecessors, capturing nonlinear relationships and feature interactions that a linear model cannot. XGBoost typically wins on raw predictive power but offers less interpretability and more opportunities for overfitting.

Hyperparameters for both models were optimized on the random training set using grid search combined with 5-fold cross-validation, a procedure that repeatedly divides the training data into five folds, trains on four, and validates on the fifth, averaging results to select the most robust parameter combination. With optimal parameters locked in, the models were retrained on both the random training set and the temporal training set, and final performance was evaluated on the random test set and the temporal validation set respectively. The evaluation went well beyond accuracy, incorporating the F1 score, the area under the receiver operating characteristic curve (AUC), the Brier score, calibration slope, and net benefit derived from decision curve analysis, a metric that asks whether acting on the model’s predictions would help patients more than it harms them.

On the random test set, both models delivered encouraging three-class performance, and, notably, no statistically significant differences emerged between the logistic regression and XGBoost across any evaluated metric, with all p-values exceeding 0.05. This parity is itself an interesting finding: the flexible, nonlinear XGBoost offered no measurable advantage over its transparent counterpart, suggesting that the eleven clinical features carry most of their predictive signal in relationships a linear model can capture. Feature importance analysis, supported by SHAP explanation plots in the supplementary material, identified history of cancer and history of trauma as the two most important predictors, which aligns with clinical expectation that prior malignancy and prior injury are the dominant forces shaping what a bone scan will reveal.

The story grew more complicated when the models faced the temporal validation set. Pre-calibration performance declined to some extent, a phenomenon the authors attribute to temporal drift, the gradual shift in patient characteristics, disease prevalence, and clinical practice that occurs as time passes. Decision curve analysis, which had provided exploratory evidence of potential net benefit on the random test set, showed that this potential benefit diminished in temporal validation. In other words, a model that looked clinically useful when tested on contemporaneous patients became less trustworthy when asked to make predictions for patients examined months later, precisely the situation any deployed model would face.

The team’s response to this drift is arguably the study’s most technically interesting contribution. They applied post-hoc calibration, a statistical correction that remaps the model’s raw probability outputs so they better reflect true observed frequencies. Platt scaling, which fits a simple logistic transformation to the model outputs, was used for the logistic regression, while isotonic regression, a more flexible non-parametric method that fits a monotonic step function, was applied to XGBoost. Crucially, the calibrators were fitted via 5-fold cross-validation on the temporal training set and then applied independently to the temporal validation set, avoiding the leakage that plagues many calibration studies. The results were clear: post-calibration probability reliability improved in the temporal cohort, even as raw discrimination did not recover. The models’ rankings of patients may have stayed roughly intact, but the probabilities attached to those rankings became honest again.

The authors are careful, almost unusually so, about what their findings do and do not establish. They state that calibration may be useful for improving probability reliability, but that the results do not establish temporal generalizability, discrimination improvement, or readiness for clinical deployment, and they emphasize that external validation is required before any real-world use. This restraint matters in a field where prediction models are frequently oversold. A model that predicts diagnostic categories before a scan could, in principle, help prioritize urgent cases, reduce unnecessary imaging, or guide the choice of additional tests, but only if its probabilities remain trustworthy across time and across hospitals, questions this single-center retrospective study cannot fully answer.

The study also carries practical lessons for anyone building clinical prediction tools. First, temporal validation should be treated as a standard requirement, not an optional extra, because random splits systematically flatter model performance. Second, calibration deserves the same attention as discrimination; a model with a respectable AUC can still produce probabilities so miscalibrated that they mislead clinical decisions. Third, the equivalence of logistic regression and XGBoost here is a reminder that simpler models remain competitive in many clinical prediction tasks, particularly when the feature set is small and largely categorical. The work, approved by the Institutional Ethics Committee of Dongzhimen Hospital and conducted under the Declaration of Helsinki with waived informed consent for anonymized retrospective data, received no specific funding and reports no competing interests. As hospitals increasingly experiment with pre-scan triage algorithms, this study offers both a template for rigorous evaluation and a warning: the future arrives one day at a time, and models must be recalibrated to meet it.

Subject of Research: Machine learning prediction of pre-scan diagnostic categories in whole-body bone scintigraphy

Article Title: Development of a prediction model for pre-scan diagnostic categories in bone scintigraphy using multinomial logistic regression and XGBoost

Article References: Ma, K., Du, Z., & Wu, P. (2026). Development of a prediction model for pre-scan diagnostic categories in bone scintigraphy using multinomial logistic regression and XGBoost. BMC Medical Imaging. https://doi.org/10.1186/s12880-026-02848-5

Image Credits: AI Generated

DOI: 10.1186/s12880-026-02848-5

Keywords: bone scintigraphy, machine learning, XGBoost, multinomial logistic regression, prediction model, calibration, temporal validation, bone metastasis, fracture, decision curve analysis, risk stratification, nuclear medicine

News Source: Ophelia Keating. (October 6, 2026). AI Models Predict Bone Scan Outcomes Before Imaging, but Time May Erode Their Accuracy. Scienmag.

Tags: bone metastasisbone scintigraphycalibrationdecision curve analysisfractureMachine Learningmultinomial logistic regressionnuclear medicineprediction modelrisk stratificationtemporal validationXGBoost
Share12Tweet7Share2ShareShareShare1

Related Posts

Landmark Study Maps the Normal Child Heart From Birth to 18 With MRI

Landmark Study Maps the Normal Child Heart From Birth to 18 With MRI

October 7, 2026
Mindfulness Meets Magic Mushrooms: USC Trials Psilocybin Therapy for Depression

Mindfulness Meets Magic Mushrooms: USC Trials Psilocybin Therapy for Depression

October 7, 2026

Coupon Clipping at the Pharmacy Counter: What Manufacturer Discounts Really Do to GLP-1 Drug Spending

October 7, 2026

Fast Walkers May Escape the Cognitive Toll of Aging, Dementia Risk Study Finds

October 7, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.