• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Monday, October 5, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Technology

AI Learns to Measure New Bone Growth in Scaffolds, But Standard Metrics Mislead

by
October 5, 2026
in Technology
Reading Time: 5 mins read
0
AI Learns to Measure New Bone Growth in Scaffolds, But Standard Metrics Mislead

AI Learns to Measure New Bone Growth in Scaffolds, But Standard Metrics Mislead

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Deep learning has quietly become one of the most powerful tools in regenerative medicine, and a new study published in Medical & Biological Engineering & Computing shows just how subtle the technology can be when the stakes are measured in micrometers of growing bone. A team led by Sohaila Aboutaleb and Amy Wagoner Johnson at the University of Illinois Urbana-Champaign trained three artificial intelligence models to segment microscopic CT images of synthetic bone scaffolds, the lattice-like implants designed to coax a patient’s own bone to regrow inside damaged defects. Their central finding is a warning for the entire field of biomedical image analysis: the metrics scientists routinely use to grade their AI can be spectacularly misleading, and the fix is to judge the model by the biological quantity it was built to measure in the first place.

The scaffolds at the heart of the study are made of hydroxyapatite, a mineral chemically similar to natural bone, arranged as an orthogonal lattice of rods with pores between 100 and 400 micrometers wide. Implanted into mandibular bone defects in Yucatan minipigs, these structures invite new bone to grow into their macropores, and quantifying how much bone has infiltrated each scaffold is the key to comparing different lattice designs and biological additives. The problem is that bone and hydroxyapatite share nearly identical mineral composition, so their X-ray attenuation values overlap almost completely in micro-CT scans. Ingrown bone also overlaps in brightness with the surrounding soft tissue and noise. Manual labeling of the roughly 1,000 images per scaffold is prohibitively time consuming across studies involving tens of scaffolds, which is exactly where the U-Net convolutional neural network enters the picture.

The U-Net architecture, first introduced in 2015, remains a favorite in biomedical imaging because it delivers strong segmentation from small training sets. Its U-shaped design works like an autoencoder with a memory: an encoding path progressively shrinks the image while doubling its feature channels to detect patterns, a bottleneck connects the two halves, and a decoding path restores the original dimensions, with skip connections threading fine spatial detail from encoder to decoder at every level. The Illinois team used a four-level network starting with eight feature channels and reaching 128 at the bottleneck, ultimately classifying every pixel in a 512-by-512 patch as scaffold, bone, or background. Rather than inventing a new architecture, the researchers held the network constant and varied only the loss function, the mathematical rule that tells the model how wrong its predictions are during training.

Three loss functions were pitted against each other under identical conditions. Cross-entropy loss, the workhorse of classification, evaluates every pixel individually regardless of class balance. Soft Dice loss instead optimizes the overlap between prediction and ground truth, a strategy favored when one class, like bone, occupies only a small fraction of the image. A third, less common option called Accuracy loss directly rewards correct classification of the entire patch. The team trained each model for 1,000 epochs on 225 patches derived from 31 manually labeled images, using four NVIDIA V100 GPUs and confirming that all three losses converged without overfitting. A second expert annotator independently relabeled a subset to gauge human variability, finding a bone Dice score of 0.78 between annotators and an average difference in bone growth fraction of just 1.2 percentage points.

Then came the surprise. Standard evaluation metrics, the numbers researchers typically report to demonstrate that a segmentation model works, produced results that seemed catastrophic but were actually meaningless. Some patches scored a Dice coefficient of zero and a Precision of zero, values usually interpreted as total failure, yet visual inspection showed those patches were segmented perfectly well. The catch: those patches contained no bone in the ground truth at all. When the model predicted a handful of spurious bone pixels in a bone-free image, the Dice and Precision formulas collapsed to zero even though the biological error was negligible. Similar distortions plagued Recall, which tanked on patches where a few missed pixels represented a large fraction of a tiny bone area. Under extreme class imbalance, the familiar metrics were measuring the imbalance, not the model.

To escape this trap, the researchers invented an application-specific endpoint they call delta BGF: simply the difference between the bone growth fraction predicted by the model and the bone growth fraction in the ground truth. This is the number that actually matters for the science, because the entire purpose of segmenting the scans is to quantify bone regeneration and compare treatment groups. Bland-Altman analysis showed the systematic bias between predicted and true bone growth fraction was less than 1.5 percentage points for all three models, comparable to the disagreement between two human experts. When the team correlated the standard metrics against the absolute value of delta BGF, Accuracy emerged as the clear winner, with correlation coefficients reaching negative 0.8 for the Accuracy-loss model and negative 0.78 for the Dice model. A Steiger test confirmed Accuracy’s correlation was significantly stronger than those of Dice, Precision, and Recall in eight of nine comparisons, and bootstrap resampling across 10,000 samples showed the relationship was stable, with confidence intervals lying entirely below zero.

The loss function also shaped where the errors landed, which matters because different regions of the image carry different scientific weight. The Accuracy-loss model tended to overestimate bone, scattering false positives inside the scaffold boundary, precisely where the ingrown bone quantification happens. The cross-entropy and Dice models instead missed bone in the dense ring surrounding the scaffold, errors that are conveniently located outside the region of interest and can be corrected automatically, since scaffold pixels cannot legitimately exist beyond the scaffold boundary. Given this geography of mistakes, the cross-entropy model emerged as the recommended choice for future in vivo studies quantifying spatial variation in bone growth, combining clean class edges within the scaffold with correctable peripheral errors.

Robustness testing added another layer of confidence. The team created augmented test sets by rotating patches 15, 30, and 45 degrees, shifting contrast in both directions, and degrading resolution by factors of 0.66 and 0.5, simulating the natural variation in micro-CT scanning parameters and scaffold orientation. None of these perturbations produced a statistically significant change in Accuracy, Dice, Precision, Recall, or delta BGF. The systematic bias remained below 2.6 percentage points across all augmentations, with contrast changes producing the widest limits of agreement and rotation the narrowest, suggesting that intensity shifts stress the model more than orientation changes. For a field where scanning protocols differ between labs and even between experiments, this resilience means the trained models can plausibly be applied to new datasets without retraining from scratch.

Perhaps the most consequential result is biological rather than computational. Comparing against the group’s earlier in vivo studies, the smallest treatment differences those studies could detect were on the order of 8 to 20 percentage points in bone fraction, vastly larger than the 1.5 percentage point error of the AI segmentation. In other words, the model’s mistakes are too small to obscure genuine biological differences between scaffold designs, meaning future comparisons of bone regeneration should remain statistically sound. The study’s limitations are acknowledged candidly: the 54 test patches derive from only six images, so patch-level statistics treat clustered data as independent, and the authors frame their correlation findings as descriptive rather than confirmatory. Still, the message radiates well beyond bone scaffolds. Whenever deep learning is deployed to extract a quantitative measurement rather than a pretty picture, the evaluation metric should be the measurement itself. A model that aces the standard report card can still flunk the science, and a model that appears to fail it may be quietly doing excellent work.

Subject of Research: Evaluating U-Net deep learning segmentation of bone ingrowth in micro-CT images of hydroxyapatite scaffolds using loss functions and application-specific metrics

Article Title: Segmenting scaffolds with ingrown bone in micro-CT images: considerations of evaluation metrics and loss functions

Article References: Aboutaleb, S., Haug, N., Keni-McCray, P., McCray, A. R. C., Phillips, H., Sharping, S., Cohen, D., Norato, J., & Wagoner Johnson, A. (2026). Segmenting scaffolds with ingrown bone in micro-CT images: considerations of evaluation metrics and loss functions. Medical & Biological Engineering & Computing. https://doi.org/10.1007/s11517-026-03686-x

Image Credits: AI Generated

DOI: 10.1007/s11517-026-03686-x

Keywords: U-Net, semantic segmentation, micro-CT, bone regeneration, hydroxyapatite scaffold, loss functions, evaluation metrics, Dice coefficient, cross-entropy loss, Bland-Altman analysis, deep learning, biomedical imaging

News Source: Denise Maddox. (October 5, 2026). AI Learns to Measure New Bone Growth in Scaffolds, But Standard Metrics Mislead. Scienmag.

Tags: Biomedical ImagingBland-Altman analysisBone regenerationcross-entropy lossdeep learningDice coefficientevaluation metricshydroxyapatite scaffoldloss functionsmicro-CTsemantic segmentationU-Net
Share12Tweet7Share2ShareShareShare1

Related Posts

Machine Learning Cracks the Vast Code of High-Entropy Catalysts

Machine Learning Cracks the Vast Code of High-Entropy Catalysts

October 5, 2026
Eight Minutes of Brahms' Lullaby a Day Boosts Brain Growth in Preterm Babies, Trial Finds

Eight Minutes of Brahms’ Lullaby a Day Boosts Brain Growth in Preterm Babies, Trial Finds

October 5, 2026

Tiny AI Model Reads Tumor Shapes to Map Brain Cancer With a Fraction of the Parameters

October 5, 2026

Tiny Particles, Big Defense: Nanoparticles Shield Concrete From Sulfate Attack

October 5, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.