• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Thursday, October 8, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Health

Flawed Benchmarks Hid AI Success in Gene Perturbation Prediction

by
October 8, 2026
in Health
Reading Time: 6 mins read
0
Flawed Benchmarks Hid AI Success in Gene Perturbation Prediction

Flawed Benchmarks Hid AI Success in Gene Perturbation Prediction

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

For the past two years, a cloud has hung over one of computational biology’s most ambitious goals: using deep learning to predict how a cell’s entire gene expression program responds when a gene is switched off, dialed up, or silenced entirely. A string of high-profile benchmarking studies concluded that sophisticated neural networks, including transformer-based foundation models trained on millions of single-cell profiles, failed to beat something almost embarrassingly simple — the mean baseline, which merely averages the perturbed expression profiles seen during training. If a model that ignores the identity of the perturbed gene performs as well as a state-of-the-art neural network, the entire enterprise of in silico genetic screens for drug discovery looks shaky. Now a team at Shift Bioscience, working with Bo Wang of the University of Toronto, argues that the models were never as bad as the benchmarks suggested. The problem, they report in Nature Biotechnology, lies in the rulers used to measure them.

The researchers’ central claim is that the metrics most commonly used to score perturbation prediction models — mean squared error (MSE), mean absolute error (MAE), and the control-referenced Pearson correlation, Pearson(Δ_ctrl) — are frequently miscalibrated. A calibrated metric should reward any predictor that captures genuine perturbation-specific signal and penalize predictors that do not. To test whether a metric meets that standard, the team borrowed a concept from experimental biology: controls. The mean baseline serves as an intuitive negative control, since it contains no perturbation-specific information. But benchmarks had lacked a positive control — a predictor that, by construction, contains real signal — leaving an ambiguity whenever models scored poorly. Low scores could mean the model failed, or simply that the metric was too insensitive to notice success.

The team’s earlier work had proposed a technical-duplicate baseline as a positive control: split the cells from each perturbation into two halves, use one half to predict the other, and see how well that works. Because both halves share the same perturbation-specific signal, the duplicate should outperform the mean baseline if a metric is working properly. On the Norman19 dataset, with roughly 102 differentially expressed genes (DEGs) per perturbation, it did. But on the Replogle22 K562 genome-wide Perturb-seq dataset, where perturbations are far weaker — only about 3.25 DEGs per perturbation on average — the technical duplicate actually underperformed the mean baseline on MSE in roughly 95 percent of individual perturbations. The culprit, the researchers found, is an artifact they call signal dilution: when perturbations touch only a handful of genes out of thousands, error metrics are dominated by the vast transcriptomic background, and a better estimate of unperturbed expression — which the mean baseline provides thanks to greater statistical power — looks more accurate than a prediction carrying real but sparse biological signal.

To build a positive control that works even for weak perturbations, the team introduced the interpolated duplicate. This baseline blends the technical duplicate and the mean baseline gene by gene, using an interpolation weight α derived from the adjusted P value of a differential expression analysis. Genes showing strong statistical evidence of being perturbed are weighted toward the technical duplicate; genes that appear unaffected are weighted toward the mean baseline. The result is a predictor that consistently outperforms the mean baseline across datasets, providing the reliable positive control the field had been missing. With negative and positive controls in hand, the researchers could finally ask the question that had gone unanswered: which benchmarking metrics can actually tell the two apart?

Their answer takes the form of a new meta-metric called the dynamic range fraction, or DRF. For any candidate metric, DRF measures how much of the theoretical gap between a perfect prediction and the negative control is actually covered by the empirical gap between the positive control and the negative control. If the positive control comfortably beats the negative control, the metric is well calibrated and DRF approaches one. If DRF hovers near zero, the metric is essentially blind to perturbation-specific signal, no matter how good the model is. Applied per perturbation, DRF turns metric evaluation from a matter of taste into a measurable quantity — and the results were striking.

Across 14 datasets and 18 metrics spanning direct reconstruction error, reference-based delta metrics, and retrieval-based metrics, the workhorse metrics of the field fared poorly. MSE and Pearson(Δ_ctrl) showed low DRF values in most perturbations of the Replogle22 K562 dataset, meaning they could barely distinguish a signal-bearing predictor from an uninformative one. The evidence was not merely statistical. When the team examined SP2, a transcription factor whose inhibition in Replogle22 K562 yields a DRF for MSE near zero, gene set enrichment analysis of the differential expression ranks still recovered SP2’s own binding targets among the top ENCODE/ChEA gene sets — clear biological signal that MSE simply could not see. In contrast, weighted and rank-based metrics, including weighted MSE (WMSE), weighted R²Δ, and the normalized inverse rank (NIR), consistently showed higher calibration. These metrics share a design principle: they upweight the small fraction of genes that actually respond to a perturbation, rather than letting the silent transcriptome drown the signal out.

With well-calibrated metrics in place, the team re-benchmarked nine models on the task of predicting responses to perturbations held out during training: scGPT, GEARS, PRESAGE, scLambda, CellFlow, and four foundation-model embedding probes built on Geneformer, ESM2, scGPT, and GenePT representations. The reversal was dramatic. Under poorly calibrated metrics such as MSE and Pearson(Δ_ctrl), earlier models like scGPT and GEARS showed no advantage over baselines, consistent with prior benchmarking reports. Under well-calibrated metrics, however, even these earlier models mostly outperformed the uninformative baselines — their perturbation-specific signal had been present all along, merely obscured by the choice of ruler. Newer architectures went further: PRESAGE, an attention-based model that encodes biological prior knowledge, and scLambda, a variational autoencoder with language-model-derived embeddings, often beat the baselines even under poorly calibrated metrics. On every metric tested, at least one deep learning model outperformed all baselines, and the findings replicated in the independent Nadig25 HepG2 dataset.

The combination-prediction task, where models must predict the transcriptomic effect of perturbing two genes at once, told a more nuanced story. On Norman19, the field’s favorite benchmark, the additive baseline — which simply sums the effects of the two single perturbations — proved nearly unbeatable, and the team explains why: about 96 percent of effects in that dataset are additive, and the training set covers only 31 of 4,950 possible gene pairs, or 0.63 percent of the combinatorial space, leaving almost nothing to learn about nonadditive interactions. The additive baseline captured roughly 88 percent of the ideal performance gap on the best-calibrated metrics. On Wessels23, a dataset with ten times the combinatorial coverage, the picture changed: multiple models surpassed the additive baseline under WMSE and weighted R²Δ, and PRESAGE beat it on 15 of 18 metrics. The team recommends that future studies stop using Norman19 as the sole test of combinatorial modeling, since it largely measures arithmetic rather than learned genetic interactions.

Importantly, the researchers checked that their conclusions reflect biology rather than metric trivia. Models that scored well under well-calibrated metrics also performed better on two downstream tasks: pathway recovery, measured by the correlation between gene set enrichment results computed on predicted versus true perturbation effects, and neighborhood-structure recovery, measured by the overlap between nearest-neighbor graphs built from predicted and observed perturbation shifts. Pathway recovery rankings correlated strongly with well-calibrated metric rankings (average Spearman ρ of 0.76 in Replogle22 K562 and 0.84 in Wessels23), and neighborhood recovery tracked retrieval-based metrics most closely. Sensitivity analyses varying cell counts, sequencing depth, and the number of highly variable genes confirmed that well-calibrated metrics retain their advantage even when data quality degrades.

The study does not declare victory for deep learning outright. The authors caution against relying on any single metric, note that the unseen-context task — predicting responses in unobserved cell types or donors — remains unaddressed, and acknowledge that any fixed set of models represents only a snapshot of a fast-moving field. But the reframing is significant. Prior benchmarks played a constructive role in exposing modeling limitations, and architectural advances have begun to address them; an independent preprint by Cole and colleagues reached complementary conclusions about the value of prior-knowledge embeddings. What this work establishes is that the debate over whether neural networks can predict genetic perturbations was, in part, a debate about measurement. With calibration-aware evaluation built on proper positive and negative controls, the models look considerably more capable than the field believed — and the path toward reliable in silico perturbation screens looks considerably shorter.

Subject of Research: Metric calibration for benchmarking deep learning models of genetic perturbation responses

Article Title: Deep learning perturbation models can outperform baselines on calibrated metrics

Article References: Miller, H. E., Mejia, G. M., Leblanc, F. J. A., Swain, B., Wang, B., & de Lima Camillo, L. P. (2026). Deep learning perturbation models can outperform baselines on calibrated metrics. Nature Biotechnology. https://doi.org/10.1038/s41587-026-03307-w

Image Credits: AI Generated

DOI: 10.1038/s41587-026-03307-w

Keywords: deep learning, genetic perturbation, Perturb-seq, benchmarking, metric calibration, single-cell RNA-seq, dynamic range fraction, weighted MSE, foundation models, drug discovery, computational biology, Nature Biotechnology

News Source: Juliet Wilcox. (October 8, 2026). Flawed Benchmarks Hid AI Success in Gene Perturbation Prediction. Scienmag.

Tags: benchmarkingcomputational biologydeep learningdrug discoverydynamic range fractionfoundation modelsgenetic perturbationmetric calibrationNature BiotechnologyPerturb-seqsingle-cell RNA-seqweighted MSE
Share12Tweet7Share2ShareShareShare1

Related Posts

Simple AI on a Wristband Predicts When Women With Chronic Pelvic Pain Will Sit Too Long

Simple AI on a Wristband Predicts When Women With Chronic Pelvic Pain Will Sit Too Long

October 8, 2026
Immune Webs of DNA May Drive Brain Bleeding After High-Sugar Stroke

Immune Webs of DNA May Drive Brain Bleeding After High-Sugar Stroke

October 8, 2026

Hepatitis B and C Remain Hidden Threats Among Men Who Have Sex With Men in Kigali

October 8, 2026

Accepted but Not Final: How Articles in Press Are Reshaping Scientific Publishing

October 8, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.