• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Friday, October 2, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Biology

AI Model Fills the Gaps in Protein Mutation Maps

Bioengineer by Bioengineer
October 2, 2026
in Biology
Reading Time: 5 mins read
0
AI Model Fills the Gaps in Protein Mutation Maps
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Deep mutational scanning has transformed the way biologists interrogate proteins, allowing thousands of amino acid substitutions to be tested in parallel and assembled into detailed maps of variant effects. Yet for all its power, the technique rarely delivers a complete picture. Low sequencing depth, limited coverage, and assay dropout mean that many mutations in a typical experiment simply have no measured score. A new study published in Molecular Systems Biology tackles this persistent problem head-on, introducing a machine learning framework called VEFill that can accurately fill in the missing values of deep mutational scanning datasets and generalize to proteins it has never seen before.

Developed by Polina Polunina and Wolfgang Maier of the University of Freiburg together with Alan Rubin of the Walter and Eliza Hall Institute of Medical Research, VEFill is built on LightGBM, a gradient boosting framework well suited to structured, high-dimensional tabular data. Rather than relying on a single source of information, the model integrates a rich set of biologically informed features: evolutionary conservation scores from the EVE generative model, amino acid substitution matrices such as BLOSUM62 and PAM250, physicochemical descriptors of the wild-type and variant residues, and contextual sequence embeddings from ESM-1v, a transformer-based protein language model trained on millions of natural sequences. The authors trained and optimized the model using Bayesian hyperparameter search and stored all underlying data in a structured PostgreSQL database to ensure reproducibility.

The training ground for VEFill was the Human Domainome 1 dataset, a standardized collection of stability measurements generated with an abundance-based protein fragment complementation assay across 522 human protein domains, using site-saturation mutagenesis to introduce every possible amino acid substitution. From this resource, the team assembled a feature-complete subset of 140 domains comprising 136,854 mutations for the full model, and used the broader set of 521 domains, totaling 562,208 mutations, to train reduced-feature versions that depend only on ESM-1v embeddings and positional mean scores.

The performance figures are striking. In a leave-protein-out evaluation, where the model must predict variant effects for protein domains entirely absent from training, the best configuration achieved a coefficient of determination of 0.64 and a Pearson correlation of 0.80. Feature ablation experiments revealed a clear hierarchy of information: models relying solely on evolutionary scores, substitution matrices, or biochemical descriptors performed poorly, with R-squared values below 0.15. Adding ESM-1v embeddings lifted performance substantially, and combining the embeddings with the mean experimental score per position proved even more powerful, indicating that learned sequence representations and data-driven positional priors capture complementary aspects of mutational tolerance.

Perhaps the most practically important finding concerns data scarcity. In controlled subsampling experiments on 28 high-quality datasets, VEFill consistently outperformed a battery of existing imputation methods, including Envision, FactorizeDMS, AALasso, and nearest-neighbor approaches, once at least 20 percent of variants had been experimentally measured. Below that threshold, predictions became markedly harder, and the authors are candid that true zero-shot prediction without any positional context remains challenging, particularly for functionally complex proteins. This matters because a survey of 841 public score sets from the MaveDB repository showed that fully complete datasets are vanishingly rare, with only 44 out of 841 achieving complete coverage of all possible substitutions. The vast majority of real-world experiments therefore fall squarely within the regime where VEFill delivers its greatest benefit.

The team also probed how few measurements are actually needed per position. When per-protein models were trained with a restricted number of substitutions per site, accuracy rose rapidly and began to plateau at roughly four mutations per position. Intriguingly, models trained exclusively on five carefully chosen amino acids, histidine, glutamic acid, asparagine, isoleucine, and glycine, each representing a distinct physicochemical class, matched the performance of models trained on randomly selected substitutions at equivalent coverage. This suggests that sparse, information-efficient mutational libraries could be designed deliberately, prioritizing a small representative set of substitutions per site without sacrificing predictive power.

Noise ceiling analyses added an important dose of realism. Because experimental variability imposes a fundamental upper bound on achievable accuracy, the researchers simulated replicate experiments using reported measurement uncertainties and found that estimated ceilings ranged from 0.68 to 0.99 across domains. VEFill’s per-protein correlations, between 0.52 and 0.91 under an 80/20 split, sat consistently below but generally close to these limits, indicating that much of the remaining discrepancy between predicted and observed scores reflects measurement noise rather than model failure. Error profiles were lowest for near-neutral variants, where data are densest, and predictions at the extreme deleterious and high-activity tails should be read as conservative approximations rather than precise estimates.

Generalization beyond the training data showed both promise and limits. When the cross-protein model was applied to eight full-length proteins from independent MaveDB datasets, including Parkin, PTEN, aspartoacylase, calmodulin, TPK1, and TP53, it performed substantially better on stability-based assays than on activity-based ones, a mismatch the authors attribute to differences between training and test assay modalities and the complexity of cellular phenotypes. A lightweight two-feature version using only ESM-1v embeddings and positional mean scores performed comparably to the full model in most settings, offering a practical alternative when evolutionary scores or other annotations are unavailable. Analysis of protein family composition revealed that a few Pfam families dominated the training data, but retraining on a Pfam-unique subset increased error only modestly, suggesting the findings are robust rather than inflated by family-level redundancy.

Fine-grained evaluations exposed the model’s blind spots in biologically meaningful ways. Proline substitutions consistently produced elevated errors, likely because their unusual conformational constraints disrupt secondary structure in ways that sequence-based features do not fully capture. In the TRIM44 zinc finger domain, histidine-to-cysteine mutations were poorly predicted when entire positions were held out but accurately recovered when single variants were withheld, hinting at context-dependent chemistry, such as zinc coordination, that only fine-grained positional information can resolve. The authors suggest that future versions could incorporate explicit structural features, cross-species transfer learning, or systematic assay harmonization to close these gaps.

Overall, VEFill arrives at a moment when the variant effect field is scaling rapidly, with MaveDB now listing more than 2,600 public datasets and over 1,100 for human proteins. By providing an interpretable, scalable tool that turns partially complete mutational maps into denser, more usable resources, the study lowers the experimental barrier for variant prioritization, protein engineering, and the clinical interpretation of missense mutations. The code and trained models have been released openly, inviting the community to apply sparse-library design strategies and, ultimately, to build the comprehensive Atlas of Variant Effects that the field has long envisioned.

Subject of Research: Machine learning-based imputation of missing scores in deep mutational scanning datasets across human protein domains

Article Title: VEFill: accurate and generalizable deep mutational scanning score imputation across protein domains

Article References: Polunina, P. V., Maier, W., & Rubin, A. F. (2026). VEFill: accurate and generalizable deep mutational scanning score imputation across protein domains. Molecular Systems Biology, 22(6), 979-1002. https://doi.org/10.1038/s44320-026-00203-y

Image Credits: AI Generated

DOI: 10.1038/s44320-026-00203-y

Keywords: deep mutational scanning, variant effect prediction, machine learning, gradient boosting, protein stability, ESM-1v, protein language models, MaveDB, Human Domainome, imputation, missense variants, bioinformatics

Cite Scienmag News
APA MLA Chicago

Drew Townsend. (October 2, 2026). AI Model Fills the Gaps in Protein Mutation Maps. Scienmag. https://scienmag.com/ai-model-fills-the-gaps-in-protein-mutation-maps/

Drew Townsend. “AI Model Fills the Gaps in Protein Mutation Maps.” Scienmag, 2 October 2026, https://scienmag.com/ai-model-fills-the-gaps-in-protein-mutation-maps/. Accessed 2 October 2026.

Drew Townsend. “AI Model Fills the Gaps in Protein Mutation Maps.” Scienmag. October 2, 2026. https://scienmag.com/ai-model-fills-the-gaps-in-protein-mutation-maps/

Copy citation Download RIS

Tags: amino acid substitution matricesbioinformaticsbioinformatics protein mutation toolsdeep mutational scanningESM-1vevolutionary conservation in proteinsgradient boostinghandling missing data in mutational datasetsHuman DomainomeimputationLightGBM protein analysisMachine learningmachine learning in protein analysisMaveDBmissense variantsprotein language modelsprotein mutation mappingprotein stabilityprotein variant effect mapsprotein variant effect predictiontransformer-based protein language modelsvariant effect predictionVEFill protein mutation prediction

Share12Tweet7Share2ShareShareShare1

Related Posts

Hidden Water Chemistry Shifts Are Quietly Sabotaging Fish Breeding in Aquaculture

Hidden Water Chemistry Shifts Are Quietly Sabotaging Fish Breeding in Aquaculture

October 2, 2026
Genes, Environment and the Epigenome: Why the Nature-Nurture Debate Refuses to Die

Genes, Environment and the Epigenome: Why the Nature-Nurture Debate Refuses to Die

October 2, 2026

Infrared Light and Machine Learning Reveal Which Animals Mosquitoes Bite

October 2, 2026

India’s Native Dogs Carry a Genetic Signature All Their Own, Landmark SNP Study Reveals

October 2, 2026

POPULAR NEWS

  • AI Model Predicts Chromosomally Normal IVF Embryos Without Genetic Biopsy

    29 shares
    Share 12 Tweet 7
  • Waterless Dyeing: Weld Plant Pigment Colors Nanofibers From the Inside Out

    29 shares
    Share 12 Tweet 7
  • Hidden Water Chemistry Shifts Are Quietly Sabotaging Fish Breeding in Aquaculture

    29 shares
    Share 12 Tweet 7
  • Carbon-Coated Trimetallic Nanosheets Push Supercapacitors Toward Rapid Charging

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

AI Model Predicts Chromosomally Normal IVF Embryos Without Genetic Biopsy

Waterless Dyeing: Weld Plant Pigment Colors Nanofibers From the Inside Out

Hidden Water Chemistry Shifts Are Quietly Sabotaging Fish Breeding in Aquaculture

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.