• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Thursday, August 27, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Biology

New Framework Reveals Unannotated Metabolites by Linking Chemical Clusters and Retention Times

Bioengineer by Bioengineer
August 27, 2026
in Biology
Reading Time: 6 mins read
0
New Framework Reveals Unannotated Metabolites by Linking Chemical Clusters and Retention Times
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Untargeted metabolomics can detect thousands of chemical signals in blood, urine or tissue, yet a large fraction cannot be assigned a definitive molecular identity. These unidentified signals, often called “metabolic dark matter,” may include biologically important compounds, products of microbial metabolism, modified nutrients, drug-related molecules or entirely unfamiliar chemicals. A new computational framework described in Metabolomics offers a way to narrow this vast uncertainty by combining molecular similarity with chromatographic behavior. Tested on plasma samples from pregnant women with obesity, the method reduced more than half a million possible structures to 418 high-confidence candidate annotations for previously unassigned metabolic features. The approach does not claim to identify unknown compounds outright, but creates a transparent shortlist that researchers can test experimentally with purified standards and additional analytical methods.

Metabolomics is designed to capture the small molecules that reflect what cells are doing at a particular moment. Unlike genomics, which describes biological potential, metabolomics records the chemical consequences of genetics, diet, exercise, disease, medication and environmental exposure. Liquid chromatography–tandem mass spectrometry, or LC–MS/MS, is one of the field’s most powerful tools. In this technique, liquid chromatography separates compounds according to their chemical interactions with a column, while mass spectrometry measures their mass-to-charge ratios and fragments selected molecules to generate structural clues. The resulting data contain thousands of peaks, each representing an ionized molecular feature. Some can be matched to reference standards or spectral libraries. Others have an accurate mass and perhaps a molecular formula, but no secure name, structure or biological interpretation.

The problem is more complicated than simply searching a larger database. Two compounds can share the same molecular formula while having different arrangements of atoms, known as structural or positional isomers. Stereoisomers can also possess the same connectivity but differ in three-dimensional orientation. These molecules may produce nearly identical masses and overlapping fragmentation patterns, while behaving differently in the chromatographic system and exerting different effects in the body. Unknown signals may also arise from chemical transformations that are poorly represented in databases, including compounds generated by gut microbes, noncanonical biochemical reactions or exposure to external chemicals. Instrumental artifacts, such as in-source fragmentation and ion rearrangement, can further create peaks that resemble genuine metabolites. As a result, a formula match alone is rarely enough to support a biologically meaningful annotation.

The researchers addressed this challenge by first defining the chemical landscape of the biological dataset itself. Their test data came from a randomized controlled trial investigating exercise during pregnancy in women with obesity. Plasma samples were collected from 119 sedentary pregnant participants, including women assigned to a combined aerobic and resistance-training program and others receiving standard care. In total, 235 samples were analyzed, including samples collected after submaximal exercise assessments. The study was not primarily intended to compare exercise and control groups in the present analysis. Instead, it provided a biologically relevant collection of plasma metabolites in which to develop and evaluate a strategy for interpreting unknown LC–MS features.

The samples were analyzed in both positive and negative electrospray ionization modes using an Orbitrap Exploris 480 mass spectrometer coupled to ultra-high-performance liquid chromatography. The instrument acquired high-resolution full-scan data across broad mass ranges and tandem spectra from pooled quality-control samples. After filtering for mass accuracy, signal quality, reproducibility and background contamination, the researchers retained 2,857 metabolite features. Of these, 1,021 had some level of annotation, ranging from relatively strong matches supported by an in-house standard and retention time to weaker assignments based only on accurate mass or molecular formula. The remaining 1,836 features lacked a chemical identity and were designated as metabolic dark matter. Rather than treating the partially annotated compounds as proven identifications, the researchers used them as landmarks defining regions of chemically plausible space.

To organize those landmarks, the team grouped the 1,021 known or partially characterized metabolites into ten structurally coherent clusters. The clustering process used molecular fingerprints, digital representations of the substructures present in each molecule. Specifically, the researchers generated Morgan circular fingerprints from SMILES chemical notation using the RDKit cheminformatics toolkit. These fingerprints encode local atomic environments as binary vectors. Structural similarity between two molecules was then measured with the Tanimoto coefficient, calculated from the number of shared fingerprint features relative to the total features present in either molecule. A score of 1 indicates identical fingerprint representations, while a score approaching 0 indicates little overlap. Importantly, the score reflects two-dimensional structural similarity and does not establish identical stereochemistry, tautomeric state or biological activity.

The next stage began with a deliberately broad search. For each of the 1,836 unannotated features, the researchers used its predicted molecular formula and molecular weight to retrieve possible structures from PubChem, allowing a relatively generous mass tolerance of plus or minus 0.5 daltons. This produced 569,115 candidate entries. After removing duplicates, malformed structures and records without usable chemical representations, 368,197 unique structures remained. The candidates were compared with the ten known-metabolite clusters, and those with Tanimoto similarity scores between 0.50 and 1.00 were retained. The lower threshold was intentionally permissive: a score of 0.5 was treated as evidence of a potentially useful chemical relationship, not as proof of identity. More stringent thresholds of 0.6 and 0.7 were also examined to show how the shortlist changed under increasingly conservative assumptions.

Similarity alone still left too many possibilities, so the researchers added a second, independent clue: retention time. In LC–MS, retention time is the point at which a compound emerges from the chromatography column. It depends on physicochemical properties such as hydrophobicity, polarity, hydrogen bonding and interactions with the stationary phase. Positional isomers with the same formula and mass may therefore separate at different times. The team used machine-learning models to predict retention times for candidate structures and compared those predictions with experimentally observed retention times for the unknown features. Candidates received higher priority when their predicted chromatographic behavior agreed with the measured signal. This step was especially valuable for distinguishing isomers that could not be separated confidently by formula matching or conventional spectral similarity.

Combining structural neighborhoods and retention-time agreement reduced the candidate space to 418 high-confidence candidate annotations. Among these, 83 were additionally supported by cross-referencing the Human Metabolome Database and the LIPID MAPS Structure Database. The final candidates were also placed in biological context through analyses involving absorption, distribution, metabolism and excretion properties, predicted protein targets, molecular docking and mapping to Kyoto Encyclopedia of Genes and Genomes pathways. Those analyses were applied to the known metabolites used to establish the reference landscape, helping the researchers interpret which chemical families and biological processes were represented in the dataset. The framework could thus prioritize not only molecules that resemble known compounds, but also candidates that fit the chemical and physiological environment of the samples.

The method is not a replacement for authentic standards, high-quality reference spectra or direct structural confirmation. A candidate structure that matches a formula, resembles a known metabolite and has a compatible retention time can still be wrong, particularly when stereoisomers or unusual ion forms are involved. The PubChem search tolerance was also broad enough to generate many chemically implausible possibilities before filtering, and the Tanimoto coefficient does not capture every aspect of molecular behavior. In addition, the test dataset came from a specific population and analytical platform, so retention-time models trained or calibrated in one laboratory may not transfer perfectly to another instrument, column or solvent gradient. Even so, the workflow provides a practical way to decide which unknown signals deserve scarce experimental resources first.

The broader significance is that metabolic dark matter is no longer being treated solely as an obstacle caused by incomplete databases. Unknown features can be interpreted as members of chemical neighborhoods, assessed according to how they travel through a chromatographic system and connected to biological pathways that may explain their origin or importance. By integrating formula-based retrieval, molecular fingerprints, retention-time prediction and biological context, the framework creates an auditable chain of reasoning between an unexplained mass-spectrometry peak and a testable molecular hypothesis. If validated with authentic standards and expanded across tissues, diseases, diets and environmental exposures, such approaches could accelerate the discovery of previously overlooked biomarkers and signaling molecules. For metabolomics, the viral idea is simple: thousands of mysterious peaks may not be noise waiting to be discarded, but clues waiting for the right combination of chemistry, computation and biology.

Subject of Research: A computational framework for prioritizing and biologically interpreting unannotated metabolites in untargeted LC–MS/MS metabolomics.

Subject of Research: Biology

Article Title: From known chemical space to unannotated metabolites: a cluster-guided retention-time driven framework for biologically informed annotation

Article References: Bhandari, D. et al., “From known chemical space to unannotated metabolites: a cluster-guided retention-time driven framework for biologically informed annotation,” Metabolomics 22, Article 146 (2026). Original publication

Image Credits: AI Generated

DOI: 10.1007/s11306-026-02520-7

Keywords: Metabolomics, metabolic dark matter, LC–MS/MS, mass spectrometry, retention time prediction, molecular similarity, Tanimoto coefficient, untargeted metabolomics, metabolite annotation, cheminformatics, lipidomics, biological pathway mapping

Tags: chemical clustering and retention time correlationchemical clustering in metabolomicschemical feature annotation in complex biological matriceschromatographic behavior in metabolite discoverycomputational framework for metabolite annotationhigh-confidence metabolite candidate shortlistidentifying unknown compounds in biological samplesLC-MS/MS metabolomics analysisLC–MS/MS metabolite profilinglinking chemical similarity with chromatographymetabolic dark matter discoverymetabolic dark matter identificationmetabolomics data analysismetabolomics of pregnancy and obesitymetabolomics of pregnant women with obesitymicrobial metabolism products detectionmicrobial metabolites in human biofluidsmolecular similarity and chromatographic behaviormolecular similarity in metabolomicsretention time-based metabolite annotationunannotated metabolites identificationuntargeted metabolomicsuntargeted metabolomics workflow

Share12Tweet7Share2ShareShareShare1

Related Posts

Review Identifies Testable Framework for Hidden Genotype D Hepatitis B Mutation Clusters

Review Identifies Testable Framework for Hidden Genotype D Hepatitis B Mutation Clusters

August 27, 2026
Bacterial and fungal diversity varies across Anopheles larval habitats by productivity

Bacterial and fungal diversity varies across Anopheles larval habitats by productivity

August 27, 2026

Norovirus Diagnostics Transform Through Innovation, Contextual Needs, and Collaborative Efforts

August 27, 2026

Gut bacterial enzyme unlocks polysaccharides’ power to boost cancer immunotherapy

August 27, 2026

POPULAR NEWS

  • Prompt-Based Knowledge Fusion Improves Faithfulness in Open-Domain Question Answering

    29 shares
    Share 12 Tweet 7
  • GDLS-D Reconstructs Medical Surfaces in Two Stages Using Diffusion and Iterative Optimization

    29 shares
    Share 12 Tweet 7
  • Study Tests MPEG-4 AAC Codec for DICOM Neurophysiology with EEG and EMG

    29 shares
    Share 12 Tweet 7
  • Pilot study explores electromyography for early detection of deep vein thrombosis risk

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Prompt-Based Knowledge Fusion Improves Faithfulness in Open-Domain Question Answering

GDLS-D Reconstructs Medical Surfaces in Two Stages Using Diffusion and Iterative Optimization

Study Tests MPEG-4 AAC Codec for DICOM Neurophysiology with EEG and EMG

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.