• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Thursday, September 24, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Chemistry

New Sum-Connectivity Descriptors Give Molecular Graphs a Sharper Predictive Edge

Bioengineer by Bioengineer
September 24, 2026
in Chemistry
Reading Time: 6 mins read
0
New Sum-Connectivity Descriptors Give Molecular Graphs a Sharper Predictive Edge
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

A molecule can be drawn as a graph: atoms become vertices, bonds become edges, and from that abstraction chemists have spent decades extracting numbers that capture how branching, connectivity and local density shape chemical behaviour. A study published in the Journal of Saudi Chemical Society by Azzam Altairi, Mohammed Alsharafi and Zaied Alhaj now proposes a systematic way to upgrade one of the most widely used families of these numbers. Their ‘sum-connectivity descriptors’ take classical degree-based topological indices and reweight every edge contribution by the inverse square root of the sum of the degrees of its two endpoint atoms. The result is a unified family of new descriptors that, in controlled benchmarks, modestly but consistently outperform the classical originals in predicting physicochemical properties, while offering complementary signal for the much harder problem of predicting antibacterial activity.

The mathematical backbone of the work is elegantly simple. Many celebrated indices in chemical graph theory, including the Zagreb indices, the Randić index, Sombor-type measures, the Albertson irregularity index, arithmetic-geometric and geometric-arithmetic indices, the Forgotten index and inverse Nirmala indices, can all be written as a sum over the edges of a molecular graph of a symmetric function of the degrees of the two atoms joined by each bond. The authors multiply each of these contributions by a normalisation factor derived from the general sum-connectivity index introduced by Zhou and Trinajstić: the weight one over the square root of the sum of the two endpoint degrees. Because this factor decreases as the local connectivity increases, it moderates the dominance of edges attached to high-degree atoms, a known source of instability when classical descriptors are applied across chemically diverse datasets.

The framework is not merely a numerical tweak. The authors prove a set of structural results that give the new family a rigorous footing. For any molecular graph whose minimum and maximum vertex degrees are delta and capital delta respectively, each sum-connectivity descriptor is provably sandwiched between its classical counterpart scaled by one over the square root of twice capital delta and the same counterpart scaled by one over the square root of twice delta. On regular graphs, where every atom has the same degree, the relationship collapses to exact proportionality with a constant of one over the square root of two r. The paper further derives explicit min-max bounds for representative descriptors in terms of the number of edges and the degree extremes, develops complement-graph identities showing how the reweighting behaves under the transformation that maps each degree to n minus one minus the original degree, and introduces a ‘hyper sum-connectivity’ index obtained when the construction is applied to the classical sum-connectivity index itself.

To test whether this theory translates into predictive power, the team assembled a carefully curated antibacterial dataset from the ChEMBL bioactivity database, restricted to minimum inhibitory concentration measurements against Escherichia coli, an organism that is both a common member of the human microbiome and a leading cause of urinary tract, bloodstream and intra-abdominal infections increasingly complicated by multidrug resistance. After standardising all MIC values to micromolar units, aggregating duplicates by median and expressing potency as the negative base-ten logarithm, the final dataset comprised 6,657 unique compounds, each characterised by eight physicochemical properties, two blocks of fourteen topological indices, the classical family and the new sum-connectivity family, and a block of ninety RDKit descriptors capturing size, branching, ring content and related structural features.

The first empirical test was a deliberately controlled one: a head-to-head comparison of the two index families alone across eight physicochemical endpoints, with everything else held constant. Under a demanding five-times-repeated five-fold cross-validation protocol, the sum-connectivity family produced higher mean coefficients of determination for all eight targets, lifting the average R-squared from 0.6951 for the classical indices to 0.7015 for the reweighted ones. The gains were most visible for topological polar surface area, lipophilicity and the spacial score, while molecular weight barely moved. The authors are careful to characterise this as a modest but systematic improvement, a claim backed by the theoretical guarantees rather than by an isolated lucky result.

A second, broader QSPR experiment concatenated everything: RDKit descriptors, both index families and the physicochemical properties, into a single ‘All Combined’ feature set. This configuration delivered the highest scores of the study, with R-squared values ranging from 0.9349 for NP-likeness to 0.9997 for molecular weight. But the team resists the temptation to oversell these numbers. Correlation analyses reveal that several endpoints are almost directly encoded by closely related features already present in the pool: molecular weight correlates with the RDKit exact molecular weight at a Spearman coefficient of 0.99999, and topological polar surface area tracks the nitrogen-oxygen count at 0.9304. The near-unity scores are therefore best read as upper-bound benchmarks within the available descriptor space, not as proof of independent mechanistic insight. A sensitivity rerun with principal component compression lowered scores across the board, showing that blind dimensionality reduction is no substitute for careful proxy-descriptor removal.

The harder test was antibacterial activity itself. Predicting the pMIC endpoint from structure alone is intrinsically difficult because potency reflects a tangle of scaffold class, charge distribution, lipophilicity, hydrogen-bonding profile and, often, a specific mechanism of action that no purely topological number can see. The best-performing QSAR model was an Extra Trees regressor trained on RDKit descriptors, reaching an R-squared of 0.5959 plus or minus 0.0246 under repeated cross-validation. Critically, the team then repeated the evaluation using a far stricter Murcko scaffold-based validation, which keeps close structural analogues out of the training folds and probes whether models generalise to genuinely new chemotypes. Performance dropped, as expected, to an R-squared of 0.4799 plus or minus 0.0252, but the ranking of feature families remained stable, indicating reproducible rather than chance-driven signal.

Within that stricter landscape, the new descriptors tell an honest story. On their own, the standalone sum-connectivity indices edged out the classical family for pMIC prediction, 0.2887 versus 0.2767 under random cross-validation and 0.2025 versus 0.2019 under scaffold validation, but both were clearly weaker than the rich RDKit representation, and the sum-connectivity advantage effectively disappeared once RDKit features were present. Diagnostics on the best model reinforced the picture of genuine but moderate predictability: a Y-randomization test with fifty permutations produced a mean R-squared of minus 0.0356, separated from the real model by a Z-score of 91.8, while permutation importance highlighted heterocycle counts, aliphatic ring content, charge-sensitive surface-area terms and electronic-state measures, a chemically plausible fingerprint of the determinants of antibacterial potency. A learning curve still rising with training-set size suggests further curated data would continue to help.

The study’s conclusions are notably measured for a field where bold claims are common. The authors state plainly that their descriptors are most strongly supported as systematically improved graph-topological descriptors in matched comparisons, with added value for the heterogeneous pMIC endpoint that is complementary and modest. They also acknowledge that every classical index is almost collinear with its sum-connectivity counterpart, with Spearman correlations between 0.9967 and 0.9986, and that a PCA rerun reduced the strongest benchmark scores while leaving constitution-driven targets comparatively easy to encode, so targeted removal of proxy descriptors and regularised feature selection remain the more promising next steps.

For the wider cheminformatics community, the work offers a template as much as a tool. It shows how a single, mathematically transparent transformation can generate a coherent descriptor family with provable bounds, how controlled matched benchmarks should be separated from headline-grabbing combined-model scores, and how scaffold-based validation should be the default standard for activity modelling on real drug-discovery data. The mathematical side also leaves a clear opening: extending the reweighting to the general chi-alpha family and studying how the exponent on the degree sum affects extremal behaviour and descriptor discrimination is flagged as an open direction. In an era when machine learning models are only as trustworthy as the features fed into them, a rigorous attempt to make one small corner of the descriptor space measurably better is a quietly consequential contribution.

Subject of Research: Sum-connectivity molecular descriptors for QSPR and antibacterial QSAR modelling

Article Title: Advancing QSPR with sum-connectivity descriptors: physicochemical and antibacterial modelling via molecular graph connectivity

Article References: Altairi, A., Alsharafi, M., & Alhaj, Z. (2026). Advancing QSPR with sum-connectivity descriptors: physicochemical and antibacterial modelling via molecular graph connectivity. Journal of Saudi Chemical Society, 30(4), Article 50. https://doi.org/10.1007/s44442-026-00105-6

Image Credits: AI Generated

DOI: 10.1007/s44442-026-00105-6

Keywords: QSPR, QSAR, topological indices, sum-connectivity descriptors, molecular graphs, antibacterial activity, Escherichia coli, ChEMBL, machine learning, MIC prediction, chemical graph theory, RDKit

Cite Scienmag News
APA MLA Chicago

Bethany Barker. (September 24, 2026). New Sum-Connectivity Descriptors Give Molecular Graphs a Sharper Predictive Edge. Scienmag. https://scienmag.com/new-sum-connectivity-descriptors-give-molecular-graphs-a-sharper-predictive-edge/

Bethany Barker. “New Sum-Connectivity Descriptors Give Molecular Graphs a Sharper Predictive Edge.” Scienmag, 24 September 2026, https://scienmag.com/new-sum-connectivity-descriptors-give-molecular-graphs-a-sharper-predictive-edge/. Accessed 24 September 2026.

Bethany Barker. “New Sum-Connectivity Descriptors Give Molecular Graphs a Sharper Predictive Edge.” Scienmag. September 24, 2026. https://scienmag.com/new-sum-connectivity-descriptors-give-molecular-graphs-a-sharper-predictive-edge/

Copy citation Download RIS

Tags: antibacterial activityantibacterial activity predictionChEMBLchemical graph theorydegree-based graph descriptorsEscherichia coligraph-theoretic molecular characterizationirregularity indexMachine learningMIC predictionmolecular graph analysismolecular graphsphysicochemical property predictionQSARQSPRRandić indexRDKitSombor measuressum-connectivity descriptorstopological indicesZagreb indices

Share12Tweet7Share2ShareShareShare1

Related Posts

Bismuth Tungstate Wrapped in Conductive Polymer Yields Supercapacitor Electrode That Barely Fades Over 1,000 Cycles

Bismuth Tungstate Wrapped in Conductive Polymer Yields Supercapacitor Electrode That Barely Fades Over 1,000 Cycles

September 24, 2026
Ball-Milled and CVD Silicon Anodes Face Off in All-Solid-State Batteries

Ball-Milled and CVD Silicon Anodes Face Off in All-Solid-State Batteries

September 24, 2026

Induction-Heating PCR Chip Cracks Worm Genes in 30 Minutes

September 24, 2026

Chiral 2D Framework Turns Rhodium Catalyst Into a Recycling Champion

September 24, 2026

POPULAR NEWS

  • Counting Energy Like Calories: The Theory Behind the Activity Calculator for Chronic Fatigue

    29 shares
    Share 12 Tweet 7
  • Skin aging gene Foxn1 revealed as master regulator of epidermal structure and redox balance

    29 shares
    Share 12 Tweet 7
  • Rapamycin Nanoparticles Rejuvenate Aging Spine Discs by Silencing a Cellular Aging Switch

    29 shares
    Share 12 Tweet 7
  • Self-Coldness, Not Self-Kindness, Amplifies Stress-Linked Anxiety and Depression in Mothers of Autistic Children

    29 shares
    Share 12 Tweet 7

About

BIOENGINEER.ORG

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Counting Energy Like Calories: The Theory Behind the Activity Calculator for Chronic Fatigue

Skin aging gene Foxn1 revealed as master regulator of epidermal structure and redox balance

Rapamycin Nanoparticles Rejuvenate Aging Spine Discs by Silencing a Cellular Aging Switch

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.