• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Sunday, October 4, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Health

AI Learns the Language of Hormones to Predict Peptide Drugs

Bioengineer by Bioengineer
October 4, 2026
in Health
Reading Time: 6 mins read
0
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Peptide hormones are among the body’s most powerful chemical messengers. These short chains of amino acids regulate metabolism, growth, development, and the delicate balance of homeostasis in organisms ranging from plants to humans. When their regulation goes awry, the consequences can be serious, which is precisely why they have become prized targets in drug development. Unlike small-molecule drugs, peptides carry more hydrogen bond donors and acceptors, allowing them to bind their targets with higher specificity and fewer off-target side effects. They are also readily degraded by enzymes in the body, show low metabolic toxicity, and exhibit good biocompatibility. Yet identifying which of the countless short peptide sequences in a genome actually function as hormones has long been a slow, expensive, and experimentally demanding task.

Now, a research team led by Chunyan Ao, Shihu Jiao, Xi Su, and Huan Yang has unveiled a deep learning framework called pLM-HP that promises to change that. Writing in the Journal of Advanced Research, the authors describe a system that combines a pre-trained protein language model, ESM2, with a bidirectional long short-term memory network, or BiLSTM, to predict whether a given peptide sequence is a hormone. On an independent test set, the model achieved a balanced accuracy of 95.64 percent, with sensitivity of 94.48 percent and specificity of 96.79 percent, a Matthews correlation coefficient of 0.824, and an area under the ROC curve of 0.991. Those numbers place it well ahead of existing tools, and they hint at a broader shift in how computational biology identifies functional molecules.

The challenge the researchers set out to solve is not new. Traditional identification of peptide hormones relies on experimental techniques such as mass spectrometry, affinity purification, and activity screening, all of which are complex and time-consuming. Computational methods emerged as a complementary approach, but early efforts leaned on sequence similarity searches or motif recognition algorithms, which falter when homologous sequences are scarce. Machine learning brought improvements: the HOPPred model, for example, integrates BLAST similarity, MERCI motifs, and logistic regression, achieving an AUROC of 0.96 and an MCC of 0.80 on an independent test set, while mHPpred adopted a multi-view feature representation and a meta-modeling strategy to build a more robust ensemble. Still, many methods remained tethered to hand-crafted features or predefined motifs, limiting their ability to represent sequences fully, and class imbalance in real data often biased models toward the majority class, dulling their sensitivity to genuine hormone peptides.

Protein language models have upended that paradigm. Models such as ProtBERT, ProtT5, and ESM2 are pre-trained on enormous collections of unlabeled protein sequences using masked language modeling, in which the model learns to predict hidden amino acids from their surrounding context. In doing so, they acquire high-dimensional embeddings that encode evolutionary, structural, and physicochemical properties of proteins. These embeddings have already powered advances in post-translational modification prediction, peptide structure modeling, and functional peptide classification. But until now, their use in peptide hormone prediction had been limited, and no systematic study had evaluated how well they perform on this specific task.

The pLM-HP framework treats a peptide sequence like a sentence and each amino acid like a word. ESM2, built on the Transformer encoder architecture and pre-trained on UniRef50 sequences, processes the sequence with a multi-head self-attention mechanism that captures long-range dependencies between residues, combined with feed-forward networks for nonlinear feature transformation. Each input is augmented with special beginning-of-sequence and end-of-sequence tokens, and token and positional embeddings are combined to capture residue order and context. Crucially, the researchers used ESM2 solely as a frozen feature extractor, without fine-tuning, and passed the residue-level embeddings directly to the downstream classifier.

One of the study’s most striking findings concerns model size. The team benchmarked four ESM2 variants, ranging from the 8-million-parameter esm2_t6_8M to the 650-million-parameter esm2_t33_650M, with output dimensions of 320, 480, 640, and 1280 respectively. All achieved balanced accuracy of roughly 95 percent in both five-fold cross-validation and independent testing, but the mid-sized esm2_t12_35M model, with 480-dimensional outputs, performed best, reaching a balanced accuracy of 95.61 percent and an MCC of 0.802 in cross-validation and 95.64 percent and 0.824 on the independent test set. The largest model, despite its far greater capacity, actually performed slightly worse, with a balanced accuracy of 95.15 percent and an MCC of 0.804, suggesting that for short peptide sequences of 11 to 41 residues, bigger is not necessarily better and may even introduce redundant features or mild overfitting.

The choice of classifier mattered just as much. When the researchers compared deep learning architectures and traditional machine learning methods on top of the ESM2 features, the BiLSTM emerged as the clear winner. By running two recurrent networks in forward and backward directions and concatenating their hidden states, the BiLSTM exploits both past and future context around each residue, capturing contextual dependencies that other architectures miss. It outperformed convolutional neural networks, multilayer perceptrons, support vector machines, logistic regression, random forests, XGBoost, and LightGBM across the board. Tree-based models fared worst: random forest achieved only about 73 percent balanced accuracy on the independent test set, with a strong bias toward predicting negative samples, while XGBoost and LightGBM reached high specificity near 98 percent but sensitivity below 81 percent, limiting their ability to spot true hormone peptides.

Class imbalance posed another formidable obstacle, since hormone peptides are vastly outnumbered by non-hormone sequences in real datasets. The team built their training data from 5,729 plant- and animal-derived peptide hormone sequences curated in the Hmrbase2 database, reduced to 1,174 non-redundant positives using CD-HIT at a 60 percent similarity threshold, paired with 11,740 non-hormone peptides drawn from PeptideAtlas. Rather than relying on resampling tricks, they incorporated class weights into a weighted binary cross-entropy loss, boosting the contribution of the rare hormone class. When they stress-tested the model on training sets with positive-to-negative ratios ranging from 1:1 to 1:10, performance remained remarkably stable, with balanced accuracy consistently above 94 percent, MCC values between 0.741 and 0.824, AUC values from 0.983 to 0.992, and AUPRC values from 0.841 to 0.915. Notably, the gap between sensitivity and specificity stayed small, at just 2.31 percentage points on the final test set, indicating a balanced classifier rather than one that simply favors the majority class.

The advantages over traditional approaches were dramatic. Handcrafted sequence descriptors such as pseudo amino acid composition, composition of k-spaced amino acid pairs, and quasi-sequence-order descriptors, when paired with classifiers like XGBoost, showed high specificity but weak sensitivity, and classical imbalance-handling strategies such as SMOTE and ADASYN often improved sensitivity at the cost of specificity, producing unstable results. In head-to-head comparisons on the same independent test set, pLM-HP outperformed HOPPred on every metric, improving balanced accuracy by 24.33 percentage points, sensitivity by 35.00 points, and specificity by 13.64 points, while raising the MCC by 0.522 and the AUC by 0.264. Against mHPpred, it improved balanced accuracy by 9.95 points, specificity by 18.51 points, MCC by 0.368, and AUC by 0.055. Visualization analyses using t-SNE and UMAP reinforced the story: while handcrafted features left positive and negative samples substantially overlapping, and raw ESM2 embeddings only partly separated them, the representations refined by the BiLSTM formed compact, well-separated clusters.

The implications reach well beyond one prediction task. As high-throughput omics and peptide-based drug development accelerate, efficiently and reliably identifying hormone-functional peptides from vast sequence spaces has become a key bridge between basic research and drug discovery. Peptide hormones, capable of probing and modulating protein-protein interactions, are considered ideal scaffolds for drug design, and a tool that can screen candidate sequences with near-99 percent AUC could dramatically narrow the experimental search space. The authors are candid about limitations: the current framework relies on linear sequence information and does not explicitly integrate structural context that shapes peptide conformation and signaling. Future work, they suggest, could incorporate structure predictions from AlphaFold2, apply geometric deep learning with graph neural networks, and explore focal loss, contrastive learning, and fine-tuning strategies to improve robustness under extreme imbalance and cross-dataset transfer. Even so, pLM-HP offers a compelling demonstration that when protein language models meet the right sequence-aware classifier, the hidden language of hormonal peptides becomes far easier to read.

Subject of Research: Computational prediction of peptide hormones using pre-trained protein language model embeddings and deep learning

Article Title: pLM-HP: Peptide hormone prediction using pre-trained protein language model representations

Article References: Ao, C., Jiao, S., Su, X., & Yang, H. (2026). pLM-HP: Peptide hormone prediction using pre-trained protein language model representations. Journal of Advanced Research. https://doi.org/10.1016/j.jare.2026.10.002

Image Credits: AI Generated

DOI: 10.1016/j.jare.2026.10.002

Keywords: peptide hormones, protein language models, ESM2, BiLSTM, deep learning, drug discovery, machine learning, bioinformatics, class imbalance, sequence analysis, peptide drugs, computational biology

Cite Scienmag News

APA
MLA
Chicago

Blake Davidson. (October 4, 2026). AI Learns the Language of Hormones to Predict Peptide Drugs. Scienmag. https://scienmag.com/ai-learns-the-language-of-hormones-to-predict-peptide-drugs/

Blake Davidson. “AI Learns the Language of Hormones to Predict Peptide Drugs.” Scienmag, 4 October 2026, https://scienmag.com/ai-learns-the-language-of-hormones-to-predict-peptide-drugs/. Accessed 4 October 2026.

Blake Davidson. “AI Learns the Language of Hormones to Predict Peptide Drugs.” Scienmag. October 4, 2026. https://scienmag.com/ai-learns-the-language-of-hormones-to-predict-peptide-drugs/

Copy citation
Download RIS

Tags: advances in peptide therapeuticsAI for hormone sequence analysisAI-driven discovery of peptide hormonesBiLSTMBiLSTM neural networks for peptide function predictionbioinformaticsclass imbalancecomputational biologycomputational biology for hormone researchdeep learningdeep learning in drug discoverydrug discoveryESM2hormone peptide identification techniquesMachine learningmachine learning models for biological sequence analysispeptide drugspeptide hormone predictionpeptide hormone regulation and signalingpeptide hormonespeptide-based drug developmentprotein language modelsprotein language models for peptide classificationsequence analysis

Share12Tweet7Share2ShareShareShare1

Related Posts

Longer Antibiotic Courses After Heart Surgery May Backfire, Study Finds

October 4, 2026

AI Reads Spine Scans to Measure Back Muscle Health Without Any Human Help

October 4, 2026

Single Mutation, Many Faces: Study Maps Mitochondrial Disease in Chinese Children

October 4, 2026

American Disability Test Norms May Not Fit German-Speaking Children, Study Warns

October 4, 2026

POPULAR NEWS

  • Emergency Embolization Halts Life-Threatening Bleeding from Advanced Breast Tumor

    29 shares
    Share 12 Tweet 7
  • AI Learns the Language of Hormones to Predict Peptide Drugs

    29 shares
    Share 12 Tweet 7
  • Hidden pH Swings at Electrodes Turn Electricity and Oxygen into Plastic Feedstocks

    29 shares
    Share 12 Tweet 7
  • Shikonin Shields the Heart From Chemotherapy Damage by Restoring Mitochondrial Balance

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Emergency Embolization Halts Life-Threatening Bleeding from Advanced Breast Tumor

AI Learns the Language of Hormones to Predict Peptide Drugs

Hidden pH Swings at Electrodes Turn Electricity and Oxygen into Plastic Feedstocks

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.