• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Monday, October 5, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Biology

AI Learns the Hidden DNA Code That Marks Enhancers Across Species

by
October 4, 2026
in Biology
Reading Time: 6 mins read
0
AI Learns the Hidden DNA Code That Marks Enhancers Across Species

AI Learns the Hidden DNA Code That Marks Enhancers Across Species

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Deep inside the genome, scattered among stretches of DNA that do not code for proteins, lie short segments with an outsized influence on life. These segments, called enhancers, act as docking sites for the proteins that switch genes on or off, and their misreading is implicated in developmental disorders, cancer, and countless other conditions. Yet despite decades of study, scientists still cannot reliably look at a piece of DNA and say whether it functions as an enhancer. Sequencing technology now produces genomes far faster than laboratories can experimentally annotate their regulatory regions, leaving a widening gap between what we can read and what we can understand. A new study published in BMC Genomics by Luis M. Solis, Geyenna Sterling-Lentsch, Marc S. Halfon, and Hani Z. Girgis tackles this problem head-on with a deep learning tool called EnhancerDetector, which the authors say can spot enhancers in organisms as distant as humans and fruit flies using nothing more than the raw DNA sequence.

The central question the researchers posed is deceptively simple: do enhancers share recurring, sequence-based features that distinguish them from the rest of the genome? The team refers to this hypothetical signature as enhancerness. If such a signature exists and can be learned by a machine, it would mean that the grammar of gene regulation is written in the sequence itself, in patterns of bases that recur across species, cell types, and experimental assays. Confirming that would be biologically fundamental, because it would suggest a universal logic underlying how regulatory information is encoded. It would also be enormously practical, since a model that recognizes enhancerness could annotate newly sequenced genomes without waiting for years of costly experiments. The study provides evidence on both fronts, combining computational benchmarks with validation in living flies.

EnhancerDetector is built on a convolutional neural network, the same family of architectures that revolutionized image recognition. Convolutional networks excel at detecting local patterns and assembling them into higher-order structures, which makes them well suited to DNA, where transcription factor binding sites and their combinations form the functional vocabulary of a regulatory element. Trained on human data, the model takes short sequence windows as input and directly outputs a probability that the window is an enhancer. This design choice matters. Many existing enhancer predictors rely on chromatin features, such as histone modifications like H3K4me1 and H3K27ac, which are measured in specific cell types and require complicated post hoc thresholding to convert into enhancer calls. By working directly from sequence, EnhancerDetector sidesteps that dependency and simplifies the discovery workflow, producing scores that can be ranked and filtered without additional processing layers.

The performance results are striking in their breadth. On human, mouse, and fly datasets, EnhancerDetector consistently outperformed existing methods in precision and F1 score, two metrics that matter greatly in genomics because the number of true non-enhancers vastly exceeds the number of true enhancers, making false positives the dominant error. The model also generalized across datasets generated with diverse experimental assays, suggesting it had not merely memorized the artifacts of one laboratory technique. An ensemble strategy, in which multiple models vote together, further improved reliability by reducing false positives. In a field where a single mislabeled enhancer can send a research group down an expensive dead end, that reduction in noise is a meaningful advance. The authors also compared their approach against sequence-based competitors such as gapped k-mer support vector machines, and the convolutional network came out ahead.

Perhaps the most practically important result concerns transfer learning. EnhancerDetector was trained on human data, yet it performed well on mouse and fly, species separated from humans by tens and hundreds of millions of years of evolution respectively. Even more remarkably, the model can be fine-tuned on a new species and retains strong performance when adapted with as few as 20,000 enhancer sequences. For newly sequenced genomes, where experimental data are scarce by definition, this is a game changer. A research consortium that has just assembled the genome of an obscure organism can adapt a pretrained model with a modest dataset and begin prioritizing candidate regulatory elements immediately, rather than starting from scratch. The authors position EnhancerDetector explicitly as a first-stage annotation tool for exactly this scenario.

Accuracy alone, however, was not the goal. The team wanted a model that biologists could interrogate, one whose predictions could be traced back to specific features of the DNA. To achieve this, they applied class activation maps, a visualization technique that highlights which regions of an input sequence contributed most strongly to the model’s decision. These maps allow researchers to see, base by base, where the network is focusing its attention when it declares a stretch of DNA to be an enhancer. Such interpretability transforms the model from a black box into a hypothesis generator: the highlighted motifs can be compared against known transcription factor binding sites, and unexplained patterns can point toward regulatory logic that science has not yet characterized. In an era when machine learning is often criticized for opacity, building interpretability into the architecture from the start is a deliberate and welcome choice.

Computational benchmarks are one thing, but the gold standard in genomics is experimental validation, and here the study delivers. The team tested candidate enhancers predicted by the model in transgenic flies, using reporter constructs in which a candidate DNA sequence is placed next to a green fluorescent protein gene, so that if the sequence acts as an enhancer, the tissue where it is active glows under the microscope. Of six candidates tested, five drove reporter expression, a hit rate that would be respectable even with extensive prior experimental support. Moreover, four of the five exhibited expression patterns consistent with those reported in the earlier literature, providing independent corroboration that the model was not merely finding sequences capable of driving expression in some generic way, but was recovering biologically meaningful regulatory elements with the correct tissue specificities.

Taken together, the computational and experimental results support the paper’s central claim: enhancers possess learnable, recurring sequence features that can be captured in one species and transferred to another. The analyses identified distinct sequence and contextual features associated with enhancer activity, the constellation the authors call enhancerness. This finding resonates with a long-standing debate in regulatory genomics. Enhancers are notoriously diffuse, with binding sites for multiple transcription factors arranged in flexible combinations rather than rigid templates, which has led some researchers to doubt whether any universal sequence signature could exist. The success of a single model trained on human sequence and applied to fly suggests that, beneath the flexibility, there is a shared statistical texture to enhancer DNA, one deep enough for a convolutional network to grasp but subtle enough to have eluded simpler analytical methods.

The implications extend well beyond the three species studied. As sequencing costs continue to fall, thousands of genomes are being assembled for organisms with no experimental annotation at all, from non-model animals to endangered species to agricultural pests. A tool that can flag candidate enhancers from sequence alone, and be fine-tuned with a few tens of thousands of examples, offers a scalable path to functional annotation of this flood of new data. The work was supported by the National Human Genome Research Institute of the National Institutes of Health under award number R21HG011507, and the article is open access, so the tool and its findings are available to any laboratory that wants to build on them. The authors acknowledge assistance with experimental validations from Jack Leatherbarrow and James Weiser, and they declare no competing interests.

There remain limits to what any sequence-only model can do. Enhancer activity depends on cellular context, and a sequence that functions in one tissue may be silent in another, so predictions from DNA alone will always need experimental follow-up. But by demonstrating that a common, learnable signature of enhancerness exists across species and assays, and by packaging that insight into an interpretable, transferable, and experimentally validated framework, the EnhancerDetector team has taken a substantial step toward making the dark matter of the genome a little less dark. For biologists staring at a freshly assembled genome and wondering where the switches are, the answer may now be one pretrained neural network away.

Subject of Research: Cross-species enhancer prediction from DNA sequence using interpretable convolutional neural networks

Article Title: EnhancerDetector: enhancer discovery from human to fly via interpretable deep learning

Article References: Solis, L. M., Sterling-Lentsch, G., Halfon, M. S., & Girgis, H. Z. (2026). EnhancerDetector: enhancer discovery from human to fly via interpretable deep learning. BMC Genomics. https://doi.org/10.1186/s12864-026-13262-0

Image Credits: AI Generated

DOI: 10.1186/s12864-026-13262-0

Keywords: enhancers, deep learning, convolutional neural network, genome annotation, regulatory genomics, cross-species prediction, interpretability, class activation maps, Drosophila, transcription factor binding sites, transfer learning, BMC Genomics

News Source: Juliet Wilcox. (October 4, 2026). AI Learns the Hidden DNA Code That Marks Enhancers Across Species. Scienmag.

Tags: BMC Genomicsclass activation mapsconvolutional neural networkcross-species predictiondeep learningDrosophilaenhancersgenome annotationinterpretabilityregulatory genomicstranscription factor binding sitesTransfer Learning
Share12Tweet7Share2ShareShareShare1

Related Posts

A Methylation Signature in Brain Fluid Predicts Who Survives a Deadly Stroke

A Methylation Signature in Brain Fluid Predicts Who Survives a Deadly Stroke

October 5, 2026
Scientists Build a 44-SNP Genetic Barcode to Catch Mislabelled Samples Across Labs

Scientists Build a 44-SNP Genetic Barcode to Catch Mislabelled Samples Across Labs

October 5, 2026

City Life Reshapes the Microbes Inside Thailand’s Brown Dog Ticks

October 5, 2026

Green Shopping Online Only Works When Consumers Know What They Are Buying, Study Finds

October 4, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.