Every messenger RNA molecule in a cell carries a molecular signature at its tail: a stretch of adenine bases, known as the poly(A) tail, that helps determine how stable the transcript is, how efficiently it is exported from the nucleus, and how readily it is translated into protein. Measuring the length of these tails across thousands of genes has long been a technical headache, particularly for researchers using Oxford Nanopore Technologies sequencers, whose raw electrical signals are notoriously difficult to interpret in regions of repetitive adenines. A new open-access study published in BMC Biology introduces PolyAnalysis, a deep learning framework designed to estimate poly(A) tail length directly from raw nanopore signal data and to profile the regulatory architecture of transcript 3′ ends with a level of confidence awareness that its developers say has been missing from earlier tools.
The challenge that motivated the work is rooted in the physics of nanopore sequencing. As an RNA or DNA molecule threads through a protein pore, the sequencer records characteristic disruptions in ionic current, and basecalling algorithms translate those disruptions into sequence. Homopolymer runs, long stretches of a single base such as the adenines in a poly(A) tail, produce nearly identical current signals base after base, making it hard to count exactly how many adenines are present. On top of that, the 3′ ends of transcripts are architecturally complex: the boundary between the end of the transcript proper and the start of the tail can be fuzzy, alternative polyadenylation sites can generate multiple transcript isoforms from a single gene, and repetitive elements, including remnants of endogenous retroviruses, can complicate the identification of where a transcript truly terminates.
PolyAnalysis tackles these problems at the signal level rather than working from basecalled sequence alone. The framework combines a convolutional neural network with a bidirectional long short-term memory network, a pairing commonly abbreviated as CNN–BiLSTM, to classify regions of the raw signal as poly(A)-proximal or otherwise. It then applies a multi-task connectionist temporal classification decoding scheme, adapted from speech recognition, to infer the boundaries of the poly(A)-proximal region and the tail length itself. Crucially, the authors built adenine-aware training and decoding into the pipeline: the loss function used during training weights adenine-related errors differently, and the decoding step is biased toward A-rich interpretations, reflecting the biological expectation that the region of interest is dominated by adenines.
Benchmarking was carried out against three established tools: Nanopolish, tailfindr, and Dorado. The evaluation used controlled datasets generated with both RNA002 and RNA004 nanopore chemistries, allowing the team to assess performance across sequencing generations. According to the study, PolyAnalysis achieved tail-length accuracy that was competitive with the existing methods while delivering high callability, meaning it produced usable estimates for a larger fraction of reads. The authors stress that tail-length estimation is the most extensively validated output of the framework, and they are careful not to overstate the maturity of the other analyses the tool offers.
To understand which components of the model actually mattered, the researchers performed component-wise ablation experiments, systematically removing or altering parts of the pipeline and measuring the effect on performance. They also ran bias-control analyses. The results indicated that the region classification module, the multi-task optimization objective, the adenine-weighted loss, and the A-rich decoding strategy each contributed to the overall performance, suggesting that the design choices were not arbitrary but jointly responsible for the accuracy gains. This kind of ablation work is increasingly expected in machine learning applied to genomics, where it is easy to attribute improvements to a single flashy component when in fact the benefit comes from an interplay of design decisions.
Beyond tail length, PolyAnalysis links its per-read estimates to analyses of polyadenylation sites, alternative polyadenylation, and repetitive-element-associated 3′ ends. When applied to human cell-line datasets and vertebrate transcriptome datasets, the framework identified variation in tail-length distributions at the dataset level and heterogeneity at the gene level. It also found only weak associations between tail length and steady-state transcript abundance, a finding that adds nuance to the widely held expectation that longer tails should reliably predict more abundant or more translated transcripts. The relationship between tail length and expression, the data suggest, is more context-dependent than simple models imply.
One of the more distinctive features of the framework is its confidence-based filtering of polyadenylation sites. Rather than presenting all detected sites as equally reliable, PolyAnalysis separates them into database-matched high-confidence sites, novel candidates flagged with high confidence, and low-confidence calls. This tiered output acknowledges a persistent problem in 3′ end profiling: novel polyadenylation sites predicted from sequencing data can reflect genuine biology or merely artifacts of alignment, basecalling, or annotation gaps. By making confidence explicit, the tool gives downstream users a principled way to decide which calls to trust for follow-up experiments.
The treatment of endogenous retroviruses illustrates the same caution. Retroviral remnants scattered through the genome can supply polyadenylation signals to nearby genes, and distinguishing true locus-level evidence of such events from family-level patterns, where many similar retroviral sequences produce ambiguous signals, is genuinely difficult. PolyAnalysis applies ambiguity-aware analysis to separate these two levels of evidence, and the authors are explicit that the sequence-composition, polyadenylation-site, alternative-polyadenylation, and endogenous-retrovirus-related results all remain dependent on model assumptions, sequencing chemistry, and annotation quality, and require further orthogonal validation before being treated as established biology.
The significance of the work lies partly in its scope. Alternative polyadenylation is a major layer of gene regulation: by choosing a proximal or distal polyadenylation site, a cell can change the untranslated region of a transcript and thereby alter its stability, localization, and translation efficiency, with documented roles in development, immunity, and cancer. Tools that can profile these choices per read, on long-read platforms that capture full transcript molecules, offer a view that short-read sequencing cannot easily provide. By coupling tail-length estimation with polyadenylation-site analysis in a single signal-level framework, PolyAnalysis aims to make that view more accessible to laboratories that already run nanopore sequencers.
The study, a software contribution from researchers at Hunan University, Hunan University of Finance and Economics, and the University of Electronic Science and Technology of China, was supported by the National Natural Science Foundation of China and involved secondary analysis of publicly available datasets. The framework is published open access under a Creative Commons Attribution license, and the authors have released supplementary figures, tables, and numerical source data underlying the main results. For a field where poly(A) tail measurement has often been a bespoke, error-prone exercise, a benchmarked, confidence-aware pipeline that works across both RNA002 and RNA004 chemistries represents a practical step forward, even as the authors themselves caution that the most exploratory parts of their analysis, particularly the retrovirus-associated findings, should be treated as hypotheses awaiting independent confirmation rather than settled results.
Subject of Research: Deep learning-based estimation of poly(A) tail length and 3′-end regulation profiling from nanopore sequencing data
Article Title: PolyAnalysis: a deep learning–based framework for nanopore-based poly(A) tail estimation and 3′-end regulation profiling
Article References: Tian, Q., Song, B., Zou, Q., & Wang, Y. (2026). PolyAnalysis: a deep learning–based framework for nanopore-based poly(A) tail estimation and 3′-end regulation profiling. BMC Biology. https://doi.org/10.1186/s12915-026-02749-7
Image Credits: AI Generated
DOI: 10.1186/s12915-026-02749-7
Keywords: poly(A) tail, nanopore sequencing, deep learning, alternative polyadenylation, polyadenylation site, long-read transcriptomics, connectionist temporal classification, endogenous retrovirus, RNA sequencing, 3′ end regulation, Oxford Nanopore Technologies, BMC Biology
News Source: Blake Davidson. (October 5, 2026). Deep Learning Tool Reads RNA Tails Straight From Nanopore Signals. Scienmag.



