A widespread blind spot in one of genomics’ most routine measurements has been brought into sharp focus by a new study comparing short-read and long-read transcriptome sequencing. Researchers report that next-generation sequencing, the short-read workhorse of modern RNA analysis, systematically underestimates the levels of adenosine-to-inosine (A-to-I) RNA editing, one of the most abundant chemical modifications in human RNA. The discrepancy, they show, is not a minor technical quirk but a pervasive bias rooted in the very length of the sequencing reads themselves, and it affects data from laboratories around the world, including widely used public datasets.
A-to-I RNA editing is catalyzed by the ADAR family of enzymes, primarily ADAR1 and ADAR2, which chemically convert adenosine bases to inosine within RNA transcripts. Because the cellular machinery reads inosine as guanosine, this single change can rewrite the meaning of a gene: it can swap amino acids in proteins, alter splice sites, create or destroy start and stop codons, reshape RNA folding, and modulate microRNA targeting. In humans, millions of editing sites dot the transcriptome, with the vast majority clustered in repetitive elements such as Alu sequences, where paired inverted repeats form the double-stranded RNA substrates that ADAR enzymes prefer. Editing in these repetitive regions also helps cells distinguish their own RNA from foreign double-stranded RNA, preventing inappropriate immune activation.
The functional stakes are high. Aberrant A-to-I editing has been implicated in amyotrophic lateral sclerosis, autism spectrum disorder, schizophrenia, epilepsy, depression and multiple cancers, and dynamic editing levels are known to vary across tissues, developmental stages and species. Because editing levels, defined as the fraction of adenosines converted to inosine at a given site, are the currency of nearly every study in this field, accurate measurement is foundational. Yet while considerable effort has gone into developing methods to identify editing sites from sequencing data, far less attention has been paid to whether the quantification of editing levels itself is trustworthy.
To test this, the team, led by Shi Cheng and Yulong Song of Guangzhou Medical University alongside colleagues at Sun Yat-sen University, performed matched short-read and long-read cDNA RNA sequencing on two human cell lines, HEK293T and U2OS. The short-read data were generated on an Illumina NovaSeq 6000 platform using both polyA-selected and rRNA-depleted library strategies, producing reads of roughly 150 nucleotides. The long-read data were produced on Oxford Nanopore Technologies PromethION sequencers, which capture full-length cDNA molecules with median clean read lengths of 812 to 878 nucleotides. Both approaches were anchored to a curated reference list of 2,858,028 human A-to-I editing sites compiled from the group’s previous work.
The results were striking. In HEK293T cells, long-read sequencing identified 3,897 editing sites measured significantly higher than by polyA-selected short-read sequencing, against only 445 sites measured higher by the short-read approach, an approximately 8.8-fold difference in the number of sites. Globally, the median editing level estimated by long reads was about 2.3 times that from short reads. The pattern held for sites in Alu elements, in repetitive non-Alu regions and in nonrepetitive regions alike. Most dramatically, at nonrepetitive sites the median short-read estimate was 0 percent, implying no editing at all, whereas long-read sequencing revealed a true median level of about 2.1 percent. Parallel comparisons using rRNA-depleted short-read libraries, and independent experiments in U2OS cells, reproduced the same directional bias, with long reads detecting roughly 6.6 to 9.0 times more significantly higher-measured sites.
To rule out artifacts, the researchers ran an extensive battery of controls. They raised coverage thresholds from 30 to 40, 50 and 60 reads per site, tightened the base-quality cutoff for long reads, and removed PCR duplicate reads; the underestimation by short reads persisted in every case. As expected, no meaningful difference emerged between polyA-selected and rRNA-depleted short-read libraries, or between the two cell lines, confirming that the bias tracks with sequencing technology rather than biology. A full-length DNA and RNA amplicon sequencing experiment provided direct validation: DNA amplicons, which cannot carry RNA edits, showed editing levels near zero, while RNA amplicons and long-read cDNA sequencing agreed with each other and both substantially exceeded the short-read estimates.
The mechanistic explanation converged on read length and its effect on genome alignment. Because short reads carry less sequence context, they are more likely to map to the wrong place in the reference genome, particularly within repetitive regions where many loci look alike. By computationally constructing artificial short-read datasets of 50, 75, 100 and 150 nucleotides from their own 150-nucleotide data, the team showed a strongly positive correlation between read length and measured editing level: longer reads yielded higher estimates, with fold differences reaching 12.1 between 150-nt and 50-nt reads. Crucially, when they truncated reads at the alignment stage rather than the sequence stage, recalculating mapping coordinates so that no re-alignment occurred, the correlation vanished. This pinpointed misalignment, not any property of the sequences themselves, as the culprit.
The alignment analysis went further. As read lengths decreased, millions of reads shifted from uniquely mapped to multiply mapped or unmapped statuses, and the counts of these misalignment categories correlated negatively with measured editing levels and with the Alu editing index, a widely used global metric computed with the RNAEditingIndexer tool. The same positive relationship between read length and editing quantification appeared in public short-read data from HEK293T, U2OS, HeLa, HepG2, K562, A549 and U87 cells, and in artificial datasets derived from long-read data segmented to shorter lengths, indicating the effect is universal across platforms and laboratories.
The generality of the finding was then tested against independent public resources. A paired short-read and long-read dataset covering eight human lung cancer cell lines showed higher editing quantification by long reads in every line, with nonrepetitive sites again registering zero percent under short reads versus roughly two percent under long reads. Cross-study comparisons using long-read data from GM12878 and HeLa cells against ENCODE and other short-read datasets revealed 23.8-fold and 6.4-fold excesses, respectively, of sites measured higher by long reads. The authors conclude that long-read RNA sequencing is preferable for accurate quantification of A-to-I editing levels, while acknowledging its higher cost; for laboratories constrained to short reads, they recommend raising coverage depth, restricting analyses to uniquely mapped reads, or validating key sites with full-length amplicon or Sanger sequencing.
Beyond editing, the study raises a caution for transcriptomics at large. Since read-length-dependent misalignment degrades any measurement that depends on precise genomic placement, splicing and gene expression estimates from short-read data may also harbor systematic errors that warrant closer scrutiny. Given that A-to-I editing shapes nearly 85 percent of the human transcriptome and is an emerging target for RNA-based therapeutics, the message is clear: the technology chosen to read RNA does not merely observe biology, it shapes what biology appears to say.
Subject of Research: Comparison of short-read and long-read RNA sequencing for accurate quantification of A-to-I RNA editing levels
Article Title: Short-read RNA-seq yields lower estimates of A-to-I RNA editing levels than long-read cDNA sequencing
Article References: Cheng, S., Qi, Y., Ya, J., Xia, L., Zhang, W., Xiong, Q., Liu, Q., Zhang, J., & Song, Y. (2026). Short-read RNA-seq yields lower estimates of A-to-I RNA editing levels than long-read cDNA sequencing. Advanced Biotechnology, 4(3), Article 27. https://doi.org/10.1007/s44307-026-00123-w
Image Credits: AI Generated
DOI: 10.1007/s44307-026-00123-w
Keywords: A-to-I RNA editing, RNA-seq, short-read sequencing, long-read sequencing, ADAR, alignment accuracy, read length, transcriptome, Alu repeats, cancer cell lines, Oxford Nanopore, quantification bias
Cite Scienmag News
APA MLA Chicago
Drew Townsend. (September 12, 2026). Short RNA Reads Underestimate A-to-I RNA Editing, Long-Read Sequencing Reveals. Scienmag. https://scienmag.com/short-rna-reads-underestimate-a-to-i-rna-editing-long-read-sequencing-reveals/
Drew Townsend. “Short RNA Reads Underestimate A-to-I RNA Editing, Long-Read Sequencing Reveals.” Scienmag, 12 September 2026, https://scienmag.com/short-rna-reads-underestimate-a-to-i-rna-editing-long-read-sequencing-reveals/. Accessed 12 September 2026.
Drew Townsend. “Short RNA Reads Underestimate A-to-I RNA Editing, Long-Read Sequencing Reveals.” Scienmag. September 12, 2026. https://scienmag.com/short-rna-reads-underestimate-a-to-i-rna-editing-long-read-sequencing-reveals/
Copy citation Download RIS
Tags: A-to-I RNA editingADARADAR enzymesalignment accuracyAlu repeatsAlu sequencescancer cell linesgenomic measurement accuracyinosine detectionlong-read sequencingOxford Nanoporequantification biasread lengthrepetitive elements in RNARNA editingRNA modificationRNA-seqsequencing technology comparisonshort-read sequencingshort-read sequencing biastranscriptometranscriptome analysis


