Long non-coding RNAs have long been treated as if each gene produces a single molecular identity: one RNA, one cellular address, and—at least in broad database annotations—one likely function. A computational study published in Molecular Genetics and Genomics challenges that simplifying view. The analysis argues that the subcellular localization of human long non-coding RNAs, or lncRNAs, is frequently an isoform-specific property that disappears when transcripts are combined at the gene level. In practical terms, different RNA molecules produced from the same gene may occupy opposite sides of the cell, with one isoform enriched in the nucleus and another predominantly found in the cytoplasm. The distinction could reshape how researchers interpret lncRNA function, disease associations, and regulatory mechanisms.
The study, led by Hidenori Tani of Yokohama University of Pharmacy, focused on a central problem in transcriptomics. A single lncRNA gene can generate multiple isoforms through alternative transcription start sites, splicing patterns, and 3′-end processing. These isoforms may share a large part of their sequence while differing at their ends or in the inclusion of particular exons. Such changes can alter RNA structure, binding sites for RNA-binding proteins, nuclear-retention signals, degradation rates, and interactions with ribosomes or other cellular machinery. Yet subcellular localization is commonly calculated using gene-level measurements, in which all isoforms are pooled together. If one isoform is nuclear and another cytoplasmic, the combined signal may misleadingly suggest that the gene produces an RNA with an intermediate or neutral distribution.
To investigate this issue, Tani reused tissue-specific dominant-isoform switch calls derived from long-read GTEx transcriptomes. Long-read sequencing is especially valuable in this context because it can read much more of an RNA molecule in a single fragment, helping distinguish full-length isoforms that short-read sequencing often confounds. The analysis began with 268 lncRNA genes showing tissue-dependent changes in their dominant isoform. The researcher then assigned isoform-level localization measurements using ENCODE RNA-sequencing data from separated nuclear and cytoplasmic fractions across 15 human cell lines. In total, 1,262 lncRNA isoforms received a cytoplasmic–nuclear relative concentration index, or CN-RCI, a quantitative measure describing whether an isoform is relatively enriched in the cytoplasm or nucleus.
The results exposed a striking pattern. Among 163 genes for which both tissue-dominant isoforms could be localized, 45—27.6 percent—appeared “masked” at the gene level. Their combined CN-RCI fell within a neutral range even though the individual isoforms occupied opposing compartments. A gene-level analysis would therefore describe these transcripts as neither strongly nuclear nor strongly cytoplasmic, while the isoform-resolved analysis would reveal two molecular populations with potentially different functions. This is not simply a technical distinction. Nuclear lncRNAs can influence chromatin organization, transcription, RNA processing, and the assembly of nuclear bodies, whereas cytoplasmic lncRNAs may regulate translation, messenger-RNA stability, signaling pathways, or interactions with ribosomes. Assigning both isoforms the same localization could obscure these biological roles.
The study also addresses an important statistical complication: a high percentage of apparent compartmental opposites can arise by chance. If isoforms are exchangeable and their localization assignments are randomly distributed, two isoforms may occupy different categories in roughly one-quarter of cases under certain two-compartment null models. The unfiltered 27.6 percent figure, therefore, cannot on its own demonstrate widespread within-gene localization heterogeneity. To distinguish meaningful patterns from random dispersion, Tani required localization calls to be reproducible across cell lines and evaluated the null models on the same genes being tested. Under these stricter conditions, 21 of 98 genes, or 21.4 percent, remained masked, compared with an expected 3.1 percent under the relevant null model. Under the most stringent criterion, eight of 54 genes—14.8 percent—were masked, while the expected rate fell to 0.3 percent. Both comparisons yielded a reported p value of 0.001.
Because subcellular fractionation can be affected by contamination, RNA degradation, transcript quantification, and annotation choices, the researcher tested whether the pattern survived independent datasets and analytical pipelines. The measurements reproduced in RKO cells using data generated by an unrelated laboratory, a different fractionation protocol, and a separate quantification method, producing a Spearman correlation of 0.71. Agreement was even higher in BLaER1 cells, with a correlation of 0.80. The localization estimates also remained consistent under an independent quantifier, suggesting that the result was not solely an artifact of one software package or a particular transcript-abundance model.
A further validation came from Halo-seq, a proximity-labeling approach that does not physically fractionate cells. In this method, RNA molecules are labeled according to their proximity to engineered cellular landmarks, offering an orthogonal way to estimate localization. The isoform-level measurements from fractionation and Halo-seq showed a correlation of 0.56. Although lower than the correlations between some fractionation datasets, the result was consistent with a shared biological signal measured by substantially different technologies. The unchanged computational pipeline also placed 48 of 49 control isoforms in compartments where previous studies had established their localization. In addition, it recovered the reported localization pattern at four of five genes whose individual isoforms had been characterized by experimental methods other than sequencing.
The findings appear not to be explained by several transcript features previously associated with RNA localization. The study reports that tissue-specific isoform switching did not preferentially cross the nuclear–cytoplasmic boundary; in other words, isoform changes between tissues were not systematically more likely to produce compartment changes than expected. Differences in localization also could not be attributed to Alu-repeat content, despite earlier work linking Alu-enriched sequences to nuclear retention. Nor were they explained by alternative 3′-end usage, structural mode, or translation potential. These negative results do not identify a single universal localization code. Instead, they suggest that multiple features—including RNA sequence, structure, processing history, and protein interactions—may combine differently for each gene and isoform.
The biological implications reach beyond lncRNA annotation. A nuclear isoform and a cytoplasmic isoform from the same locus could respond differently to cellular stress, developmental signals, or disease-associated changes in transcription. A tissue switch that replaces one isoform with another may alter not only RNA abundance but also the compartment in which the transcript can act. This possibility is relevant to cancer biology, where lncRNA expression is frequently used as a biomarker without distinguishing transcript variants, and to therapeutic design, where antisense oligonucleotides or RNA-targeting drugs may affect some isoforms more strongly than others. It may also help explain why studies connecting a lncRNA gene to a phenotype sometimes produce inconsistent functional conclusions: different experiments may unknowingly measure different isoforms.
The work does not claim that every lncRNA gene produces functionally segregated nuclear and cytoplasmic molecules, nor that computational localization replaces direct imaging or biochemical validation. The analysis relies on public datasets, transcript annotations, and RNA-sequencing measurements that carry intrinsic limitations. Fractionation data provide relative enrichment rather than a perfect census of individual molecules, and low-abundance isoforms remain difficult to quantify reliably. Nevertheless, the combination of long-read transcript identification, isoform-level localization indices, reproducibility filters, null-model controls, independent quantifiers, and orthogonal validation strengthens the central conclusion. For lncRNA research, the message is direct: a gene-level localization label may be too coarse to represent the biology of the transcripts it produces. As transcript-resolved sequencing becomes more accessible, subcellular localization may need to be recorded not as a fixed attribute of a gene, but as a property of individual RNA isoforms.
Subject of Research: Isoform-specific subcellular localization of human long non-coding RNAs
Article Title: Subcellular localization of human long non-coding RNAs is an isoform-resolved property masked by gene-level analysis
Article References: Tani, H. “Subcellular localization of human long non-coding RNAs is an isoform-resolved property masked by gene-level analysis.” Molecular Genetics and Genomics 301, article 177 (2026).
Image Credits: AI Generated
DOI: https://doi.org/10.1007/s00438-026-02510-3
Keywords: Long non-coding RNA; subcellular localization; isoform; alternative transcription; subcellular fractionation; CN-RCI
Tags: alternative transcription and splicing in lncRNAscomputational methods for lncRNA localizationgene-level vs isoform-level RNA analysisimpact of lncRNA isoforms on cellular functionimplications for lncRNA-related disease mechanismslong noncoding RNA isoform-specific localizationnuclear versus cytoplasmic lncRNA localizationregulation of lncRNA subcellular targetingRNA-binding protein interactions with lncRNA isoformssubcellular distribution of lncRNA isoformsTranscriptomics


