Cancer has long been studied through the lens of gene expression, the simple question of whether a gene is switched on or off in a tumor compared with healthy tissue. But genes do not speak in a single voice. Each one can produce multiple transcript isoforms, distinct RNA molecules assembled through alternative splicing, and these variants can carry dramatically different functions. A new pan-cancer study published in the International Society for Health Data Science’s journal has now mapped this hidden layer of transcriptomic chaos across ten human organs, and its central finding is striking: the way genes change their expression during tumorigenesis follows fundamentally different rules from the way their isoforms are remodeled. The work, based on Nanopore long-read RNA sequencing of 144 paired tumor-normal tissue sets, was published on 30 September 2026 under DOI 10.1002/imm3.70061.
The technical foundation of the study is what sets it apart. Conventional short-read RNA sequencing, the workhorse of most large cancer atlases including TCGA, chops RNA molecules into fragments of a few hundred bases and then computationally reconstructs transcripts by aligning those fragments to a reference genome. This approach excels at quantifying aggregated gene-level expression but routinely collapses the identity of individual full-length transcripts. Long-read sequencing, by contrast, reads entire RNA molecules from end to end, preserving the complete exon chain of each transcript. The researchers applied Nanopore-based long-read RNA sequencing to their tumor-normal pairs and subjected the resulting data to stringent quality filtering to remove outlier specimens. Validation was multi-pronged: the transcript read-length distribution confirmed full-length capture, the 5-prime ends of reads aligned with known CAGE peaks marking transcription start sites, the 3-prime ends showed expected poly-A motif occupancy, and cross-platform correlation against matched short-read RNA-seq data confirmed that gene-level quantification was technically sound.
From this carefully validated dataset, the team detected nearly 300,000 isoforms, and a substantial proportion of them represented previously unannotated entities. These included novel splice variants never cataloged in reference annotations as well as tumor-associated fusion transcripts, hybrid molecules created when genomic rearrangements join parts of different genes. Across the cancer types examined, the researchers recovered thousands of tumor-specific isoforms covering all seven major classes of alternative-splicing events, alongside hundreds of recurrent fusion transcripts. Critically, the authenticity of a subset of these fusion transcripts was supported by partial experimental validation using PCR followed by Sanger sequencing, a gold-standard confirmation method that lends credibility to the computational discoveries.
The study’s headline discovery emerges from pan-cancer comparative analysis, which exposed two uncoupled layers of transcriptomic dysregulation. At the gene expression level, tumors arising from diverse organs displayed convergent transcriptional signatures. In other words, a liver tumor and a lung tumor, despite their radically different origins, tend to switch similar sets of genes up or down. This convergence is tightly linked to oncogenic dedifferentiation, the cellular reversion toward stem-like phenotypic states that is a hallmark of aggressive malignancy. Many of the genes convergently down-regulated across multiple tumor types participate in tissue-specific physiological metabolism and ion homeostasis, functions that specialized cells abandon as they lapse into a primitive, proliferative program.
Isoform-level remodeling told a completely different story. Rather than converging, isoform composition shifted in a highly divergent, organ-restricted manner, with each tissue type exhibiting its own characteristic splicing perturbations. The most compelling evidence came from cross-organ correlation analysis: the correlations between isoform ratios across different tumor types resembled those observed in normal tissues, contradicting the heightened inter-tumor correlation seen for gene expression. This key observation demonstrates that widespread isoform shifts occur largely independently of bulk messenger-RNA abundance changes. Because conventional short-read workflows mostly quantify aggregated gene-level counts rather than individual full-length transcripts, this entire dimension of cancer biology has been systematically invisible to most previous pan-cancer investigations.
What drives this extensive isoform reprogramming? One major force is the dysregulation of splicing factors, the proteins that decide which exons are included or excluded during RNA processing. The researchers built a random-forest regression model trained on paired tumor-normal expression profiles to quantify how predictive the expression of more than four hundred splicing-factor genes was for genome-wide isoform-composition shifts. The analysis pinpointed a small subset of high-impact regulators, including HNRNPC and RBFOX2, both of which are already known to govern alternative-splicing circuits in cancer. The machine-learning approach thus provides a quantitative ranking of which splicing regulators exert the strongest influence on the isoform landscape in each tumor context.
Perhaps the most clinically provocative finding concerns genes whose isoform ratios change dramatically even when their total expression remains perfectly stable. The gene DUT, which encodes an enzyme involved in nucleotide metabolism, exemplified this pattern: pronounced isoform compositional remodeling took place in liver tumors while total gene transcript abundance stayed flat. Such cases underscore a critical limitation of short-read-oriented pan-cancer research. Many oncogenic transcriptomic events manifest exclusively through altered transcript-isoform proportions and structural variants, and they are simply invisible when analytical pipelines collapse all sequencing reads onto aggregated gene models. A cancer study that only measures gene expression is, in effect, listening to the average volume of an orchestra while ignoring which instruments are playing.
The authors are appropriately careful about the limits of bulk sequencing. All long-read measurements in this study represent population-averaged signals derived from heterogeneous tissue biopsies containing malignant cells, stromal cells, and immune infiltrates. Consequently, detected isoform alterations and fusion transcripts cannot be definitively mapped solely to cancer cells; some shifts may partly reflect variations in tumor purity and cellular composition within specimens. Even so, sensitivity analyses using deconvolution approaches confirmed that the core isoform-ratio trends persisted irrespective of purity fluctuations. The exact cellular origin of most detected transcript-level aberrations awaits future single-cell long-read RNA sequencing experiments, which would resolve intratumoral heterogeneity and assign molecular events to defined cell subpopulations.
To translate these multi-layered transcriptomic dimensions into biologically meaningful priorities, the team constructed a custom pan-cancer scoring framework. The framework merges gene-level differential expression, convergent-divergent variance patterns, tumor-specific isoform occurrence, fusion-transcript evidence, and isoform-ratio perturbations, supplemented with weighted evidence retrieved from TCGA and the Cancer Gene Census database. Each gene received an organ-specific score capped at one, and these were aggregated across all ten organs to compute a final pan-cancer susceptibility score. More than ten thousand genes were scored and stratified into low-, medium-, and high-score groups. High-score entries included well-characterized canonical cancer drivers such as TP53, RB1, and FAT4, providing an internal sanity check for the method.
More notably, the scoring system unearthed under-appreciated susceptibility genes whose dysregulation is dominated by isoform-level defects rather than total gene-expression change. Representative examples include IBSP, involved in extracellular-matrix organization and metastasis, and FANCL, which participates in DNA-damage-repair pathways; both display recurrent tumor-specific isoforms across multiple cancer types without prominent gene-level dysregulation. Kaplan-Meier survival analysis demonstrated that selected high-score genes carried significant prognostic implications, whereas near-zero-score genes showed no survival association, confirming the framework’s capacity to nominate prognosis-relevant molecules. The authors caution that the resource is not intended for early-cancer screening or treatment-response prediction, given the cohort’s lack of complete clinical follow-up metadata. Even with that caveat, the study delivers a public atlas for isoform-centered cancer research and a shortlist of actionable candidates for downstream functional validation, biomarker development, and therapeutic targeting, while making a forceful case that the next generation of cancer genomics must read transcripts in full.
Subject of Research: Pan-cancer long-read transcriptomics of gene and isoform expression changes in tumorigenesis
Article Title: Long-read pan-cancer transcriptomics unravel distinct alteration trends between gene and isoform expression in tumorigenesis
Article References: Long-read pan-cancer transcriptomics unravel distinct alteration trends between gene and isoform expression in tumorigenesis. (n.d.). Original publication
Image Credits: AI Generated
DOI: Not provided
Keywords: long-read sequencing, Nanopore, pan-cancer, isoform expression, alternative splicing, fusion transcripts, tumorigenesis, splicing factors, transcriptomics, cancer biomarkers, dedifferentiation, RNA sequencing
News Source: Juliet Wilcox. (October 9, 2026). Long-Read RNA Sequencing Reveals Hidden Isoform Chaos That Gene-Level Cancer Studies Miss. Scienmag.



