CRISPR experiments have become remarkably precise at the level of individual DNA letters, yet proving exactly what happened inside a genome remains a difficult analytical challenge. A new study in Nature Biomedical Engineering introduces CRISPRLungo, a computational approach designed to analyse long-read sequencing data from genome-editing experiments. Developed by researchers including GH Hwang, B. Vyshedskiy and T. Barry, the platform addresses a problem that is becoming increasingly important as CRISPR moves from laboratory demonstrations toward clinical and agricultural applications: edited DNA can be far more complicated than a short sequence around the intended cutting site suggests.
Most CRISPR analyses have traditionally relied on short-read sequencing. These methods generate millions of relatively small DNA fragments, often a few hundred bases long, which are then aligned to a reference genome and classified as unedited or edited. Short reads are powerful for measuring small insertions and deletions, but they can struggle when an editing event involves a large deletion, an insertion of foreign or donor DNA, an inversion, repeated sequences or several changes occurring on the same DNA molecule. Because the fragments are short, the analytical software may be unable to determine whether distant changes belong to one chromosome, to different copies of the gene or to separate cells in the sample.
Long-read sequencing offers a different view. Instead of reconstructing a genomic region from many small fragments, it can read thousands or even tens of thousands of DNA bases in a single molecule. In principle, this allows researchers to see an entire target locus, including the guide-RNA binding site, the cleavage region, nearby regulatory sequences and structural changes extending far beyond the intended edit. It can also reveal how multiple variants are physically linked. Yet long-read data introduces its own obstacles: individual reads can contain sequencing errors, the same edit may appear in multiple molecular forms, and conventional alignment tools may not represent complex rearrangements accurately. CRISPRLungo is presented as a response to this gap between what long-read sequencing can capture and what standard analysis pipelines can reliably interpret.
At the heart of the approach is the effort to classify complete editing outcomes rather than reducing every event to a single number. In a typical CRISPR experiment, a nuclease such as Cas9 is directed by a guide RNA to a chosen genomic sequence. The resulting break may be repaired through error-prone end joining, producing small insertions or deletions, or through a template-directed mechanism that introduces a planned sequence. Repair can also generate unexpected outcomes, including deletions that extend over large distances, duplications, inversions and complex combinations of these events. A long-read workflow must therefore determine where a molecule begins and ends relative to the target, identify mismatches and gaps, and distinguish genuine biological changes from technical errors.
CRISPRLungo is designed to organize these observations into an interpretable picture of the edited population. Rather than examining only a narrow window around the expected cut site, the software can analyse reads spanning a broader region and compare their structures with an unedited reference. This makes it possible to separate molecules that retain the original sequence from those carrying small indels, larger rearrangements or inserted material. The method can also preserve the connection between variants found on the same read, an important feature known as phasing. Phasing tells researchers whether two sequence changes occur together on one DNA molecule, a distinction that short-read data often cannot make without statistical inference.
That capability matters because apparently successful editing can conceal unwanted molecular diversity. A sample may show a high proportion of correctly modified reads near the target while also containing a smaller population with extended deletions or rearrangements. Such events could remove regulatory elements, disrupt neighbouring genes or create new junctions between genomic regions. In a therapeutic setting, these rare products may be biologically significant even if they represent only a small fraction of all sequenced molecules. By following the full length of individual reads, long-read analysis can make these products visible and provide a more detailed estimate of editing outcomes.
The platform is also relevant to experiments involving targeted insertion. Researchers may use CRISPR to place a therapeutic gene, a reporter sequence or another designed payload into the genome. Determining whether the inserted sequence is present is only the first step. Investigators must also establish its orientation, copy number, boundaries and relationship to the surrounding genome. A payload can integrate as intended, but it may also appear in truncated form, in multiple copies or alongside unintended rearrangements. A tool built specifically for CRISPR long-read data can inspect the junctions between the insert and the host genome, helping distinguish precise integration from partial or structurally altered products.
Another important feature of long-read analysis is its potential to connect editing outcomes with genomic context. Human cells often contain two copies of each chromosome, and the two alleles may differ naturally through single-nucleotide variants, insertions or larger structural changes. An edit may occur on one allele but not the other, or the same guide may cut the two alleles with different efficiencies. When a read spans both the editing site and informative genetic markers, researchers can determine which allele was modified. This information is particularly valuable for studying dominant mutations, allele-specific editing and diseases in which the clinical effect depends on the precise genetic background surrounding a target.
The arrival of CRISPRLungo reflects a broader shift in genome engineering, from asking whether an edit occurred to asking exactly what the edited genome looks like molecule by molecule. More complete characterization will be essential as researchers compare editing enzymes, guide RNAs and delivery systems, and as they evaluate the safety of therapies intended for patients. Long-read sequencing is not a universal replacement for short-read methods: it can be more expensive, may have lower throughput for some applications and remains sensitive to platform-specific errors. Nevertheless, by combining long-range molecular information with specialized computational interpretation, CRISPRLungo could help turn complex sequencing output into a clearer map of intended and unintended genome changes.
The study’s significance lies less in adding another sequencing pipeline than in recognizing that genome editing has outgrown simple measurements of cut efficiency. A single percentage cannot fully describe a population of cells containing precise edits, small repair scars, large deletions, inversions, insertions and unanticipated rearrangements. Tools such as CRISPRLungo offer a way to catalogue that diversity and preserve the physical relationships between mutations on individual DNA molecules. As CRISPR technologies become more powerful, the ability to inspect their consequences at this resolution may prove just as important as the ability to make the edit in the first place.
Subject of Research: Long-read sequencing analysis of CRISPR genome-editing experiments
Article Title: Analysing long-read CRISPR experiments with CRISPRLungo
Article References: Hwang, GH., Vyshedskiy, B., Barry, T. et al. Analysing long-read CRISPR experiments with CRISPRLungo. Nat. Biomed. Eng (2026). https://doi.org/10.1038/s41551-026-01776-7
Image Credits: AI Generated
DOI: https://doi.org/10.1038/s41551-026-01776-7
Keywords: CRISPR, CRISPRLungo, genome editing, long-read sequencing, structural variants, genetic engineering, DNA repair, bioinformatics, precision medicine
Tags: agricultural genome editingbioinformatics for CRISPRclinical CRISPR applicationscomputational genome editing toolsCRISPR experiment analysisCRISPRLungoDNA editing complexitygenome editing validationlong-read sequencing analysislong-read sequencing challengessequencing data interpretationstructural variation detection


