Whole-genome sequencing has long been promoted as the foundation of a new era in personalized preventive medicine, a tool capable of scanning more than six billion DNA letters to reveal an individual’s susceptibility to disease long before the first symptom appears. Yet researchers at Semmelweis University in Budapest warn that the entire enterprise rests on a reference map that may be quietly misleading the automated software on which modern genomics depends. In a study published in the journal GeroScience, the team showed that when the whole genomes of twenty healthy Hungarian individuals were compared against the internationally accepted human reference sequence, the automated analysis flagged the same severe genetic variant in every single case. Expert curation later revealed that none of these variants was real. The finding exposes a bioinformatics vulnerability that, if left unaddressed, could become a serious bottleneck for population-wide genomic screening programs of the future.
The scale of the challenge becomes clear when one considers what whole-genome sequencing actually involves. Reading a person’s genetic information means capturing roughly 40,000 genes together with the vast regulatory regions that control their function. Dr. Gyula Richárd Nagy, a clinical geneticist and associate professor in the Department of Obstetrics and Gynecology at Semmelweis University, offers a vivid illustration of the volume: if all of this genetic text were laid out side by side, it would fill approximately 4,000 books of 500 pages each. The international Human Genome Project, launched in 1990 to map this enormous text, produced its first draft in 2001, but it was not until 2022 that researchers finally assembled a substantially complete, gap-free human genome. Even now, every new genome sequenced in a clinic or research laboratory must be interpreted by comparing it against a reference, and that comparison is performed almost entirely by automated software processing massive quantities of data.
The reference sequence currently in use worldwide is known as GRCh38. It represents the standard coordinate system against which variations in any individual’s genetic makeup are assessed. The problem, as the Semmelweis researchers point out, is that GRCh38 is not an average human genome that applies equally to everyone. It is built on a single, linear reference coordinate system, and although it incorporates genetic data from multiple individuals, roughly seventy percent of its content derives from the genome of a single male donor from Buffalo in the United States. When the software encounters a healthy genetic variant in a patient that happens to be listed in the reference as an exceptionally rare minor allele, it may trigger a false alarm. In other words, if the map fails to capture the true breadth of human genetic diversity, healthy traits can easily be mistaken for defects, particularly when the interpretation is left solely to automated pipelines without expert oversight.
The Hungarian study demonstrated exactly how this vulnerability plays out in practice. The whole genomes of twenty healthy volunteers were aligned against GRCh38, and in all twenty cases the automated analysis identified the same severe genetic variant. Only after expert curation, which cross-checked the findings against additional representative population regions, was the variant shown to be erroneous. None of the twenty individuals carried a medically relevant variant at all. Dr. Balázs Győrffy, Head of the Department of Bioinformatics at Semmelweis University and co-lead author of the paper, describes the result as the demonstration of an important bioinformatics vulnerability, one that arises not from faulty sequencing technology but from the inadequacy of the reference against which the sequences are judged.
At present, this weakness does not translate into clinical harm, because the discrepancies flagged by software are routinely reviewed and analyzed manually by specialists who consult other population databases before any conclusion is drawn. Human expertise acts as a safety net, catching the false positives that the automated systems generate. However, that safety net has a cost. Dr. Nagy warns that if whole-genome screening were expanded to large populations, the volume of manual corrections required would create an unsustainable bottleneck. Each false alarm demands expert time, and multiplying those demands across millions of screened individuals could overwhelm the very workforce needed to deliver the promised benefits of preventive genomics. In this sense, the reference genome problem is not merely a technical curiosity but a potential obstacle to the future expansion of population-wide preventive screening.
The solution proposed by the researchers is a graph-based pan-genome, a fundamentally different way of representing human genetic variation. Instead of describing the DNA sequence of a single person as one linear string, a pan-genome graph depicts the common and divergent segments of many human genomes in a branching, network-like structure. Such a computational reference already exists, and it offers a much richer genetic mapping system that covers human diversity more comprehensively. Crucially, it would eliminate the class of errors observed in the Hungarian study by design, because a variant common in some populations would appear in the graph as a known branch rather than as an alarming rarity. Comparing every person’s genome to a diverse network of genomes, rather than to a single reference sequence, provides a far more accurate picture of what is genuinely rare and what is simply part of normal human variation.
Yet the existence of a pan-genome graph does not mean it is ready for routine clinical use. Dr. Győrffy identifies the key question for the future as the extent to which information technology will be able to keep pace with advances in medicine. Current everyday IT systems, he notes, simply cannot support the daily use of such massive databases. Graph genomes are computationally far more demanding than linear references, requiring more storage, more processing power and more sophisticated alignment algorithms. Bridging the gap between a working computational prototype and a system that hospitals and screening programs can rely on every day is a formidable engineering challenge, and one that the researchers argue must be tackled proactively rather than reactively.
Dr. Nagy emphasizes that actively exploring the limits of the technology is essential if this branch of preventive medicine is to provide the highest level of safety without compromises. Genetic testing is already embedded in modern medicine in more targeted forms. For certain cancers, tumor samples are analyzed to identify genetic abnormalities that indicate which targeted therapy is likely to be effective. Genetic testing also plays an important role in diagnosing and treating rare conditions, in detecting hereditary diseases, and in assessing risk among affected family members, enabling early detection and preventive interventions in fields such as cardiogenetics. These applications demonstrate the clinical value of reading genetic information, but they operate on a far smaller scale than the whole-population vision that whole-genome sequencing promises.
The fundamental approach of future preventive and predictive medicine goes beyond treating disease once it has developed. The researchers argue that its core premise is knowing the risk before symptoms appear, and knowing who needs timely attention and for what. Whole-genome sequencing could give every person a personalized guide to their disease risks and predispositions, opening the door to genuinely individualized prevention strategies. That future, however, depends on a reference map worthy of the name. Until a graph-based pan-genome can be deployed routinely, and until the IT infrastructure catches up with the medical ambition, every automated genome analysis will carry the risk of mistaking healthy human diversity for disease, a risk that only expert human judgment currently stands between. The study, published in GeroScience under DOI 10.1007/s11357-026-02436-z, is a timely reminder that the tools we use to read the book of life must be as sophisticated as the book itself.
Subject of Research: Limitations of the GRCh38 human reference genome in automated whole-genome sequencing and the case for a graph-based pan-genome
Article Title: We may need a new map to properly interpret our genome
Article References: We may need a new map to properly interpret our genome. (n.d.). Original publication
Image Credits: AI Generated
DOI: Not provided
Keywords: whole-genome sequencing, reference genome, GRCh38, pan-genome, bioinformatics, preventive medicine, genetic variants, false positives, personalized medicine, GeroScience, Semmelweis University, genomic screening
News Source: Juliet Wilcox. (October 7, 2026). A Flawed Genetic Reference Could Trigger False Alarms in Whole-Genome Screening. Scienmag.



