Every day, in hospitals around the world, laboratories report a number that can set off a cascade of anxiety: the serum level of CA125, a protein marker most closely associated with ovarian cancer. For decades, that number has been judged against a single, largely universal upper limit, typically around 35 units per milliliter, a threshold inherited from small studies conducted in populations very different from the patients being tested today. A new study published in BMC Cancer argues that this one-size-fits-all approach is overdue for revision. Drawing on real-world health-checkup data from more than 27,000 individuals, researchers at the First Affiliated Hospital of Zhejiang University School of Medicine, working with colleagues at Dalian University, have established sex- and age-stratified reference intervals for CA125 using five different statistical algorithms, and then tested how well those new intervals perform in actual clinical practice.
The scale of the dataset is what sets the study apart. The establishment set comprised 27,426 healthy health-checkup participants examined between January 2024 and December 2025 at the Zhejiang hospital. Rather than relying on a small, hand-picked cohort of presumed healthy volunteers, the team mined the laboratory’s routine data stream, an approach known as an indirect method because it infers the healthy distribution of a biomarker from a mixed population without requiring explicit confirmation that every individual is disease-free. Indirect methods have gained momentum in laboratory medicine precisely because modern hospital information systems generate enormous volumes of test results that would be impossible to replicate with traditional, labor-intensive reference interval studies.
Technically, the researchers applied five algorithms spanning two generations of methodology. The traditional pair, designated EP28-NP and EP28-P, follows the nonparametric and parametric approaches outlined in the Clinical and Laboratory Standards Institute’s EP28 guideline, the long-standing standard for defining reference intervals. The modern trio consisted of refineR, Kosmic, and TMC, algorithms that use sophisticated mixture modeling to separate the underlying healthy distribution from pathological values buried within routine data. Each algorithm estimates where the central 95 percent of healthy values lie, and the upper boundary of that range, the 97.5th percentile, becomes the proposed upper reference limit. By running all five algorithms on the same data and comparing their outputs, the team could assess how sensitive their conclusions were to the choice of statistical method, a form of robustness testing that few single-algorithm studies can offer.
The results confirm something gynecologists have long suspected: CA125 does not behave the same way in everyone. The marker shows significant sex differences and dynamic variation around menopause. Using the 97.5th percentile derived from the EP28-NP algorithm as an example, the study proposes three distinct upper limits. For males aged 18 and older, the limit is 19.2 units per milliliter. For women of reproductive age, between 18 and 49 years, it rises to 31.9 units per milliliter, reflecting the well-documented tendency of CA125 to fluctuate with the menstrual cycle and to be mildly elevated in conditions such as endometriosis. Strikingly, for postmenopausal women aged 50 and older, the proposed limit drops sharply to 19.4 units per milliliter, nearly matching the male value. In other words, a postmenopausal woman whose result sits comfortably below the traditional 35-unit threshold may in fact be well above what is normal for her age group, a distinction that could matter enormously in the early detection of ovarian cancer.
Consistency among the five algorithms proved generally strong. The modern algorithms, refineR, Kosmic, and TMC, tended to produce slightly higher reference limits than the traditional methods, but the absolute bias between modern and traditional estimates remained below 0.375 in all but eight comparisons, seven involving TMC and one involving Kosmic. That level of agreement, the authors report, indicates good inter-algorithm consistency and suggests that the stratified intervals are not an artifact of any single statistical approach. For laboratory directors weighing whether to adopt data-driven reference intervals, this kind of multi-algorithm cross-validation offers a practical template: if several independent methods converge on similar limits, the case for updating local reference ranges becomes far more persuasive.
To evaluate clinical applicability, the researchers assembled two additional datasets. The evaluation set included health-checkup participants, outpatients, and inpatients tested for CA125 between January and April 2026, and was used to measure the discrepancy rate between the newly established intervals and the laboratory’s preset ranges. Those discrepancy rates came in at 6.74 percent for outpatients, 7.01 percent for inpatients, and 8.41 percent for the disease set, calculated as medians. In practical terms, roughly one in every fifteen results would be interpreted differently under the new stratified limits than under the old ones, a substantial fraction given how many CA125 tests are ordered annually. The disease set comprised patients with eight CA125-related diseases diagnosed between January 2022 and December 2025, allowing the team to characterize how well the marker discriminates between disease and health using receiver operating characteristic analysis.
The ROC analysis delivered a nuanced verdict on CA125’s diagnostic power. For ovarian cancer overall, the marker performed well, achieving an area under the curve of 0.858 with an optimal cut-off above 29.5 units per milliliter, a value notably lower than the conventional 35-unit limit and closer to the stratified reproductive-age limit of 31.9. However, the study also quantified a persistent weakness: for stage I to II ovarian cancer, the area under the curve fell to 0.740, indicating limited discriminative ability for early-stage disease. This finding echoes a long-standing frustration in gynecologic oncology, where CA125’s modest sensitivity in early-stage tumors has limited its usefulness as a standalone screening test. The marker also provided clinically useful indications for lung cancer, pancreatic cancer, and endometriosis, with areas under the curve of at least 0.711 and an optimal cut-off above 22.1 units per milliliter, underscoring that CA125 is far from ovarian-specific and that interpretation must always account for clinical context.
What does this mean for the average patient? The most immediate implication concerns the borderline zone, the gray band of results that are neither clearly normal nor clearly alarming. Under the current single-threshold system, a postmenopausal woman with a CA125 of 30 units per milliliter might be reassured that her value is below 35, even though it exceeds the proposed postmenopausal limit of 19.4 by more than 50 percent. Conversely, a premenopausal woman with a value of 33 might be flagged as abnormal when the new data suggest she is within the normal range for her age. Stratified reference intervals, the authors conclude, may help refine the interpretation of borderline results and support more individualized risk assessment, though they are careful to note that prospective validation is still needed before the new limits should be adopted into routine practice.
The study also carries a broader message about how laboratory medicine is evolving. The combination of laboratory big data with multi-algorithm evaluation, the researchers argue, is a reliable approach for optimizing reference intervals not just for CA125 but potentially for the hundreds of other analytes whose reference ranges rest on decades-old studies. Hospitals already generate the raw material; what has been missing is a validated framework for extracting trustworthy normal ranges from it. By demonstrating agreement across five algorithms, quantifying discrepancy rates in three distinct patient populations, and linking the new limits to disease-specific diagnostic performance, the Zhejiang team has built exactly such a framework. If prospective studies confirm their findings, the humble CA125 report may soon come with context that has been missing since the marker was first introduced: an understanding of what normal actually means for the person being tested, whether a young menstruating woman, a man in midlife, or a woman past menopause for whom a lower threshold could one day mean an earlier diagnosis.
Subject of Research: Establishment of sex- and age-stratified reference intervals for serum CA125 using big data and five statistical algorithms
Article Title: Establishment of stratified reference intervals for serum CA125 using five algorithms based on real-world big data and evaluation of clinical performance
Article References: Tang, Z., Qi, X., Fan, L., & Yang, D. (2026). Establishment of stratified reference intervals for serum CA125 using five algorithms based on real-world big data and evaluation of clinical performance. BMC Cancer. https://doi.org/10.1186/s12885-026-17036-5
Image Credits: AI Generated
DOI: 10.1186/s12885-026-17036-5
Keywords: CA125, reference intervals, ovarian cancer, biomarkers, laboratory medicine, big data, indirect methods, refineR, Kosmic, ROC analysis, diagnostic markers, menopause
News Source: Nathaniel Bowman. (October 10, 2026). Big Data Study Redraws the Normal Range for Ovarian Cancer Marker CA125. Scienmag.



