Medical imaging is entering a new era in which scans are no longer used only to produce pictures for physicians to interpret. Increasingly, images are being converted into precise numerical measurements that can describe tumors, quantify how tissues function and track how a disease responds to treatment. Artificial intelligence is accelerating this transformation, generating new tools designed to extract information that may be invisible to the human eye. Yet the promise of these technologies has created a difficult scientific problem: How can researchers determine whether an imaging tool is reliable when the true value of what it is measuring is unknown?
A team at Washington University in St. Louis has developed a method intended to solve that problem. The technique, called NGSE-Corr, allows researchers to compare quantitative imaging methods without relying on a gold standard—the definitive, independently verified value against which other measurements are judged. The approach could help scientists evaluate AI-powered imaging systems, assist physicians in selecting more dependable tools and give regulators a way to assess emerging technologies before they enter widespread clinical use. The findings were reported in IEEE Transactions on Medical Imaging in a study led by Abhinav Jha, an associate professor of biomedical engineering at the McKelvey School of Engineering and of radiology at WashU Medicine’s Mallinckrodt Institute of Radiology.
The need for an alternative to the gold standard is widespread in medicine. Consider the challenge of measuring a cancerous tumor inside a living patient. The actual size, biological activity or treatment response of the tumor cannot always be known with complete certainty. A biopsy may sample only a small portion of a complex mass, while a highly accurate reference measurement may require invasive procedures, surgery or extensive follow-up. In other situations, the desired quantity may not be directly observable at all. Researchers may therefore have several imaging tools that measure the same clinical property but no absolute benchmark that reveals which one is closest to reality.
Without such a benchmark, conventional validation can become difficult. A method may agree with another method without either being accurate, or it may appear inconsistent because of variations in image quality, patient anatomy or scanning conditions. Quantitative imaging systems also contain measurement noise: random fluctuations that arise from the scanner, reconstruction algorithms, biological motion and other sources. When several tools are applied to the same patient or tumor, however, their errors are not necessarily independent. Because the instruments observe the same underlying anatomy and are often affected by common conditions, their fluctuations can be correlated.
That insight is central to NGSE-Corr. The method builds on an earlier mathematical formulation for evaluating quantitative imaging, but modifies it to account explicitly for correlated noise among measurements. In simplified terms, the technique examines how different imaging methods vary across repeated or related observations and uses those patterns to estimate their relative precision. Rather than asking whether a measurement matches a known truth, NGSE-Corr asks which method produces the most dependable information under the same clinical circumstances. This distinction is important because precision—the consistency of a measurement—is often assessable even when absolute accuracy cannot be directly established.
Jha and his collaborators, including first author Yan Liu, tested the method through numerical experiments designed to mimic different measurement conditions. Their results indicated that NGSE-Corr could correctly rank imaging methods according to precision, even when the researchers withheld knowledge of the simulated ground truth. The goal was not simply to identify whether an individual measurement was correct, but to determine which of several competing methods was best suited to the task. That ranking capability could be especially useful in fields where new algorithms appear rapidly and where performance may vary depending on the disease, the imaging protocol or the clinical question.
The researchers then moved from mathematical simulations to a virtual imaging trial. They generated computer-based patients with bone-metastatic castration-resistant prostate cancer who were treated with radium-223, a radioactive therapy used in certain cases of advanced disease. The trial compared three quantitative single-photon emission computed tomography, or SPECT, methods for measuring regional activity uptake. In this context, the imaging systems were being evaluated for how precisely they could quantify the distribution of radioactivity in different regions of the body—information that may help researchers understand treatment delivery and response.
The virtual trial produced striking results. When the methods were evaluated in groups of 50 computer-generated patients, NGSE-Corr correctly ranked the imaging approaches in 91% of the trials without being given the underlying ground truth. It identified the most precise method in 95% of the trials. Increasing the number of virtual patients improved performance further, suggesting that larger clinical datasets may allow the method to distinguish between competing tools with greater confidence. Although virtual trials cannot replace carefully designed studies involving real patients, they provide a controlled environment in which researchers can test whether an evaluation strategy behaves as expected.
The implications extend beyond SPECT or prostate cancer. Quantitative imaging is being used to estimate tumor volume, blood flow, tissue composition, metabolic activity and other clinically relevant properties across radiology and nuclear medicine. AI systems are also being developed to transform images into risk scores and treatment predictions. Such systems may be highly sensitive to the data used to train them, the characteristics of the patient population and the technical details of image acquisition. A method that performs well in one hospital may be less reliable elsewhere. By enabling comparisons without requiring a perfect reference measurement, NGSE-Corr could offer a practical way to monitor and rank these tools across diverse settings.
The researchers say the technique could ultimately strengthen confidence in medical imaging technologies while reducing the cost and time associated with traditional validation. For developers, it may provide an objective framework for comparing algorithms during innovation. For physicians, it could clarify which measurements are most dependable when several tools are available. For regulators, it may offer additional evidence when assessing AI-backed products whose outputs cannot easily be checked against an unquestionable biological truth. The method does not eliminate the need for clinical validation or establish accuracy by itself, but it addresses a major gap: evaluating relative performance when the gold standard is unavailable. As medical images increasingly become sources of numerical data rather than pictures alone, tools such as NGSE-Corr could help ensure that those numbers are precise enough to support decisions that affect patient care.
Subject of Research: A method for evaluating the precision and reliability of quantitative medical imaging tools, including AI-based systems, without a gold standard.
Article Title: NGSE-Corr: A Technique for Objective Clinical Evaluation of Quantitative-Imaging Methods Without a Gold Standard
Web References: Washington University in St. Louis; Abhinav Jha profile; IEEE Transactions on Medical Imaging article: https://doi.org/10.1109/TMI.2026.3707743
References: Y. Liu et al., “NGSE-Corr: A technique for objective clinical evaluation of quantitative-imaging methods without a gold standard,” IEEE Transactions on Medical Imaging. DOI: 10.1109/TMI.2026.3707743
Keywords
Medical imaging, artificial intelligence, quantitative imaging, NGSE-Corr, correlated noise, image analysis, SPECT, prostate cancer, virtual clinical trials, machine learning, radiology, medical technology
Tags: AI in medical diagnosticsAI-powered medical imaging evaluationbiomedical engineering in medical imagingdisease response trackingimaging tool validation without gold standardmedical imaging measurement reliabilitymedical imaging technology regulationNGSE-Corr imaging assessment methodquantitative imaging in healthcarereliability assessment of emerging imaging toolstissue function quantificationtumor measurement accuracy




