Predicting whether a cancer patient will respond to therapy has long been one of oncology’s most consequential gambles. For people with metastatic non-small cell lung cancer, the most common and deadliest form of the disease, first-line chemoimmunotherapy can produce dramatic tumor shrinkage in some patients while leaving others with little benefit and substantial toxicity. Now, a research team at the University of Washington and Fred Hutchinson Cancer Center has built a prototype clinical decision support system that fuses three very different biological signals into a single prediction, and, crucially, tells clinicians exactly how much it trusts each answer.
The study, published in the Journal of Translational Medicine, drew on data from the PET-BRIGHT clinical trial, in which thirty-five patients with metastatic non-small cell lung cancer received the standard combination of carboplatin, pemetrexed, and the immunotherapy pembrolizumab. Each patient underwent a FDG-PET/CT scan and a blood draw at two moments: before treatment began, and again after just three weeks, a single cycle of therapy. From these timepoints the researchers extracted three families of biomarkers: imaging metrics that quantify how much glucose metabolically active tumor tissue consumes, measures of T-cell receptor diversity that capture the breadth of the immune system’s anti-cancer repertoire, and a panel of inflammatory cytokines circulating in the blood.
The choice of these three modalities reflects a growing recognition that no single biomarker can capture the complexity of immunotherapy response. FDG-PET imaging measures uptake of a radioactive glucose analog, revealing where and how aggressively tumor cells are metabolizing energy, with summary metrics such as standardized uptake value and total lesion glycolysis summarizing whole-body tumor burden. T-cell receptor sequencing reads out the diversity of immune recognition, on the theory that a richer repertoire of tumor-targeting T cells signals a stronger response to immune checkpoint blockade. Cytokines, the small signaling proteins through which immune cells communicate, offer a systemic window into inflammatory states that may either support or suppress anti-tumor immunity.
Against these signals, the researchers benchmarked the current clinical standard: PD-L1 tumor proportion score, an immunohistochemistry measurement that guides treatment selection in lung cancer today. The results were sobering for the status quo. PD-L1 achieved an area under the receiver operating characteristic curve of just 0.58, barely better than a coin flip at distinguishing responders from non-responders. This weak discrimination, the authors note, is precisely what motivates a multimodal, uncertainty-aware approach, because patients and clinicians currently must commit to months of therapy armed with a biomarker that carries limited predictive weight.
The machine learning pipeline itself was deliberately conservative, a design choice appropriate for a small cohort. Using nested leave-one-out cross-validation with automated feature selection, the team identified a single representative biomarker per modality, guarding against the overfitting that plagues models trained on dozens of features and only a handful of patients. Classification was performed with class-balanced logistic regression, an interpretable method chosen over more opaque architectures, and performance was assessed both by discrimination, how well the model separated responders from non-responders, and by calibration, whether predicted probabilities matched observed outcomes as measured by the Brier score.
The headline innovation is the application of conformal prediction, a statistical framework that converts a model’s raw output into prediction sets with formal guarantees. Rather than declaring that a patient will respond or not, the conformal framework targets a specified reliability level, in this case eighty percent, and returns either a confident single-class prediction or a set containing both possibilities, signaling that the model cannot confidently classify that patient. The proportion of confident singleton predictions therefore becomes a direct, interpretable measure of uncertainty at the level of the individual patient, not merely the population average that conventional accuracy metrics provide.
The performance results chart a coherent story about how biomarker information evolves over treatment. Before therapy began, T-cell receptor diversity was the strongest single modality, reaching an AUROC of 0.80, followed by PET imaging at 0.76, while cytokines lagged at 0.54. Combining PET and T-cell data through late fusion, in which separately trained modality-specific models are merged, pushed pre-treatment discrimination to 0.85, and raised the proportion of confident singleton predictions from 78 percent to 89 percent. After the three-week blood draws and scans were added, the hierarchy shifted: cytokines became the strongest single modality at 0.78 while T-cell metrics were attenuated at 0.66, and late fusion of PET and cytokines achieved the highest overall discrimination of 0.86. Thirteen of twenty-two model and timepoint combinations beat a permutation-based null at the conventional significance threshold, and empirical coverage landed at 78 percent and 82 percent, close to the eighty percent reliability target the conformal framework was set to honor.
That dynamic shift in which modality carries the most predictive weight is itself scientifically interesting. It suggests that pre-treatment immune repertoire breadth anticipates who will benefit from checkpoint blockade, but that once therapy begins, the inflammatory milieu measured in plasma becomes the more informative readout of whether the treatment is taking hold. A static biomarker panel frozen at baseline would miss this evolution entirely, and the study’s longitudinal design, with its week-three reassessment, demonstrates how a decision support system could update its confidence as new evidence accumulates over the first critical weeks of therapy.
The authors are appropriately careful about the limitations of what they have built. Thirty-five patients from a single center constitute a proof-of-concept feasibility study, and the findings are explicitly framed as hypothesis-generating rather than practice-changing. Small cohorts make any machine learning model vulnerable to optimistic bias, which is why the permutation testing and conservative feature selection matter so much here. Before such a system could inform real treatment decisions, it would need validation in larger, multi-institution cohorts spanning different tumor types, treatment regimens, and demographic populations, along with calibration checks to confirm that the conformal guarantees hold outside the development data.
Even so, the prototype points toward a future in which clinical AI systems do something more honest than issuing confident verdicts. By pairing multimodal biomarker fusion with formal uncertainty quantification, the framework identifies not only which patients are likely to respond, but also, and perhaps more valuably, those for whom no confident prediction can be made, flagging them for closer monitoring or alternative strategies. In a disease where the current standard biomarker performs barely better than chance, knowing the limits of prediction may prove as clinically important as the predictions themselves. The interface prototype developed alongside the statistical framework shows how such uncertainty could be surfaced directly to oncologists, turning a black-box risk score into a transparent, caveat-aware recommendation that clinicians can weigh against the realities of each patient’s care.
Subject of Research: Uncertainty-aware multimodal prediction of chemoimmunotherapy response in metastatic non-small cell lung cancer
Article Title: Toward uncertainty-aware clinical decision support for treatment response prediction in metastatic NSCLC: integrating FDG-PET, T-cell repertoire, and cytokines with conformal prediction
Article References: Yaseen, F., Hippe, D. S., Cui, S., Fu, J., Kim, Y., Grassberger, C., Deng, L., Ye, T., Kinahan, P. E., Zeng, J., Gennari, J. H., & Bowen, S. R. (2026). Toward uncertainty-aware clinical decision support for treatment response prediction in metastatic NSCLC: integrating FDG-PET, T-cell repertoire, and cytokines with conformal prediction. Journal of Translational Medicine. https://doi.org/10.1186/s12967-026-09030-z
Image Credits: AI Generated
DOI: 10.1186/s12967-026-09030-z
Keywords: metastatic non-small cell lung cancer, FDG-PET, T-cell receptor repertoire, cytokines, conformal prediction, clinical decision support, immunotherapy, PD-L1, biomarkers, machine learning, uncertainty quantification, treatment response prediction
Cite Scienmag News
APA MLA Chicago
Nathaniel Bowman. (September 27, 2026). AI Tells Doctors When It Is Unsure About Cancer Treatment Success. Scienmag. https://scienmag.com/ai-tells-doctors-when-it-is-unsure-about-cancer-treatment-success/
Nathaniel Bowman. “AI Tells Doctors When It Is Unsure About Cancer Treatment Success.” Scienmag, 27 September 2026, https://scienmag.com/ai-tells-doctors-when-it-is-unsure-about-cancer-treatment-success/. Accessed 27 September 2026.
Nathaniel Bowman. “AI Tells Doctors When It Is Unsure About Cancer Treatment Success.” Scienmag. September 27, 2026. https://scienmag.com/ai-tells-doctors-when-it-is-unsure-about-cancer-treatment-success/
Copy citation Download RIS
Tags: AI confidence in medical predictionsassessing treatment efficacy in lung cancerbiomarker data integration in cancer therapyBiomarkerscancer treatment response predictionclinical decision supportconformal predictioncytokinesearly response indicators in cancer therapyFDG PETimmune system biomarkers in cancer treatmentImmunotherapyMachine learningmachine learning for cancer treatment outcomesmetastatic non-small cell lung canceroncology clinical decision support systemsPD-L1personalized cancer therapy decision toolsPET imaging in cancer prognosisT-cell receptor repertoiretreatment response predictiontumor response prediction using AIuncertainty quantification


