For decades, the fate of a clinical trial has hinged on a single, almost ritualistic question: did the p-value fall below 0.05? If it did, the trial was positive and the treatment declared a success; if it did not, the trial was negative and the treatment, however promising, was consigned to the statistical graveyard. A new study published in BMC Medicine argues that this binary logic is quietly distorting medicine, and it offers a strikingly simple alternative: tell doctors and patients, in plain language, the probability that a treatment actually works. The research, led by Orestis Efthimiou of the University of Bern together with colleagues from Bern and Geneva, re-analyzed hundreds of published trials using Bayesian methods and found that the conventional labels of positive and negative often conceal the truth about whether a therapy helps patients.
The technical heart of the approach lies in the difference between frequentist and Bayesian inference. A p-value answers a rather convoluted question: how unlikely are the observed data if the treatment has no effect at all? It does not answer the question anyone actually cares about, namely whether the treatment works. Bayesian analysis flips the logic. By combining the trial data with a prior distribution, a mathematical summary of what was believed before the trial, it produces a posterior distribution that describes the uncertainty around the true treatment effect. From that posterior, researchers can compute directly interpretable probabilities: the chance that a new drug beats placebo, the chance that it reduces the risk of death by at least twenty percent, or the chance that it causes a specific harm. A Bayesian 95% credible interval can be read as there being a 95% probability that the true effect lies within it, whereas a frequentist confidence interval technically refers to the behavior of intervals across infinite hypothetical repetitions of the study, a subtlety that even seasoned researchers routinely get wrong.
To test what this means in practice, the team re-analyzed two large collections of published trials. The first comprised 169 primary outcomes from 130 trials published in six leading general medical journals in 2021, all of which had been deemed negative because their p-values exceeded 0.05. The second covered 234 primary outcomes from 225 oncology trials published between 2009 and 2019, all statistically significant and all reporting hazard ratios for survival. For each outcome, the researchers fitted a Bayesian model combining the published effect estimate and its standard error with an uninformative prior, then calculated the posterior probability that the experimental treatment was superior to the control. In sensitivity analyses they repeated everything with a strongly skeptical prior, giving 95% prior probability that the true effect ratio lay between 0.41 and 2.42, to check that the conclusions were not artifacts of a permissive starting assumption.
The results were startling. Among the supposedly negative trials, 26 of 169 outcomes, about 15 percent, showed a greater than 90 percent probability that the experimental treatment was genuinely better than the control, and for 12 outcomes, 7 percent, that probability exceeded 95 percent. Even under the heavily skeptical prior, 20 of the 113 ratio-based outcomes retained a probability above 90 percent. When the analysis was restricted to survival outcomes among the negative trials, the picture grew starker still: in 5 of 18 cases, 28 percent, the probability that the treatment improved survival was at least 90 percent. In other words, evidence that most clinicians would read as absence of benefit was, in a meaningful fraction of cases, fairly strong evidence of benefit that had simply failed to cross an arbitrary statistical threshold.
Two case studies illustrate how consequential this can be. The first involved a trial of blinatumomab, an immunotherapy substituted for multiagent chemotherapy during consolidation treatment in children, adolescents, and young adults with relapsed B-cell acute lymphoblastic leukemia. The reported hazard ratio for disease-free survival was 0.70 with a 95% confidence interval of 0.47 to 1.03, and because the interval included 1, the authors concluded there was no statistically significant difference. The Bayesian re-analysis, however, estimated a 96 percent probability that blinatumomab was the better treatment, and a 74 percent probability that it reduced the hazard of relapse by at least 20 percent. Treating such evidence as null, the researchers argue, could deprive young patients of a potentially life-extending therapy.
The second case was arguably even more dramatic. A fetal surgery trial compared fetoscopic endoluminal tracheal occlusion, or FETO, against standard prenatal management for moderate left congenital diaphragmatic hernia. The primary outcomes, survival to neonatal intensive care discharge and survival without oxygen supplementation at six months, both favored FETO but were statistically non-significant, with relative risks of 0.79 and 0.81 respectively. The trial also found that FETO increased the risks of preterm prelabor rupture of membranes and preterm birth, which were significant. The Bayesian re-analysis estimated a 97 percent probability that FETO increases survival at discharge and a 96 percent probability that it increases survival without oxygen at six months, with 86 and 84 percent probabilities of absolute survival gains exceeding five percentage points. The authors of the new study warn that dismissing such evidence as non-significant, and continuing to withhold a procedure that appears to save babies lives because it also raises preterm birth risk, is highly questionable, and that dichotomizing evidence by p-value thresholds may contribute to avoidable deaths.
The flip side of the problem emerged from the oncology trials. Although all 234 outcomes were statistically significant, with posterior probabilities above 97 percent that the experimental treatments beat the controls, statistical superiority is not the same as clinical worth. Using a smallest worthwhile effect of a hazard ratio of 0.80, a benchmark recommended by the American Society of Clinical Oncology in 2014 for metastatic solid tumors, the team found that in 25 of 234 outcomes, 11 percent, there was less than a 50 percent probability that the treatment effect reached that threshold. Under the skeptical prior this rose to 13 percent. A detailed example involved adding cetuximab to FOLFIRI chemotherapy as first-line treatment for metastatic colorectal cancer. The combination produced a significant survival improvement, with a hazard ratio for death of 0.88 and a Bayesian probability of 98 percent that it was better than chemotherapy alone. But median survival improved only from 18.6 to 19.9 months, and the combination raised the overall risk of side effects by 15 to 20 percent, including neutropenia, diarrhea, rash, and dermatitis. Against a hazard ratio threshold of 0.80, the probability that the combination was genuinely worthwhile was just 8 percent, a sobering counterpoint to the trial’s positive label.
The team also demonstrated the approach on a contemporary neurology question, re-analyzing the ELAN trial, which tested whether direct oral anticoagulants should be started early, within 48 hours of a minor or moderate ischemic stroke and on days 6 to 7 after a major stroke, in patients with atrial fibrillation. Early treatment might prevent recurrent stroke but was feared to increase intracranial hemorrhage. Using uninformative priors, the Bayesian analysis found a 93 percent probability that early treatment was superior for the composite outcome at 30 days and 97 percent at 90 days, and a 96 percent probability of superiority for preventing recurrent ischemic stroke at both time points. Crucially, the method can quantify the feared harm: there was a 97.6 percent probability that early treatment does not increase symptomatic intracranial hemorrhage risk by more than 0.5 percentage points, and a 99.995 percent probability that any increase stays below 1 percentage point. Even when the researchers imposed a strongly skeptical prior encoding clinicians’ pre-trial fears that early anticoagulation raises bleeding risk, the posterior probability of a clinically relevant increase was only 7 percent. Combining benefit and harm, the analysis estimated an 81 percent probability that early treatment reduces recurrent stroke risk by at least 0.5 percent while raising hemorrhage risk by no more than 0.5 percent, and a Bayesian decision analysis assigning equal severity weights to the two outcomes found a 95 percent probability that early treatment is preferable, a conclusion that remained robust even when hemorrhage was given double or triple weight.
The authors are careful about the limits of their work. They used trial-level rather than individual patient data, made simplifying assumptions such as normality of estimated effects, and chose the hazard ratio of 0.80 threshold largely for illustration, noting that the smallest worthwhile effect is context-specific and should ideally reflect treatment toxicities and patient preferences. They also caution against a subtle misreading: statements about average treatment effects must not be confused with statements about individual patients, so the phrase on average matters. Perhaps most importantly, they do not advocate simply swapping one threshold for another, replacing the 0.05 p-value cutoff with, say, a 95 percent posterior probability cutoff. Instead, they argue that probabilistic statements should complement, not replace, conventional estimates, credible intervals, and confidence intervals, and that formal decision or cost-effectiveness analyses considering multiple outcomes, patient preferences, and costs remain the ideal basis for treatment choices.
To lower the barrier to adoption, the researchers released a free web application that takes a trial’s event counts or effect estimates and returns the posterior distribution along with the probability that the true effect exceeds any chosen reference value, though it currently implements only uninformative priors. All data and code are publicly available on GitHub, and the paper outlines methods for eliciting informative priors from panels of experts, an approach the authors recommend in practice since researchers almost never begin from genuine ignorance. With the United States Food and Drug Administration having issued draft guidance on Bayesian methodology in drug trials, and prior surveys showing clinicians find Bayesian results more useful than p-values when correctly explained, the study lands at a moment of genuine regulatory and cultural openness. Whether medicine’s communication habits change may ultimately depend on involving patients in deciding what probabilities matter to them, but this analysis makes a compelling case that the question doctors should be answering is not whether the data were unlikely under a null hypothesis, but how likely it is that the treatment helps.
Subject of Research: Using Bayesian probabilistic statements to communicate and interpret clinical trial results
Article Title: Bayesian probabilistic statements for communicating clinical trial results
Article References: Efthimiou, O., Chalkou, K., Siontis, G. C., Fischer, U., Perneger, T., Salanti, G., & Gayet-Ageron, A. (2026). Bayesian probabilistic statements for communicating clinical trial results. BMC Medicine, 24(1), Article 554. https://doi.org/10.1186/s12916-026-05208-w
Image Credits: AI Generated
DOI: 10.1186/s12916-026-05208-w
Keywords: Bayesian statistics, clinical trials, p-values, posterior probability, statistical significance, credible intervals, hazard ratio, ELAN trial, oncology, shared decision-making, smallest worthwhile effect, BMC Medicine
News Source: Ophelia Keating. (October 9, 2026). When Negative Trials Hide Real Benefits: Bayesian Probabilities Rethink Clinical Evidence. Scienmag.



