Acute appendicitis is the most common cause of the acute abdomen, and every emergency physician knows the stakes: miss it, and the inflamed appendix can rupture, flooding the abdomen with bacteria; overcall it, and a patient undergoes surgery they never needed. For decades, clinicians have leaned on the Alvarado score, a simple checklist of symptoms, signs and blood values, to tip the balance. Now a retrospective diagnostic accuracy study from Mansoura University Hospital in Egypt suggests that a large language model, ChatGPT-5, can read the same clinical data and outperform that venerable score in a striking way, particularly in its ability to rule the disease out.
The research team, led by Amr A. Elgharib and colleagues, assembled a cohort of 162 patients aged 15 and over who presented with abdominal pain and were clinically suspected of having appendicitis between March 2024 and March 2025. All of them underwent appendectomy, and the final word on their diagnosis came from histopathological examination of the removed appendix, the definitive reference standard. The cohort was deliberately enriched: roughly 70 percent of patients turned out to have confirmed appendicitis, a prevalence far higher than in a general emergency department, which shapes how the results must be interpreted.
The researchers fed ChatGPT-5 structured clinical narratives for each patient, including age, sex, symptoms such as fever, nausea, vomiting and right lower quadrant pain, physical examination findings like rebound tenderness, and laboratory values including white blood cell count and neutrophil percentage. The prompt cast the model as a general surgery physician in an emergency department and, crucially, forbade it from using any established appendicitis scoring system such as Alvarado, AIR, RIPASA or AAS. Instead, the model had to rely on its own clinical reasoning, outputting either a diagnosis of acute or non-acute appendicitis and, for positive cases, a percentage probability stratified into low, intermediate or high risk.
The study was designed in two phases following the classic machine learning convention of a 70-30 split. In Phase 1, 114 cases were presented without any prior exposure to labeled examples, simulating how a clinician might use a readily available chatbot off the shelf. After the model’s predictions were recorded, the true diagnoses were revealed, and in Phase 2 the remaining 48 cases were used to test whether this in-context learning improved performance. The prompt format was fixed across all cases, and the model’s memory was cleared before each evaluation round to test its default behavior.
The headline numbers are eye-catching. In Phase 1 without randomization, ChatGPT-5 achieved 100 percent sensitivity, meaning it caught every single case of appendicitis with zero false negatives, alongside 81.6 percent specificity, a positive predictive value of 91.6 percent, a negative predictive value of 100 percent, and overall accuracy of 93.9 percent. The Alvarado score, by contrast, managed only 67.1 percent sensitivity, missing 25 true cases, though its specificity was higher at 92.1 percent, and its overall accuracy came to 75.4 percent. Statistical comparison using McNemar’s test for paired proportions confirmed that the model’s sensitivity and accuracy advantages were highly significant, with p-values below 0.001.
But the story took a twist when the researchers examined a subtle methodological trap known as order leakage. When cases are presented in a sequence, a language model can inadvertently exploit the ordering of the data rather than genuine clinical reasoning, artificially inflating its apparent accuracy. To probe this, the team re-ran the evaluations with the case order randomized. Sensitivity barely budged, remaining at 98.7 percent in Phase 1, but specificity collapsed dramatically from 81.6 percent to 36.8 percent, a difference that was highly statistically significant. In other words, part of the model’s apparent precision in ruling appendicitis in may have been an artifact of how the cases were sequenced in the prompt.
Phase 2 told a more balanced story. After in-context learning, the non-randomized model reached 100 percent sensitivity, 93.8 percent specificity, 97.0 percent positive predictive value, 100 percent negative predictive value and 97.9 percent accuracy, statistically indistinguishable from the Alvarado score’s 96.9 percent sensitivity and 93.8 percent specificity. Agreement between the two methods, measured with Cohen’s kappa, climbed from a moderate 0.503 in Phase 1 to an almost perfect 0.952 in Phase 2. The authors note that in-context learning produced a modest, non-significant gain in specificity, while sensitivity was largely unaffected, suggesting the model’s core ability to exclude appendicitis was robust from the start.
The findings sit within a growing but cautionary literature. Previous evaluations have found that ChatGPT’s answers to surgeon-designed appendicitis management questions were clinically pertinent but inconsistent, and that ChatGPT-3.5 performed significantly worse than clinicians for some acute abdominal conditions such as cholecystitis and diverticulitis. A systematic review of 29 studies on artificial intelligence in appendicitis highlighted wide heterogeneity in inputs, validation strategies and metrics. Imaging-based machine learning models reading CT scans have achieved sensitivities around 77 percent and specificities around 86 percent. The new study’s authors also point to broader evidence that data leakage has inflated reported performance in fields ranging from Parkinson’s disease detection to brain MRI classification, reinforcing their insistence on randomization as a guard against over-optimistic estimates.
The authors are careful about what their results do and do not license. Because the cohort consisted exclusively of surgically confirmed patients with high pre-test probability, the estimates cannot simply be generalized to the unselected emergency department population in whom appendicitis must be ruled in or out, and the study was limited to a single center, a single model and a single comparator, with newer scores like AIR and RIPASA omitted because the required data were not consistently available. ChatGPT-5, they conclude, cannot yet be relied upon to confirm acute appendicitis, given the Alvarado score’s higher specificity, and the findings should be regarded as hypothesis-generating pending prospective validation.
What the study does establish is a proof of concept with real clinical texture: a general-purpose language model, guided by a carefully engineered prompt and given nothing more than the routine clinical data already sitting in a patient’s chart, can match or exceed a century-old scoring heuristic at the task that matters most for safety, excluding the disease. The authors envision ChatGPT-5 as a clinical decision-support tool to assist professionals during early assessment, particularly in resource-limited settings, always under physician supervision and never as an autonomous diagnostician. They also warn explicitly against patient self-diagnosis, noting that false reassurance from a chatbot could delay timely medical evaluation. As large language models continue their march into medicine, this Egyptian study offers both an encouraging data point and a methodological warning: before trusting an AI’s diagnostic brilliance, check whether the cases were shuffled.
Subject of Research: Diagnostic accuracy of ChatGPT-5 compared with the Alvarado score for acute appendicitis
Article Title: Diagnostic performance of ChatGPT for acute appendicitis compared with the Alvarado score: a retrospective diagnostic accuracy study
Article References: Elgharib, A. A., Ghanem, M. I., Ibrahim, R. S., Elwakeel, N., & Shemes, A. (2026). Diagnostic performance of ChatGPT for acute appendicitis compared with the Alvarado score: a retrospective diagnostic accuracy study. Discover Artificial Intelligence, 6(1), Article 1352. https://doi.org/10.1007/s44163-026-02332-7
Image Credits: AI Generated
DOI: 10.1007/s44163-026-02332-7
Keywords: ChatGPT-5, large language models, acute appendicitis, Alvarado score, diagnostic accuracy, artificial intelligence, clinical decision support, histopathology, order leakage, in-context learning, emergency medicine, prompt engineering
News Source: Denise Maddox. (October 5, 2026). ChatGPT-5 Outperforms Classic Alvarado Score in Detecting Appendicitis, Study Finds. Scienmag.



