A team of ophthalmologists in China has put some of the world’s most powerful artificial intelligence chatbots through one of the most demanding real-world tests yet devised for medical AI: deciding which type of refractive surgery, if any, suits a patient’s eyes. The results, published in BMC Medicine, show that five leading large language models could match or exceed the performance of an intermediate-level ophthalmologist when triaging patients for procedures such as LASIK and SMILE, reaching accuracies above 98.5 percent in some tasks. But the study also draws a careful line: the machines excel at straightforward yes-or-no judgments yet remain noticeably shakier when asked to choose among multiple surgical options, a nuance the researchers say should shape how clinics deploy these tools.
Refractive surgery is one of the most common elective procedures in medicine, encompassing techniques that reshape the cornea with lasers or implant a corrective lens inside the eye. Choosing the right procedure is far from trivial. Surgeons must weigh corneal thickness, refractive error, anterior chamber depth, pupil size, lifestyle factors and a catalogue of contraindications. A patient who is an ideal candidate for Small Incision Lenticule Extraction, or SMILE, might be a poor fit for transepithelial photorefractive keratectomy, known as TransPRK, and vice versa. Misjudging that calculus can lead to complications, retreatments or long-term visual problems, which is why surgical suitability decisions are traditionally reserved for trained specialists.
The study, led by Qi Wan, Ran Wei, Jing Tang, Ying-ping Deng and Ke Ma of the Department of Ophthalmology at West China Hospital of Sichuan University, drew on preoperative data from 11,966 consecutive patients evaluated at their institution. That sheer scale is what makes the work unusual. Most AI-in-medicine benchmarks rely on a few hundred curated cases; this one threw nearly twelve thousand messy, real-world clinical records at the algorithms. To establish ground truth, a panel of three senior refractive surgeons independently reviewed every case, assigning each a recommendation score from 0 to 100 and a suitability classification across four procedures: Femtosecond LASIK, SMILE, TransPRK and the Implantable Collamer Lens, or ICL. The panel’s agreement was excellent, with a kappa statistic exceeding 0.85, meaning the experts rarely disagreed about what the right answer was.
Five large language models faced the exam: DeepSeek-Chat, GLM-4.7, GPT-4o, Kimi-K2-Thinking and Qwen-Max. Each was fed the same structured, expert-mimicking prompt for every case and asked to produce the same recommendation scores and classifications the human panel had given. Alongside the machines, an intermediate-level physician, representing a mid-career ophthalmologist, independently evaluated all 11,966 cases, providing a human benchmark that is arguably more realistic than comparing AI to world-renowned professors. The study then scored everyone, human and machine alike, on a battery of standard metrics: accuracy, the area under the receiver operating characteristic curve for binary decisions, Cohen’s kappa for multi-class agreement, correlation coefficients for the numeric scores, and the root mean square and mean absolute errors measuring how far predictions strayed from expert judgment.
The headline result is that the top-performing language models were remarkably good at the binary question of whether a patient is suitable for a given procedure. For LASIK and SMILE, the best models achieved accuracies above 98.5 percent and areas under the curve exceeding 0.96, figures that place them in territory generally considered excellent for clinical classification. Perhaps more striking was the performance of the intermediate physician, who scored significantly lower than the leading models, particularly on complex classifications. That comparison matters because it reframes the debate about AI in medicine. The question is no longer simply whether machines can match elite specialists, but whether they can consistently outperform the average clinician making routine decisions, and on this evidence, in this narrow domain, they can.
The picture grows more complicated when the task shifts from two options to four. In the multi-class setting, where a model must pick the single best-suited procedure among LASIK, SMILE, TransPRK and ICL, agreement with the expert panel was more modest. Qwen-Max led this metric, achieving a Cohen’s kappa of up to 0.743, which statisticians generally label substantial agreement but which falls well short of the panel’s own internal consistency. The researchers interpret this gap candidly: the models are best understood as decision-support tools rather than autonomous decision-makers. In other words, an AI can reliably flag that a patient is a candidate for corneal refractive surgery, but a human expert should still make the final call about which procedure, especially in borderline cases where corneal topography, pupil characteristics and patient preference interact in subtle ways.
Beyond raw accuracy, the study examined how closely the models’ numeric recommendation scores tracked the experts’ 0-to-100 ratings. Here Qwen-Max and DeepSeek-Chat showed strong correlation with the panel’s scores, suggesting the models grasp not just the category of recommendation but its gradations, recognizing, say, that a patient with thin corneas and moderate myopia deserves a lower suitability score for LASIK than one with abundant corneal tissue. Regression errors captured the same story: the strongest models deviated from expert scores by margins small enough to be clinically useful, while weaker models and the intermediate physician showed wider scatter. This graded, score-like behavior matters for real deployment, because a tool that merely says yes or no discards the nuance surgeons use to counsel patients about risk.
One of the study’s most practically significant contributions is its analysis of cost and speed, dimensions rarely quantified in medical AI evaluations. DeepSeek-Chat emerged as the efficiency champion, combining the lowest cost per query with the fastest response times while still delivering strong accuracy. GPT-4o, by contrast, was the most expensive of the five. These economics are not trivia. In high-volume screening scenarios, where thousands of preoperative assessments must be triaged before a surgeon ever sees the patient, a model that is nearly as accurate but a fraction of the price can transform workflow economics. The authors point to resource-limited settings, including regions with few refractive surgeons, as the clearest beneficiaries, since a cheap, fast, accurate pre-screening layer could extend specialist-grade triage to populations that currently lack access.
Technically, the study also offers a template for how such evaluations should be run. The structured expert-mimicking prompt, the massive consecutive-patient dataset, the multi-surgeon gold standard with quantified inter-rater agreement, and the multi-dimensional metric battery together form a benchmarking blueprint that other specialties can copy. Too many published AI evaluations rest on convenience samples and single-metric reporting; this work demonstrates what a more rigorous standard looks like, including the honest acknowledgment that kappa values in multi-class tasks remain moderate. It is worth noting that the study is retrospective: the models reviewed recorded data rather than live patients, and real clinical deployment would raise additional questions about data privacy, liability, and how surgeons integrate algorithmic advice into consultations.
What emerges is a measured but genuinely exciting picture of where medical language models stand. On narrow, well-defined classification tasks grounded in structured clinical data, they now perform at or above the level of mid-career physicians, at pennies per case and in seconds. On the harder judgment calls that define expert practice, they still lag behind senior specialists and need human oversight. The West China Hospital team frames the technology exactly as the evidence supports: a powerful assistive layer for screening, triage and decision support, positioned to augment rather than replace the surgeon’s judgment. As these models continue to improve, and as prospective trials validate retrospective results, the routine preoperative assessment of refractive surgery candidates may become one of the first places where patients routinely benefit from an AI second opinion, whether or not they ever know it is there.
Subject of Research: Evaluation of large language models for refractive surgery recommendation and clinical decision support
Article Title: Benchmarking advanced large language models for refractive surgery recommendation: a multi-model, real-world evaluation
Article References: Wan, Q., Wei, R., Tang, J., Deng, Y.-P., & Ma, K. (2026). Benchmarking advanced large language models for refractive surgery recommendation: a multi-model, real-world evaluation. BMC Medicine. https://doi.org/10.1186/s12916-026-05262-4
Image Credits: AI Generated
DOI: 10.1186/s12916-026-05262-4
Keywords: large language models, refractive surgery, LASIK, SMILE, ophthalmology, artificial intelligence, clinical decision support, GPT-4o, DeepSeek, cost-effectiveness, machine learning, BMC Medicine
Cite Scienmag News
APA
MLA
Chicago
Ophelia Keating. (September 26, 2026). AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds. Scienmag. https://scienmag.com/ai-chatbots-match-human-doctors-in-recommending-laser-eye-surgery-study-finds/
Ophelia Keating. “AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds.” Scienmag, 26 September 2026, https://scienmag.com/ai-chatbots-match-human-doctors-in-recommending-laser-eye-surgery-study-finds/. Accessed 26 September 2026.
Ophelia Keating. “AI Chatbots Match Human Doctors in Recommending Laser Eye Surgery, Study Finds.” Scienmag. September 26, 2026. https://scienmag.com/ai-chatbots-match-human-doctors-in-recommending-laser-eye-surgery-study-finds/
Copy citation
Download RIS
Tags: AI chatbotsAI limitations in complex medical choicesAI triage for eye proceduresAI vs human ophthalmologistsArtificial IntelligenceBMC Medicineclinical decision supportCost-effectivenessDeepSeekGPT-4olarge language modelslarge language models in healthcarelaser eye surgery decision-makingLASIKLASIK and SMILE surgical recommendationsMachine learningmedical AI performance comparisonmedical decision support AIophthalmologyophthalmology AI diagnosisreal-world AI validation in ophthalmologyrefractive surgeryrefractive surgery AI accuracySMILE


