• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Saturday, October 3, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Biology

AI Takes the Prosthodontics Board Exam: Chatbots Flirt With Resident-Level Scores

Bioengineer by Bioengineer
October 3, 2026
in Biology
Reading Time: 5 mins read
0
AI Takes the Prosthodontics Board Exam: Chatbots Flirt With Resident-Level Scores
Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

Artificial intelligence has already passed bar exams, medical licensing tests, and chess grandmasters, but a new study has now put four of the world’s most prominent chatbots through one of dentistry’s most demanding knowledge assessments: the National Prosthodontic Resident Examination, the annual test taken by dentists training to become specialists in rebuilding and replacing teeth. The results, published in the journal Heliyon, are both striking and sobering. The best-performing models scored within striking distance of the average human resident, yet their answers shifted between test sessions in ways that reveal how fragile machine reasoning remains when confronted with genuinely specialized clinical knowledge.

The research team, led by Marwa Shembesh and Cortino Sukotjo and spanning institutions in the United States and abroad, assembled 300 multiple-choice questions drawn from the 2023 and 2024 examinations administered by the American College of Prosthodontists. Each year’s exam contained 150 questions, written by prosthodontics program directors and fellows of the college, and the correct answers were verified against the official answer key. Four large language models were put to the test: ChatGPT-3.5 and ChatGPT-4 from OpenAI, Microsoft’s Bing Copilot, and Google’s Gemini. The investigators created fresh accounts for each model to eliminate any influence of stored conversation history, used standardized zero-shot prompts instructing each system to pick the single best answer, and set the sampling temperature to zero for API-based queries to favor deterministic output.

The examination itself is structured around a three-tier hierarchy of topics that mirrors the written section of the American Board of Prosthodontics certification. Tier 1 covers the core of the specialty: complete dentures, fixed and removable partial dentures, dental biomaterials, occlusion, implant prosthodontics, and esthetics. Tier 2 spans adjacent disciplines such as periodontics, pharmacology, wound healing, and temporomandibular disorders, while Tier 3 collects the broader periphery, from biostatistics and diagnostic imaging to sleep disorders, oral pathology, and maxillofacial prosthodontics. In 2023, more than 61 percent of questions fell into Tier 1, with implant prosthodontics and removable partial dentures the most heavily represented topics. In 2024, dental materials and fixed prosthodontics dominated, and Tier 1 still accounted for nearly 57 percent of the exam.

The headline numbers are remarkable. In the 2023 examination, ChatGPT-4 answered 72.7 percent of questions correctly on the first testing round, edging out Gemini at 67.3 percent, Bing at 62.7 percent, and ChatGPT-3.5 at 62.0 percent. The average human resident that year scored 64 percent. In other words, the most advanced chatbot of its generation outperformed the typical prosthodontics resident on the same questions, at least descriptively. A year later the picture shifted: Gemini led the first round with 66.7 percent, followed by ChatGPT-4 at 64.0 percent, Bing at 60.0 percent, and ChatGPT-3.5 at 44.7 percent, against a resident average of 61 percent. Because the human and machine scores were not compared with paired statistical tests, the authors caution that these comparisons are descriptive rather than proof of superiority.

Statistical analysis told a more nuanced story. In 2023, the differences among the four models never reached statistical significance at either testing point. In 2024, however, the overall difference was significant at the first time point, with ChatGPT-4, Bing, and Gemini each significantly outperforming ChatGPT-3.5 after Bonferroni correction; Gemini’s advantage over the older model amounted to 22 percentage points. By the second round two weeks later, the omnibus difference was still significant, but no individual pairwise comparison survived correction, underscoring how unstable model rankings can be across sessions and examination years.

Perhaps the most unsettling finding concerns consistency. When the researchers repeated the entire examination two weeks after the baseline session, three of the four models held steady in 2023, with observed agreement between sessions ranging from 78.7 to 85.3 percent and Cohen’s kappa values between 0.518 and 0.647. But in 2024, ChatGPT-4’s accuracy dropped from 64.0 percent to 51.3 percent, a statistically significant decline of 12.7 percentage points. The authors note that unobserved changes on the provider side, such as backend model updates between testing sessions, cannot be excluded as a source of this variability, since exact platform-level technical details were not retained during data collection. For educators, the message is clear: the same chatbot asked the same question two weeks apart may give a different answer, even under tightly controlled conditions.

The topic-tier analysis added another layer of insight. Descriptively, all four models found Tier 3 questions, the broad peripheral subjects, easiest, pooling to 74.2 percent accuracy in 2023 and 69.9 percent in 2024, while Tier 1, the heart of prosthodontics, proved hardest at 61.1 percent and 53.1 percent respectively. After accounting for repeated responses to the same questions using generalized estimating equations, the tier effect was statistically significant only in 2024, when Tier 3 questions carried roughly twice the odds of a correct answer compared with Tier 1. Within the core Tier 1 material, ChatGPT-4 and Gemini both significantly outperformed ChatGPT-3.5, with odds ratios of 1.80 and 1.92 respectively, while Bing’s advantage did not survive statistical adjustment.

Counterintuitively, newer was not always better. In the 2023 exam, the older ChatGPT-3.5 actually beat ChatGPT-4 on Tier 2 questions at baseline, 77.8 percent to 70.4 percent, and on Tier 3 as well. In 2024, ChatGPT-3.5’s second-round total accuracy of 52.7 percent narrowly exceeded ChatGPT-4’s 51.3 percent. The authors point to similar reversals reported elsewhere, including cases where GPT-3.5 outperformed GPT-4 on classification tasks because its simpler pattern recognition avoided overgeneralization. Model performance, they conclude, depends on the interplay of architecture, training data, and question type rather than a simple hierarchy of model generations.

The study is the first to evaluate large language models on the National Prosthodontic Resident Examination, filling a gap in a literature that has mostly focused on medicine, periodontology, and endodontics. Its implications reach beyond dentistry. The findings suggest that while chatbots can serve as supplementary study aids for residents reviewing implant dentistry or dental materials, they cannot yet be trusted as authoritative sources, particularly on the specialized core knowledge that defines a specialty. The authors also flag concerns about examination integrity if candidates gain access to such tools during secure testing, and they emphasize that model-generated content must always be verified against reliable sources because these systems can fabricate plausible-sounding but incorrect information, a phenomenon known as hallucination.

The researchers acknowledge limitations that temper the conclusions. The question bank is proprietary and accessible only to members of the American College of Prosthodontists, the dataset consisted solely of multiple-choice items that may not capture the complexity of real clinical decision-making, and the two-week testing window cannot predict how performance might drift with future model updates. Future work, they suggest, should test open-ended and case-based questions that better mirror chairside judgment, extend the evaluation across longer periods and more specialties, and probe how training strategies shape performance across different question types. For now, the study stands as a vivid snapshot of a fast-moving frontier: machines that can nearly match the specialists of tomorrow on their own exam, yet still wobble from one week to the next, and that therefore belong beside the textbook, not in place of the clinician.

Subject of Research: Performance of large language models on the American College of Prosthodontists national resident examination

Article Title: The performance of different large language models in the national prosthodontic resident examination in 2023 and 2024

Article References: Shembesh, M., Koseoglu, M., Fang, Q., Gheisarifar, M., Kattadiyil, M. T., Yuan, J. C.-C., Barao, V. A., & Sukotjo, C. (2026). The performance of different large language models in the national prosthodontic resident examination in 2023 and 2024. Heliyon, 12(15), Article e45465. https://doi.org/10.1016/j.heliyon.2026.e45465

Image Credits: AI Generated

DOI: 10.1016/j.heliyon.2026.e45465

Keywords: artificial intelligence, large language models, ChatGPT, Google Gemini, prosthodontics, dental education, medical licensing exams, test-retest reliability, hallucination, American College of Prosthodontists, machine learning, healthcare education

Cite Scienmag News

APA
MLA
Chicago

Blake Davidson. (October 3, 2026). AI Takes the Prosthodontics Board Exam: Chatbots Flirt With Resident-Level Scores. Scienmag. https://scienmag.com/ai-takes-the-prosthodontics-board-exam-chatbots-flirt-with-resident-level-scores/

Blake Davidson. “AI Takes the Prosthodontics Board Exam: Chatbots Flirt With Resident-Level Scores.” Scienmag, 3 October 2026, https://scienmag.com/ai-takes-the-prosthodontics-board-exam-chatbots-flirt-with-resident-level-scores/. Accessed 3 October 2026.

Blake Davidson. “AI Takes the Prosthodontics Board Exam: Chatbots Flirt With Resident-Level Scores.” Scienmag. October 3, 2026. https://scienmag.com/ai-takes-the-prosthodontics-board-exam-chatbots-flirt-with-resident-level-scores/

Copy citation
Download RIS

Tags: AI accuracy and variability in specialized knowledgeAI and human comparison in dental examsAI in prosthodontics board examAmerican College of ProsthodontistsArtificial Intelligenceartificial intelligence in dental specialization assessmentschatbot performance in dental licensing testsChatGPTChatGPT-3.5 and 4 in prosthodonticsclinical knowledge testing with chatbotsdental educationevaluation of AI capabilities in healthcareGoogle Geminihallucinationhealthcare educationimpact of AI on dental education and licensinglarge language modelslarge language models in medical examsMachine learningmachine learning models in medical certificationmachine reasoning in clinical dentistrymedical licensing examsprosthodonticstest-retest reliability

Share12Tweet7Share2ShareShareShare1

Related Posts

New AI Method Ranks Metabolites by Their Impact on Graph Neural Network Predictions

New AI Method Ranks Metabolites by Their Impact on Graph Neural Network Predictions

October 3, 2026
Transplanted Mitochondria Rejuvenate the Aging Heart by Restarting Cellular Cleanup

Transplanted Mitochondria Rejuvenate the Aging Heart by Restarting Cellular Cleanup

October 3, 2026

Dynamic Land-Use Model Reveals Extinction Risk May Spike Abruptly, and Protected Areas Can Bend the Curve

October 3, 2026

New computational tool maps the cell-state crosstalk that decides survival in IDH-mutant glioma

October 3, 2026

POPULAR NEWS

  • Prediabetes May Quietly Weaken Bone Quality in Men Even When Density Looks Normal

    29 shares
    Share 12 Tweet 7
  • Diamond Steps Up as the Ultimate Heat Shield for Next-Generation Chips

    29 shares
    Share 12 Tweet 7
  • Two-Stage Surgery With Biologic Mesh Offers New Hope for Contaminated Hernia Repair

    29 shares
    Share 12 Tweet 7
  • New AI Method Ranks Metabolites by Their Impact on Graph Neural Network Predictions

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Prediabetes May Quietly Weaken Bone Quality in Men Even When Density Looks Normal

Diamond Steps Up as the Ultimate Heat Shield for Next-Generation Chips

Two-Stage Surgery With Biologic Mesh Offers New Hope for Contaminated Hernia Repair

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.