• HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
Tuesday, October 6, 2026
BIOENGINEER.ORG
No Result
View All Result
  • Login
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
  • HOME
  • NEWS
  • EXPLORE
    • CAREER
      • Companies
      • Jobs
        • Lecturer
        • PhD Studentship
        • Postdoc
        • Research Assistant
    • EVENTS
    • iGEM
      • News
      • Team
    • PHOTOS
    • VIDEO
    • WIKI
  • BLOG
  • COMMUNITY
    • FACEBOOK
    • INSTAGRAM
    • TWITTER
No Result
View All Result
Bioengineer.org
No Result
View All Result
Home NEWS Science News Health

AI Sits a Dermatology Board Exam: Five Chatbots Put to the Test

by
October 6, 2026
in Health
Reading Time: 5 mins read
0
AI Sits a Dermatology Board Exam: Five Chatbots Put to the Test

AI Sits a Dermatology Board Exam: Five Chatbots Put to the Test

Share on FacebookShare on TwitterShare on LinkedinShare on RedditShare on Telegram

When a patient presents with an unusual rash, the dermatologist across the table is drawing on years of training, thousands of clinical images, and the hard-won experience of board certification. Now a team at Rutgers Robert Wood Johnson Medical School has asked a provocative question: how would the newest generation of artificial intelligence chatbots fare against the same examination standards? In a research letter published in the Archives of Dermatological Research, Rose Rasty, Elyse Mackenzie, Shaunt Mehdikhani and Babar Rao report their evaluation of five prominent large language models on board-style dermatology questions drawn from the American Board of Dermatology’s testing tradition, offering one of the most direct head-to-head comparisons yet attempted in the specialty.

The five systems examined represent the current cutting edge of consumer-facing artificial intelligence. The study tested ChatGPT in its GPT-5.2 configuration, xAI’s Grok 4.1, Google’s Gemini 3, Perplexity running Sonar and GPT-5-series backends, and Microsoft Copilot built on GPT-5. This lineup matters because these are not obscure research prototypes buried in a laboratory; they are the tools that medical students, residents and practicing physicians already consult daily for quick answers. Understanding how they perform on rigorous, specialty-specific questions is therefore not an academic curiosity but a matter of practical patient safety, since patients too are increasingly turning to these chatbots with their own dermatological concerns.

Board-style questions are a particularly demanding benchmark for any question-answering system. The American Board of Dermatology’s examinations are designed to distinguish competent specialists from merely knowledgeable ones, and the questions typically present nuanced clinical vignettes in which several answer choices are plausible. A candidate must weigh the patient’s age, lesion morphology, distribution, history of prior treatments, and sometimes subtle histopathological findings before selecting the single best response. This format tests something closer to clinical reasoning than to factual recall, which is precisely why it has become a popular yardstick for evaluating medical artificial intelligence. A model can memorize that psoriasis affects roughly two to three percent of the population, but it must reason through a vignette to recognize a psoriasiform drug eruption masquerading as the idiopathic disease.

The technical architecture underlying these models helps explain both their promise and their pitfalls. Large language models are neural networks, generally built on the transformer architecture, trained on vast text corpora to predict the next token in a sequence. Through this deceptively simple objective, they absorb grammar, factual knowledge, and patterns of reasoning encoded in their training data. Medical knowledge enters the picture because textbooks, review articles, clinical guidelines and discussion forums form a substantial portion of the public internet. However, the models do not store facts as a database would; knowledge is distributed across billions of numerical parameters, which means retrieval is probabilistic rather than deterministic. This is why a model can answer an obscure question flawlessly one moment and hallucinate a nonexistent drug interaction the next, a behavior that examiners would immediately fail in a human candidate.

Dermatology poses distinctive challenges for text-based models, even in a question-answer format that strips away the visual element. The specialty sits at the intersection of internal medicine, immunology, infectious disease, oncology and pathology, and its vocabulary is notoriously dense. Distinguishing lichen planus from lichenoid drug eruption, or dermatitis herpetiformis from linear IgA bullous dermatosis, depends on pattern recognition honed through exposure to thousands of cases. Board-style questions compress that pattern recognition into prose, which arguably favors language models, yet the compression also removes the contextual cues a clinician would use at the bedside. The Rutgers team’s choice of this format follows a growing research tradition: earlier studies had already tested ChatGPT on board-style dermatology questions, including image-based items, and Liu and colleagues in 2025 assessed dermatological knowledge and image analysis using specialty certificate examinations as their benchmark.

What makes the new study notable is its breadth of comparison. Most prior evaluations examined a single model, usually a ChatGPT variant, against a question bank. By running the same board-style items through five different systems, each with distinct underlying architectures and training pipelines, the researchers could probe whether strong medical performance is a general property of frontier language models or something specific to particular products. The differences among these systems are not trivial. GPT-5.2, Grok 4.1, Gemini 3 and the GPT-5-based Copilot each reflect different training data mixtures, different reinforcement learning strategies, and different approaches to reasoning. Perplexity adds a further wrinkle: it couples a language model with live web search, meaning its answers may draw on retrieved documents rather than purely on parametric memory, a hybrid architecture that could behave very differently on questions with recently updated guidelines.

The distinction between parametric knowledge and retrieval-augmented answering is one of the most technically interesting aspects of this kind of evaluation. A pure language model answers from what is baked into its weights during training, frozen at a cutoff date. A retrieval-augmented system like Perplexity queries the live web, which brings freshness but also vulnerability: it can absorb misinformation, outdated guidelines or commercially biased content from the open internet. For dermatology, where treatment recommendations evolve and where the internet is saturated with anecdotal skin-care advice, that distinction could materially change performance. A benchmark like the one used in this study therefore measures not just a model’s medical knowledge but its entire answer-generation pipeline, including how it filters and prioritizes whatever sources it consults.

The authors of the research letter are appropriately measured about what such testing can and cannot establish. Answering multiple-choice questions correctly is not the same as managing a patient. A board examination candidate who selects the right option for a bullous disorder vignette may still struggle to perform a safe skin biopsy or to counsel a frightened patient about a new diagnosis of melanoma. Conversely, models can exploit statistical cues in question wording, patterns that human test-takers learn to recognize as well, without genuinely understanding the underlying medicine. There is also the ever-present risk of data contamination: if board-style questions or close paraphrases circulate online, a model trained on the internet may have effectively seen the answer key. Rigorous benchmarking in medical artificial intelligence must therefore always ask not only how well a model scored, but whether the score reflects reasoning or recall.

Nevertheless, the trajectory across successive studies is striking. When ChatGPT first burst into public awareness in late 2022, early medical evaluations produced mixed results, with the model often falling short of passing thresholds on professional examinations. Each subsequent generation has closed the gap, and by the mid-2020s frontier models were routinely performing at or near the level of human test-takers across a range of specialties. The Rutgers study extends that line of evidence into dermatology with the newest available models, and its very framing, published as a research letter rather than a lengthy original article, reflects how quickly the field is moving: the findings needed to reach the literature before the models they evaluated were superseded. The authors report that all data were generated using the five named systems, with the dataset available from the corresponding author on reasonable request, and the work received no external funding while the authors declared no conflicts of interest.

For clinicians and patients alike, the practical message is one of cautious engagement rather than either alarm or celebration. These systems are already embedded in clinical workflows, from drafting patient messages to summarizing literature, and their demonstrated competence on board-style questions suggests they can serve as powerful adjuncts for education and decision support. A resident reviewing for boards might legitimately use a chatbot to generate practice vignettes or explain the immunopathology of pemphigus vulgaris. But the same probabilistic architecture that enables fluent expertise also produces confident errors, and no benchmark, however demanding, substitutes for clinical supervision. The Rutgers team’s comparison of five frontier models gives the field a clearer map of where artificial intelligence stands in dermatology today, and a reminder that the map will need redrawing with every new model release.

Subject of Research: Evaluation of large language models on American Board of Dermatology-style examination questions

Article Title: Performance of five large language models on board-style dermatology questions from the American Board of Dermatology

Article References: Rasty, R., Mackenzie, E., Mehdikhani, S., & Rao, B. (2026). Performance of five large language models on board-style dermatology questions from the American Board of Dermatology. Archives of Dermatological Research, 318(1), Article 462. https://doi.org/10.1007/s00403-026-04894-z

Image Credits: AI Generated

DOI: 10.1007/s00403-026-04894-z

Keywords: large language models, dermatology, artificial intelligence, board examination, ChatGPT, Gemini, Grok, Microsoft Copilot, Perplexity, medical education, clinical reasoning, benchmarking

News Source: Dean Parker. (October 6, 2026). AI Sits a Dermatology Board Exam: Five Chatbots Put to the Test. Scienmag.

Tags: Artificial Intelligencebenchmarkingboard examinationChatGPTclinical reasoningDermatologyGeminiGrokLarge Language ModelsMedical EducationMicrosoft CopilotPerplexity
Share12Tweet7Share2ShareShareShare1

Related Posts

When Bypass Grafts Slow Down: New Study Separates Competitive Flow from True Blockage

When Bypass Grafts Slow Down: New Study Separates Competitive Flow from True Blockage

October 6, 2026
Longer Community Exercise Programs Sharpen Balance, Strength and Endurance in Older Adults

Longer Community Exercise Programs Sharpen Balance, Strength and Endurance in Older Adults

October 6, 2026

Lifting Weights May Rejuvenate the Aging Brain’s Metabolism and Deepen Sleep

October 6, 2026

Three-Year Trial Shows Mediterranean Diet Eases Food Addiction Symptoms but Only for Some

October 6, 2026

POPULAR NEWS

  • Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

    29 shares
    Share 12 Tweet 7
  • Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

    29 shares
    Share 12 Tweet 7
  • Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

    29 shares
    Share 12 Tweet 7
  • New Scale Measures How Ready Nurse Educators Really Are for the AI Era

    29 shares
    Share 12 Tweet 7

About

We bring you the latest biotechnology news from best research centers and universities around the world. Check our website.

Follow us

Recent News

Alloys That Shrink Their Own Grains: New PIX Mechanism Refines Metals With Heat Alone

Endurance Exercise Reshapes the Liver in Males and Females Through Distinct Molecular Routes

Single Transcription Factor PU.1 Rapidly Converts Fibroblasts into Macrophage-Lineage Cells

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm' to start subscribing.

Join 85 other subscribers
  • Contact Us

Bioengineer.org © Copyright 2023 All Rights Reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Homepages
    • Home Page 1
    • Home Page 2
  • News
  • National
  • Business
  • Health
  • Lifestyle
  • Science

Bioengineer.org © Copyright 2023 All Rights Reserved.