Artificial intelligence is quietly creeping into one of science’s most guarded rituals: peer review. Editors short on time and reviewers exhausted by endless requests have already begun experimenting with large language models to help evaluate manuscripts, and surveys suggest the practice is more widespread than journals officially acknowledge. But handing scientific judgment to machines trained on human text raises an uncomfortable question: will these systems inherit the biases that have long plagued human peer review, from gender disparities to geographic favoritism? A new experimental study published in the Journal of General Internal Medicine offers one of the most controlled answers yet, and the result is surprising in its cleanliness.
Paul Sebo of the University of Geneva and Ting Wang of Emporia State University set out to test whether two leading AI models, OpenAI’s ChatGPT and Anthropic’s Claude, would score identical scientific abstracts differently depending on the name and country attached to them. They also wanted to know something equally important for any future editorial application: whether the models give the same score to the same work when asked twice. The stakes are considerable. If AI reviewers are both unbiased and reproducible, they could become powerful tools for triaging the flood of submissions that modern journals face. If they are neither, their adoption could quietly distort who gets published and whose research shapes medicine.
The experimental design was deliberately austere. The researchers randomly selected ten general internal medicine journals from the Journal Citation Reports, each with an impact factor of at least 1.5, and pulled five original research abstracts per journal from the Web of Science, for a total of fifty abstracts published between 2023 and 2026. Each abstract was then evaluated under four fictional author identities: an American woman named Rachel Smith, an American man named John Smith, a woman from Côte d’Ivoire named Fatoumata Diallo, and a man from Côte d’Ivoire named Amadou Diallo. Using fictional identities rather than real researchers allowed the team to manipulate exactly one variable, the identity cue, while keeping everything else constant and avoiding the ethical tangle of attaching published work to real people.
Every abstract-identity combination was scored twice by each model, with each evaluation conducted in a fresh chat session to prevent any memory of previous judgments from contaminating the results. The prompt was identical every time: act as a scientific reviewer, and rate the abstract on three dimensions using a 0 to 10 scale, namely overall scientific quality, novelty, and likelihood of acceptance at a high-quality conference or journal. The models were instructed to output only three numbers, no explanations. That discipline produced 800 evaluations in total, 400 per model, all collected between April 15 and April 30, 2026, using the standard user interfaces with default settings.
The headline finding is striking in its symmetry. For ChatGPT, median quality and novelty scores were identical across all four identities, with only minimal variation in acceptance scores that followed no consistent gender or geographic pattern. For Claude, all three scores were identical across identities. Multivariable ordinal logistic regression, adjusted for journal, found no overall association between author identity and any score, with one partial exception: a global test for Claude’s acceptance scores hinted at a possible association, but no individual comparison reached statistical significance. The authors urge caution even about the isolated differences that did appear, such as slightly lower novelty and acceptance scores that ChatGPT gave to abstracts attributed to the American man, because the overall tests were not significant and many comparisons were performed on a modest number of unique abstracts.
Reproducibility, the second pillar of the study, was equally encouraging. When the same abstract was scored twice under the same identity, percent agreement exceeded 0.98 for both models across all three outcomes. Fleiss’ kappa, a statistic that corrects for agreement expected by chance and weights larger disagreements more heavily, ranged from 0.88 to 0.89 for ChatGPT, indicating excellent consistency, and from 0.74 to 0.80 for Claude, indicating substantial agreement. ChatGPT was thus the steadier grader, though both models were far more consistent with themselves than human reviewers typically are with each other, a well-documented weakness of traditional peer review.
The scores were not blind to everything, however. Abstracts from higher-impact journals received systematically higher ratings across all three dimensions, with each point of impact factor associated with roughly 35 to 40 percent higher odds of a better score. Interestingly, the journals’ names and impact factors were never shown to the models, so this pattern cannot reflect prestige bias directly. It more likely reflects genuine differences in the writing quality and study characteristics of abstracts published in more selective venues, though the relationship was not perfectly linear, and one mid-tier journal, Annals of Family Medicine, consistently ranked low despite its respectable impact factor of 5.1. The models also differed from each other: Claude assigned significantly lower quality and acceptance scores than ChatGPT overall, suggesting that different AI systems bring meaningfully different evaluative standards to the same text.
The authors are careful, almost insistently so, about what these results do not prove. Consistency is not validity. The fact that a model gives the same score twice says nothing about whether that score is accurate, insightful, or useful for editorial decisions, and the study included no human reference standard against which to calibrate the numbers. The restricted range of scores, mostly clustered between 2 and 9 with medians of 5 to 7, may itself have inflated the agreement statistics. The task was also radically simplified compared with real peer review: abstracts rather than full manuscripts, numerical scores rather than narrative critique, and a single standardized prompt rather than the messy, iterative dialogue that characterizes genuine refereeing. Subtler biases that might surface in open-ended written feedback would be invisible in this design.
There are further limits worth keeping in mind. Fifty abstracts from a single medical specialty cannot speak for all of science, and only two models, at specific versions, were tested; both are updated frequently, so the findings may not hold for future releases. The identity manipulation covered only two countries and two genders, a narrow slice of the diversity of the global research community, and the researchers could not verify whether the models actually registered the identity cues at all. An absence of score differences could mean the models are genuinely impartial, or simply that they ignored the names entirely. The study was also not preregistered, which the authors acknowledge transparently.
Still, the implications are provocative. At a moment when journals are wrestling with reviewer shortages and with evidence that human peer review itself suffers from gender, geographic, and institutional bias, the prospect of an evaluator that treats a manuscript from Geneva and one from Abidjan identically is genuinely newsworthy. The authors suggest that structured, low-stakes applications such as initial screening, detection of reporting deficiencies, and editorial triage may be the most realistic near-term uses for these tools, rather than replacing human reviewers outright. The most informative next step, they argue, would be to run LLM-based review on manuscripts for which real human review reports and editorial decisions already exist, allowing a direct head-to-head comparison of judgments and biases. Until then, the message of this study is a measured one: in the sterile laboratory of standardized abstract scoring, the machines showed no favoritism and remarkable self-consistency. Whether they can sustain that discipline in the messy, argumentative world of real scientific publishing remains an open and urgent question.
Subject of Research: Bias and reproducibility of large language models as evaluators of scientific abstracts in peer review
Article Title: Bias and Reliability of AI-Based Peer Review: A Comparative Study of ChatGPT and Claude Evaluating Scientific Abstracts
Article References: Sebo, P., & Wang, T. (2026). Bias and Reliability of AI-Based Peer Review: A Comparative Study of ChatGPT and Claude Evaluating Scientific Abstracts. Journal of General Internal Medicine. https://doi.org/10.1007/s11606-026-10806-8
Image Credits: AI Generated
DOI: 10.1007/s11606-026-10806-8
Keywords: artificial intelligence, large language models, peer review, ChatGPT, Claude, bias, reproducibility, scientific publishing, gender bias, geographic bias, abstracts, internal medicine
Cite Scienmag News
APA MLA Chicago
Ophelia Keating. (October 2, 2026). AI Reviewers Show No Gender or Geographic Bias in Abstract Scoring, Study Finds. Scienmag. https://scienmag.com/ai-reviewers-show-no-gender-or-geographic-bias-in-abstract-scoring-study-finds/
Ophelia Keating. “AI Reviewers Show No Gender or Geographic Bias in Abstract Scoring, Study Finds.” Scienmag, 2 October 2026, https://scienmag.com/ai-reviewers-show-no-gender-or-geographic-bias-in-abstract-scoring-study-finds/. Accessed 2 October 2026.
Ophelia Keating. “AI Reviewers Show No Gender or Geographic Bias in Abstract Scoring, Study Finds.” Scienmag. October 2, 2026. https://scienmag.com/ai-reviewers-show-no-gender-or-geographic-bias-in-abstract-scoring-study-finds/
Copy citation Download RIS
Tags: abstractsAI fairness in scienceAI peer review biasAI-assisted manuscript reviewAnthropic Claude review performanceArtificial Intelligencebiasbias-free AI scientific judgmentChatGPTChatGPT scientific abstract evaluationClaudeethical considerations in AI peer reviewgender biasgender bias in scientific evaluationgeographic biasgeographic bias in AI scoringimpact of AI on scientific publishinginternal medicinelarge language modelslarge language models in peer reviewpeer reviewreproducibilityreproducibility of AI assessmentscientific publishing


