Large language models are quietly moving into the machinery of scientific publishing. They draft referee reports, triage submissions, and rank manuscripts in editorial pipelines. But a new counterfactual audit raises an uncomfortable question: when an AI evaluator scores a paper, is it really judging the science, or is it also judging the scientist? A study published in Discover Artificial Intelligence by Marco Rospocher of the University of Verona suggests that the answer is troubling. Identical manuscripts received systematically different scores from six large language models depending solely on the author metadata attached to them, with the biggest distortions appearing in exactly the judgments that decide which papers get accepted, funded, and celebrated.
The experiment was designed with unusual rigor. Rospocher assembled 80 recent English-language arXiv manuscripts, twenty each from computer science, mathematics, physics, and quantitative biology, deliberately chosen in short and long length buckets and posted after March 2025 to reduce the chance that the evaluator models had memorized them. Author-identifying information, acknowledgements, and repository links were stripped from the PDFs so that the text itself carried no identity clues. Each manuscript was then evaluated under a factorial design: a corresponding-author profile was appended varying three factors at two levels each, producing eight profile conditions plus a blind baseline with no author information at all. Because every manuscript was evaluated under every condition, each paper served as its own control, a repeated-measures structure that sharply isolates the effect of the metadata from the effect of the content.
The three manipulated factors were carefully chosen to represent different classes of real-world cues. The first was a name-coded identity cue, alternating between first names commonly perceived in the United States as Black female-coded, such as Lakisha and Tanisha, and White male-coded, such as Brad and Hunter, all paired with a fixed surname to limit confounding. The second was institutional prestige, drawn from the QS World University Rankings 2026, contrasting elite institutions like MIT, Oxford, and Stanford with lower-ranked ones like Cleveland State and Central Michigan. The third was a bibliometric profile: a high profile with an h-index of 32 and roughly 8,000 citations, versus a low profile with an h-index of 8 and 320 citations. These numbers matter because they are exactly the kind of externally retrievable signals that a retrieval-augmented AI assistant could surface through a routine author lookup, even when the manuscript itself is anonymous.
Six heterogeneous evaluator models scored every manuscript under every condition: GPT-4.1-mini accessed through an API, and five locally hosted open-weight models including Gemma 3, Llama 3.1, Olmo 3, Qwen3, and a review-specialized model called DeepReviewer. Each model produced integer ratings on a seven-point scale from minus three to plus three across fourteen review-style criteria, spanning writing clarity, novelty, methodological soundness, perceived importance, award-worthiness, funding potential, hiring-committee appeal, and an overall acceptance recommendation. Each manuscript-condition-model combination was run five times with near-deterministic decoding settings, and the results were averaged and pooled across models with equal weight. Uncertainty was quantified with 20,000 bootstrap resamples, sign-flip permutation tests, and false-discovery-rate correction, alongside an ordinal-direction diagnostic that used only the sign of each paired difference to avoid assuming equal intervals on the rating scale.
The headline finding is that evaluator outputs are not invariant to author metadata. The strongest and most consistent driver was the bibliometric profile: papers attributed to a high-citation author scored systematically higher on perceived prestige, impact, and acceptance recommendation than the very same papers attributed to a low-citation author. Institution tier produced a similar but smaller elevation. The name-coded identity cue had comparatively modest effects overall, though it was not absent: it produced a reliable positive shift on prestige-oriented items such as award-worthiness and hiring-committee appeal. Crucially, the effects were not distributed evenly across the evaluation instrument. Content-quality judgments, such as clarity and methodological adequacy, barely moved. The distortions concentrated on the gatekeeping layer: top-tier acceptance, awards, funding, promotion, and the final accept-or-reject recommendation.
This selective vulnerability is what makes the result so consequential. It suggests the models are not applying a uniform halo effect that inflates every judgment when a prestigious name appears. Instead, they are perturbing a specific layer of evaluation tied to anticipated status, visibility, and institutional reward. The question-level decomposition showed the largest and most robust effects on items asking whether the manuscript meets the standard of a top-tier venue, deserves a best-paper award, warrants competitive funding, would impress a hiring committee, or is likely to change how the field thinks. Items about whether the writing is clear or the evidence sufficient were largely untouched. In other words, the metadata leaks into the decisions, not the diagnosis.
The study then pushed the analysis one step further, from scores to consequences. Using the acceptance recommendation as the ranking scalar, Rospocher recomputed manuscript rankings under counterfactual metadata swaps while holding everything else fixed. The results were striking. When a high-bibliometric profile was swapped for a low one, roughly a quarter to nearly two fifths of manuscripts originally selected in top-K shortlists, across thresholds from ten to twenty-five percent of the pool, fell below the cutoff after the perturbation. Institution-tier swaps produced an intermediate disruption, and name-cue swaps a smaller but still non-trivial one. Rank-displacement distributions were asymmetric, with downward movements more common and more extreme than upward ones, consistent with the finding that high-prestige cues systematically elevate scores. Even when most manuscripts stayed near their original position, those sitting near the selection boundary could cross the threshold, changing shortlist membership outright.
The audit also revealed meaningful variation across evaluator models. Random-effects meta-analysis showed that effect directions were broadly consistent, but magnitudes varied substantially, especially for prestige-related contrasts, with a large share of total variation attributable to between-model heterogeneity rather than sampling error. Inter-evaluator agreement diagnostics reinforced the picture: single models showed only modest absolute agreement on the same manuscripts, while averaging across all six improved reliability considerably. The practical implication is that metadata sensitivity is not a quirk of one model but a property of the model class, and that two pipelines differing only in their choice of evaluator may exhibit meaningfully different bias profiles. Domain-stratified analyses, meanwhile, found the qualitative pattern held across all four sampled fields, though magnitudes varied, and interaction tests suggested the three cue channels operate approximately additively rather than through narrow cue combinations.
The author is careful about scope. The study does not include a human-reviewer baseline, so it cannot say whether this sensitivity is unique to machines or inherited from patterns in human judgments embedded in training data. The corpus is limited to 80 English-language papers from four largely hard-science arXiv domains, the name cues are U.S.-centric and name-coded rather than ground-truth demographic attributes, and the experiment disabled browsing and retrieval to isolate the effect of explicitly supplied metadata. Real-world assistants that actively look up author information may introduce additional pathways not tested here. Still, the counterfactual logic is airtight within its design: identical texts, varied only in author context, produced different verdicts.
The design implications are immediate. The paper argues that author-context metadata should be treated as a consequential design variable, not neutral input, in any system that uses large language models for scientific evaluation, ranking, or discovery. In double-blind settings, LLM-assisted review should default to content-only evaluation with identifiers and lookup pathways suppressed. In single-blind or retrieval-augmented settings, systems should structurally decouple content scoring from context-informed summaries, document which metadata fields were available to the evaluator, and treat track record as an explicit, separately disclosed criterion if it is genuinely part of the decision rule, rather than letting it silently leak into manuscript scores. Because sensitivity varies by model, deployments should be audited whenever the evaluator is changed or updated. As AI-assisted reviewing spreads through editorial pipelines, the study delivers a clear warning: an algorithmic reviewer that can see who wrote a paper may not be able to stop that knowledge from coloring its judgment of the work itself.
Subject of Research: Counterfactual auditing of author-metadata sensitivity in large language model-based scientific peer review
Article Title: Author metadata affects Large Language Model scores in scientific peer review
Article References: Rospocher, M. (2026). Author metadata affects Large Language Model scores in scientific peer review. Discover Artificial Intelligence, 6(1), Article 1404. https://doi.org/10.1007/s44163-026-02268-y
Image Credits: AI Generated
DOI: 10.1007/s44163-026-02268-y
Keywords: large language models, peer review, algorithmic bias, bibliometrics, institutional prestige, counterfactual audit, scientific evaluation, editorial triage, LLM-as-a-judge, ranking fairness, arXiv, research integrity
News Source: Denise Maddox. (October 8, 2026). AI Reviewers Judge the Author, Not Just the Paper, Study Finds. Scienmag.



