When a deadly pathogen begins to spread, policymakers increasingly turn to computer simulations to answer urgent questions: How high will the infection peak climb? When will it arrive? How severe will subsequent waves be? Among the modeling approaches used for these questions, agent-based simulation occupies a special place, because it represents every individual in a population as a discrete entity that can become infected, recover, lose immunity, and become susceptible again. But a new study published in Complex & Intelligent Systems raises an uncomfortable question about this practice: if several competent teams independently build the same epidemic model from the same written specification, will their simulations agree? The answer, according to a large collaborative experiment led by Benjamin Antunes and colleagues, is a resounding no — and the reasons behind that disagreement carry lessons for the entire field of computational science.
The study set out to test replicability, a concept that is closely related to but distinct from reproducibility. Reproducibility typically asks whether the same code, run again, produces the same numbers. Replicability asks something harder: whether an independent implementation of the same model, built from the same specification, produces the same results. This distinction matters enormously in agent-based modeling, where a model is rarely a single equation but rather a set of rules describing how autonomous agents interact, move through disease states, and drive emergent population-level dynamics. Because so many discretionary choices are involved in turning written rules into running code, the researchers suspected that the act of implementation itself might inject variability that no amount of careful specification could eliminate.
To find out, the team designed a rigorous and unusually ambitious experiment. They chose a SEIRS epidemiological model — a classic framework in which individuals cycle through Susceptible, Exposed, Infectious, and Recovered states, with the crucial twist that recovered individuals eventually lose their immunity and return to the susceptible pool, allowing epidemic waves to recur. The model’s dynamics are well understood: an initial sharp infection peak followed by a series of progressively damped oscillations as the population settles toward equilibrium. The researchers then wrote a common formal specification of the model using the ODD protocol — Overview, Design concepts, and Details — which is the closest thing the agent-based modeling community has to a standardized blueprint for describing models.
With the specification in hand, seven independent modelers each implemented the identical model on a different platform or in a different programming language. The roster spans the full breadth of the simulation ecosystem: hand-coded implementations in C++, Julia, and Python; the dedicated agent-based modeling platforms NetLogo, GAMA, and Cormas; and PythonPDEVS, a framework rooted in the Discrete Event System Specification formalism. Each implementation was developed without coordination among the modelers, mimicking the real-world conditions under which scientific software is typically produced. Every implementation was then run thirty times with different random seeds, generating a total of 210 simulation runs whose outputs could be compared statistically across implementations and within each implementation’s own stochastic variation.
The first, reassuring finding was qualitative. All seven implementations reproduced the expected epidemic trajectory: a sharp initial infection peak followed by damped oscillations converging toward a stable equilibrium. In other words, every team had correctly captured the essential structure of the model. Had the researchers stopped there, the story would have been one of success. But when they examined quantitative metrics — the amplitude of the infection peak, the timing of that peak, and other key characteristics of the epidemic curve — statistically significant differences emerged between implementations. Simulations that were supposed to be computationally equivalent were, in numerical terms, telling measurably different stories about the same hypothetical outbreak.
The magnitude of those differences is what elevates the finding from a technical footnote to a genuine concern. The variation between implementations turned out to be roughly 7.6 times larger than the stochastic variation within a single implementation across its thirty replications. To appreciate what this means, consider that stochastic noise is the variation modelers routinely expect and account for: run the same simulation many times, and random events produce a spread of outcomes. The divergence between independently written implementations was nearly eight times that expected, well-understood noise. In practical terms, an uncertainty band that a modeling team would report from their own repeated runs would dramatically understate the true uncertainty surrounding the model’s predictions, because it would exclude the variability introduced by implementation choices.
One immediate objection is that the experiment confounds two factors: the modeling platform and the person using it. Because each implementation was produced by a different modeler on a different technology, the observed variance reflects the combined effect of tool and developer, and the study’s design cannot separate the two. The researchers addressed this partially with a complementary analysis: several independent developers worked in a single language, either C++ or Java, removing the platform variable. Even under these controlled conditions, the individual developer alone induced statistically significant and large differences between implementations. This confirms that the human element — the countless small interpretive decisions each programmer makes when translating a specification into code — is itself an important source of divergence. Whether the technology contributes an additional, separable effect remains an open question that future work will need to disentangle.
The most actionable part of the study came from a controlled ablation analysis, in which the researchers systematically traced the divergence back to its origins. Much of it, they found, could be attributed to just two under-specified execution details. The first concerns how residence times — the durations individuals spend in the Exposed, Infectious, or Recovered states — are discretized into whole days. A specification that says individuals remain infectious for, say, five days on average leaves open whether that duration is drawn from a continuous distribution and rounded, truncated, or sampled as an integer directly, and each choice subtly shifts the epidemic dynamics. The second detail concerns whether the state-transition test is strict: when a model checks whether an individual should move from one disease state to another, small differences in how that probabilistic test is implemented can accumulate across thousands of agents and hundreds of time steps.
Neither of these details is exotic or obscure. They are precisely the kinds of decisions that fall through the cracks of even careful model documentation, because they seem too minor to mention. Yet the ablation analysis showed that simply specifying them explicitly in the ODD protocol would remove much of the observed divergence between implementations. This is a genuinely hopeful conclusion. It suggests that the replicability problem in agent-based modeling is not an intractable consequence of complexity, but rather a specification gap that the community can close. The practical prescription is clear: model descriptions should pin down not only the conceptual rules of a model but also the execution-level details — how continuous quantities are discretized, how probabilistic transitions are evaluated, and at what point in the simulation loop each operation occurs.
The broader implications extend well beyond epidemiology. Agent-based models are used to study economics, ecology, urban systems, and social dynamics, and their results increasingly inform real decisions. The reproducibility crisis that has shaken psychology, medicine, and other empirical fields now has a computational analogue, and this study provides some of the clearest quantitative evidence of its scale within simulation science. The authors emphasize that reproducibility is not merely a technical virtue but a cornerstone of scientific credibility, and their findings argue for rigorous practices at every stage: precise specifications, independent replication as a standard check, and open sharing of code and data. In a welcome demonstration of that principle, the team has made all source code, raw simulation outputs, and analysis scripts publicly available on GitHub, allowing anyone to scrutinize, reuse, or extend the experiment. As simulation models take on ever greater roles in policy and prediction, this study stands as both a warning and a roadmap: divergence between implementations is real and large, but with sufficiently precise specifications, it is largely preventable.
Subject of Research: Replicability of agent-based epidemiological simulations across programming platforms and languages
Article Title: Replicability of agent-based simulations: a case study with a SEIRS epidemiological model
Article References: Antunes, B., Havion, H., Delay, É., Foucher, C., Hill, D. R. C., Mazel, C., Bisgambiglia, P.-A., Scriban, A., Zaitsev, O., & Duboz, R. (2026). Replicability of agent-based simulations: a case study with a SEIRS epidemiological model. Complex & Intelligent Systems. https://doi.org/10.1007/s40747-026-02546-3
Image Credits: AI Generated
DOI: 10.1007/s40747-026-02546-3
Keywords: agent-based modeling, replicability, reproducibility crisis, SEIRS model, epidemiological simulation, ODD protocol, NetLogo, GAMA, Cormas, simulation platforms, computational science, model specification
News Source: Kristina Jarvis. (October 9, 2026). Seven Teams, One Epidemic Model, Diverging Results: Inside Simulation’s Replicability Problem. Scienmag.



