Every mathematical model of a biological system begins with an act of judgment. Which molecular interactions matter? Which reactions can be safely ignored? Which hypotheses deserve to be encoded in equations at all? These questions sit at the heart of systems biology, and answering them has long been a slow, error-prone, and deeply idiosyncratic process. A team of researchers now reports a solution that could change how the field works: PEtab Select, a new specification standard and software package for automated model selection, published in PLOS Computational Biology. The tool promises to bring the same rigor and reproducibility to the problem of choosing between competing models that earlier standards brought to parameter estimation itself.
The problem it addresses is deceptively simple to state and notoriously hard to solve. When scientists model a biological process—say, a signaling cascade inside a cell or the dynamics of a metabolic pathway—they typically face a thicket of competing hypotheses. Each hypothesis yields a different candidate model, and each candidate model contains unknown parameters that must be fitted to experimental data before the model’s performance can be judged. Comparing candidates therefore requires not one optimization problem but many, sometimes an enormous number of them. Until now, there has been no standard way to specify such model selection problems, no common format that different tools could read, and no straightforward route to evaluating a broad spectrum of selection approaches on equal footing.
PEtab Select fills that gap with two tightly linked components. The first is a file format standard: a concise, machine-readable specification of a model selection problem and its associated calibration problems. The second is a supporting software package that implements the standard and connects it to established modeling and calibration workflows. The design builds directly on PEtab, an existing and widely adopted standard for specifying parameter estimation problems in systems biology. By extending PEtab rather than starting from scratch, the developers ensured that the enormous ecosystem of tools and datasets already compatible with PEtab can be leveraged immediately, lowering the barrier to adoption for research groups around the world.
What makes the standard remarkable is its scalability. Model selection problems in biology can balloon to staggering sizes, and the PEtab Select format is designed to represent even very large model spaces compactly. In one demonstration highlighted by the authors, the format handles a problem comprising billions of model alternatives—far too many to enumerate by hand or to describe with conventional ad hoc scripts. The compact representation means that a researcher can define an entire universe of candidate models in a small set of files, and any PEtab Select-compliant tool can then interpret that universe identically. This eliminates a whole class of errors that arise when selection problems are re-implemented from scratch in each laboratory, each script, each paper.
Interoperability is the second pillar of the project. PEtab Select enables the use of state-of-the-art modeling and calibration workflows through connections to several widely used software platforms, including COPASI, Data2Dynamics, PEtab.jl, and pyPESTO. Each of these tools brings its own strengths—different optimization algorithms, different sampling methods, different language ecosystems—and each can now consume the same standardized problem description. For the working scientist, this means a model selection problem defined once can be attacked with multiple independent toolchains, and results can be cross-checked across implementations. For the community at large, it means that methods developers can benchmark their new algorithms against established ones using identical problem definitions, a practice that has been frustratingly rare in this corner of computational biology.
The standard also takes a clear position on how models should be judged. PEtab Select supports the model selection criteria most commonly used in the field, including the Akaike information criterion and the Bayesian information criterion, both of which balance goodness of fit against model complexity to avoid overfitting. Crucially, the framework is not locked into these choices: it can be easily extended to accommodate other criteria as the field’s statistical practices evolve. This flexibility matters because the debate over the right way to compare models—whether through information criteria, likelihood ratios, or Bayesian evidence—is far from settled, and a standard that hard-coded one philosophy would quickly become an obstacle rather than an enabler.
Equally important is the way candidate models are explored. PEtab Select implements several model space exploration strategies, ranging from the elementary to the sophisticated. At the simple end sit brute-force approaches that evaluate every candidate in the space, along with forward selection, which starts from a minimal model and iteratively adds components, and backward selection, which starts from a maximal model and prunes them away. Beyond these classics, the software supports advanced and flexible selection methods that can navigate vast model spaces more intelligently. The availability of multiple strategies within a single standardized framework allows researchers to match the exploration method to the problem at hand—and to compare strategies directly, something that previously required stitching together incompatible tools.
The authors describe PEtab Select as the first standardization of model selection tasks, and the claim carries weight. Parameter estimation, the sibling problem, has benefited from standards for years, and that standardization has paid dividends in reproducibility, software reuse, and the speed of method development. Model selection, by contrast, has remained a wilderness of bespoke pipelines, each constructed for a single study and rarely reused. By filling this critical gap in existing computational pipelines, PEtab Select extends the reach of standardization one crucial step further upstream, to the very point where hypotheses are formulated and compared. The result, the researchers argue, is an essential contribution to FAIR research software—software that is findable, accessible, interoperable, and reusable—in systems biology.
The practical implications reach well beyond convenience. Reproducibility crises in the life sciences have repeatedly implicated computational methods, and model selection is a stage where researcher degrees of freedom multiply silently: which candidates were tried, which criterion was applied, which exploration strategy was used, and how the winner was chosen can all shape the final published model in ways that are difficult to audit after the fact. A standardized, machine-readable specification makes every one of those choices explicit and shareable. A reviewer can inspect the exact model space a paper explored; a rival group can rerun the selection with a different criterion or algorithm; a methods developer can test a new optimizer against a published problem without reverse-engineering someone’s scripts. In this sense, PEtab Select does not merely automate a tedious task—it converts an opaque judgment call into a transparent, testable procedure.
For a field increasingly awash in high-dimensional omics data and increasingly ambitious in the mechanistic models it attempts to build, the timing could hardly be better. As models grow and hypothesis spaces expand into the billions, manual curation of candidate models becomes not just impractical but conceptually untenable. Automated, standardized, and interoperable model selection offers a path forward in which the choice among competing descriptions of life’s machinery is made systematically, at scale, and in the open. With PEtab Select, the systems biology community has taken a decisive step toward that future—and, in doing so, has reminded the broader scientific world that the humble file format, done right, can be one of the most powerful instruments in the laboratory.
Subject of Research: A standardized file format and software package for automated model selection in systems biology
Article Title: PEtab Select: Specification standard and supporting software for automated model selection
Article References: Pathirana, D., Bergmann, F. T., Doresic, D., Lakrisenko, P., Persson, S., Neubrand, N., Timmer, J., Kreutz, C., Binder, H., Cvijovic, M., Weindl, D., Hasenauer, J., & with the PEtab Select community (2026). PEtab Select: Specification standard and supporting software for automated model selection. PLOS Computational Biology, 22(9), e1014774. https://doi.org/10.1371/journal.pcbi.1014774
Image Credits: AI Generated
DOI: 10.1371/journal.pcbi.1014774
Keywords: PEtab Select, systems biology, model selection, parameter estimation, computational biology, FAIR research software, Akaike information criterion, Bayesian information criterion, COPASI, pyPESTO, open standards, reproducibility
News Source: Drew Townsend. (October 8, 2026). New Standard PEtab Select Brings Order to the Chaos of Choosing Biological Models. Scienmag.



