A new study suggests that one of the biggest challenges in artificial intelligence—teaching machines to understand science rather than merely recognize patterns—may have a powerful solution: train neural networks on simulations built from the laws and mechanisms governing the real world.
In research published in Scientific Reports, Dudley, Magdaleno, Harding and colleagues show how mechanistic simulations can be used to improve scientific inference. The approach links two rapidly advancing fields: physics-based or process-based modeling, which attempts to reproduce how a system actually works, and machine learning, which excels at finding complex relationships in data. Together, they offer a route toward artificial intelligence that can estimate hidden scientific quantities, test explanations and make predictions even when direct measurements are incomplete.
Most neural networks learn by examining large collections of examples. During training, the system adjusts millions of internal parameters until its predictions become increasingly accurate. This strategy has transformed image recognition, language processing and automated decision-making, but it has an important weakness in scientific applications: a network can identify statistical correlations without understanding whether those correlations reflect genuine causation. If the data change, or if the network encounters conditions that were not represented during training, its predictions can fail dramatically.
Mechanistic simulations address that weakness by generating data from an explicit description of a system. Such a simulation may encode equations for physical forces, biological interactions, chemical reactions, population changes or other processes. Instead of simply recording what happened, the model represents why it happened. Researchers can then vary the underlying parameters, produce thousands or millions of synthetic observations and label each example with the hidden conditions that generated it. A neural network trained on those examples learns to reverse the process: given observable data, it estimates the unseen mechanisms or parameters responsible for them.
This is an inverse problem, and inverse problems are among the most difficult tasks in science. Measurements are often noisy, incomplete or indirect. Several different combinations of parameters may produce similar observations, a problem known as non-identifiability. A neural network can provide rapid estimates, but speed alone does not guarantee that the answer is scientifically meaningful. By training with mechanistic simulations, the researchers aim to constrain the network with information about the structure of the system, reducing the chance that it will rely on accidental features in a particular dataset.
The method also changes the economics of scientific discovery. Generating real-world observations can require expensive experiments, long-term monitoring or access to rare events. Simulations, by contrast, can produce controlled datasets at a fraction of the cost and can explore conditions that would be difficult or unethical to reproduce experimentally. A model can be exposed to extreme values, unusual combinations of variables and carefully designed perturbations, allowing researchers to investigate how its conclusions behave beyond the narrow range of existing measurements.
The technical advantage is especially important when a neural network must estimate parameters quickly. Traditional inference often involves repeatedly running a computationally demanding simulator while searching for the parameter values that best match observed data. This process can require thousands of simulations for a single case. Once trained, a neural network can approximate that inverse mapping in a fraction of the time. The simulator supplies the scientific structure during training, while the neural network provides speed during analysis.
The study’s central message is not that simulations can replace experiments, but that they can make machine-learning systems more disciplined and useful. A simulation is itself a hypothesis about how nature operates, and an inaccurate model can teach an AI system the wrong lessons. The researchers therefore emphasize the importance of testing whether conclusions learned from simulated data remain reliable when the network encounters observations from the real world. Differences between simulated and real measurements—known collectively as the simulation-to-real gap—can arise from noise, missing processes, calibration errors or assumptions built into the model.
Addressing that gap requires more than increasing the size of a neural network. Researchers may need to incorporate realistic measurement noise, compare predictions across multiple mechanistic models, use experimental data to recalibrate simulations and quantify uncertainty in the network’s output. Uncertainty is essential because a scientific estimate is not simply a number; it must also indicate how strongly the evidence supports that number. A well-designed system should be able to distinguish between a confident inference based on informative observations and an apparently precise answer produced when the available data cannot uniquely determine the underlying cause.
The growing interest in this approach reflects a broader change in scientific computing. Artificial intelligence is moving from tools that classify existing information toward systems that can assist with hypothesis testing, parameter estimation and model comparison. Neural networks trained on mechanistic simulations could help researchers analyze complex datasets in areas where direct interpretation is difficult, including biological systems, environmental processes and physical phenomena. Their greatest value may come not from producing an instant answer, but from narrowing the range of plausible explanations and revealing which additional measurements would most effectively discriminate between them.
The work also highlights a principle that is increasingly shaping the future of scientific AI: predictive performance should be connected to scientific validity. A model that performs well on a benchmark can still be misleading if it exploits shortcuts or fails under changed conditions. Mechanistic training offers a way to build those conditions into the learning process from the beginning. By combining equations, simulated experiments and data-driven algorithms, researchers can create systems that are faster than traditional inference methods while remaining anchored to explicit assumptions about the world.
As these methods mature, the critical test will be whether they improve reproducibility and discovery across independent datasets, laboratories and scientific disciplines. The most trustworthy systems will likely be hybrid: mechanistic simulations will provide structure, neural networks will provide computational efficiency, and real experiments will determine whether the resulting inferences survive contact with nature. The study by Dudley, Magdaleno, Harding and colleagues adds momentum to that vision, suggesting that the path to more capable scientific AI may not lie in abandoning scientific models, but in teaching machines to learn through them.
Subject of Research: Training neural networks on mechanistic simulations to improve scientific inference
Article Title: Training neural networks on mechanistic simulations improves scientific inference
Article References: Dudley, C., Magdaleno, R., Harding, C. et al. Training neural networks on mechanistic simulations improves scientific inference. Sci Rep (2026). https://doi.org/10.1038/s41598-026-64106-6
Image Credits: AI Generated
DOI: 10.1038/s41598-026-64106-6
Keywords: neural networks, mechanistic simulations, scientific inference, machine learning, inverse problems, computational modeling, simulation-to-real gap, uncertainty quantification, artificial intelligence
Tags: artificial intelligence for understanding physical lawscausation-aware AI modelsenhancing generalization in scientific AIimproving scientific inference with neural networksintegrating physics simulations with machine learningmechanistic simulation in AIovercoming data limitations in scientific AIphysics-driven machine learning modelspredictive modeling using mechanistic simulationsprocess-based modeling for scientific discoveryScience-based neural network trainingsimulation-based neural network training



