UNIVERSITY PARK, Pa. — Artificial intelligence could dramatically accelerate the search for technologies that transform carbon dioxide into useful fuels, but a new multi-laboratory study shows that even the most advanced model can be undermined by something as simple as inconsistent stirring. Researchers from four U.S. laboratories have demonstrated that experimental variability can substantially alter measurements of catalyst performance, creating a serious challenge for scientists who want to use real-world laboratory data to train artificial intelligence and machine-learning systems. The work, led by scientists at SLAC National Accelerator Laboratory with contributions from Penn State, Stanford University and the University of California, Santa Barbara, was published in Nature Catalysis.
The study focused on a rhodium-based catalyst used in carbon dioxide hydrogenation, a chemical process in which carbon dioxide reacts with hydrogen to produce carbon-containing compounds. In the experiments, the catalyst was designed to promote the formation of carbon monoxide, an important intermediate that can be converted into synthetic fuels and other industrial chemicals. Methane was also produced as an undesirable side product. By comparing results from laboratories that followed a shared experimental protocol, the researchers hoped to measure how reproducible catalyst testing would be when performed by independent teams. The experiment was structured as a “round-robin” study, in which the same material and testing instructions are distributed among multiple laboratories so that differences in performance can be traced to procedures, equipment or operating conditions rather than to the catalyst itself.
The motivation is closely tied to the growing use of artificial intelligence in chemistry and chemical engineering. A well-trained model can examine relationships among catalyst composition, temperature, pressure, reaction time and product yield, then predict which combinations are most likely to work. Instead of testing thousands of possibilities experimentally, researchers could use the model to identify a smaller set of promising candidates and validate those predictions in the laboratory. Such an approach could reduce the time, cost and energy required to develop catalysts for carbon capture, fuel production and industrial chemical manufacturing. However, machine-learning systems do not inherently understand whether two measurements were obtained under truly equivalent conditions. If data from different laboratories contain systematic differences, an algorithm may interpret those differences as meaningful chemical trends.
The researchers initially expected the four laboratories to produce broadly comparable results because each team had agreed to use the same catalyst and experimental framework. Instead, the data showed noticeable differences in the amounts of carbon monoxide and methane generated, as well as in how catalyst activity changed over time. These discrepancies meant that a machine-learning model trained on the combined dataset could struggle to distinguish genuine chemical behavior from laboratory-specific effects. A catalyst might appear more selective in one dataset and less selective in another, even though the material itself was nominally identical. The outcome highlighted a central problem in data-driven science: large datasets are not necessarily high-quality datasets, particularly when measurements are assembled from multiple instruments, facilities and operating cultures.
To find the source of the disagreement, the teams systematically examined their methods and equipment. They compared reactor designs, gas-delivery systems, analytical instruments, temperature control, catalyst preparation and operating procedures. One of the most important contributors to variability was the way the reacting mixture was shaken or stirred. Mixing affects how efficiently gases dissolve, how reactants reach the catalyst surface and how quickly products leave the reaction environment. Small differences in agitation can therefore change the local concentrations surrounding the catalyst and alter the apparent reaction rate and product distribution. In heterogeneous catalysis, where a solid catalyst interacts with gases or liquids at an interface, these transport effects can be just as important as the catalyst’s intrinsic chemical properties.
The issue becomes even more significant when researchers attempt to model catalyst deactivation. Catalysts rarely maintain their initial performance indefinitely. During extended operation, active sites can become blocked by impurities, the catalyst structure can change, or high temperatures can cause particles to grow and lose surface area. In practical systems, deactivation may occur over months or years, while laboratory studies often last only days. If one laboratory records faster apparent deactivation because of differences in mass transfer, reactor geometry or temperature stability, an AI model may learn an inaccurate prediction of long-term behavior. This could lead researchers to reject a promising material, select an unsuitable catalyst for scale-up or overestimate the lifetime of a process intended for industrial deployment.
After identifying the main sources of mismatch, the researchers introduced additional standardization across the participating laboratories. They refined reactor configurations, clarified operating protocols and tightened control over experimental conditions. The results became more consistent, demonstrating that reproducibility can be improved when laboratories pay close attention not only to chemical composition but also to the physical details of an experiment. The findings support a broader shift in how scientific data are collected for machine learning. Rather than treating experimental measurements as interchangeable once basic conditions have been reported, researchers may need to record a far more complete description of equipment, mixing, calibration, catalyst handling and reaction history.
Robert Rioux, Friedrich G. Helfferich Professor of Chemical Engineering at Penn State and a co-author of the study, said the work represents an initial effort to quantify uncertainty in heterogeneous catalysis experiments performed across laboratories. The study’s authors argue that uncertainty should be treated as a central component of AI-ready scientific data, not as an inconvenient error term added at the end of an analysis. Selin Bac, a postdoctoral researcher at the University of California, Santa Barbara and the paper’s first author, emphasized that seemingly minor variations in experimental design can influence the reliability of machine-learning outcomes. Adam Hoffman, a staff scientist at SLAC and the study’s senior author, said the experience revealed practical difficulties that may be overlooked when real-world data are incorporated into predictive models.
The implications extend beyond carbon dioxide conversion and rhodium catalysts. Laboratories around the world are building databases for batteries, solar fuels, pharmaceuticals, biomaterials and chemical manufacturing, often by combining results generated under different conditions. If those datasets are not harmonized, AI systems may produce predictions that appear precise but fail when tested outside the laboratory that generated the original data. The researchers recommend stronger coordination among facilities, more consistent reactor and protocol designs, improved reporting of experimental parameters and deliberate round-robin testing before data are merged for modeling. Their message is not that artificial intelligence is unsuitable for catalyst discovery, but that its success depends on understanding the experimental uncertainty embedded in every measurement.
The study was supported in part by the U.S. Department of Energy’s Office of Science under award FWP 101064. Testing equipment was supplied in part by Co-ACCESS, part of the SUNCAT Center for Interface Science and Catalysis, a joint research center supported by SLAC National Accelerator Laboratory and Stanford University. Greg Barber, assistant professor of chemistry at Penn State Altoona and an affiliate researcher in Penn State’s Institute of Energy and the Environment, also contributed to the research. By showing how protocol and equipment standardization can determine whether independent laboratories obtain comparable catalyst results, the team has provided a practical warning for the rapidly expanding field of AI-driven science: before algorithms can reliably discover the next generation of fuels and chemical technologies, scientists must ensure that the data they learn from are genuinely comparable.
Subject of Research: Not applicable
Article Title: Quantifying uncertainty in catalyst activity and deactivation during CO2 hydrogenation via round-robin testing for data-driven modelling
News Publication Date: 31-Jul-2026
Web References: SLAC National Accelerator Laboratory, https://www6.slac.stanford.edu/ ; Penn State research profile, https://www.che.psu.edu/department/directory-detail-g.aspx?q=RMR189 ; Original news release, https://www6.slac.stanford.edu/news/2026-08-03-ai-drive-science-discoveries-highly-reproducible-data-key
References: Nature Catalysis, DOI: 10.1038/s41929-026-01559-y
Image Credits: Greg Stewart/SLAC National Accelerator Laboratory
Keywords
Artificial intelligence, machine learning, catalysis, carbon dioxide conversion, carbon monoxide, hydrogenation, catalyst deactivation, heterogeneous catalysis, experimental reproducibility, round-robin testing, chemical engineering, synthetic fuels, data-driven modeling, SLAC National Accelerator Laboratory, Penn State, Nature Catalysis
Tags: accuracy of AI models in predicting catalytic reactionsAI model reliability in laboratory catalyst testingchallenges in standardizing experimental procedures across labseffects of laboratory differences on carbon dioxide conversion efficiencyimpact of experimental variability on catalyst performance measurementsimplications for training AI and machine learning in catalysis researchimportance of reproducibility for industrial chemical process developmentinfluence of stirring and protocol differences on laboratory resultsmulti-laboratory study on chemical process consistencyreproducibility challenges in CO2 hydrogenation experimentsrole of multi



