Diacylglycerol, or DAG, is one of those quiet molecules that food scientists have long wanted more of. It occurs naturally in vegetable oils at levels below ten percent, yet studies have linked it to beneficial effects on lipid metabolism and a reduced risk of obesity and related metabolic disorders. Because its natural abundance is too low to deliver functional benefits, producers must synthesize it, typically by reacting triglycerides with glycerol under the guidance of an enzyme. A new study published in Food Chemistry: X has now tackled this task with an unexpected tool: a machine learning framework that actively chooses its own experiments, and in doing so outperformed the industry-standard optimization method by a wide margin.
The research team, led by Haoran Liu of Wuhan Polytechnic University, focused on sunflower oil as their starting material. Sunflower oil is an attractive feedstock because it is more than 96 percent digestible and rich in tocopherols, phytosterols, and trace elements. The researchers used immobilized Candida antarctica lipase B, sold commercially as Novozym 435, to catalyze the glycerolysis reaction, in which triglycerides exchange fatty acid groups with glycerol to form diacylglycerol. Four process variables govern the outcome: reaction time, temperature, enzyme loading, and the weight ratio of oil to glycerol. The challenge was finding the combination of these four factors that maximizes DAG content.
For decades, the standard answer to such problems has been response surface methodology, or RSM. The approach is elegant in its simplicity: run a structured set of experiments, such as a central composite design, fit a second-order polynomial equation to the results, and then use mathematics to locate the predicted optimum. In this study, the team ran a 30-experiment central composite design spanning reaction times of 9 to 13 hours, temperatures of 50 to 90 degrees Celsius, lipase loadings of 6 to 10 percent, and oil-to-glycerol ratios of 9 to 13. The fitted model was statistically strong, explaining over 96 percent of the variance, and analysis of variance confirmed that temperature, enzyme loading, and time all had significant effects.
But RSM has a fundamental weakness: it is static. Every experiment is planned before any data is collected, so if the true optimum lies outside the initial design space, or if the response surface is more twisted and nonlinear than a parabola can capture, the method has no way to adapt. Worse, when replicate measurements are averaged before modeling, RSM discards information about local variability that could reveal how trustworthy its predictions actually are. As optimization problems in food engineering grow more complex, these limitations compound.
The alternative tested here is called active learning sequential surrogate-based optimization, or AL-SSBO. Instead of committing to a fixed experimental plan, the method works in a closed loop: predict, recommend, experiment, update. An XGBoost model, a gradient-boosted ensemble of decision trees known for its computational efficiency and built-in regularization, serves as the surrogate that stands in for the real chemistry. In each of five rounds, the algorithm proposes two new experimental conditions, the lab runs them in triplicate, and the results feed back into the model before the next round begins.
The clever part lies in how the algorithm chooses what to test. Its scoring function balances exploitation, meaning conditions where the model predicts high DAG yield, against exploration, meaning conditions far from any previously sampled point, using a distance-based bonus scaled to the average spacing of the training data. A penalty discourages candidates that stray beyond a predefined distance threshold, and K-means clustering ensures the two recommendations in each round are spatially diverse rather than near-duplicates. The result is a search that leaps efficiently between distant regions of the parameter space early on, then converges tightly on the most promising zone.
The numbers tell a striking story. RSM predicted an optimum of 53.13 percent DAG at 11.78 hours, 70.96 degrees Celsius, 9 percent lipase, and an oil-to-glycerol ratio of 11.23, but experimental validation delivered only 51.70 percent, an error of 1.43 percentage points. AL-SSBO, working from the very same starting dataset, recommended conditions of 11.93 hours, 68.56 degrees Celsius, 8.93 percent lipase, and a ratio of 11.71, predicting 53.87 percent. The measured value was 53.77 percent, an error of just 0.10 percentage points, roughly fourteen times smaller than RSM’s. The yield improvement over RSM was 2.07 percentage points, achieved in a near-optimal region where the polynomial model had essentially exhausted its capability.
The team also compared surrogate models, testing Gaussian process regression and random forest alongside XGBoost under identical rules. Random forest achieved the lowest cross-validation error on the shared dataset, and Gaussian process performed best across its own sequential trajectory, but XGBoost delivered the lowest final-round prediction error and the highest experimentally achieved DAG content, so it was retained. A bootstrap-based statistical comparison, chosen because sequential data violates the independence assumptions of classical tests, showed that the 95 percent interval for the AL-SSBO optimum, spanning 52.77 to 54.16 percent, sat entirely above the RSM result, with a standardized effect size of 2.88 and zero probability of practical equivalence to the older method.
Perhaps most compelling is that the algorithm’s search path made chemical sense. Principal component analysis of the sampling trajectory revealed that early rounds explored high-temperature regions, which were quickly deprioritized once experiments showed they underperformed, while later rounds concentrated on lower temperatures, longer times, and higher enzyme and substrate ratios. This matches the underlying enzymology: moderate heat accelerates molecular collisions, but lipase is a thermally sensitive protein that degrades past a critical threshold, and longer reaction times allow the reversible glycerolysis network to approach thermodynamic equilibrium. The machine was not wandering blindly; it was rediscovering reaction kinetics from data alone, with a final-round spatial dispersion 56.6 percent lower than the initial design.
The authors are careful to note the limits of their claims. The five-round budget was fixed in advance rather than terminated by a convergence rule, retrospective checks confirmed stabilization only within the studied domain, and the results represent a laboratory-scale reference rather than an industrial operating specification. Future work must test the workflow with sparser starting datasets, during scale-up, and across multiple objectives including byproduct formation and enzyme reuse. Still, the message is clear: when a machine can learn from each experiment and decide what to try next, even a mature, well-characterized process like sunflower oil glycerolysis holds untapped potential, and the fifty-year reign of the response surface in the food lab may finally be facing a serious challenger.
Subject of Research: Machine learning-based optimization of enzymatic diacylglycerol synthesis from sunflower oil
Article Title: Active learning sequential surrogate-based optimization (AL-SSBO) versus response surface methodology in sunflower oil diacylglycerol synthesis: Performance and insight into AL-SSBO
Article References: Liu, H., Dai, Y., Liu, H., Zhang, S., Yu, H., Liao, Q., Chen, D., Jiang, X., & He, D. (2026). Active learning sequential surrogate-based optimization (AL-SSBO) versus response surface methodology in sunflower oil diacylglycerol synthesis: Performance and insight into AL-SSBO. Food Chemistry: X, 39, Article 104516. https://doi.org/10.1016/j.fochx.2026.104516
Image Credits: AI Generated
DOI: Not provided
Keywords: diacylglycerol, sunflower oil, machine learning, XGBoost, active learning, response surface methodology, enzymatic glycerolysis, lipase, food chemistry, process optimization, surrogate modeling, Novozym 435
Cite Scienmag News
APA MLA Chicago
Bethany Barker. (October 2, 2026). AI Learns to Cook Healthier Sunflower Oil, Beating a 50-Year-Old Chemistry Method. Scienmag. https://scienmag.com/ai-learns-to-cook-healthier-sunflower-oil-beating-a-50-year-old-chemistry-method/
Bethany Barker. “AI Learns to Cook Healthier Sunflower Oil, Beating a 50-Year-Old Chemistry Method.” Scienmag, 2 October 2026, https://scienmag.com/ai-learns-to-cook-healthier-sunflower-oil-beating-a-50-year-old-chemistry-method/. Accessed 2 October 2026.
Bethany Barker. “AI Learns to Cook Healthier Sunflower Oil, Beating a 50-Year-Old Chemistry Method.” Scienmag. October 2, 2026. https://scienmag.com/ai-learns-to-cook-healthier-sunflower-oil-beating-a-50-year-old-chemistry-method/
Copy citation Download RIS
Tags: active learningadvanced experimental design in food scienceAI-driven food processingbiotechnology in oil refiningdiacylglycerolenzymatic glycerolysisenzymatic glycerolysis optimizationenzyme catalysis in food industryfood chemistryfood chemistry innovationlipaselipid metabolism enhancementMachine learningmachine learning in lipid synthesisnatural diacylglycerol productionNovozym 435obesity risk reductionprocess optimizationresponse surface methodologysunflower oilsunflower oil health benefitssurrogate modelingsustainable oil modification methodsXGBoost


