Concrete is the most consumed manufactured material on Earth, and it comes with an enormous carbon bill. Portland cement, the glue that holds concrete together, is responsible for a substantial share of global carbon dioxide emissions, driving engineers to search for greener alternatives. One of the most promising strategies is to replace a portion of cement with supplementary cementitious materials, or SCMs, such as fly ash, ground granulated blast furnace slag, and silica fume. These industrial by-products reduce clinker content, improve durability, and enhance resistance to aggressive environments, making them attractive for marine structures, transportation infrastructure, and large-scale foundations. But there is a catch: the mechanical behaviour of SCM-based concrete is governed by a tangled web of nonlinear, interdependent interactions between binders, water, aggregates, and chemical admixtures, and conventional empirical formulas struggle to capture it.
A new study published in Results in Engineering tackles this complexity head-on with a comprehensive machine learning framework that predicts the compressive strength of SCM-based concrete and, crucially, explains its predictions. A team of researchers led by Abdul Qadir Bhatti, Muhammad Nasir Amin, and Ayaz Ahmad compiled a heterogeneous dataset of 329 SCM-blended concrete mixtures drawn from published experiments across multiple regions, originally assembled by Liu and colleagues and publicly available through the Mendeley Data repository. The inputs are the seven quantities engineers know before a single batch is mixed: cement, water, fine aggregate, coarse aggregate, fly ash, slag, and superplasticizer dosages. The target is the cylinder compressive strength in megapascals, which ranged from 10.26 to 78 MPa across the database, with a mean of roughly 41 MPa.
Rather than inventing a new algorithm, the researchers systematically benchmarked four established modelling strategies under identical conditions: CatBoost, LightGBM, Random Forest, and Gaussian Process Regression. CatBoost builds decision trees sequentially with an ordered boosting mechanism that reduces prediction bias from information leakage. LightGBM grows trees leaf-wise, splitting whichever branch yields the greatest loss reduction, and uses histogram-based learning to slash computational cost. Random Forest aggregates many decision trees trained on bootstrap samples, averaging their outputs to suppress variance. Gaussian Process Regression, by contrast, is a probabilistic approach that defines a distribution over possible functions using a covariance kernel, offering smooth trend representation and built-in uncertainty estimates, though at higher computational cost and with strong sensitivity to kernel choice.
Hyperparameter optimisation played a central role in the workflow. Instead of relying on default settings, the team employed a randomised search to explore configurations governing tree complexity, learning rates, sampling strategies, and regularisation strength, tuning against unseen data to mitigate overfitting. Model performance was then assessed with three complementary metrics: the coefficient of determination, which captures the proportion of explained variance; mean absolute error, which measures the average magnitude of deviations; and root mean square error, which penalises larger errors. Ten-fold cross-validation and residual analysis were layered on top to verify that predictions were stable across data subsets and free of systematic bias.
The results delivered a clear winner. LightGBM achieved coefficients of determination of 0.91, 0.70, and 0.80 on training, testing, and validation subsets under the original random partition, with a mean absolute error of 4.74 MPa and a root mean square error of 7.08 MPa. Random Forest followed closely with training, testing, and validation values of 0.95, 0.71, and 0.77, while CatBoost and Gaussian Process Regression trailed with validation scores of 0.67 and 0.65 respectively. Residual distributions for all models were centred near zero, with LightGBM’s mean residual at just 0.21 MPa, indicating minimal systematic bias. The gap between mean absolute and root mean square errors suggested that extreme deviations were rare and did not dominate the error structure.
One of the study’s most instructive findings concerns data partitioning rather than algorithms. When the team repeated the analysis using an SPXY-based representative partition, which allocates samples to ensure coverage of the full mixture and strength domains, performance improved dramatically across the board. LightGBM’s coefficients of determination jumped to 0.949, 0.952, and 0.955 for training, testing, and validation, while Random Forest reached 0.907, 0.938, and 0.911. The lesson is striking: the apparent weakness of the models under random splitting was largely an artefact of unrepresentative subsets in a heterogeneous, multi-source dataset, not a flaw in the learning methods themselves. For anyone applying machine learning to compiled literature data, representativeness of the data split may matter as much as the choice of model.
The ten-fold cross-validation results reinforced this hierarchy. LightGBM posted the highest mean coefficient of determination at 0.751 with the lowest standard deviation of 0.070, followed by CatBoost at 0.745, Random Forest at 0.733, and Gaussian Process Regression at 0.732 with the widest fold-to-fold variability of 0.111. The kernel-based approach, while elegant for smooth functional relationships, proved least able to adapt to the jagged, interaction-driven behaviour of blended concrete mixtures, whereas the ensemble methods thrived on it.
Perhaps the most scientifically valuable contribution is the explainability analysis. Using SHAP, a technique from explainable artificial intelligence that quantifies each feature’s contribution to individual predictions, the researchers ranked the drivers of compressive strength. Cement content emerged as the most influential parameter, with higher cement quantities producing positive contributions through increased hydration and matrix densification. Water content ranked second and behaved in the opposite direction: increasing water shifted predictions downward, reflecting the well-established damage that excess water inflicts on porosity and mechanical performance. Coarse aggregate, slag, fly ash, fine aggregate, and superplasticizer completed the ranking.
The SHAP dependence plots went further, revealing coupled interactions that single-variable analyses miss entirely. The positive contribution of cement was strongest at low water contents, confirming the coupled effect of binder and mixing water on strength development. The negative influence of water intensified as cement content decreased. Slag and fly ash interacted noticeably with each other, meaning their benefit depends on combined replacement levels rather than individual dosages. Even the aggregates and superplasticizer showed both positive and negative contributions across different regions of the mixture space, demonstrating that strength in SCM-based concrete is governed by the joint composition of the mix, not by any ingredient in isolation. This physically meaningful interpretation directly addresses the black-box criticism that has limited machine learning adoption in civil engineering.
To translate the framework into practice, the team built a graphical user interface in Python’s Tkinter library that lets engineers enter mixture proportions and receive an immediate compressive strength estimate without any programming expertise. The tool is designed as a decision-support aid for preliminary mixture design, allowing rapid comparison of candidate mixes before committing to batching, casting, curing, and destructive testing, processes that consume material, laboratory time, and money over extended curing periods. The authors are candid about the framework’s limits: the dataset lacks curing age, oxide composition, particle size distribution, and microstructural descriptors, and the models do not yet include formal uncertainty quantification. Future work will expand the database, incorporate richer physicochemical variables, validate against independent field data, and add probabilistic prediction intervals. Even so, the study marks a meaningful step toward sustainable construction, showing that interpretable machine learning can compress months of trial-and-error mix development into seconds of computation while keeping engineers, rather than opaque algorithms, firmly in control of the design decisions that shape the built environment.
Subject of Research: Explainable machine learning prediction of compressive strength in supplementary cementitious material-based concrete
Article Title: Predicting and interpreting compressive strength of SCM-based concrete using explainable machine learning
Article References: Bhatti, A. Q., Amin, M. N., Ahmad, A., Qadir, M. T., Tanoli, W. A., & Faraz, M. I. (2026). Predicting and interpreting compressive strength of SCM-based concrete using explainable machine learning. Results in Engineering, 32, Article 113184. https://doi.org/10.1016/j.rineng.2026.113184
Image Credits: AI Generated
DOI: 10.1016/j.rineng.2026.113184
Keywords: machine learning, concrete, compressive strength, supplementary cementitious materials, LightGBM, SHAP, fly ash, ground granulated blast furnace slag, sustainable construction, Random Forest, Gaussian Process Regression, explainable artificial intelligence
News Source: Denise Maddox. (October 9, 2026). AI Cracks the Code of Green Concrete, Predicting Strength Before a Single Batch Is Poured. Scienmag.



