Casting is one of the oldest manufacturing processes on Earth, yet it remains stubbornly difficult to master. When molten metal meets a mold, dozens of interacting variables—chemical composition, sand properties, temperatures, humidity—combine in nonlinear and often unpredictable ways to determine whether a finished component is sound or scrap. For large, thick-walled castings used in energy equipment, rail transportation, and construction machinery, the stakes are especially high, because a single hidden defect can compromise a part weighing many tons. Now, a study published in Results in Engineering describes a data-driven framework that replaced trial-and-error intuition with machine learning, and then proved itself on the factory floor with a 98.5 percent qualification rate across four consecutive production batches.
The research, conducted by Yutong Guo and Chao Yang, began with a problem familiar to every foundry engineer: pouring-process parameters have traditionally been set by experience and adjusted through repeated, costly iterations. When several parameters interact, operating conditions fluctuate, and multiple defect types appear simultaneously, this conventional approach suffers from slow adjustment cycles and difficulty pinpointing which factors actually drive defects. Industrial digitalization, however, has quietly accumulated vast archives of shop-floor records, and the authors recognized that this data could serve as the foundation for a closed-loop workflow: clean the data, predict quality, screen for likely defect causes, optimize the adjustable parameters, and then verify the recommendations in real production.
The dataset came from actual pouring records at a cooperating foundry producing large, complex castings. The raw template contained 65 fields: 49 input variables covering chemical composition, molding-sand properties, and temperature conditions; a product qualification rate; pouring and scrap quantities; and 13 categories of scrap-cause labels drawn from routine quality-inspection records. Of 354 raw records, four incomplete entries were removed, leaving 350 valid records for modeling. Because full-volume radiographic or industrial CT inspection was impractical for such large components, the factory relied on destructive sectioning at designated positions followed by visible-color liquid penetrant testing, in line with ISO 4987:2020 and ISO 3452-1:2021, exposing internal discontinuities such as sand holes and shrinkage porosity on the freshly cut surfaces.
With the data cleaned, the team compared four candidate regression models for predicting qualification rate: a feedforward neural network, k-nearest neighbors regression, Gaussian process regression, and a random forest. Under five-fold cross-validation, the 150-tree random forest achieved the lowest root mean square error, at 3.6865 percentage points, along with a mean absolute error of 2.5832. A sensitivity analysis varying the ensemble size from 10 to 500 trees confirmed that 150 was sufficient; larger forests produced no further error reduction. Random forest offered an additional practical advantage: it provides explicit feature-relevance rankings, which the researchers needed to identify which of the 49 variables deserved attention, and it could be evaluated rapidly thousands of times during a subsequent optimization search.
Feature relevance pointed to chemical-composition elements such as calcium, magnesium, carbon, sulfur, and silicon, along with AFS sand fineness and temperature-related variables. But the six variables ultimately selected for direct adjustment were not simply the six highest-ranked. The authors jointly weighed model relevance, shop-floor adjustability, and engineering feasibility, arriving at moisture content, hot-wet tensile strength, tapping temperature, AFS fineness, permeability, and sand temperature—factors that foundry crews can realistically control. Two-dimensional conditional-response surfaces showed that predicted qualification rate changed in interval-dependent, partially interacting ways across variable combinations, and Friedman H statistics quantified measurable but moderate non-additive interactions, the largest between tapping temperature and permeability.
The second computational stage tackled a subtler question: could routine scrap-cause records, rather than just the single qualification-rate number, inform the optimization? Because several defect causes can be recorded under the same process conditions, the team formulated the task as a multi-label screening problem rather than a mutually exclusive classification. A deep feedforward neural network with a 49–128–64–13 architecture—two ReLU-activated hidden layers feeding a 13-neuron sigmoid output—was trained on the 49 process parameters to produce scores for all 13 scrap-cause categories simultaneously. Given the extreme sparsity of the labels, only three categories had at least five positive records: sand hole, sand adhesion, and shrinkage porosity. Decision thresholds were calibrated on a validation subset instead of being fixed at 0.5, and the network was evaluated with standard multi-label metrics including Micro-F1, Hamming loss, and subset accuracy.
Crucially, the researchers did not treat the neural network as a second optimization objective. Instead, its label scores served as feasibility constraints during candidate screening: a proposed parameter setting was accepted only if the context-averaged scores for sand hole, sand adhesion, and shrinkage porosity did not increase relative to the historical baseline. This design kept the random forest as the primary evaluator of quality while the DNN acted as a guardrail against settings that might trade a higher predicted qualification rate for a higher risk of specific defects. It is an elegant division of labor—prediction, screening, and search each handled by a model suited to its role, all operating on the same production context.
The optimization itself used an evolutionary search implemented in the DEAP framework, with a population of 100, crossover probability of 0.7, and mutation probability of 0.3, searching within bounds set by observed historical values. A supplementary verification running the search for 100 generations across 10 independent runs showed that the best feasible solution at generation 2 was already within roughly 0.008 percentage points of the eventual plateau, evidence of rapid convergence. The random-forest surrogate estimated that replacing the six variables with optimized-center values would raise the context-averaged qualification rate from 97.9441 percent to 98.7799 percent—a model-based counterfactual gain of about 0.84 percentage points. Importantly, the authors present this as a model estimate supporting parameter recommendation, not a causal claim of field improvement.
Equally notable is how the recommendations were expressed. Rather than a single theoretical optimum—impractical in a factory where raw materials, equipment states, and ambient conditions constantly shift—the team reported bounded ranges for each variable. Moisture content was to decrease slightly from a historical mean of 3.028 percent toward 2.866 to 3.023 percent; hot-wet tensile strength was to increase from 4.90 kPa into the 5.51 to 5.94 kPa window; AFS fineness was to rise from 57.54 toward 58.30 to 63.45; and permeability was to fall from 172.4 into the 140.0 to 165.9 range. Tapping and sand temperatures changed only marginally. Read together, these coupled adjustments point toward a more stable mold condition: finer sand with lower permeability and moisture, but stronger hot-wet tensile resistance against erosion and scabbing at the mold–metal interface.
The final proof came from the foundry itself. A pilot trial of five products produced under the recommended settings yielded five qualified castings. Scale-up followed across four consecutive batches of 50 products each, with qualification rates of 98.0, 100.0, 98.0, and 98.0 percent—197 of 200 products overall, or 98.5 percent, compared with a historical aggregate baseline of 98.17 percent. Sectioned penetrant inspection added physical corroboration: before the adjustment, localized red penetrant bleed-out indications were visible near the junction region of exposed sections; after implementation, no comparable indication appeared in the inspected region. Combined with the lower modeled shrinkage-porosity label score, the evidence suggests genuinely improved local defect conditions. The study stands out precisely because so many data-driven manufacturing papers stop at offline modeling—here, prediction, multi-label defect screening, constrained evolutionary search, and staged field validation form one continuous loop, offering foundries a realistic template for turning accumulated production data into executable process windows.
Subject of Research: Data-driven quality prediction and process-parameter optimization for large complex castings using machine learning and evolutionary search
Article Title: Data-driven quality prediction and process optimization for large complex castings
Article References: Data-driven quality prediction and process optimization for large complex castings. (n.d.). https://doi.org/10.1016/j.rineng.2026.113269
Image Credits: AI Generated
DOI: 10.1016/j.rineng.2026.113269
Keywords: casting, machine learning, random forest, deep neural network, multi-label classification, evolutionary optimization, process optimization, defect prediction, foundry, liquid penetrant testing, qualification rate, manufacturing
News Source: Blake Davidson. (October 4, 2026). AI Learns to Pour Perfect Metal: Machine Learning Boosts Casting Quality in Real Factory Trial. Scienmag.



