Artificial intelligence has quietly become the load-bearing infrastructure of global commerce. Machine learning models predict what shoppers will want weeks before they order it, deep learning networks scan port terminals and warehouse floors, and algorithms reroute freight around storms, strikes, and congested customs queues. Yet as companies pour staggering sums into these technologies, a deceptively simple question has gone largely unanswered: how do you actually measure whether machine learning and deep learning are good for the business? A new study published in the Journal of Big Data offers one of the most concrete answers yet. Mahmoud M. A. AbdEllatif of the University of Jeddah in Saudi Arabia and Samah Ibrahim Abdel Aal of Zagazig University in Egypt have built a decision-making framework that translates the messy, uncertain judgments of supply chain stakeholders into rigorous, comparable scores—revealing which artificial intelligence techniques genuinely pay off and which merely look impressive on a benchmark.
The timing is hardly accidental. Supply chains now generate torrents of data—point-of-sale records, sensor readings from trucks and containers, weather feeds, supplier dashboards—and the complexity of the decisions made on top of that data has grown in step. Machine learning, in which algorithms learn statistical patterns from historical data, has become central to demand forecasting, inventory optimization, supplier selection, and dynamic pricing. Deep learning, its more powerful cousin, stacks artificial neurons into deep neural networks capable of digesting unstructured inputs such as images, audio, and raw text, enabling tasks like visual inspection of products, natural-language processing of contracts and demand signals, and anomaly detection across sprawling logistics networks. Researchers have reported gains across every link of the chain, from factory floor to last-mile delivery. But the authors of the new paper argue that the field’s obsession with a particular family of performance statistics has created a blind spot at precisely the moment executives need clarity most.
That blind spot involves the metrics everyone cites. Accuracy, precision, sensitivity, recall, and the F1 score are the standard currencies of machine learning evaluation. Accuracy measures the share of predictions a model gets right; precision captures how many of its positive predictions were actually correct; sensitivity—also called recall—reveals how many true cases the model managed to catch; and the F1 score blends precision and recall into a single harmonic mean. These numbers are excellent for comparing two algorithms against the same dataset. What they cannot do, AbdEllatif and Abdel Aal contend, is tell a procurement director, a logistics manager, or a shareholder whether deploying a given technique improved the organization in ways those stakeholders actually care about. A model can post a stellar F1 score while failing to reduce costs, accelerate deliveries, or ease workloads. The study therefore shifts the question from “how well does the algorithm perform?” to “how much benefit does the integration of machine learning or deep learning with supply chain management tasks deliver, according to the people who must live with the results?”
To capture those human judgments, the researchers reached for one of the most exotic toolkits in modern decision science: neutrosophic numbers. Fuzzy sets, introduced by Lotfi Zadeh in 1965, let an observation belong partially to a category—for instance, judging a forecasting model’s benefit as “0.7 good.” Krassimir Atanassov’s intuitionistic fuzzy sets added a second degree of freedom, allowing experts to state both how true and how false a claim feels, with the two summing to at most one. Neutrosophic logic, proposed by mathematician Florentin Smarandache in the mid-1990s, goes further still: truth, indeterminacy, and falsehood become three fully independent membership degrees, each ranging between zero and one. That third channel—indeterminacy—is the crucial one. A supply chain expert asked whether a deep learning system will deliver benefits might genuinely not know, and neutrosophic mathematics can encode that hesitancy instead of forcing artificial precision. The study employs single-valued trapezoidal neutrosophic numbers, which attach a four-point interval to each judgment, letting evaluators express ranges of belief rather than single brittle values.
The centerpiece of the framework is an aggregation operator the authors call the Single-Valued Trapezoidal Neutrosophic Number Weighted Arithmetic Average, or SVTNNWAA. In essence, the method gathers evaluations from multiple stakeholders, each expressed as a trapezoidal neutrosophic number carrying its own truth, indeterminacy, and falsity components, and fuses them into a single collective rating. Because the averaging is weighted, judgments tied to more important criteria exert greater influence on the final score. The mathematics preserves all three neutrosophic components throughout the computation, so the uncertainty voiced by experts is not laundered away in the process. Once the aggregated values are computed, techniques drawn from the fuzzy-numbers literature—including the graded mean integration representation, a standard defuzzification approach that converts an interval judgment into a crisp, comparable figure—translate the results into benefit rates that can be ranked and visualized. The output is a league table of machine learning and deep learning techniques ordered not by benchmark accuracy but by their expected organizational payoff.
Assigning those weights is where the second half of the framework comes in: the Full Consistency Method, or FUCOM. Classical weighting techniques such as the Analytic Hierarchy Process require experts to compare every criterion against every other one, producing n(n−1)/2 judgments for n criteria—a burden that balloons quickly and invites inconsistency. FUCOM, introduced in 2018 by operational researchers led by Dragan Pamučar, slashes the workload to roughly n−1 comparisons. Experts rank the criteria in order of priority and then compare the top-ranked criterion against each of the others, yielding a compact set of ratio judgments. The method then solves an optimization problem that minimizes the maximum deviation from perfect consistency, exploiting the transitivity of the comparisons to guarantee mathematically coherent weights. In the new study, FUCOM supplies the consistent weighting structure that the SVTNNWAA operator requires, ensuring that the aggregated stakeholder scores rest on a defensible foundation rather than on arbitrarily chosen priorities. Fewer comparisons also mean less fatigue for busy executives—a practical virtue in corporate settings where evaluation panels have limited patience for lengthy questionnaire exercises.
To demonstrate that the machinery works outside of theory, the researchers applied the method in a practical case study. The results show that it enables decision-makers to assess and visualize the benefit rates of machine learning and deep learning techniques, indicating which approach is most suitable for more effective supply chain management. Just as importantly, the evaluation incorporates stakeholders’ viewpoints directly: the benefit of integrating a given AI technique with supply chain tasks is scored through the eyes of the people responsible for the chain’s performance, rather than inferred from technical benchmarks alone. Because the entire calculation runs in a neutrosophic environment, the framework absorbs the uncertainty and indeterminacy that inevitably accompany judgments about emerging technology—situations where experts hold partial knowledge, evidence conflicts, or outright indecision reigns. The case study thus functions as a proof of concept that neutrosophic multi-criteria decision-making can move from academic journals into the meeting rooms where technology investments are actually argued over.
The broader significance lies in bridging two communities that often talk past each other. Data scientists publish benchmark results; operations managers ask what those results mean for costs, service levels, and resilience. By fusing FUCOM-derived weights with neutrosophic aggregation, the new framework gives both sides a shared language. It belongs to the growing field of multi-criteria decision-making in supply chain management, a discipline that has previously transformed how firms select suppliers and prioritize risks with tools such as the Analytic Hierarchy Process and TOPSIS. What distinguishes this contribution is its explicit targeting of AI integration decisions—an area where enthusiasm routinely outruns evidence. Frameworks like this one could help executives determine where a deep learning investment beats a simpler machine learning model, and where neither justifies the disruption. They also create an auditable record of why a technology was chosen, which matters as organizations face growing scrutiny over algorithmic procurement decisions.
The study, published open access in the Journal of Big Data, a Springer Nature title, appeared on 5 August 2026 after a peer-review journey that began with submission on 18 July 2025 and acceptance on 12 July 2026. Springer is releasing the paper early as a citable, peer-reviewed accepted version carrying a permanent DOI, ahead of the final Version of Record. The work was funded by the University of Jeddah under grant number UJ-23-DR-67, with the authors thanking the university for its technical and financial support. AbdEllatif is affiliated with the university’s College of Business, while Abdel Aal is based at the Faculty of Computers and Informatics at Zagazig University in Egypt. Published under a Creative Commons license that permits sharing with appropriate credit, the paper falls squarely within the journal’s research area of multi-criteria decision-making in supply chain management. Both authors declare no competing interests, and the study required no ethical approval.
As global supply chains strain under geopolitical shocks, climate disruption, and relentless consumer expectations, corporations are expected to keep escalating their AI spending—and every one of those dollars will eventually face a boardroom reckoning. Tools that convert human judgment into transparent, uncertainty-aware rankings may prove as consequential as the algorithms they evaluate. The authors’ approach is not limited to logistics; any domain where experts must weigh emerging technologies under uncertainty—from healthcare informatics to smart manufacturing—could, in principle, adopt the same neutrosophic machinery. For now, the study stands as a reminder that the hardest part of artificial intelligence is not teaching machines to learn. It is teaching organizations to know, with confidence, whether the machines are actually helping. With this framework, the answer arrives as a number—one that finally accounts for doubt.
Subject of Research: A neutrosophic multi-criteria decision-making framework combining the Single-Valued Trapezoidal Neutrosophic Number Weighted Arithmetic Average (SVTNNWAA) with the Full Consistency Method (FUCOM) to assess the organizational benefits of integrating machine learning and deep learning into supply chain management tasks from stakeholders’ viewpoints under uncertainty.
Subject of Research: Technology and Engineering
Article Title: Machine learning and deep learning techniques for effective supply chain management
Article References: AbdEllatif, M. M. A., & Aal, S. I. A. (2026). Machine learning and deep learning techniques for effective supply chain management. Journal of Big Data. https://doi.org/10.1186/s40537-026-01516-3
Image Credits: AI Generated
DOI: 10.1186/s40537-026-01516-3
Keywords: Supply chain management, Machine learning, Deep learning, Multi-criteria decision-making, Neutrosophic numbers, SVTNNWAA, Full Consistency Method (FUCOM), Decision-making under uncertainty, Stakeholder evaluation, Organizational benefits of AI
Cite Scienmag News
APA MLA Chicago
Blake Davidson. (August 30, 2026). Machine and deep learning reshape modern supply chain management. Scienmag. https://scienmag.com/machine-and-deep-learning-reshape-modern-supply-chain-management/
Blake Davidson. “Machine and deep learning reshape modern supply chain management.” Scienmag, 30 August 2026, https://scienmag.com/machine-and-deep-learning-reshape-modern-supply-chain-management/. Accessed 30 August 2026.
Blake Davidson. “Machine and deep learning reshape modern supply chain management.” Scienmag. August 30, 2026. https://scienmag.com/machine-and-deep-learning-reshape-modern-supply-chain-management/
Copy citation Download RIS
Tags: AI for port and warehouse operationsAI performance evaluation in commerceAI-driven freight routingAI-driven supply chain optimizationartificial intelligence in logisticsbig data in supply chain managementdata-driven supply chain optimizationdecision-making frameworks for supply chainsdeep learning for freight routingdeep learning network applicationsimpact measurement of AI in logisticsmachine learning in supply chainsmeasuring AI impact on businesspredictive analytics for inventorypredictive analytics in supply chainsensor data in logisticssupply chain data analyticsSupply Chain Managementtechnological transformation in global tradetechnology evaluation in supply chain performance


