Artificial intelligence has moved from laboratory curiosity to critical infrastructure at astonishing speed. It approves loans, screens medical images, steers autonomous vehicles, and moderates the information billions of people consume. Yet the frameworks meant to guarantee these systems are safe, fair, and accountable have not kept pace with the technology they are supposed to govern. A new peer-reviewed review published in the journal Artificial Intelligence Review by Anita Khadka and Carsten Maple of WMG at the University of Warwick delivers a sobering diagnosis: AI assurance, despite a flurry of international activity, remains fragmented, under-specified, and poorly matched to the messy realities of deployed systems. The paper, published open access on 29 September 2026, argues that the field needs a fundamental rethink of how high-level governance ambitions are translated into auditable, maintainable engineering practice.
The core problem, the authors explain, lies in the unusual properties of modern machine learning systems. Traditional software assurance grew up around deterministic programs: given the same inputs, the same outputs follow, and engineers can trace behavior through explicit lines of code. AI systems break these assumptions. They are non-deterministic, meaning identical inputs may yield different outputs across runs or model versions. They are deeply data-dependent, so their behavior shifts with the distributions of the data they are trained and operated on. And they carry evolving vulnerabilities, from adversarial attacks to silent model degradation, that static, one-off certification approaches were never designed to catch. Assurance methods borrowed from safety-critical engineering, such as the certification regimes used in aviation, assume a stability that learning systems simply do not possess.
Khadka and Maple frame their review around a crucial conceptual move. Rather than treating safety, fairness, robustness, privacy, explainability, and accountability as separate silos, each with its own framework, checklist, and community, they argue these should be understood as assurance goals that only become meaningful when operationalised through five cross-cutting commitments. The first is testable claims: any assurance statement about an AI system must be phrased in a way that can, in principle, be verified or falsified. The second is evidence obligations: whoever makes the claim must be able to produce the artifacts, test results, documentation, and monitoring data that support it. The third is lifecycle validity, the recognition that evidence gathered before deployment decays as models, data, and operating contexts change. The fourth is accountability and recourse, ensuring that when systems fail, affected parties can identify responsible actors and obtain redress. The fifth is interoperability, so that evidence and assurance artifacts can move across jurisdictions, sectors, and toolchains without being rebuilt from scratch.
Using these five commitments as an analytical lens, the researchers examined the leading AI assurance initiatives now shaping global practice. Chief among them is the NIST AI Risk Management Framework, the United States benchmark that organises AI risk work into govern, map, measure, and manage functions. They also assessed international guidance including the toolkits developed under the Organisation for Economic Co-operation and Development and the G7, along with practitioner-facing resources and representative research-led approaches. The verdict is nuanced. These initiatives have been valuable in establishing shared vocabulary and elevating AI risk to boardroom and ministerial attention, but the review finds that they frequently stop short of specifying the deployment-facing practices that credible assurance in real-world settings demands. High-level principles abound; concrete, auditable procedures for a hospital, a bank, or a transport operator remain scarce.
This gap between regulatory intent and operational reality is the review’s central finding. A regulator may mandate that an AI system be fair, but fairness itself fractures into multiple incompatible mathematical definitions, demographic parity, equalised odds, calibration across groups, and choosing among them is a context-sensitive judgment, not a technical formality. Similarly, requirements for explainability or robustness often lack the specified metrics, evidence formats, and thresholds that would allow an auditor to determine compliance. The result is a landscape in which organisations can declare alignment with principles while the underlying assurance work remains shallow, inconsistent, or unverifiable. The authors warn that this under-specification is not a cosmetic problem; it undermines the very credibility that assurance regimes exist to provide.
To make these abstractions concrete, the paper develops illustrative assurance cases across three high-impact domains. One sketches how explainability might be assured for an AI-enabled skin cancer detection system, tracing the chain from clinical claims about model outputs to the evidence a clinician and a regulator would need. Another addresses fairness in an AI-driven financial decision-making system, where lending outcomes must be demonstrably free of unjustified bias across protected groups. A third considers safety for an AI-enabled transportation system, where the consequences of failure are measured in lives. These worked examples show how the five commitments interact in practice: a fairness claim is only as strong as the evidence behind it, that evidence is only valid across the system’s lifecycle if monitoring continues after deployment, and recourse mechanisms must exist for individuals harmed by erroneous decisions.
The review also surveys the growing ecosystem of toolkits, platforms, benchmarks, and evaluation suites that practitioners increasingly treat as assurance evidence. Catalogued in the paper’s appendices, these resources range from model documentation templates to adversarial robustness benchmarks. The authors’ assessment is cautious: such artifacts are useful building blocks, but benchmarks alone do not constitute assurance. A model that scores well on a static test set may still fail catastrophically under distribution shift, and a completed documentation template may describe a system that has since been retrained. Evidence must be tied to specific, testable claims and kept current, or it becomes assurance theater, paperwork that simulates accountability without delivering it.
What emerges from the analysis is a call for assurance strategies that are coherent, context-sensitive, and adaptive. Coherent means that safety, fairness, privacy, and the other goals are pursued within a single integrated structure rather than as competing compliance exercises. Context-sensitive means that the appropriate depth and rigor of assurance scales with the risk profile of the deployment: a recommendation engine and an autonomous vehicle should not face identical burdens. Adaptive means that assurance is treated as a continuous lifecycle activity, with monitoring, re-evaluation, and evidence refreshment built into operations rather than bolted on at certification time. This vision aligns with, but pushes beyond, the direction of travel in emerging regulation, including the European Union’s risk-based approach to AI systems.
The practical implications reach every stakeholder in the AI ecosystem. For policymakers, the message is that principles without operationalisation create a compliance culture that rewards paperwork over performance; standards bodies must invest in specifying measurable requirements and acceptable evidence formats. For developers, the five commitments offer a design discipline: build systems so that claims about them can be tested, evidence can be generated automatically, and lifecycle changes are tracked. For regulators and auditors, the review highlights the need for interoperable assurance artifacts so that evidence produced in one jurisdiction or sector can be evaluated elsewhere, reducing duplication and closing loopholes. The authors also point toward the insurance industry as an underappreciated actor, noting that research programs on assurance and insurance for AI suggest financial markets may soon demand the kind of verifiable evidence that assurance frameworks provide.
The Warwick study arrives at a decisive moment. Governments across the world are standing up AI safety institutes, regulators are drafting enforcement rules, and enterprises are racing to demonstrate trustworthiness to customers and courts. Yet as Khadka and Maple demonstrate, the connective tissue between ambition and verification is still missing. Their contribution is less a new framework than a rigorous audit of the frameworks we already have, measured against a simple standard: can a claim about an AI system be tested, evidenced, maintained across its lifecycle, contested through recourse, and trusted across boundaries? By that measure, most current initiatives fall short. The path forward they chart, treating assurance as a living, integrated engineering practice rather than a static compliance checkbox, offers the clearest route yet to AI systems that earn, rather than merely assert, the public’s trust.
Subject of Research: AI assurance frameworks, their gaps, and the operationalisation of trustworthy AI governance
Article Title: Artificial Intelligence (AI) assurance: challenges, gaps, and the path forward
Article References: Khadka, A., & Maple, C. (2026). Artificial Intelligence (AI) assurance: challenges, gaps, and the path forward. Artificial Intelligence Review. https://doi.org/10.1007/s10462-026-11719-y
Image Credits: AI Generated
DOI: 10.1007/s10462-026-11719-y
Keywords: AI assurance, trustworthy AI, NIST AI RMF, machine learning, AI regulation, assurance cases, AI safety, fairness, explainability, accountability, risk management, AI governance
Cite Scienmag News
APA MLA Chicago
Denise Maddox. (October 1, 2026). Why AI Assurance Is Failing in the Real World and How to Fix It. Scienmag. https://scienmag.com/why-ai-assurance-is-failing-in-the-real-world-and-how-to-fix-it/
Denise Maddox. “Why AI Assurance Is Failing in the Real World and How to Fix It.” Scienmag, 1 October 2026, https://scienmag.com/why-ai-assurance-is-failing-in-the-real-world-and-how-to-fix-it/. Accessed 1 October 2026.
Denise Maddox. “Why AI Assurance Is Failing in the Real World and How to Fix It.” Scienmag. October 1, 2026. https://scienmag.com/why-ai-assurance-is-failing-in-the-real-world-and-how-to-fix-it/
Copy citation Download RIS
Tags: accountabilityAI assuranceAI assurance failureAI deployment risksAI governanceAI regulationAI regulation challengesAI safetyAI safety frameworksAI system auditing practicesAI system transparencyassurance casesautonomous vehicle safetyethical AI governanceExplainabilityfairnessfairness and accountability in AIgoverning AI in real-world applicationsinternational AI assurance policiesMachine learningmachine learning safety standardsNIST AI RMFrisk managementtrustworthy AI

