After adopting AI, enterprises need the capability to absorb what AI brings.
In 2025, Deloitte delivered a report worth approximately AUD 440,000 to an Australian government department. Problems were later identified in the report: it contained non-existent academic references, mismatched footnotes, and fabricated citations to court judgments. The revised version disclosed that parts of the report had used Azure OpenAI’s GPT-4o. Following the controversy, Deloitte agreed to refund the final payment and resubmit a revised version of the report (AP News; The Guardian).
This was not a minor episode of “AI getting a few citations wrong.” A AUD 440,000 professional services engagement ultimately turned into public correction, rework, refund, and trust repair. The direct loss was not limited to the amount refunded. Secondary review, redelivery, client explanation, and damage to brand credibility can all consume the time and cost that AI was originally expected to save.
More importantly, this case is difficult to explain simply as a problem of an “outdated model” or “insufficient model capability.” GPT-4o was already a leading model at the time, and Deloitte is not an organization unfamiliar with professional review. The problem occurred at another level: AI had been inserted into the delivery workflow, but fact verification, human review, transparent disclosure, and responsibility boundaries had not kept pace.
This is exactly the part that many enterprises underestimate when introducing AI. Model capability determines whether AI can generate an output that looks plausible. But what enterprises actually pay for, bear risk for, and hand over to clients or employees is whether that output can be verified, traced, corrected, and integrated into business processes.
In other words, what determines the ROI of enterprise AI is not only how intelligent the model is, but whether what the enterprise embeds into its business chain is a genuinely verifiable, controllable, and accountable Responsible AI system. Responsible AI, or RAI as used below, is not about adding a polished layer of governance language to AI. It is about whether, when model output enters customer service, contracts, quotations, approvals, and internal workflows, the enterprise knows where it came from, what it can do, and who catches it when it goes wrong.
Stanford University’s AI Index Report 2026 places Responsible AI in a dedicated chapter. In the report, RAI is defined as the practices and governance mechanisms designed to ensure that AI systems are safe, fair, beneficial, and operating as intended.
Behind this definition is one of the report’s core judgments: RAI-related infrastructure is expanding, with evaluations, policies, regulation, and AI safety institutions all developing, but they are not keeping pace with the speed of AI deployment. This judgment is not an abstract industry intuition; it is an evidence-backed insight. The AI Incidents Database recorded 362 enterprise AI incidents in 2025. Before 2022, that number had consistently remained below 100. The enterprise survey conducted for the second consecutive year by AI Index and McKinsey & Company also provides a score for corporate Responsible AI maturity. In 2025, the global average was only 2.3 out of 4, meaning that most organizations are still integrating relevant rules and processes internally, rather than operating a mature RAI system.
RAI maturity can be understood as the maturity with which an enterprise embeds Responsible AI into day-to-day operations. A score of 1 usually means that the enterprise has only basic rules or initial awareness. A score of 2 means that relevant practices are being integrated into organizational processes. A score of 3 means that the necessary governance, risk control, and monitoring mechanisms are largely in place. A score of 4 means that the enterprise has formed a comprehensive, proactive, and sustainably operating RAI system. The higher the score, the more capable the enterprise is of managing AI output quality, data boundaries, access controls, accountability, and incident response within real business operations.
Enterprises are not unaware of the risks. In the report, 74% of enterprises identify inaccuracy as a relevant risk, 72% cite cybersecurity, and 63% cite regulatory compliance. But there is still a long distance between knowing that risks exist and being able to govern those risks. This is especially true when enterprises begin discussing AI agents that can autonomously call tools and execute actions, automated approvals, intelligent customer service, internal knowledge bases, and contract or quotation assistance. At that point, the question is no longer “Can the model answer?” but “Can its answer enter the business?”

The Deloitte incident is one slice of this gap: AI had already entered professional delivery, but the supporting mechanisms for verification, audit, and responsibility were not equally mature.
The report breaks RAI into three layers.
The first layer is Core Function and Behaviors. This layer focuses on what the AI system itself should achieve. The report emphasizes several dimensions: factuality and truthfulness, validity and reliability, privacy, data stewardship, fairness and bias, transparency and auditability, and explainability. In an enterprise context, this layer corresponds to output quality: whether the conclusions generated by AI can be verified, whether citations are real, whether data sources are clear, and whether the model turns uncertain content into seemingly certain facts.
Deloitte’s problem first occurred at this layer. AI generated citations that were complete in format and looked professional, but they included non-existent academic references and fabricated court judgment citations. The issue was not only model hallucination. It was also that the delivery process had not treated factuality, source verification, and auditability as hard gates.
The second layer is System Integrity and Risk Controls. This layer focuses on how risks are managed through technical and operational mechanisms. The report highlights dimensions including security, safety, and robustness. For enterprises, this layer determines whether AI can enter real workflows: what data AI can access, which systems it can call, which actions it can execute, which steps require human confirmation, and how the system rolls back after failure.
When an AI agent can search email, query inventory, modify orders, and send messages, an error is no longer just an incorrect answer. It becomes a real operation. Access control, logging, anomaly alerts, rollback mechanisms, and human takeover must be designed together with functionality.
The third layer is Governance, Accountability, and Enforcement. This layer focuses on responsibility, oversight, and remedy. The report emphasizes accountability and liability, as well as human oversight and contestability. Enterprises need to answer: Who approves AI launch? Who is responsible for daily monitoring? Who can suspend the system? Can employees override AI judgments? Do customers have complaint and human escalation channels? After an incident, can the enterprise trace the model version, prompt version, data source, and tool invocation records?
Together, these three layers point to the same question: Has AI moved from functional capability into operational capability? A model can generate output, but the enterprise bears the consequences once that output enters the business. Rework, misjudgment, complaints, compliance pressure, customer churn, and employee distrust all enter the same ROI calculation.
The premise of AI-driven efficiency is that the enterprise has the capability to manage the output, permissions, risks, and responsibilities that AI brings. Without that capability, AI merely accelerates part of the work while also allowing errors to enter the business faster.
RAI is often misunderstood as simply imposing stricter controls. The focus of page 170 of the AI Index report is precisely the opposite: there are tradeoffs among multiple RAI objectives. Safety, accuracy, privacy, fairness, explainability, speed, and cost often cannot all be maximized at the same time.
The report notes that differential privacy reduces the risk of identifying personal data by adding noise during training. In some experimental settings, it improved privacy protection but reduced explainability, fairness, and accuracy, with accuracy falling by as much as 33 percentage points.
Federated learning presents similar issues. It allows multiple institutions to train a model together without directly exchanging raw data, instead exchanging model updates. The report notes that in one Alzheimer’s disease detection scenario, stronger privacy protection reduced accuracy by 14.8 percentage points; for hospitals with smaller datasets, the false negative rate increased by 21.4%. Other encrypted alternatives can make fairness more stable, but computing costs may increase by two to three times.
The enterprise implication behind these technical examples is direct: stronger privacy may bring higher cost or lower effectiveness; more conservative safety strategies may affect user experience; higher automation can improve efficiency, but it can also increase the speed at which errors spread.
The report also notes that the industry currently lacks a unified framework for measuring and comparing these tradeoffs. This is especially important for enterprises. Standards can provide baselines, but when applied to customer service, contracts, sales, recruitment, risk control, and AI agents, the risk priorities differ by scenario.
Enterprises implementing RAI need a tradeoff mechanism: where speed should be prioritized, where accuracy should be prioritized, where automation rates must be sacrificed, and where human judgment must be retained. If these tradeoffs are not explicitly recorded, they become ad hoc decisions and accountability gaps after the project goes live.
The enterprise survey by AI Index and McKinsey also shows that organizations are adding policies, but remain far from operationalizing them. The share of companies with no RAI policies fell from 24% in 2024 to 11% in 2025. This shows that enterprises are indeed writing RAI into their formal rules.
New bottlenecks have also emerged. Fifty-nine percent of enterprises cite knowledge and training gaps, 48% cite resource or budget constraints, 41% cite regulatory uncertainty, and 38% cite technical limitations. For fully scaled agentic AI, the largest barrier becomes security and risk concerns, cited by 62%.

This set of data shows that enterprises do not lack a principle. They lack the concrete capability to translate principles into projects.
First, standardized metrics. Customer service, sales, contracts, knowledge bases, and internal process automation should all have basic monitoring metrics: factual error rate, hallucination rate, rework rate, human review ratio, the share of AI outputs overridden by employees, sensitive data trigger count, unauthorized access attempts, complaint and appeal rate, incident response time, and records of model versions, prompt versions, tool calls, and human edits.
The metrics do not need to be complex at the beginning, but they cannot be absent altogether. Without metrics, enterprises can only see how much time AI has saved; they cannot see how much hidden cost it has created.
Second, scenario-specific judgment. Marketing copy can allow room for stylistic experimentation, but it cannot make false claims, infringe rights, or mislead customers. Customer service responses should be fast, but they cannot exceed commitment boundaries. Quotations and contracts can use AI to assist with organization and drafting, but key clauses, amounts, versions, and sources still require human confirmation. Recruitment, risk control, credit, and similar scenarios also require consideration of fairness, explainability, and appeal mechanisms. When AI agents enter business systems, permissions, logs, rollback, and human takeover should receive high priority. Once business rules are written into automated systems, they will be executed consistently, quickly, and repeatedly. If enterprises have not checked rules, metrics, and appeal mechanisms in advance, the more efficient the system becomes, the greater the impact of its errors.
Third, accountability mechanisms. Once an AI system enters the business, the enterprise needs to know who approved launch, who reviews monitoring metrics, who handles incidents, who can suspend the system, and which issues must be escalated to humans. If accountability is unclear, AI projects easily remain functionally usable but fail to truly enter business operations.
RAI does not need to start as a heavy governance system. A more practical approach is to conduct a self-assessment under the report’s three-layer framework before an AI project enters real workflows.
Output layer:
System layer:
Accountability layer:
The value of this set of questions is that it turns RAI from a principle into a go-live condition. Enterprises do not need to first build a large governance department, but they must know which gates every AI system must pass before entering the business.
Models will continue to become stronger, and usage costs will continue to fall. The gap between enterprises will ultimately not only be reflected in who has connected to the newest model, but in who can place model capabilities into the business in a stable way.
This cannot rely only on post-launch remediation. RAI needs to enter the AI solution scope, system design, and post-launch operating mechanism from the beginning of the project.
At the technical layer, enterprises need to define how models are called, how permissions are allocated, how logs are recorded, how anomalies are alerted, how systems roll back, and which deployment approaches can meet data and security requirements. Multi-model management, secure invocation, access control, log auditing, private deployment, and failure fallback are not appendices to governance documentation. They are the foundation of whether enterprise AI systems can operate reliably.
At the business layer, enterprises need to define which type of ROI the AI project is meant to improve, where risks arise in the workflow, which metrics need to be monitored continuously, which actions must remain with humans, which errors are acceptable, and which errors must be intercepted in advance. If these judgments are not written into the solution scope, they will become ad hoc coordination after the project goes live.
This is where the value of AI consulting lies: placing the business problem, risk boundary, and system delivery on the same map. What enterprises truly need is not an AI demo that can be shown in a presentation, but an AI system that can be deployed, maintained, and governed.
Whether a company uses AI well should not be measured only by how new the model it has connected to is. The more important question is: when AI enters real business operations, can the enterprise steadily manage the efficiency, risk, and responsibility that it brings?
Cite as · Deep Reads · 21 July 2026
If you want both columns delivered together, four times a year, in one quiet email — leave an address. Otherwise just bookmark this page.