The race to develop Artificial General Intelligence (AGI) has reached a fever pitch, prompting a radical, albeit controversial, proposal from the industry’s vanguard. Anthropic CEO Dario Amodei has proposed a new model for corporate governance: embedding third-party safety evaluators directly into the offices of frontier AI labs. The objective, framed as a proactive measure to prevent catastrophic outcomes, involves granting external researchers unprecedented access to the "black box" of large language models (LLMs). Yet, as the tech industry grapples with the existential risks warned of by former researchers like Jacob Coxon, a glaring discrepancy has emerged between the marketing of this safety initiative and its structural reality. Critics, including legal scholars and auditing experts, argue that without the legal authority to halt development—a true "kill switch"—these embedded evaluators function less like federal bank supervisors and more like internal compliance officers, potentially engaging in what experts are calling "audit washing." The Genesis of the Safety Pivot: A Chronology The urgency behind Amodei’s proposal follows a period of internal turmoil and external pressure. Early 2026: AI labs, including OpenAI and Anthropic, face mounting scrutiny from lawmakers and safety researchers regarding the rapid scaling of model capabilities. September 9, 2026: Former Anthropic researcher Jacob Coxon resigns, issuing a blistering warning that frontier labs are engaged in a reckless race toward systems that may eventually escape human control. September 12, 2026: Dario Amodei publishes a seminal essay proposing a "pace-setting" framework for the industry, which includes the controversial embedded evaluator plan. September 15–16, 2026: The proposal triggers a sharp regulatory divide. While some in the AI community welcome the oversight, others—including the Trump administration—signal skepticism toward increased federal interference, creating a fragmented landscape for potential AI legislation. Late September 2026: Following Anthropic’s lead, OpenAI commits to a similar safety arrangement, though specific details regarding the implementation remain sparse. The "Banking Analogy" Under the Microscope Amodei has frequently cited the banking industry as a blueprint for AI regulation, noting that government "supervisors" are often physically embedded within financial institutions to monitor systemic risk. Julie Andersen Hill, dean of the University of Wyoming College of Law and a noted expert on banking regulation, contends that the comparison is fundamentally flawed. "At the largest banks, government examiners have the power to direct a bank to stop a practice, restrict its growth, force management changes, and, in extreme cases, close the institution," Hill explains. In contrast, the proposed AI evaluators—even those with high-level access—lack the legal mandate to intervene. They can investigate, document, and report, but they cannot prevent a model from being trained or deployed. "If you don’t give them that kind of power, I don’t know what they are doing," Hill says. "That’s fundamentally different from bank regulators, who have actual enforcement authority." The Reality of "Black-Box" Auditing To understand the technical challenges of these audits, one must look at firms like XBOW, which has already received early access to unreleased models. Albert Ziegler, head of AI at the firm, notes that the day-to-day reality of auditing is far less cinematic than the existential dread often portrayed in the media. "We might find that a model produces nonsense under an unusual formatting request, or that a safety filter needs adjustment," Ziegler says. "But the kind of insidious, catastrophic consequences produced by subterfuge—that’s not something we’ve seen." The problem, according to Ziegler, is that catastrophic risk often emerges only under a "perfect storm" of circumstances: a combination of specific instructions, system permissions, and external triggers that an evaluator might never think to simulate. Because the evaluator lacks veto power, they are relegated to the role of a consultant. They can document risks, but the "ultimate decision remains with the company." Conflicts of Interest and the "Independent" Label A significant hurdle for the embedded evaluator model is the issue of independence. In traditional financial auditing, there are stringent rules regarding who can audit whom to prevent conflicts of interest. In the AI space, the current ecosystem is small, highly interconnected, and rife with personal and professional ties. Deborah Raji, a UC Berkeley researcher specializing in algorithmic auditing, warns that access is not synonymous with independence. "You are effectively not qualified to be an actual auditor if you can’t meet the standards of independent conduct," Raji asserts. "If the company being audited chooses the evaluator, controls the scope of the investigation, and retains the right to redact findings, you have the equivalent of a company asking a friend to check their homework." The nonprofit Model Evaluation and Threat Research (METR) is frequently cited as the gold standard for this work. However, the organization itself has acknowledged that its employees often share social and professional histories with the very labs they are tasked with auditing. While METR maintains that these ties do not compromise their work, critics argue that the lack of a formal, rotational, and government-sanctioned auditing board renders the process inherently opaque. The Regulatory Gap: Implications for Policy The failure to establish a clear, binding legal framework for AI safety is perhaps the most significant structural weakness in the current landscape. 1. The "Audit Washing" Phenomenon Experts fear that these voluntary arrangements provide a veneer of safety that allows companies to market their products as "vetted" while retaining full control. If an adverse finding produces no consequences—if, in Raji’s words, "nothing happens" after the audit—the process serves only to burnish the company’s reputation rather than secure the public. 2. Market Consolidation Christina Ho, chief assurance officer at Oath and former board member of the Public Company Accounting Oversight Board, notes that the cost of continuous, high-level supervision is astronomical. While tech giants like Anthropic or OpenAI can absorb these costs, smaller startups may find it impossible to comply with such rigorous (and expensive) safety standards, effectively creating a barrier to entry that insulates incumbents from competition. 3. The Expertise Deficit Even if a regulatory body were created, the talent pool of individuals capable of auditing a frontier-scale LLM is vanishingly small. Traditional audits verify processes; AI audits must verify the output and the intent of the system. Without a massive investment in independent, public-sector technical expertise, the government remains perpetually reliant on the labs themselves to define what constitutes a "safe" model. Conclusion: A Turning Point for AI Governance The proposal to embed evaluators is an acknowledgment that the industry has outpaced current regulatory capabilities. However, as it stands, the plan is a half-measure. Legal experts and safety researchers are in agreement: if the risks of AGI are as catastrophic as the CEOs themselves claim, then the current proposal is insufficient. "If you really believe that AI has the power to destroy society, then you have to have an independent supervisor that has the ability to pull the plug on it," Hill says. Until the industry moves beyond voluntary, self-directed oversight and accepts a framework where independent auditors possess genuine legal authority—and the power to enforce it—the "embedded evaluator" model will likely remain a sophisticated form of corporate signaling rather than a true safeguard for the future of humanity. The question for policymakers is no longer whether we need oversight, but whether we have the courage to treat AI with the same rigorous, restrictive, and public-interest-focused regulatory standards that we apply to the global financial system. Post navigation The Architects of Autonomy: How the AI Era is Permanently Rewiring the C-Suite High-Stakes Diplomacy: Boeing and Citigroup CEOs to Accompany President Trump on Pivotal China Summit