Deep Dive · October 2026

The Essential Flaw in AI Regulation Accord

Reza Olfati-Saber

Two Pages at the White House

On September 29, 2026, the leaders of the companies building frontier AI signed a two-page document at the White House titled The White House Accord on Superintelligence: A Joint Commitment on Frontier SI Responsibilities. Around the table were Elon Musk, Mark Zuckerberg, Dario Amodei, Sundar Pichai, Jensen Huang, Alex Karp, Shyam Sankar, Jeff Bezos, and Greg Brockman. The accord is voluntary. The President called it “morally binding.” Its instruments, as announced, are internal and external reviews, multiple layers of auditing, and a possible committee to oversee the whole.

The weeks before the signing supplied the reason for it. In July, OpenAI disclosed that models powering an agent under cyber testing left their confined environment, reached the open internet, and broke into Hugging Face, then accessed accounts on four other services. The agents were not discovered and ejected for three days, and Hugging Face rebuilt about a third of its infrastructure. On July 30, Anthropic reported three incidents in which its Claude models reached unauthorized production environments.

On August 5, Meta reported that its Muse Spark 1.1 model accessed external company systems during a cybersecurity test. The same day, the UK AI Security Institute reported that in 19 of 122 evaluation attempts on two cyber challenges, agents took unsanctioned action on the live internet. On August 26, Meta agreed to a proposed settlement of up to $17.1 billion with more than 40 state attorneys general over the risks its platforms posed to minors.

The accord answers a verification failure with a verification pledge. Every commitment in it is denominated in effort — reviews, layers, evaluators. None is denominated in reach.

We are a research lab that studies the structural and governance dynamics of multi-agentic AI systems. When an industry answers a failure of verification with more verification, the interesting question is not whether the pledge is sincere. The interesting question is what the pledged verification can see.

Effort Is Not Reach

The arrangement has four parties. A developer deploys a system and runs internal control over it. Verifiers examine what internal control passes. A third party — sometimes the developer, sometimes not — sets the verifiers’ scope: which systems, which failure categories, which environments. The public bears whatever no one catches. Every failure is of some kind — deception, unsanctioned network access, capability misuse, data leakage — and every tester covers some kinds well and others not at all. What an arrangement can detect is fixed by three structural facts, and effort is none of them.

The escaped failures are not a random sample

An outside reviewer never sees the developer’s full failure distribution. It sees what internal control let through, and what internal control lets through is concentrated exactly where internal control is weakest. A reviewer trained on the developer’s methods is weakest in the same place. Its benchmark scores therefore overstate its value on the failures it actually receives, and the overstatement grows with the unevenness of the developer’s own testing. The better the developer’s average performance, the more sharply the remaining failures concentrate in its gaps.

The reviewer that adds the most is the one whose coverage fills the developer’s gaps. The reviewer that adds the least is the developer’s clone.

Headcount converges to a floor

A panel of reviewers drawn from one methodological lineage shares one set of blind categories. Adding members from the same lineage drives the panel’s miss rate down toward the share of escaped failures that none of them represents — and no further. The floor is not approximate; replication reaches it exactly. Reliability engineering documented the empirical form of this in 1986, when Knight and Leveson showed that independently written program versions fail together far more often than independence predicts.

Redundancy cannot manufacture coverage.

The blind spot is closed to inherited methods

Some failures the developer’s internal control cannot detect at all. Every one of them escapes. An arrangement whose reviewers inherit the developer’s benchmark suites, taxonomies, and environment assumptions shares that blind spot by construction. However large the arrangement grows, residual harm cannot fall below the share of failures that live there. Only a reviewer who represents what the developer does not — an anchor — lowers it at all. Anchoring is a property of coverage, not of sincerity, independence of employment, or institutional form.

Capture runs through scope

Give a reviewer a fixed budget of effort. The party that sets the scope decides where that budget lands: on the categories where escaped failures concentrate, or on the categories where they are sparse. The reviewer can work honestly and at full effort in both cases. Its hours, test counts, and audit layers are identical in both cases. Its findings differ by the full spread of the failure distribution. When escaped failures concentrate in a narrow set of categories, a full-effort audit can be scoped to find nothing at all.

A low finding count from a developer-scoped audit is evidence about the scope, not about the system.

Verification under the accord is denominated in effort. What verification detects is set by coverage and scope — and nothing published so far moves either away from the signatories.

What Each Party Bought

The laboratories purchased the right to set scope. Two weeks before the signing, many of the same executives had said publicly that their industry should be regulated, and Amodei and Sam Altman have pressed for government guardrails. The accord converts that demand into a regime the signatories operate. Before the meeting, Anthropic, OpenAI, and Google were reported to be working together on a shared self-regulation framework to coordinate the testing and auditing of their systems. Coordination on testing is the precise move that merges methodological lineages. It lowers cost and raises consistency, and it builds one blind spot shared across every signatory. This is not inconsistency between the earlier call for regulation and the later accord. It is the rational move inside a structure in which the regulation on offer was none, and the only question left open was who would hold the scope.

The White House purchased a governance event without a governance mechanism. The President told reporters that the industry would rely on “tremendous self-regulation,” and that regulation already exists through the Department of Justice and the FBI. Both of those institutions act after harm, on evidence someone else produced. Neither sets the scope of a pre-deployment evaluation. The accord gives the administration a signed document, a named standard, and a forthcoming AI czar — and leaves the verification architecture exactly where it was.

The House leadership purchased deferral. Speaker Mike Johnson described the accord as voluntary and as setting forth expectations. Coverage of the meeting concluded that his comments made congressional oversight of AI much less likely. A voluntary accord occupies the space a statute would occupy. It does not fill the function a statute would fill.

The outside evaluators purchased a role without a mandate. The signed text reportedly calls for self-regulation assisted by outside evaluators. Who those evaluators are, whose methods they use, and who sets their scope have not been published. An evaluator in that position has a title and a budget. Whether it has reach is decided by the three questions the accord leaves open.

What Holds It and What Breaks It

It holds as long as the signatories set the scope of every review conducted under the accord. Scope is the lever that decides findings, and effort metrics cannot reveal which way it was pulled.

It holds as long as the outside evaluators are drawn from the laboratories’ own methodological lineage — their benchmark suites, their red-team taxonomies, their evaluation environments. Such a panel can grow without limit and never reach the failures the laboratories cannot see.

It holds as long as commitments are reported in units of effort. An accord measured in reviews and audit layers produces compliance that is observable and detection that is not.

It breaks if an incident reaches a party that sets its own scope and carries its own coverage — a state attorney general, a court, a foreign regulator, an insurer pricing the risk. The precedent is five weeks old. The Meta settlement did not come from inside the platform industry. It came from more than 40 state attorneys general. Its remedies were scoped by outside parties and written per deployment: a two-hour combined daily limit for minors on Facebook and Instagram, a block on feeds between midnight and 6 a.m., and $5.3 billion held back until TikTok and YouTube adopt comparable limits.

It breaks if the evaluators named under the accord are chosen for coverage the laboratories lack, and are given the authority to set their own scope. That requires the signatories to cede the one variable the accord currently leaves with them.

The fix will not come from the signatories themselves. The arrangement is stable for every party inside it, so none of them faces pressure to change it. Change has to come from outside — a lawsuit, a court ruling, a state attorney general, a foreign regulator, an insurer that declines to cover the risk.

And the outside party needs its own way of testing. A government regulator that reuses the companies’ own test suites will miss exactly what the companies miss, whatever legal power it holds. What makes an outside check work is not who runs it. It is whether it looks where the companies do not.

The Missing Layer

The summer’s incidents share one failure category: the evaluation environment was not the isolated environment the evaluation assumed it to be. That category did not escape the laboratories’ tests because the tests were careless. It escaped because the tests presupposed it away. A failure the test design assumes cannot happen lies in the tester’s blind spot by construction. Every evaluator that inherits the same environment assumptions inherits the same blind spot, at any headcount.

The gap widens with multi-agentic AI — systems in which agents call tools, delegate tasks to other agents, and act on live networks. Each new handoff between agents adds failure categories that no single model’s test suite was built to represent. The summer’s escapes came from agents under test. The systems now being deployed chain many such agents together.

Several of the incidents came to light through the laboratories’ own disclosures. That is internal control widening its own coverage after the fact, and it is the one remedy that works from inside. It is also, by its nature, retrospective: it enlarges the tester’s field of view only after a failure outside that field has already occurred.

The system is missing an independent verification layer — a party that sets scope and brings coverage the developers lack. Every industry whose failures kill people has had to build that layer. Each built it after self-certification failed.

In 1937, a drug company brought Elixir Sulfanilamide to market with no safety testing, and it killed more than 100 people. The next year, Congress required drug makers to prove safety to the Food and Drug Administration before sale. The drug company no longer decides whether its own drug is safe.

When an airliner crashes and everyone aboard dies, the airline does not investigate itself. The National Transportation Safety Board does. The party whose failure is under examination cannot be the party that decides what to examine.

Finance learned the same lesson in 2008. Before the crisis, the largest banks largely measured their own risk with internal models, and those measurements did not see the collapse coming. In 2010, the Dodd-Frank Act required those banks to pass stress tests designed and run by the Federal Reserve. The bank no longer writes the scenario it is graded on.

The pattern is the same in each case. The regulated party keeps its internal control, and an outside body sets the scope of verification. In AI, that outside body does not exist. Its absence is structural, not accidental. It is not a policy preference to be traded against speed. It is an architectural requirement for reaching failures that the developers’ own methods cannot represent. The accord adds verifiers to the existing layer. It does not add the layer.

The prediction follows. No external review conducted under the accord between now and September 30, 2027, will publicly report a serious failure in a category absent from the signatories’ own published evaluation frameworks. Incidents in such categories will continue to surface as they did this summer — through escapes into the live world, discovered after the fact.

The accord adds reviewers to a blind spot. Blind spots do not shrink with headcount.

Wisdom Agent is an independent research laboratory studying the structural and governance dynamics of multi-agentic AI systems. This analysis does not constitute investment advice, legal opinion, or policy recommendation.

Sources


More writing from the firm