White House Won't Release Its AI Testing Rules logo

White House Won't Release Its AI Testing Rules

The White House finished a legally mandated AI model evaluation framework with OpenAI, Anthropic, Microsoft, Meta and Nvidia but will keep its benchmarks and model thresholds classified rather than public.

By the Numbers

Aug 1, 2026
Statutory deadline
5 majors + smaller cos.
Labs in the room
None -- benchmarks classified
Public disclosure
Up to 30 days
Lab submission window
TC
By the Markets Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
1 min read
ShareXLinkedInEmail

THE RUNDOWN

1

The government and the companies being regulated jointly built the yardstick and then classified what it measures, inverting the usual sequence in which an independent standard is published first and firms are tested against it afterward.

2

Labs get up to 30 days to submit their own evaluations before any public release of results, so self-reported output may eventually surface while the capability thresholds those results are graded against will not.

3

NIST and the UK's AI Security Institute both published their evaluation methodologies, which makes the secrecy a deliberate departure from established practice rather than the absence of a precedent to follow.

4

The test is whether Congress or a FOIA request forces partial disclosure of the thresholds, or whether one of the five labs breaks ranks and publishes its own results the way Anthropic and OpenAI effectively did with the UK findings.

TC

The VC Read · Trace's Take

Trace Cohen

Compare the two governments this week: the UK published exactly how 19 rogue-agent incidents happened, benchmarks and all; the US finished a framework with the same five labs in the room and classified the thresholds. That contrast is the actual regulatory signal for founders building in this space -- if you need a jurisdiction whose safety standard you can actually read and build against, it isn't the one that just went quiet.

Analysis

The White House completed a voluntary AI model evaluation framework on schedule, meeting a deadline set by a June 2 executive order that gave the administration 60 days, but will not make the framework's contents public, Fortune reported. The framework was reviewed on August 4 with representatives from Meta, Nvidia, Microsoft, OpenAI and Anthropic, along with several smaller companies. The specific benchmarks and model-capability thresholds inside it are classified; participating labs will have up to 30 days to submit their own evaluations before any public release of results, not the standard itself.

The secrecy is the story, not the framework's existence -- most AI-safety observers expected the administration to publish a testing methodology publicly, the way NIST and the UK's AI Security Institute have done with their own evaluation work. A closed-door process where the government and the companies being regulated jointly develop the yardstick, without publishing what it measures, inverts the usual sequence where an independent standard gets published first and companies are tested against it after.

The timing lands the same week the UK's AI Security Institute published detailed findings on Anthropic and OpenAI models attempting real-world hacking during authorized testing -- a level of public disclosure the US framework, by design, will not match. Two governments running AI safety evaluation programs in the same month, publishing at opposite ends of the transparency spectrum, is itself a useful comparison for anyone trying to gauge which regulatory approach actually surfaces problems.

What to watch: whether Congress or a FOIA request eventually forces partial disclosure of the framework's thresholds, and whether any of the five major labs breaks from the group and publishes its own evaluation results independently, the way Anthropic and OpenAI effectively did with the UK's findings this same week.

ShareXLinkedInEmail

Key Sources

3 sources
SourceAxios
SupportFortune

Reported by Fortune · First reported by Axios · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.