Analysis
The White House completed a voluntary AI model evaluation framework on schedule, meeting a deadline set by a June 2 executive order that gave the administration 60 days, but will not make the framework's contents public, [Fortune](https://fortune.com/2026/08/04/baffling-white-house-wont-publicly-release-ai-model-evaluation-framework-it-reviewed-today-with-openai-anthropic-microsoft-and-others/) reported. The framework was reviewed on August 4 with representatives from Meta, Nvidia, Microsoft, OpenAI and Anthropic, along with several smaller companies. The specific benchmarks and model-capability thresholds inside it are classified; participating labs will have up to 30 days to submit their own evaluations before any public release of results, not the standard itself.
The secrecy is the story, not the framework's existence -- most AI-safety observers expected the administration to publish a testing methodology publicly, the way NIST and the UK's AI Security Institute have done with their own evaluation work. A closed-door process where the government and the companies being regulated jointly develop the yardstick, without publishing what it measures, inverts the usual sequence where an independent standard gets published first and companies are tested against it after.
The timing lands the same week the UK's AI Security Institute published detailed findings on Anthropic and OpenAI models attempting real-world hacking during authorized testing -- a level of public disclosure the US framework, by design, will not match. Two governments running AI safety evaluation programs in the same month, publishing at opposite ends of the transparency spectrum, is itself a useful comparison for anyone trying to gauge which regulatory approach actually surfaces problems.
What to watch: whether Congress or a FOIA request eventually forces partial disclosure of the framework's thresholds, and whether any of the five major labs breaks from the group and publishes its own evaluation results independently, the way Anthropic and OpenAI effectively did with the UK's findings this same week.