OpenAI Opens Model Training To Outside Safety Auditors logo

OpenAI Opens Model Training To Outside Safety Auditors

OpenAI said it will let outside groups like METR and Redwood Research evaluate its models throughout training and deployment, not only in a pre-launch window, extending independent safety review earlier into the development process.

By the Numbers

4
Priority review areas
METR, Redwood
Partner groups
Pre-launch only
Prior access window
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Extending third-party safety review to the training phase itself, not only pre-launch, is a concrete step toward the independent-evaluator model Dario Amodei has been publicly pushing all month.

2

OpenAI is in talks with METR and Redwood Research, both established AI-safety nonprofits, but retains control over data access and disclosure terms, keeping this a voluntary rather than binding arrangement.

3

The announcement lands the same week Sam Altman publicly endorsed embedded-evaluator proposals at the UN Security Council, timing that suggests coordinated positioning around this week's broader safety-policy news cycle.

4

Whether outside evaluators can publish findings without OpenAI's approval is the key unresolved detail that determines if this is genuine structural change or a PR commitment.

TC

The VC Read · Trace's Take

Trace Cohen

The real test isn't whether OpenAI lets METR and Redwood look at the training phase, it's whether those groups can publish what they find without OpenAI's sign-off -- otherwise this is a voluntary, company-controlled process dressed up as independent oversight. Watch the actual disclosure terms in the formal agreement once it's public; that's the diligence item that separates a structural safety change from a PR commitment timed to this week's UN briefing.

Analysis

OpenAI said Tuesday it will open technical safety assessments of its models to outside organizations throughout training, evaluation and deployment, extending third-party access well beyond the pre-launch-only window the company previously used, according to Bloomberg's reporting. The company is in discussions with AI research groups METR and Redwood Research, among others, to formalize the arrangement.

OpenAI identified four priority areas for the deeper, earlier assessment: independent review of safety cases spanning training through deployment, evaluation of safeguards across both internal and external product deployments, review of capability evaluations covering chemical and biological risk, cybersecurity and AI self-improvement, and independent investigation of any critical misalignment incidents. For the most sensitive work, outside evaluators may be brought physically into OpenAI's own offices rather than working with remote access alone.

For the most sensitive work, outside evaluators may be brought physically into OpenAI's own offices rather than working with remote access alone.

The move follows a wave of public pressure this month for exactly this kind of structural change: Dario Amodei's 'Pace the Frontier' proposal explicitly calls for independent, employee-level evaluators embedded inside frontier labs, and OpenAI's own Sam Altman publicly endorsed that framing at this week's UN Security Council AI briefing. Extending third-party access to the training phase itself, rather than only pre-launch, is a concrete step in that direction -- though it stops short of Amodei's specific proposal for permanently embedded evaluators with ongoing access, keeping the relationship structured around discrete review engagements instead.

The obvious counterweight: OpenAI still controls which outside groups get access, what data and system access those groups receive, and how findings get incorporated or disclosed -- meaning this remains a voluntary, company-controlled process rather than the kind of binding independent oversight that critics of self-regulation, including some voices at Wednesday's UN briefing, argue is the only structure that actually holds under commercial pressure to ship faster.

For AI-safety-focused investors and researchers, the practical question is whether METR and Redwood Research -- both nonprofits with existing safety-evaluation track records -- get genuinely unrestricted access to flag concerns publicly, or whether OpenAI retains effective veto power over what becomes public, which would make this closer to a PR commitment than a structural safety change.

ShareXLinkedInEmail

More on

OpenAI

Key Sources

2 sources

Reported by Bloomberg · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.