OpenAI, Anthropic Near Deal To Stress-Test Rival AI logo

OpenAI, Anthropic Near Deal To Stress-Test Rival AI

OpenAI is negotiating a legally binding agreement with Anthropic under which each company would run safety stress tests on the other's models, a formal follow-on to an informal cross-evaluation the two labs ran in August 2025.

By the Numbers

Aug 2025 cross-eval
Prior test
~70% (uncertain queries)
Claude refusal rate
Matched/outperformed
o3 vs Opus 4
Neared, not final
Deal status (Sept 21)
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

Unlike the August 2025 cross-evaluation, which was an informal one-off where each lab handed over access to its best model, this new agreement is reportedly being negotiated as legally binding -- a meaningfully higher commitment level between two companies that also compete directly for the same enterprise customers and the same funding pool.

2

The prior round of testing produced a real capability delta worth remembering: Anthropic's Claude models refused roughly 70% of uncertain queries versus a lower refusal rate from OpenAI's o3, while o3 matched or outperformed Claude Opus 4 on core alignment metrics -- meaning neither lab's model was simply 'safer' across the board.

3

The talks are unfolding alongside a broader, separate industry push -- Google DeepMind's Demis Hassabis, OpenAI's Chris Lehane and Sam Altman have been working since summer on a proposed FINRA-style self-regulatory body -- suggesting frontier labs increasingly prefer negotiated, industry-run safety mechanisms over waiting on binding federal rules.

4

As of September 21, The Information reports the deal is 'neared' but not finalized -- the gap between a negotiated framework and an actually-signed, binding commitment is exactly where past industry safety pledges have stalled before.

TC

The VC Read · Trace's Take

Trace Cohen

The tell here is 'legally binding' -- that's a materially bigger commitment than the informal August 2025 swap, and it's happening between two companies that also compete head-on for the same enterprise contracts. The diligence item for anyone with exposure to either lab: ask whether this produces a public joint report like last year's, or stays a private arrangement neither side has to answer for if it finds something ugly. A pact with no disclosed enforcement mechanism is a PR framework until proven otherwise.

Analysis

OpenAI is negotiating a legally binding agreement with Anthropic under which each company would run safety stress tests on the other's AI models, according to The Information. As of September 21, the two labs were described as having "neared" a deal, though it remained unclear whether terms were fully finalized.

A Formal Sequel To An Informal Test

This isn't the first time OpenAI and Anthropic have opened their models to each other. In August 2025, the two labs ran an informal cross-evaluation, handing each other access to their best models to run safety evaluations on a rival's technology. That earlier round produced a genuinely useful data point: Anthropic's Claude models refused roughly 70% of uncertain or potentially harmful queries, reflecting a more conservative design philosophy, while OpenAI's o3 model matched or outperformed Claude Opus 4 on core alignment benchmarks despite refusing less often. Neither company came out of that round with a clean "safer" label -- the two labs simply optimize differently, and the exercise made that difference visible rather than resolving it.

In August 2025, the two labs ran an informal cross-evaluation, handing each other access to their best models to run safety evaluations on a rival's technology.

Why Make It Binding Now

Moving from an informal, one-time swap to a reportedly legally binding arrangement is a meaningfully higher level of commitment between two companies that compete directly for the same enterprise customers, the same developer mindshare and, increasingly, the same pool of late-stage capital. A binding agreement implies recurring testing cycles and some mechanism for acting on what each side finds in the other's model, rather than a single published comparison that both companies can cite selectively in their own marketing.

Part Of A Wider Industry Push

The talks aren't happening in isolation. Pulse previously covered weeks of parallel safety coordination among OpenAI, Anthropic and Google DeepMind, including a proposal -- reportedly first floated by DeepMind's Demis Hassabis, developed further by OpenAI's Chris Lehane, and endorsed by Sam Altman -- for an industry-funded, FINRA-style self-regulatory body to test powerful AI systems before release. A bilateral stress-testing pact between the two most valuable AI labs reads as a smaller, faster-moving piece of that same broader effort: frontier labs increasingly prefer negotiated, industry-run safety mechanisms they help design, over waiting for binding federal rules that don't yet exist.

What The Cooperation Doesn't Resolve

What this kind of pact doesn't address is who arbitrates a disagreement if OpenAI's stress test flags a serious issue in an Anthropic model, or vice versa -- there is no disclosed enforcement mechanism, no independent referee, and no public reporting requirement comparable to what California's SB 53 already imposes on both companies. A voluntary bilateral agreement between two competitors, however well-intentioned, is still self-policing by the companies with the most to lose from an unfavorable public result being disclosed at all.

What To Watch

The gap between "neared a deal" and an actually signed, binding agreement is exactly where past industry safety commitments -- including earlier voluntary AI pledges from 2023 -- have stalled or been quietly narrowed before signing. Whether this pact produces a published joint report, similar to the August 2025 cross-evaluation, or stays a private arrangement between the two companies, will determine whether it functions as real accountability or as a talking point for both labs' next round of safety messaging.

ShareXLinkedInEmail

Key Sources

2 sources

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.