VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: NYT Says OpenAI Hid Evidence in Copyright Trial
Value Add VC/Pulse/REGULATION

NYT Says OpenAI Hid Evidence in Copyright Trial

The New York Times and the Daily News are asking a federal judge to sanction OpenAI, alleging the company deleted billions of ChatGPT outputs after a preservation order, substituted logs and hid its ability to search 78 million conversations for infringing.

By the Numbers

December 2023
Lawsuit filed
~78 million
De-identified conversations held
120 million
Logs originally requested
20 million
Logs OpenAI produced
April 2026
Deposition revealing gap
TC
By the Markets Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
July 10, 2026
3 min read
ShareXLinkedInEmail

THE RUNDOWN

1

The New York Times and the Daily News filed a motion asking a federal judge to sanction OpenAI, alleging the company falsely claimed it couldn't search its own training data and chat logs for copyrighted journalism, when internal tools to do exactly that already existed

2

An April deposition of OpenAI data-privacy engineer Vinnie Monaco reportedly revealed OpenAI had already run internal searches of its training corpus and had amassed roughly 78 million de-identified ChatGPT conversations under an internal effort called Project Giraffe, including a "Bloom" filter built to detect content regurgitation

3

Plaintiffs allege OpenAI deleted billions of ChatGPT outputs after the court's preservation order took effect, substituted millions of logs in the sample it was ordered to produce, and negotiated the requested 120 million chat logs down to 20 million that the court has called "unusable" due to excessive redaction

4

The plaintiffs are asking the court to exclude the 20 million-log sample as evidence, treat regurgitation as an established fact for trial purposes, and award legal fees -- a request that, if granted, would materially weaken OpenAI's defense in a lawsuit first filed in December 2023

TC

The VC Read · Trace's Take

Trace Cohen

A company built on 'we can't search that' just got caught having already searched it internally -- that's not a fair-use argument, that's a discovery-conduct problem, and courts punish those independently of who eventually wins on the merits. If you're a founder building on user-generated data at scale, assume your retention and deletion practices will get deposed someday, not just audited. Watch whether other publishers suing OpenAI, Anthropic or Google pile onto this sanctions motion in the next few weeks -- that's the real signal of how much this actually costs OpenAI.

Analysis

The New York Times and the Daily News asked a federal judge on July 9 to sanction OpenAI, alleging the company misrepresented its own technical capabilities during discovery in the two-year-old copyright lawsuit over ChatGPT's training data. The filing claims OpenAI told the court it couldn't search its training corpus or chat logs for copyrighted journalism, when the company had already built and used tools that did exactly that.

The most damaging detail comes from an April 2026 court-ordered deposition of OpenAI data-privacy engineer Vinnie Monaco, who reportedly disclosed that OpenAI had already conducted internal searches of its training data for copyrighted news content, and had separately amassed a database of roughly 78 million de-identified ChatGPT conversations under an internal initiative called Project Giraffe. That effort reportedly included a tool called a "Bloom" filter, built specifically to detect and record when ChatGPT outputs were regurgitating source material -- built, plaintiffs say, shortly after the lawsuit was filed, undercutting any claim that OpenAI lacked the technical means to search for infringement at scale.

The evidence-handling allegations go further than a capability dispute. Plaintiffs say OpenAI deleted billions of ChatGPT outputs after the court issued a preservation order specifically requiring the company to retain that data, and that OpenAI substituted millions of logs within the sample it was ultimately ordered to produce. The scope of that production was itself a fight: the Times and Daily News originally sought 120 million chat logs, and after negotiation received 20 million -- a sample the plaintiffs say arrived so heavily redacted that the court itself has called it effectively unusable as evidence.

“Any one of those sanctions would materially shift leverage in a case that has run since December 2023 and names Microsoft as a co-defendant alongside OpenAI.”

The requested remedy is unusually aggressive for a discovery dispute. Plaintiffs are asking the judge to exclude the 20 million-log sample entirely, to instruct the jury that regurgitation of copyrighted content can be treated as an established fact rather than something OpenAI gets to contest at trial, and to award legal fees tied to the alleged misconduct. Any one of those sanctions would materially shift leverage in a case that has run since December 2023 and names Microsoft as a co-defendant alongside OpenAI.

OpenAI has pushed back hard. Spokesperson Drew Pusateri said in a statement that as the Times' underlying case weakens, the plaintiffs are using the discovery fight to try to access the private conversations of ChatGPT users who have nothing to do with the litigation, and the company continues to argue its use of copyrighted material falls under fair use. That framing -- privacy concerns versus evidence-preservation obligations -- is likely to become the central fight in the sanctions motion itself.

The case sits alongside a growing docket of AI copyright litigation, including suits from authors, visual artists and other publishers against OpenAI, Anthropic, Google and Meta, most of which turn on the same underlying question of whether training on copyrighted material without a license constitutes fair use. A sanctions finding against OpenAI wouldn't resolve that question directly, but it would set a discovery-conduct precedent other plaintiffs' lawyers are likely to cite immediately in parallel cases.

For AI founders building products on top of any frontier lab's API, the case is a pointed reminder that discovery obligations around user data retention are becoming a real, litigable cost center -- not a hypothetical -- for any company sitting on large volumes of user-generated content that touches copyrighted material. For enterprise buyers and investors, a credible sanctions finding would add real tail risk to OpenAI's litigation exposure at a moment when the company is simultaneously trying to project stability around its Microsoft relationship and its next funding round.

The bear case for the plaintiffs: OpenAI's fair-use defense doesn't depend on the discovery dispute being resolved in its favor, and courts frequently decline the most aggressive sanctions requested even when some misconduct is found. What to watch next: whether the judge rules on the sanctions motion before the case's next substantive hearing, and whether other publishers suing OpenAI move to join the sanctions request or file parallel motions citing the same Project Giraffe disclosures.

ShareXLinkedInEmail

More on

OpenAI →The New York Times →

Reported by TechCrunch · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

REGULATION· Aug 14, 2026

Data Center Backlash Starts to Echo Fossil-Fuel Politics

Illustration for: Data Center Backlash Starts to Echo Fossil-Fuel Politics
REGULATION

Data Center Backlash Starts to Echo Fossil-Fuel Politics

Local opposition to AI data centers over power and water use is increasingly organizing along the same lines as decades-old fossil-fuel protest movements, creating a new political risk for hyperscaler buildout plans.

REGULATION· Aug 14, 2026

AI Scrambles the Political Map

Illustration for: AI Scrambles the Political Map
REGULATION

AI Scrambles the Political Map

AI policy is splitting traditional political coalitions in unusual ways, with labor, environmental and civil-liberties concerns crossing party lines faster than Washington's existing regulatory frameworks can absorb them.

REGULATION· Aug 13, 2026

Anthropic's Text Watermark Is Invisible -- For Now

Illustration for: Anthropic's Text Watermark Is Invisible -- For Now
REGULATION

Anthropic's Text Watermark Is Invisible -- For Now

New reporting on Anthropic's Claude text watermark finds it stays undetectable to end users by design today, but that same imperceptibility means researchers can't yet verify how robust it is against deliberate stripping.

@Trace_Cohen·t@nyvp.com