VC
Value Add VC
โšกHomePulseโšกHelpful Apps๐Ÿ“Blog๐ŸคPartner
Illustration for: NYT Says OpenAI Hid Evidence in Copyright Trial
Value Add VC/Pulse/REGULATION

NYT Says OpenAI Hid Evidence in Copyright Trial

The New York Times and the Daily News are asking a federal judge to sanction OpenAI, alleging the company deleted billions of ChatGPT outputs after a preservation order, substituted logs and hid its ability to search 78 million conversations for infringing.

By the Numbers

December 2023
Lawsuit filed
~78 million
De-identified conversations held
120 million
Logs originally requested
20 million
Logs OpenAI produced
April 2026
Deposition revealing gap
OpenAIThe New York TimesNYT
TC
By the Markets Desk
Edited by Trace Cohen ยท Early-stage VC & angel ยท Founder, New York Venture Partners
July 10, 2026
3 min read
ShareXLinkedInEmail

THE RUNDOWN

1

The New York Times and the Daily News filed a motion asking a federal judge to sanction OpenAI, alleging the company falsely claimed it couldn't search its own training data and chat logs for copyrighted journalism, when internal tools to do exactly that already existed

2

An April deposition of OpenAI data-privacy engineer Vinnie Monaco reportedly revealed OpenAI had already run internal searches of its training corpus and had amassed roughly 78 million de-identified ChatGPT conversations under an internal effort called Project Giraffe, including a "Bloom" filter built to detect content regurgitation

3

Plaintiffs allege OpenAI deleted billions of ChatGPT outputs after the court's preservation order took effect, substituted millions of logs in the sample it was ordered to produce, and negotiated the requested 120 million chat logs down to 20 million that the court has called "unusable" due to excessive redaction

4

The plaintiffs are asking the court to exclude the 20 million-log sample as evidence, treat regurgitation as an established fact for trial purposes, and award legal fees -- a request that, if granted, would materially weaken OpenAI's defense in a lawsuit first filed in December 2023

TC

The VC Read ยท Trace's Take

Trace Cohen

A company built on 'we can't search that' just got caught having already searched it internally -- that's not a fair-use argument, that's a discovery-conduct problem, and courts punish those independently of who eventually wins on the merits. If you're a founder building on user-generated data at scale, assume your retention and deletion practices will get deposed someday, not just audited. Watch whether other publishers suing OpenAI, Anthropic or Google pile onto this sanctions motion in the next few weeks -- that's the real signal of how much this actually costs OpenAI.

Analysis

The New York Times and the Daily News asked a federal judge on July 9 to sanction OpenAI, alleging the company misrepresented its own technical capabilities during discovery in the two-year-old copyright lawsuit over ChatGPT's training data. The filing claims OpenAI told the court it couldn't search its training corpus or chat logs for copyrighted journalism, when the company had already built and used tools that did exactly that.

The most damaging detail comes from an April 2026 court-ordered deposition of OpenAI data-privacy engineer Vinnie Monaco, who reportedly disclosed that OpenAI had already conducted internal searches of its training data for copyrighted news content, and had separately amassed a database of roughly 78 million de-identified ChatGPT conversations under an internal initiative called Project Giraffe. That effort reportedly included a tool called a "Bloom" filter, built specifically to detect and record when ChatGPT outputs were regurgitating source material -- built, plaintiffs say, shortly after the lawsuit was filed, undercutting any claim that OpenAI lacked the technical means to search for infringement at scale.

The evidence-handling allegations go further than a capability dispute. Plaintiffs say OpenAI deleted billions of ChatGPT outputs after the court issued a preservation order specifically requiring the company to retain that data, and that OpenAI substituted millions of logs within the sample it was ultimately ordered to produce. The scope of that production was itself a fight: the Times and Daily News originally sought 120 million chat logs, and after negotiation received 20 million -- a sample the plaintiffs say arrived so heavily redacted that the court itself has called it effectively unusable as evidence.

โ€œAny one of those sanctions would materially shift leverage in a case that has run since December 2023 and names Microsoft as a co-defendant alongside OpenAI.โ€

The requested remedy is unusually aggressive for a discovery dispute. Plaintiffs are asking the judge to exclude the 20 million-log sample entirely, to instruct the jury that regurgitation of copyrighted content can be treated as an established fact rather than something OpenAI gets to contest at trial, and to award legal fees tied to the alleged misconduct. Any one of those sanctions would materially shift leverage in a case that has run since December 2023 and names Microsoft as a co-defendant alongside OpenAI.

OpenAI has pushed back hard. Spokesperson Drew Pusateri said in a statement that as the Times' underlying case weakens, the plaintiffs are using the discovery fight to try to access the private conversations of ChatGPT users who have nothing to do with the litigation, and the company continues to argue its use of copyrighted material falls under fair use. That framing -- privacy concerns versus evidence-preservation obligations -- is likely to become the central fight in the sanctions motion itself.

The case sits alongside a growing docket of AI copyright litigation, including suits from authors, visual artists and other publishers against OpenAI, Anthropic, Google and Meta, most of which turn on the same underlying question of whether training on copyrighted material without a license constitutes fair use. A sanctions finding against OpenAI wouldn't resolve that question directly, but it would set a discovery-conduct precedent other plaintiffs' lawyers are likely to cite immediately in parallel cases.

For AI founders building products on top of any frontier lab's API, the case is a pointed reminder that discovery obligations around user data retention are becoming a real, litigable cost center -- not a hypothetical -- for any company sitting on large volumes of user-generated content that touches copyrighted material. For enterprise buyers and investors, a credible sanctions finding would add real tail risk to OpenAI's litigation exposure at a moment when the company is simultaneously trying to project stability around its Microsoft relationship and its next funding round.

The bear case for the plaintiffs: OpenAI's fair-use defense doesn't depend on the discovery dispute being resolved in its favor, and courts frequently decline the most aggressive sanctions requested even when some misconduct is found. What to watch next: whether the judge rules on the sanctions motion before the case's next substantive hearing, and whether other publishers suing OpenAI move to join the sanctions request or file parallel motions citing the same Project Giraffe disclosures.

Related Deep Dives

  • 40% Enterprise LLM โ€” Anthropic Passes OpenAI โ†’
  • OpenAI vs Anthropic Revenue Dispute: $74B Gross ARR vs $4... โ†’
  • EU AI Act 2026 โ€” High-Risk Deadline Pushed โ†’
ShareXLinkedInEmail

More on

OpenAI โ†’The New York Times โ†’

Prior Pulse Coverage

OpenAIAnthropic and OpenAI's Parallel Paths to Going PublicOpenAIMassachusetts AI Bill Pits Anthropic Against OpenAIOpenAINvidia's Investment Pullback Is a Signal for 2026's IPO ClassOpenAIChatGPT Can Now Read and Send Your iMessagesOpenAIBroadcom Seeks Up to $100B in AI Chip Debt

Key Sources

2 sources
SourceTechCrunch
AnalysisValue Add Pulse

Reported by TechCrunch ยท Analysis by Value Add Pulse.

โ† Back to Pulse

THE WIRE in your inboxโ€” Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

REGULATIONยท Aug 21, 2026

Meta Loses Landmark Social Media Addiction Trial

Illustration for: Meta Loses Landmark Social Media Addiction Trial
REGULATION

Meta Loses Landmark Social Media Addiction Trial

Meta lost a closely watched trial over claims Instagram and Facebook are designed to be addictive to teenagers, a verdict plaintiffs' lawyers say could reshape how every major platform designs its feed and notification systems.

REGULATIONยท Aug 21, 2026

DOJ, TikTok Settle Child Privacy Suit for $400M

Illustration for: DOJ, TikTok Settle Child Privacy Suit for $400M
REGULATION$400M settlement

DOJ, TikTok Settle Child Privacy Suit for $400M

The Justice Department and TikTok agreed to a $400 million settlement resolving a children's-privacy lawsuit first filed under the Biden administration, one of the largest child-privacy payouts a tech platform has made.

REGULATIONยท Aug 21, 2026

Massachusetts AI Bill Pits Anthropic Against OpenAI

Illustration for: Massachusetts AI Bill Pits Anthropic Against OpenAI
REGULATION

Massachusetts AI Bill Pits Anthropic Against OpenAI

Massachusetts is on track to pass the strictest state-level AI safety law yet, and Anthropic and OpenAI have taken opposite public positions on whether it should become law.

Deep Dives

40% Enterprise LLM โ€” Anthropic Passes OpenAIOpenAI vs Anthropic Revenue Dispute: $74B Gross ARR vs $4...EU AI Act 2026 โ€” High-Risk Deadline Pushed
@Trace_Cohenยทt@nyvp.com