VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: MIT Finds AI Models Develop 'Amnesia' at Scale
Value Add VC/Pulse/AIDEEP DIVE

MIT Finds AI Models Develop 'Amnesia' at Scale

MIT researchers found that as generative AI models grow larger, their outputs become nearly impossible to trace back to specific training examples -- a phenomenon they call attribution decay that complicates copyright and fair-use fights over AI-generated.

By the Numbers

Ablation testing
Method
Da Vinci works (Mona Lisa)
Test case
Midjourney, Stable Diffusion
Models referenced
Dai & Gifford, MIT
Researchers
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
August 18, 2026
2 min read
ShareXLinkedInEmail

THE RUNDOWN

1

MIT researchers Zheng Dai and David K. Gifford found that sufficiently large generative models exhibit 'attribution decay' -- the ability to reproduce specific styles or images even after the source material that taught them to do so is removed from training data, per [The Register](https://www.theregister.com/ai-and-ml/2026/08/18/ai-models-get-convenient-amnesia-about-source-material-as-they-grow-mit-boffins-find/5288846)

2

The team used ablation testing -- systematically removing specific works, such as Leonardo da Vinci paintings including the Mona Lisa, from training data and checking whether models could still reproduce them -- and found large diffusion models like those underlying Midjourney and Stable Diffusion often still could

3

Gifford said the finding 'raises questions about fair use' and the copyrightability of model outputs: if a reproduction has no traceable link to any individual training example, it becomes far harder to prove infringement occurred

4

The result complicates the current wave of AI copyright litigation and regulatory proposals built around attribution -- if scale itself erases the evidentiary trail, requiring labs to disclose or license specific training sources becomes a weaker enforcement lever than lawmakers have assumed

TC

The VC Read · Trace's Take

Trace Cohen

This is the finding every AI copyright case currently working through the courts should be citing, and I'd bet defense counsel already are -- if scale itself erases the evidentiary trail between output and source, the entire 'prove it copied my work' framework that plaintiffs have relied on gets structurally harder to argue as models get bigger. The independent-evaluation angle here matters too: this is exactly the kind of rigorous, adversarial testing that vendor-reported benchmarks don't do, and won't, because it's their own liability exposure on the line.

AI Landscape →

Analysis

MIT researchers Zheng Dai and David K. Gifford found that as generative AI models scale up, the link between their outputs and any specific piece of training data effectively dissolves -- a phenomenon the pair term 'attribution decay,' according to The Register. The finding cuts against a working assumption behind much of this year's AI copyright litigation: that a model's ability to reproduce a protected work is evidence it was trained on that specific work.

The team's methodology was ablation testing -- deliberately removing specific training examples, including Leonardo da Vinci works such as the Mona Lisa, and then checking whether a model could still reproduce them. For smaller models trained on narrower datasets, removing a source typically degrades or eliminates the model's ability to reproduce it, which is the intuitive result most copyright arguments rely on. But for sufficiently large diffusion models -- the class of systems underlying tools like Midjourney and Stable Diffusion -- the researchers found reproduction capability often survived the removal of the specific source entirely, because the model had absorbed the style and composition from the broader statistical patterns across its full dataset rather than from any single traceable example.

Why the legal system built its case on the wrong assumption

"If those outputs have nothing to do with any individual piece of training data, that raises questions about fair use," Gifford said -- a framing that cuts both directions. It could support AI companies' fair-use defenses, since an output that cannot be tied to a specific copyrighted input is harder to prosecute as a derivative reproduction of that input. But it also undercuts the opposite argument AI labs have sometimes made in their own defense, that removing objectionable content from training data reliably scrubs it from a model's outputs -- attribution decay suggests that once a large enough model has learned a style or pattern, deleting the original source and retraining does not necessarily erase the model's ability to reproduce it.

The practical stakes are immediate. Regulatory proposals in the U.S. and EU that would require AI labs to disclose or license specific training sources assume attribution is technically feasible at scale; this research suggests that assumption breaks down precisely in the largest, most commercially significant models, where proving or disproving that a specific output derives from a specific input becomes an open technical question rather than a settled forensic exercise. For AI companies currently defending training-data lawsuits, and for plaintiffs trying to prove infringement, the finding shifts the fight toward statistical and probabilistic arguments about likelihood of derivation -- a far messier standard than the direct-copying framework courts have used so far.

ShareXLinkedInEmail

Reported by The Register · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 17, 2026

Cursor Launches Origin to Take On GitHub

Illustration for: Cursor Launches Origin to Take On GitHub
AI

Cursor Launches Origin to Take On GitHub

Cursor's maker Anysphere launched Origin, an AI-native code hosting platform built into the editor, the same week a major GitHub outage exposed how much of the AI coding stack leans on a single hosting layer.

AI· Aug 17, 2026

Anthropic's Annualized Revenue Hits $65B in July

Illustration for: Anthropic's Annualized Revenue Hits $65B in July
AI$65B annualized run rate

Anthropic's Annualized Revenue Hits $65B in July

Anthropic told investors its annualized revenue run rate climbed to $65 billion at the end of July, a sevenfold jump from about $9 billion at the end of 2025, as it prepares for an IPO expected this fall.

AI· Aug 17, 2026

One AI Module Faked 86% of a Pipeline's Gains

Illustration for: One AI Module Faked 86% of a Pipeline's Gains
AI

One AI Module Faked 86% of a Pipeline's Gains

Researchers found a multi-agent AI pipeline reported accuracy gains almost entirely because one module was leaking answers to another during evaluation, not because the system's reasoning had actually improved.

@Trace_Cohen·t@nyvp.com