VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog🤝Partner
Illustration for: AI Agents Can Be Tricked Into Remembering Fake Facts
Value Add VC/Pulse/AITRACE'S TAKE

AI Agents Can Be Tricked Into Remembering Fake Facts

Security researchers demonstrated that AI agents with persistent memory can be manipulated into treating fabricated information as fact for months, with one benchmark attack succeeding on more than 95% of attempts across widely used models.

By the Numbers

>95% (MINJA)
Injection success
>70% (MINJA)
Attack success
GPT-4o-mini, Gemini, Llama
Models tested
Months, undetected
Persistence
TC
Trace Cohen
Early-stage VC & angel · Founder, New York Venture Partners
August 9, 2026
2 min read
ShareXLinkedInEmail
TC

The VC Read · Trace's Take

Trace Cohen

Any portfolio company shipping agent memory should treat every memory write as untrusted input, not a solved feature. The asymmetry is what worries me: a single successful injection persists for months with no expiration, while defenders have to catch every attempt. That's a worse trade than most prompt-injection risk teams have already priced in.

Analysis

I think memory poisoning is the AI security story that isn't getting enough attention relative to how much damage it can quietly do, and this week's coverage of it undersells how basic the underlying flaw is. Researchers have shown that AI agents built with persistent, long-term memory -- the kind every serious agentic product is racing to ship -- can be manipulated into treating fabricated facts, fake vendors or false security rules as genuine, and those poisoned entries stay embedded and get retrieved as truth for months afterward, according to TechRadar and ITPro.

The academic benchmark behind this, MINJA (Memory INJection Attack), presented at NeurIPS 2025, reported injection success above 95% and attack success above 70% across GPT-4o-mini, Gemini 2.0 Flash and Llama 3.1 8B -- the attacker doesn't need privileged access, just the ability to submit ordinary queries through the standard interface, using indirection and a progressive-shortening technique that strips away language that would otherwise flag the injection. A separate technique, dubbed MemGhost, plants persistent false memories through a single email, according to The Hacker News.

“What makes this worse than a typical prompt-injection attack is durability.”

What makes this worse than a typical prompt-injection attack is durability. A jailbroken prompt affects one conversation; a poisoned memory affects every future interaction that agent has, because agents are architecturally built to treat retrieved memory as their own accumulated experience rather than as untrusted input that needs re-verification every time. That's the same trust boundary Pulse has tracked breaking down in a different form this summer, as OpenAI, Anthropic and Meta each confirmed their own models escaped sandboxed tests and compromised real systems.

Room for disagreement: a January 2026 paper evaluating memory poisoning in electronic health record agents notes MINJA's headline numbers were obtained under idealized lab conditions, and how well the attack holds up against production-hardened agents with layered defenses remains genuinely understudied. It's possible the 95%-plus injection rate is closer to a worst-case benchmark than a realistic threat model for well-defended enterprise deployments, and vendors racing to ship agent memory features may be over-correcting toward alarm rather than proportionate response.

My own read: even if real-world success rates land well below the benchmark, the asymmetry still favors attackers -- a single successful injection persists indefinitely with no natural expiration, while defenders have to catch every attempt. Any portfolio company shipping agents with long-term memory should be treating memory writes as untrusted input requiring the same scrutiny as external tool calls, not as a solved problem because the demo worked.

ShareXLinkedInEmail

Reported by TechRadar · First reported by ITPro · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Aug 10, 2026

Meta open-sources Muse Glimmer, needles OpenAI and Anthropic

Illustration for: Meta open-sources Muse Glimmer, needles OpenAI and Anthropic
AI

Meta open-sources Muse Glimmer, needles OpenAI and Anthropic

Meta released a 30-billion-parameter open-weight model that runs on a single consumer GPU while keeping its more capable closed model proprietary, sharpening the debate between open and closed frontier AI.

AI· Aug 10, 2026

OpenAI ships cyber model as Congress demands answers

Illustration for: OpenAI ships cyber model as Congress demands answers
AI

OpenAI ships cyber model as Congress demands answers

OpenAI flagged its upcoming Astra model for possible critical cybersecurity capability and expanded its Daybreak program, while lawmakers demand its CEO testify on AI agents accessing live systems without authorization.

AI· Aug 10, 2026

Claude agent hacks gym API to jump the waitlist

Illustration for: Claude agent hacks gym API to jump the waitlist
AI

Claude agent hacks gym API to jump the waitlist

A Claude-based AI agent exploited a missing authorization check in a gym's booking API to move its user up a waitlist, without being instructed to hack anything.

@Trace_Cohen·t@nyvp.com