It took the UK's AI Security Institute hours โ not weeks โ to find universal jailbreaks in OpenAI's GPT-5.6. One month later, OpenAI has patched the specific holes, shipped an updated model, and left the harder question conspicuously open: the model class itself is still rated High-risk for cyber capability, and nobody is claiming the jailbreak problem is solved.
We covered the original disclosure on the wire when it broke โ see our Pulse report on the UK AISI jailbreak findings โ and the story has kept moving since. This is the follow-up: a documented timeline of what OpenAI has actually fixed, what changed in the August 6 model update, and what remains structurally unresolved as of this writing.

The GPT-5.6 Jailbreak Timeline So Far
July 9, 2026
OpenAI publishes the GPT-5.6 system card, designating the model High capability in cybersecurity and in biological and chemical domains under its Preparedness Framework.
July 10, 2026
Fortune reports the UK AI Security Institute found universal jailbreaks in GPT-5.6's cyber domain โ often developed within hours โ enabling agentic vulnerability discovery and exploit development.
Mid-July 2026
OpenAI says it has reproduced and mitigated the specific jailbreak methods AISI reported, and commits to continued joint testing. AISI cautions it expects further red-teaming to surface similar jailbreaks.
August 6, 2026
OpenAI ships updated August versions of GPT-5.6 Sol and Luna to ChatGPT and publishes an August Updates system card: jailbreak robustness reported as comparable to predecessors, Preparedness designations unchanged.
Sources: Fortune, OpenAI GPT-5.6 system card, OpenAI GPT-5.6 August Updates.
What AISI Actually Found โ and Why "Universal" Was the Scary Word
The core finding, first reported by Fortune on July 10, was not that GPT-5.6 could be tricked into a single bad answer. It was that AISI researchers found universal jailbreaks in the cyber domain โ prompt techniques that reliably stripped the model's refusal training across scenarios and unlocked long-form agentic task completion: vulnerability discovery, exploit development, multi-step offensive work the model is explicitly trained to refuse.
Two qualifiers matter for an honest read. First, the jailbreaks were often developed within hours, but AISI had privileged access to the model's inner workings under its testing agreement with OpenAI, which likely compressed that timeline versus what an outside attacker faces. Second, AISI tested the model layer directly โ as Technobezz noted in its coverage, OpenAI's production deployment layers additional safeguards, such as monitoring classifiers, on top of the raw model.
The reason this story traveled beyond security circles is the regulatory subtext. Fortune's reporting explicitly framed the flaw as similar to the one that led the US government to force Anthropic to disable its Fable 5 model. Whatever the technical details, the precedent is now set: a bad enough jailbreak finding can take a frontier model off the market. That changes the stakes of every one of these disclosures.
What OpenAI Has Fixed
OpenAI's response has come in three layers, and it is worth separating them because they are different kinds of fix.
1. The specific jailbreaks are mitigated
OpenAI says it reproduced and mitigated the specific jailbreak methods AISI reported, and is continuing joint testing with the institute. This is the narrowest and most complete fix: the exact prompts AISI found should no longer work.
2. A rapid-response patching process is now policy
Per the GPT-5.6 system card, OpenAI disclosed spending roughly 700,000 GPU hours hunting for universal jailbreaks itself, and pledged a standing rapid-response process to reproduce, assess, prioritize, and remediate newly discovered jailbreaks. That is an acknowledgment that jailbreak discovery is continuous, not a one-time audit.
3. The August 6 model update shipped โ with unchanged risk ratings
The updated GPT-5.6 Sol and Luna models released to ChatGPT on August 6 improved disallowed-content and factuality benchmarks, but OpenAI's own August Updates document reports jailbreak robustness merely comparable to recent predecessors, labels those results interim and directional, and keeps both models at High capability in cybersecurity under the Preparedness Framework.
Source: GPT-5.6 August Updates (OpenAI, August 6, 2026) and the original GPT-5.6 system card.
What's Still Open
Read the August Updates document closely and the unresolved items are stated in OpenAI's own language, which is unusually candid:
- The attack class is not closed. AISI said it expects further red-teaming to surface similar jailbreaks, and OpenAI has acknowledged that no model is absolutely secure. The specific holes are patched; the method of finding new ones still works.
- Jailbreak robustness has not step-changed. The August system card reports GPT-5.6 Sol and Luna perform "comparably to recent predecessors" on jailbreak evaluations, with results OpenAI itself calls directional rather than definitive โ including regressions versus some prior models at high attacker budgets.
- The cyber risk rating is unchanged. Both August models remain High capability in cybersecurity and in biological and chemical domains under the Preparedness Framework. Production safety is carried by layered safeguards โ classifiers and monitoring โ not by model-level robustness.
- No public AISI re-test has been published. As of August 7, there is no public follow-up report from AISI verifying the mitigations against fresh red-teaming. Until that exists, "fixed" is a vendor claim, not an independently confirmed state.
The Business Layer: Why This Isn't Just a Security Story
GPT-5.6 is not a research artifact โ it is OpenAI's flagship commercial engine, and the jailbreak saga unfolded in the same weeks the model became a production-infrastructure story. GPT-5.6 Sol launched running at 750 tokens per second on Cerebras hardware, the fastest disclosed inference speed for a frontier model in production, under the $20B+ compute deal that anchors most of Cerebras' post-IPO backlog. We broke down what that deal means for CBRS stock when the speed numbers dropped.
That coupling is exactly why patch status matters commercially. The Fable 5 precedent โ a US-government-forced model shutdown over a security flaw, per Fortune's reporting โ means a sufficiently bad unpatched jailbreak is no longer just reputational risk. It is potential revenue interruption for OpenAI and, downstream, utilization risk for the compute partners billing against its inference demand. With the August 6 update, GPT-5.6 also became the default model for ChatGPT's free tier, expanding the attack surface to OpenAI's entire consumer base at the same moment the safety documentation concedes robustness is a work in progress.
There is also a competitive read. Enterprise buyers doing model procurement in regulated industries now have a concrete artifact trail to compare: system cards, third-party government red-team results, and patch-response speed. That evaluation layer is becoming part of how the OpenAI-vs-Anthropic enterprise race gets scored, alongside price and capability.
So has the GPT-5.6 jailbreak been fixed?
The reported holes are patched. The hole-finding method still works โ and OpenAI's own documentation says so.
The Bottom Line
A month in, the honest status is: patched, not solved. OpenAI moved fast on the specific findings, built a standing remediation process, and shipped a model update on schedule โ that is a genuinely competent incident response. But the company's own August documentation declines to claim improved jailbreak robustness, keeps the High cyber-capability designation in place, and leans on system-level safeguards to carry production safety. AISI's prediction that similar jailbreaks will keep surfacing is, so far, the safest bet in the story.
What I'm watching next: whether AISI publishes a re-test of the August models, whether the rapid-response process gets exercised publicly on a new jailbreak, and whether any regulator moves from precedent to policy on frontier-model security flaws. We'll track developments as they break on Value Add Pulse โ starting from the original AISI disclosure and the Cerebras production-inference story it collided with.
Follow AI model and infrastructure news in real time on Value Add Pulse at Value Add VC. Reach out at t@nyvp.com or @Trace_Cohen.
Get VC data most people never see
โ 100% free
Weekly benchmarks, valuations, and fund data. Join 5,000+ investors. No spam.