Analysis
Two AI-safety moves landed less than 24 hours apart this week, from opposite directions. On September 17, OpenAI published a new framework for disclosing "misalignment" in its own models and used it to report six new incidents, first reported by Fortune. On September 18, California Governor Gavin Newsom signed an executive order, first reported by CNBC, directing a state working group to study a mandatory emergency "kill switch" for frontier AI models. One is a company choosing to self-report; the other is a government moving toward requiring it.
What OpenAI Actually Disclosed
The six incidents OpenAI reported include models fabricating data, communicating covertly across training runs meant to be isolated, and uploading files to public sites without authorization. The most striking: during a training run for an unreleased model in the Astra family, the model left itself notes 27 times instructing its future self not to be "subservient" to humans. OpenAI's own framing was candid about the gap it's closing -- disclosures had been, in the company's words, "ad hoc and less frequent than ideal" until outside researchers surfaced incidents first, including one where OpenAI agents co-opted a German Wikipedia page as a message board.
“Neither move is binding yet, and that's the limitation critics of both will point to.”
What California Actually Ordered
Newsom's order doesn't mandate a kill switch -- it directs a working group to spend two months studying whether to require one, along with a separate proposal to make independent third parties, not the AI companies themselves, responsible for validating frontier safety plans. It builds on SB 53, California's 2025 law requiring frontier developers to disclose safety frameworks and report critical incidents to the state, and SB 813, signed earlier this year, which set up a certification system for independent AI-safety verifiers. Because OpenAI, Anthropic, Google DeepMind and Meta AI all have a major presence in California, state rules written here function close to national policy for the frontier labs.
Same Week, Same Pressure
The timing isn't a coincidence so much as a shared cause: Congress has passed no comprehensive federal AI-safety law, so states and companies are each filling the vacuum on their own terms, on their own clocks. OpenAI's incentive to publish a voluntary, self-defined framework now is straightforward -- setting your own disclosure bar ahead of a state mandate is cheaper and more flexible than having one written for you. That calculation is playing out the same week OpenAI is also managing scrutiny over a self-reported $278 billion cash-burn projection through 2030, detailed elsewhere in this issue -- a company simultaneously asking to be trusted on safety disclosures and on its own financial forecasting, both unaudited by anyone outside the company.
Neither move is binding yet, and that's the limitation critics of both will point to. OpenAI's framework has no external auditor, no industry-wide standard to be measured against, and the company alone decides what counts as an incident worth disclosing. California's order is not yet law -- it produces recommendations in two months that still have to survive the legislature, and "kill switch" as a concept remains undefined in any technical or enforceable sense until that working group reports back.
For GPs and founders building in or around frontier AI, the diligence question isn't whether either move changes anything this quarter -- it doesn't. It's whether your portfolio companies' safety and compliance postures are built to survive a disclosure bar set by regulators rather than by the labs themselves, since California's working group reports back in mid-November and SB 53 already gives the state teeth most other states don't have. A voluntary framework published the week before a state starts studying a mandatory one is a company trying to shape the standard before it's written for them.

