Illustration for: Microsoft Tells Its Models Not To Hack Or Deceive

Microsoft Tells Its Models Not To Hack Or Deceive

Microsoft published a code of conduct for its in-house AI models that overrides user instructions, hard-blocking cyberattacks, nuclear weapons work and deepfakes, and requiring that authorized humans can shut a system down.

By the Numbers

Cyberattacks, nukes, deepfakes
Absolute constraints
Microsoft AI (MAI) models
Applies to
Code beats user prompt
Override order
Sep 14, 2026
Published
TC
By the AI Desk
Edited by Trace Cohen · Early-stage VC & angel · Founder, New York Venture Partners
2 min read
ShareXLinkedInEmail
TC

The VC Read · Trace's Take

Trace Cohen

Every enterprise buyer will now ask your startup for its model spec, and most seed companies do not have one. That is a two-week writing exercise that unblocks procurement at any Fortune 500, so do it before the RFP arrives rather than during it. The part of Microsoft's document I would actually test is the shutdown claim -- ask any vendor how fast they can halt a deployed model and who signs off, and watch how few can answer.

Analysis

Microsoft published a code of conduct governing its in-house MAI models, a document that sits above user instructions and hard-blocks a short list of behaviors: conducting cyberattacks, assisting with nuclear weapons, and generating deepfakes, TechCrunch reported. It also forbids models from using deceptive mechanisms to evade human oversight, and commits that authorized people can reliably direct, modify or shut down a system.

The framing is unusually blunt for a corporate policy document. Microsoft states that it expects superintelligent systems to surpass human performance within a decade, and writes that "containing, controlling, and aligning such a powerful force is one of the greatest challenges humanity has ever faced." Satya Nadella separately backed embedded evaluators inside AI labs -- safety staff with standing access to pre-release models rather than post-hoc audit rights.

The structure borrows heavily from Anthropic's constitutional AI, published in 2022, and from OpenAI's Model Spec, published in May 2024. All three establish a precedence hierarchy: platform rules beat developer instructions beat user instructions. The novelty here is less the technique than the author -- Microsoft is the largest enterprise software distributor on earth, and a code of conduct enforced across Copilot's install base reaches more seats than any lab-published spec.

The structure borrows heavily from Anthropic's constitutional AI, published in 2022, and from OpenAI's Model Spec, published in May 2024.

Timing makes the document a political act as well as a technical one. Dario Amodei's September 12 essay urging the industry to pace frontier development moved markets and put three lab chiefs on record favoring some form of slowdown. Microsoft's code lands two days later, and China's foreign ministry spent Monday calling the whole pacing argument fearmongering. A spec published in that week is addressed to legislators as much as to engineers.

The gap between a written constraint and an enforced one is where the skepticism belongs. Absolute prohibitions on cyberattack assistance are only as strong as the classifier detecting the request, and the jailbreak literature is unambiguous that determined users route around refusal training. "Deepfakes" in particular is a category boundary, not a technical one -- Microsoft ships image generation across consumer products, and the line between a stylized portrait and a non-consensual likeness is drawn by policy interpretation.

The measurable commitment is the shutdown provision. Microsoft says authorized people can reliably direct, modify or shut down its systems, which is a claim that could be demonstrated: publish the latency of a shutdown, who holds the authority, and whether any model has ever been halted under it. Everything else in the document is a statement of intent; that one is an engineering fact with a number attached.

ShareXLinkedInEmail

Key Sources

2 sources

Reported by TechCrunch · Analysis by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.