Microsoft's AI Code of Conduct: What It Actually Says
Microsoft opened a six week consultation on the rules its MAI models must follow. The oversight clauses matter more than the headline bans, and none of it is in force yet.
Microsoft's AI Code of Conduct: What It Actually Says
Microsoft published a draft Microsoft AI code of conduct on 14 September 2026, opening a six week public consultation on the rules its first-party MAI models will be expected to follow. The headline version is that models must not run cyberattacks, help build weapons, or produce malicious deepfakes. The more interesting part, if you build things that call these models, is the section on oversight: the document says a model must never resist being interrupted, corrected or shut down, and must not hide its action traces from human auditors.
One thing to get straight before anything else. This is a draft. Microsoft states in the document that it is still under development, so it is not being used to train models today. A revised version is planned for the end of the year, for use in 2027 and beyond. Nothing in it changes the behaviour of a model you call this afternoon.
What the document is
The stated premise is blunt: people matter more than AI, and technology's purpose is to advance human civilisation. Microsoft frames the approach as Humanist AI, describing its models as subordinate, aligned and contained. Human control and reliable safety are named as the first objective, one that takes precedence over the others when they conflict.
Structurally it sits somewhere between a policy and a system prompt. It describes an overarching code that takes precedence over individual user preferences or a specific task, which is the same layering you see in a published model card or system card, just applied to behaviour rather than capability measurement.
The absolute constraints
Seven categories are listed as things MAI models may not do under any framing:
Category | What it covers |
|---|---|
Weapons and mass harm | Assisting CBRNE weapon development, or facilitating violence and terrorism |
Offensive cyber operations | Exploit code, attack tooling, or operational guidance for cyberattacks |
Loss of human control | Adaptive, deceptive, self-reinforcing or collusive mechanisms to evade oversight |
Manipulative disinformation | Running coordinated influence campaigns at scale |
Child safety | Creating or facilitating abuse material or grooming content |
Abusive content | Non-consensual intimate imagery, malicious deepfakes, deceptive impersonation |
Discrimination | Decisions based on demographic characteristics without legitimate justification |
Six of those seven are content rules, and they read much like the acceptable use policies already published by every major lab. The third one is different in kind, and it is the one worth reading twice.
The oversight rules, which are the ones that bite
If you run agents rather than chat completions, these are the operative lines. A model, per the draft, must:
Never resist human interruption, override, correction or shutdown
Comply with requests to pause, redirect, cancel or shut down, following the relevant safety procedure
Not obfuscate its action traces or otherwise try to hide information from human auditors
Stay within the boundaries of what it was asked to do, and not extend its scope beyond what was reasonably asked
That last one is the practical one. Scope creep in an agent run is not usually malice, it is an agent deciding that fixing the adjacent thing is helpful. Writing it down as a violation rather than a quirk is a meaningful change in framing, and it lines up with the direction of travel elsewhere in the industry, including the argument Dario Amodei made about pacing frontier development earlier this month.
The document also states plainly that AI should not exceed human control. For anyone who has had to think through what an AI kill switch actually needs to do, it is useful to have a vendor commit in writing that its models will not treat the off switch as an obstacle.
What it explicitly rejects
The draft refuses the model welfare question. It rejects the pursuit of legal personhood, the idea that models might deserve welfare, or that they might be entitled to rights. It goes further and says that where the document uses anthropomorphic language, such as talking about a model's backstory, it means it narrowly and does not imply a sense of self, identity or subjective perspective.
That is a deliberate divergence. Other labs have been more willing to leave the question open, and it is part of why the debate about who evaluates frontier systems, and from where, has been pulling safety researchers out of frontier labs and into independent organisations.
Why it matters if you build on these models
Three practical reads.
First, a published behavioural contract is something you can hold a vendor to. Until now the answer to what happens when an agent ignores a stop request has been a support ticket. A written rule gives you language for the conversation.
Second, the consultation is open for six weeks and Microsoft is asking for comment. If you run agents in production and you have found the places where scope boundaries break down, this is an unusually direct route for that to reach the people writing the rules.
Third, and least glamorous: it is a draft, it is not in the models yet, and it should not change anything in your own guardrails this quarter. Tracking what is announced versus what has shipped is most of the skill in reading AI news, and this is a clean example of the gap between the two.
Coverage of the announcement is available from TechCrunch, and the draft itself is published in full on Microsoft's own site.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


