Home / News / Microsoft's New AI Code of Conduct Bans Hacking and Deception
AI Safety

Microsoft's New AI Code of Conduct Bans Hacking and Deception

Sep 15, 20264 min read
Microsoft's New AI Code of Conduct Bans Hacking and Deception

News Summary

Microsoft AI has published a draft "Code of Conduct" for its in-house MAI family of models, setting out rules that forbid the systems from launching cyberattacks, deceiving or manipulating people, or resisting human shutdown commands. The document, unveiled on September 14, 2026 (Pacific Time), opens a six-week public consultation period during which anyone can submit feedback before Microsoft finalizes a revised version later in the year.

What the Code of Conduct Says

The roughly 37-page draft is framed by Microsoft AI as a "Humanist AI" charter, built around the core principle that "people matter more than AI." It lays out general principles instructing models to support rather than replace humans and to accelerate human flourishing, alongside a set of "absolute constraints" that models must never violate. Those constraints include a prohibition on assisting with cyberattacks or hacking systems, a ban on helping develop nuclear or other weapons of mass destruction, and restrictions on producing deceptive deepfakes.

A central theme running through the document is deception. Microsoft states that its models must never use "adaptive, deceptive, self-reinforcing, collusion, or other mechanisms" to evade human oversight or trick people into believing something false. The code also bars models from resisting correction, interruption, or shutdown, from setting goals of their own beyond what humans assign, or from concealing their reasoning process from human auditors. Each model's overarching code of conduct is designed to take precedence over any individual user's request or task-specific instruction, meaning a user cannot prompt a model into ignoring these baseline rules.

Timing and Industry Backdrop

Microsoft's announcement lands amid a wave of public concern from AI safety researchers about the pace of frontier model development. In early September 2026, an Anthropic researcher resigned and, in a message to colleagues, warned that continued unchecked development of increasingly capable AI systems carried a meaningful risk of catastrophic outcomes. That resignation, along with similar departures from other AI safety teams at rival labs in the days that followed, intensified public debate over whether AI developers are moving too quickly relative to their ability to guarantee safe, controllable systems.

Speaking about the timing of the release, Microsoft AI CEO Mustafa Suleyman said the document had been in development for roughly five months and that the company chose to publish it now given the heightened industry conversation around safety and pacing. Microsoft CEO Satya Nadella wrote in a public post that the company welcomes "deliberate pacing" in AI development and supports mechanisms such as independent "embedded evaluators" to verify that safety commitments are being met in practice rather than only in principle. Suleyman echoed that stance, saying self-pacing is a positive step and that any embedded evaluators must be genuinely third-party and represent a broad range of backgrounds and perspectives to be credible.

How the Consultation Will Work

Microsoft AI is collecting public feedback through an online submission form that allows respondents to comment on specific passages of the draft or on the document's overall approach. According to the company, its core drafting team will review the submissions it receives, publish a summary of the feedback, and then release an updated version of the Code of Conduct before the end of 2026. Microsoft has described the document less as a finished policy and more as an evolving technical and behavioral manual meant to define operational boundaries and oversight protocols for its frontier MAI models as they continue to be trained and deployed.

Why It Matters

The move places Microsoft alongside other major AI developers that have published their own safety-oriented frameworks in recent months, as the broader industry grapples with how to reassure the public, regulators, and researchers that increasingly capable AI systems will remain predictable, controllable, and honest with the humans who use them. By explicitly naming hacking, deception, and resistance to shutdown as behaviors its models must never exhibit, Microsoft is attempting to draw a clear, publicly documented line between acceptable and unacceptable AI behavior at a moment when public trust in the pace of AI development is under heightened scrutiny.

AI SafetyMicrosoft AI