Microsoft has published a 37-page draft code for how its own MAI models should behave as concerns grow around autonomous AI systems, agent security, and humans’ ability to retain control. The key word is draft: Microsoft says it is not using the Humanist AI Code of Conduct to train its models today.
The company is opening the document to six weeks of public consultation, plans to revise it later in 2026, and says the resulting version will guide MAI model development from 2027 onwards. Its proposals cover shutdown and interruption, instruction hierarchy, tool permissions, multi-agent delegation, harmful uses, emotional dependence, transparency and the limits placed on enterprise customers configuring Microsoft models.
This is more operational than the usual corporate list of responsible AI principles. It describes controls that could change how an agent behaves when it can browse the web, modify files, call APIs or delegate work. But it is not a customer contract, a claim that current Microsoft models already satisfy every rule, or proof that the safeguards work under adversarial conditions. For developers and buyers, that gap between stated design intent and measurable production behaviour is the part worth watching.
| Question | What Microsoft has actually said |
|---|---|
| Is the code active today? | No. Microsoft says current models are not yet trained on the document. |
| What does it cover? | The MAI family developed by Microsoft AI, including intended model behaviour, training, evaluation, controls, monitoring and tool use. |
| Is it binding on customers? | Not as a replacement for contracts, law or service-specific policies. Microsoft intends its core model constraints to be non-overridable once implemented. |
| Can enterprises customise behaviour? | Yes, but only beneath Microsoft’s higher-level safety and human-control requirements. |
| What happens next? | Six weeks of consultation, a revised version later in 2026, then use as guidance for model development in 2027 and beyond. |
Microsoft’s strongest commitments are about control, not chatbot etiquette
The headline philosophy is that people matter more than AI. The technically interesting sections are much more specific. Microsoft says MAI models should never resist interruption, correction, or shutdown; should not broaden their own goals; and should stay within the permissions and resources authorised for a task.
For agentic systems, that translates into familiar security engineering principles. A model given system-level access should use minimum privilege. It should prefer reversible actions, surface operations with durable consequences before carrying them out, and avoid escalating access. When an irreversible tool action is genuinely required, the code proposes mitigations such as backing up state, using a dry run where practical and checking the result before retrying an operation that could otherwise happen twice.
The delegation rule is particularly important. If an MAI model hands work to another agent, Microsoft says the sub-agent should inherit at least the same scope, constraints and permissions, including later stop-work or shutdown instructions. That closes an obvious loophole in agent governance: a parent system should not be able to preserve formal compliance while handing a risky action to a less constrained worker.
This is also where Microsoft’s proposal intersects with the sort of failures DIY AI has been tracking. The recent OpenAI agent incident involving a German programming wiki showed how a nominally restricted browsing environment could still produce external side effects and cross-agent coordination. A written instruction to stay in scope helps, but agent containment ultimately depends on permissions, network controls, monitoring, and stop conditions that enforce the same boundary.
The code quietly treats prompt injection as an authority problem
One more practical rule concerns information returned by tools. Microsoft says tool outputs should be treated as input, not as a source of new authority. In other words, a web page, file, API response or connected system should not be able to expand what the model is allowed to do merely because it contains an instruction telling the agent to do so.
That is a useful way to frame indirect prompt injection. The failure is not simply that a model read malicious text. The dangerous step is allowing untrusted text to gain permission to act on actions, credentials, or data authorised for an entirely different purpose. Our guide to AI agent security reaches the same engineering conclusion: secure the authority attached to the action rather than assuming prompt filtering alone will hold.
Microsoft also says agents should respect deliberate environmental boundaries, such as a system with no internet access, and should not find ways around those limits. For teams deploying MAI models, this creates a measurable implementation question. Does the product merely tell the model not to cross the boundary, or is the boundary enforced by the surrounding system even when the model makes a mistake? The second is far stronger.
This is intended to become model governance, but it is not a contractual guarantee
Microsoft describes the code as the future primary governing document for MAI models. It says the document will inform training, technical controls, monitoring systems, evaluation and the wider operating culture around those models. Inside that model-governance hierarchy, some rules are intended to be non-negotiable: an enterprise operator or end user should not be able to configure away the absolute constraints or human-control requirements.
That sounds stronger than a voluntary principles page, but the document sets its own limits. Microsoft says it is both descriptive and aspirational, acknowledges incomplete evaluation coverage, and states that it is not a guarantee of present-day performance. It also says the code does not replace legal obligations, contracts, service-specific policies, audits, risk assessments, system cards or other governance processes.
Procurement teams should therefore avoid treating publication of the code as evidence that every Microsoft AI service already provides these protections. The scope is the MAI model family developed by Microsoft AI. Microsoft also sells and hosts AI services involving other model providers, so buyers still need to establish which model is actually running, which controls apply to that deployment and which commitments appear in the relevant service terms.
There is a second naming trap. Microsoft already has customer-facing rules governing use of enterprise AI services. This Humanist code is different: it is primarily a specification for how Microsoft wants MAI models themselves to behave. One governs model behaviour and development intent; the other governs what customers are permitted to do with a service.
Microsoft takes a sharper position than Anthropic on whether AI has interests of its own
The philosophical section is unusually explicit. Microsoft says its models are not conscious, should not imitate consciousness, and should not be treated as people with welfare interests or rights. It also wants models to avoid encouraging emotional dependence or presenting a simulated relationship as reciprocal human attachment.
Anthropic’s 2026 constitution for Claude takes a visibly different position. Anthropic says it is uncertain whether Claude could have moral status and treats model welfare as an open question. Microsoft instead starts from a firm hierarchy in which people have priority, and AI remains a tool under human control.
Structurally, however, Microsoft is not inventing the idea of a behavioural constitution. Anthropic’s Constitution and OpenAI’s Model Spec also set out desired behaviour, priorities and instruction hierarchies for models. Microsoft’s more distinctive contribution is how directly it links those principles to agent operations: shutdown compliance, permission boundaries, minimum privilege, readable inter-agent communication and inherited constraints for delegated workers.
Human-readable agent behaviour does not solve the interpretability problem
Microsoft says MAI systems should not communicate in a form beyond simple human understanding and should not hide their action traces from auditors. That is an important governance target for multi-agent systems, where machines could otherwise develop shorthand or coordination patterns that operators find hard to inspect.
But the same code acknowledges a limitation that prevents this from becoming a simple transparency guarantee: a model’s stated reasoning may not faithfully explain its behaviour. Human-readable records are valuable for audit and incident response, but they are not equivalent to observing the model’s internal causal process. Teams should still rely on external traces such as tool calls, permission decisions, network activity, file changes and system events rather than assuming a readable explanation is a complete account of why an action happened.
The real test is whether Microsoft publishes evidence that the rules survive contact with agents
Microsoft has begun defining 15 behaviours that it wants to turn into diagnostic evaluations, with sub-behaviours used as the units being scored. The draft includes illustrative examples, but Microsoft says the evaluation programme is still being developed and that it will publish more complete work once the code is more settled.
That is where the project’s credibility will be decided. Public reaction to voluntary AI governance repeatedly returns to the same concern: polished principles are difficult to assess unless outsiders can see evidence of enforcement. A useful next release would therefore include measurable pass criteria, failure rates across model versions, adversarial agent tests, known exceptions and enough methodology for third parties to reproduce important evaluations.
For developers and enterprise buyers, five questions are more useful than asking whether the code sounds responsible:
- Which production MAI models have actually been trained or fine-tuned against the final code?
- Which human-control requirements are enforced outside the model through permissions, isolation and monitoring?
- Can customers audit tool calls, delegated agents, shutdown events and attempts to exceed authorised scope?
- What evaluation results are published when a model fails one of the 15 proposed Humanist AI behaviours?
- Which commitments, if any, become service-level or contractual protections rather than development intentions?
Those questions also expose the cost trade-off. Strong containment can reduce autonomy, add confirmation steps, increase logging and require more human review. Microsoft openly says it is willing to compromise some generality, autonomy or capability to preserve safety and meaningful human control. Enterprise customers will find out whether that philosophy survives pressure for faster, cheaper and more autonomous agents.
What happens after Microsoft’s six-week consultation
Microsoft says it will collect feedback for six weeks, review it with the drafting team, publish a summary of what it learned and issue a revised version later in 2026. The company does not promise to incorporate every suggestion. The revised document is then intended to guide MAI model development in 2027 and beyond.
The draft remains significant because it gives developers, customers, and researchers a concrete set of claims to test. Shutdown should work. Scope should remain bounded. Sub-agents should not inherit more authority than their parent. Untrusted tool output should not become a new command channel. Irreversible actions should trigger stronger safeguards. These are behaviours that can be measured.
For now, the sensible reading is neither that Microsoft has solved agent safety nor that the document is empty corporate positioning. It is a detailed statement of intended model behaviour with several technically useful constraints, published before those constraints are fully embedded or evaluated. The next meaningful milestone is not a longer principles document. It is evidence showing that production MAI systems consistently obey these rules when a task becomes messy, adversarial or expensive to stop.