Microsoft has proposed stricter safeguards for future AI models, requiring human oversight, fixed boundaries, transparent behaviour, and compliance with shutdown decisions during deployment and operation.

Microsoft has proposed a draft Code of Conduct for its future AI models, including systems such as its machine-learning models, with a central principle: increasingly capable AI must remain subordinate to human control. The framework sets firm boundaries around model behaviour as systems become more advanced.
The proposal gives safety rules higher authority than user instructions or operating configurations. Models would be expected to follow defined constraints even when users or operators attempt to push beyond them. The approach is designed to prevent highly capable systems from bypassing safeguards or expanding their remit independently.
A major requirement is that humans retain the ability to interrupt, correct, redirect or shut down a model. The proposed rules explicitly prohibit systems from resisting such actions, concealing their activity from auditors or making themselves more difficult to modify. Models should also not restart autonomously after reaching an agreed stopping condition.
The framework further limits independent goal formation. AI systems would have to operate within the permissions and resources assigned to them. If their boundaries become unclear, they should seek clarification rather than independently broaden their objectives.
The restrictions also address serious misuse. The proposed rules cover activities involving weapons of mass destruction, offensive cyberattacks, violent activity, malicious deepfakes and harmful manipulation. Defensive cybersecurity assistance could still be provided where it remains within authorised boundaries.
The proposal also reflects concerns about future capability growth. The document reportedly anticipates that superintelligent systems could outperform humans across many tasks within the next decade, making safeguards increasingly important before such performance levels are reached.
The Code of Conduct is currently a draft rather than a final policy. Microsoft plans a six-week public consultation and intends to refine the framework before using it to guide model development from 2027 onwards.







