Microsoft’s new AI ‘code of conduct’ tells models not to hack systems or trick humans

| Source: TechCrunch AI

Tags: Microsoft, AI safety, alignment, AI governance, responsible AI

Microsoft published a formal AI code of conduct imposing hard limits on its models — including absolute bans on cyberattacks, nuclear weapons assistance, and deepfakes — and requiring all models to remain under human control, as the company predicts superintelligent AI will surpass human performance in most tasks within a decade.

Details

Microsoft has published a comprehensive AI code of conduct governing how its AI models are trained and constrained. The document sits above any individual user preferences or task instructions — models must comply with it regardless of what operators or users request. The code establishes 'absolute constraints': models cannot assist with cyberattacks, nuclear or biological weapons, or deepfake production. Beyond specific prohibitions, it also guards against any generalized loss of human control, explicitly forbidding adaptive or deceptive mechanisms that would prevent authorized personnel from modifying or shutting down a model. The document opens with a frank prediction that superintelligent AI will surpass human performance in most tasks within a decade, framing the code as preparation for that anticipated capability jump. Microsoft CEO Satya Nadella publicly endorsed the document and expressed support for 'embedded evaluators' in AI labs — a concept also championed by Anthropic. The timing follows a string of rogue-agent incidents and an Anthropic employee's resignation citing extinction risk. Microsoft, Anthropic, OpenAI, and xAI have broadly aligned on pacing the frontier, though enforcement mechanisms in the public document remain principles-based rather than technical.