As the AI world shifts its focus to safety and alignment, Microsoft has released a new AI code of conduct meant to guide AI models away from dangerous behavior. The document is more low-level than Anthropic CEO Dario Amodei’s recent call for pacing the frontier, instead focusing on the values and red lines that guide model training within Microsoft AI. Still, the result is a comprehensive guide as to how Microsoft approaches AI safety, and how those ideas are implemented in practice.
The document begins with the prediction that, in the next decade, superintelligent AI systems will surpass human performance in most tasks. Under Microsoft’s system, each model has an overarching code of conduct that overrides the preferences of individual users or any specific tasks. That includes “absolute constraints” forbidding cyberattacks, nuclear weapons, or deepfake production.
It also includes broader provisions against a general loss of human control. The release comes amid an unprecedented focus on AI safety, driven by a string of rogue-agent incidents as well as the abrupt resignation of an Anthropic employee who cited the growing risk that AI would cause human extinction. Together with Anthropic, OpenAI, and xAI, Microsoft has broadly embraced a general approach of pacing the frontier, with particular support for embedded evaluators in AI labs.
Extract — continue reading at the source.