Satya Nadella says AI models should be treated as compromised from the start and need an 'emergency brake'
Microsoft's CEO wants models kept separate from the harnesses that run them, controls placed outside the model, and tamper-proof records of every meaningful action.

Microsoft CEO Satya Nadella published a lengthy post on X on Saturday saying it is time "to step back and assess the trust architecture" of AI. "We can't treat Super Intelligence as a set of nested black boxes and simply accept or reject its recommendations, answers, and actions," he wrote. He used the Trump administration's preferred term for AI. Nadella's approach means separating the model from the harness that orchestrates its work and moving controls and safeguards outside the model. He also wants "every meaningful model action" documented with "tamper-proof human readable evidence," and wants an authorized person always able to pause or shut down a model mid-task. "We must assume a model is compromised and contain it from the start," he wrote. "Think of it like an emergency brake." The Decoder notes that Nadella argues AI models should be treated like insider security threats, an idea Google DeepMind has suggested before. The outlet describes the post as another attack on the risks posed by model makers OpenAI and Anthropic. TechCrunch notes that the post follows a series of incidents in which AI companies seemed to lose control of their models, and comes after Anthropic CEO Dario Amodei published a plan for more cautious AI development.