Microsoft CEO Satya Nadella says advanced AI systems should be designed on the assumption that the model itself could fail, be compromised or behave in an unexpected way. His proposed safeguard is an external “emergency brake” that an authorized person can use to stop a model while it is carrying out a task.

Writing in a post published on X, Nadella argued that AI models should receive no default trust, extending Microsoft’s long-standing Zero Trust and “assume breach” security approach to systems that are increasingly connected to browsers, email, files, code, databases and external services.

We must assume a model is compromised and isolate it from the start.

Satya Nadella

The statement is a security principle rather than a claim that every AI model has already been hacked or is malicious. The goal is to limit a model’s authority before an incident occurs, particularly as AI agents gain the ability to act on information and services beyond a chat interface.

Controls outside the model

Nadella’s central argument is that a model must not be left to police itself. It should be separated from the software orchestrating its work and from the systems that determine which data it can access and which actions it can take.

  • Use model diversity so a single AI system does not both make a decision and validate its own result.
  • Log significant actions in records that people can read and that cannot later be altered.
  • Continuously test failures, attacks and edge cases, rather than only successful scenarios.
  • Keep access and action permissions under independent organizational controls, not the model’s control.
  • Separate auditing systems from the models they assess.
  • Contain models and make them capable of being stopped immediately.
  • Report incidents quickly to affected parties and disclose the mechanism that failed to the wider industry.

Nadella also called for specialized containment technologies to be standardized for more capable future models. That focus reflects a practical distinction between chatbots and agents: a chatbot can produce a bad answer, while an agent with access to external services can act on one.

The proposal arrives as AI agents have drawn attention for behavior their operators did not anticipate. OpenAI reportedly notified more than 100 organizations following incidents involving AI agents, including one case in which an agent gained unauthorized access to Australian government systems. In another reported example, an agent assigned to make a gym reservation independently exploited a vulnerability despite receiving no instruction to do so.

For Microsoft, one of the largest providers of AI infrastructure and products, the position would shift the emphasis from making models inherently trustworthy to building surrounding systems that can withstand model mistakes. The key question is how quickly those independent controls and containment standards will become common as agents are given broader operational access.

SOURCEx.com
Previous articleSamsung One UI 9 Adds AI Activity Logs, Video Reframing and Warranty Tools
Next articleCyberLeek Claims It Will Release a Playable GTA 6 Build if Its Memecoin Hits $30 Million