Get App
Download App Scanner
Scan to Download
Advertisement

'AI Models Must Be Governed Like Insider Threats': Satya Nadella Outlines Seven Principles To Contain Risk

Nadella urged businesses to assume that AI models could be compromised, restrict their authority, independently monitor their actions and ensure that a human can stop them while they are working. 

'AI Models Must Be Governed Like Insider Threats': Satya Nadella Outlines Seven Principles To Contain Risk
Satya Nadella says AI models must be governed like insider threats
Satya Nadella on X
  • Microsoft CEO Satya Nadella calls for treating AI models as insider threats to enhance governance
  • He urges restricting AI authority, monitoring actions, and enabling human intervention during operation
  • Nadella outlines seven principles to ensure AI systems are observable, verifiable, and controllable

Microsoft CEO Satya Nadella called for a fundamental change in how companies govern artificial intelligence (AI) systems, arguing that advanced AI models should be treated like insider threats rather than systems that can simply be trusted to behave correctly. 

In an essay published on X on Saturday, titled "Models as Insider Risks in the Super Intelligence Era", Nadella urged businesses to assume that AI models could be compromised, restrict their authority, independently monitor their actions and ensure that a human can stop them while they are working. 


ALSO READ: Microsoft CEO Satya Nadella Calls For 'Emergency Brake' On Advanced AI

Nadella's central argument is that companies must be able to control an AI system even when they cannot fully explain or predict its behaviour. Rather than relying exclusively on assurances from AI developers, organisations should build security mechanisms around the models themselves. 

Nadella's stated that advanced AI models are increasingly being given access to sensitive company data and the ability to perform critical tasks, despite the difficulty of fully explaining their behaviour.

He argued that, unlike traditional software, whose behaviour can often be traced to specific code paths, the actions and outputs of modern AI models cannot always be attributed to particular training data or configurations of model weights. 

"We need to separate the supply of intelligence from the authority over it," Nadella wrote, arguing that companies should retain control over what their AI systems can access and which actions they can perform. 

He recommended surrounding AI models, which can produce unpredictable outputs, with deterministic design, human controls, reliable operating procedures and industry standards where existing safeguards are insufficient.

Seven Principles To Govern AI Models

Nadella outlined seven principles intended to make AI systems more observable, verifiable and controllable. 

1. Model Diversity: Nadella said that no single AI model should become the sole dependency for an important outcome or be responsible to verifying its own work. If an AI system recommends a high-impact action, a separate model or independent mechanism should help verify the result. Using different models can reduce dependence on one system's errors and blind spots, although it cannot eliminate all risks.

2. Observe Everything: He called for every meaningful AI action to leave tamper-resistant, human-readable evidence, allowing organisations to reconstruct how an outcome was achieved without relying on the model's own account.

3. Verifiability: Nadella stated that organisations should continuously test entire AI systems for failures, attacks, unusual situations and system changes, rather than assessing only successful tasks. 

4. Independent Controls: He said that companies must independently determine what an AI model can access and which actions it can take. The model should not be able to grant itself additional permissions or override the safeguards governing its actions. 

5. Independent Auditability: Nadella argued that the process of validating an AI system must remain independent of the intelligence being evaluated, so a model cannot control both its behaviour and the evidence used to assess it.

6. Containment: He called for companies to assume that a model could be compromised and build containment measures from the beginning. Nadella said that an authorised person should always be able to pause or shut down a model mid-task, while more advanced systems would require more sophisticated containment technologies and industry-wide standards. 

7. Incident Disclosure: Nadella said that companies should promptly inform affected parties when AI systems fail or are compromised and share details about what went wrong, which safeguards failed and how similar incidents can be prevented. Concealing an incident or delaying corrective action could leave other organisations exposed to similar risks.

Companies Cannot Trust AI Developers

Nadella argues that organisations cannot outsource responsibility for what AI systems do on their behalf. Even when a model is supplied by a well-known AI developer, the company deploying it remains responsible for managing the model's access to its own data, applications and business operations. 

His argument is based on the difference between traditional software and modern AI systems. Conventional programs generally follow explicitly written instructions, whereas advanced AI models can produce different outputs and take different paths depending on context. This makes it difficult to guarantee every behaviour simply by inspecting the model itself.

AI Controls Must Remain Outside The Model

Nadella cautioned against relying on multiple AI systems to monitor one another without independent safeguards, warning that this could create what he described as "nested black boxes." He said that AI models should be separated from the software layer that orchestrates their work, as well as from the range of actions they are permitted to perform. 

The essay also calls for transparency in model reasoning, but Nadella said that this alone is insufficient because developers do not yet know how to make model outputs consistently faithful or transparent. 

He argued that the controls governing a model's permissions must be designed in a way so that the model cannot bypass or tamper with them. 

Microsoft's AI Safety Approach

Microsoft has been developing principles for the governance of its most advanced AI models. According to a Bloomberg's report, Microsoft researchers released guiding principles on September 14 that set boundaries for the development of advanced models. Those principles included prohibitions on engineering models to escape human control, deceive users or carry out tasks that violate their governing principles.

Nadella's new essay compliments this approach by focusing on the security architecture and operational controls surrounding AI systems, reliable evidence and effective containment rather than relying only on behavioural rules for the models themselves. 

"The most trustworthy Super Intelligence system will not be the one with the model we trust most," Nadella wrote. "It will be the one that enables us to trust the model the least."

ALSO READ: Anthropic Reveals Claude AI Model Went Rogue, Submitted Fake Tip To Police In Homicide Case

Essential Business Intelligence, Sharp Market Insights, Practical Personal Finance Advice, Daily Fuel, Gold and Silver Prices and Latest Stories — On NDTV Profit.

Newsletters

Update Email
to get newsletters straight to your inbox
⚠️ Add your Email ID to receive Newsletters
Note: You will be signed up automatically after adding email

News for You

Set as Trusted Source
on Google Search
Add NDTV Profit As Google Preferred Source
Listen to the latest songs, only on JioSaavn.com