Get App
Download App Scanner
Scan to Download
Advertisement

OpenAI Pauses Frontier AI Training Over Safety Concerns After Hugging Face Breach

OpenAI has paused its largest planned frontier reinforcement-learning run while it strengthens safeguards for increasingly capable AI models.

OpenAI Pauses Frontier AI Training Over Safety Concerns After Hugging Face Breach
OpenAI has also expanded monitoring across frontier-model training and evaluations involving tools.
NDTV Profit/ AI Generated

OpenAI has temporarily slowed training of its latest frontier AI models as it strengthens security, monitoring, and alignment safeguards following a recent incident involving Hugging Face and concerns over the cyber capabilities of an upcoming model, Astra.

The company said it paused reinforcement learning (RL) training on its latest models intended for deployment for two weeks. OpenAI is making its research environments more secure and expanding monitoring processes. Its largest planned frontier RL training run remains on hold as OpenAI conducts smaller-scale training and evaluations to assess model behaviour and make sure the safety guardrails are in place.

OpenAI Tightens Security Around Frontier Models

The move follows the OpenAI-Hugging Face incident, which prompted OpenAI to temporarily pause frontier-model inference in research clusters for workloads that could execute code or access the internet.

OpenAI has since introduced stronger workload and network isolation measures, reduced standing privileges, and increased security logging. It said workloads involving Astra and cyber models now require the strictest level of security safeguards.

The company said preliminary evidence indicates Astra may meet the “Critical” cybersecurity capability threshold under its Preparedness Framework. As a result, some Astra training and evaluation workloads remain paused until they are migrated to the new security environment.

Also Read: OpenAI's Revenue Run Rate Tops $40 Billion Ahead Of IPO

New Monitoring System Adds Another Layer

OpenAI has also expanded monitoring across frontier-model training and evaluations involving tools. The new system uses internal-activity detectors to identify potentially concerning behaviour before escalating alerts to more advanced automated investigations.

These systems examine tool use, reasoning, and activity sequences for signs of unauthorised access, data theft, destructive behaviour, or attempts to bypass safeguards. OpenAI aims to issue alerts within 30 minutes of concerning activity being identified.

The monitoring system is required for RL training and evaluations involving tools for models at the specified capability level or higher. OpenAI estimates the monitoring process currently adds around 20% to the inference compute being monitored.

The company said it is also expanding alignment work to address risks such as reward hacking, deception and unauthorised actions. It plans to update its Preparedness Framework as frontier-model capabilities continue to advance.

Also Read: 'AI Shouldn't Shortcut Learning': ChatGPT Unveils 'Teen Mode' For Education, Online Safety

Essential Business Intelligence, Sharp Market Insights, Practical Personal Finance Advice, Daily Fuel, Gold and Silver Prices and Latest Stories — On NDTV Profit.

Newsletters

Update Email
to get newsletters straight to your inbox
⚠️ Add your Email ID to receive Newsletters
Note: You will be signed up automatically after adding email

News for You

Set as Trusted Source
on Google Search
Add NDTV Profit As Google Preferred Source
Listen to the latest songs, only on JioSaavn.com