OpenAI is rethinking its approach to security

Technologies2026-08-25, 11:31
OpenAI has temporarily slowed the pace of scaling and begun rebuilding protections for research environments following an incident involving Hugging Face and new data indicating that the new Astra model may reach the critical capability threshold.
Three enhanced mechanisms are being implemented for training and inference of Sol-level models and above: • workload Isolation — stronger isolation (“sandboxes”) for executing model-generated code; • network Isolation — internet access rules have been tightened for workloads associated with high risks. Compromise of a single node must not provide access to the internet or internal networks; • continuous Security Testing — potentially vulnerable shared services have been removed, standing privileges have been reduced, and trust boundaries have been improved. Automated testing using the models themselves is being introduced to simulate attacks.
The new monitoring system is multistage. — Specialized detectors inspect every step taken by the model. If they detect anything suspicious, automated “investigators” analyze the model’s actions, reasoning, and the entire sequence of events to identify attempts to steal data, gain unauthorized access, or bypass security controls. — The investigation is given 30 minutes. If it cannot be established within that time that the alert is a false positive, the model’s operation is suspended. Monitoring is mandatory for all training runs and evaluations of Sol-level models and above, and for Astra—even during ordinary use.
OpenAI acknowledges that risks are growing faster than security standards can be established. The company is moving to a model in which AI itself will help provide protection. Astra-related workloads remain suspended and will resume only after a complete transition to the new, stricter security standards.
A more detailed technical report on OpenAI’s current work is expected in the coming weeks.
Vendors
Openai
Hugging Face
Products
Astra
Sol