Safety overview: GPT-6 Astra

| Source: OpenAI Blog

Tags: GPT-6, GPT-6 Astra, OpenAI, Preparedness Framework, cybersecurity, AI safety, model release

GPT-6 Astra — OpenAI's most capable broadly deployed model — becomes the first AI system to reach the 'Critical' cybersecurity capability threshold under OpenAI's Preparedness Framework, triggering mandatory safety mitigations before public rollout.

Details

OpenAI has published a safety overview for GPT-6 Astra, describing it as both their most capable broadly deployed model and the first to cross the Critical threshold for cybersecurity capability under the Preparedness Framework. This designation is significant: Critical is the highest risk tier in OpenAI's internal classification system, meaning the model was deemed capable enough in cybersecurity contexts to require special mitigations before deployment. The Preparedness Framework was designed as OpenAI's internal tripwire for when AI capability reaches levels that demand heightened scrutiny. Crossing the Critical bar in cyberoffense and cyberdefense suggests GPT-6 Astra can meaningfully assist with tasks like vulnerability discovery or exploit development at a level OpenAI considers potentially dangerous without controls. The 'broadly deployed' framing is crucial — this is not a research preview but a production rollout to users, with safety mitigations applied. This sets a precedent: GPT-6 is the first widely available model where OpenAI formally acknowledged Critical-tier risk and deployed it anyway, relying on mitigations rather than keeping it locked. How those mitigations work in practice, what security researchers can access via the API, and whether regulators will scrutinize the Critical deployment decision are open questions the source does not fully address.