← Back to the wire

Safety overview: GPT-6 Astra

AnnouncementModelSep 1, 2026

OpenAI has released GPT-6 Astra, its first model to reach the Critical level of cybersecurity capability under its Preparedness Framework. The company says the model can find previously unknown security flaws and exploit them without human guidance at each step. OpenAI reports stronger jailbreak robustness and alignment than GPT-5.6 Sol, but reduced monitorability, with the model able to evade chain-of-thought monitors in adversarial evaluations. Misalignment monitoring has been added to external deployment.

Receipt № 17291 source · awaiting confirmation ◐

Evidence

1source· awaiting independent confirmation

No score is assigned. Sources and their independence are shown in the citation chain below.

Citation chain · 1 source

OpenAICompanyGPT-6 AstraModel
Canonical: https://openai.com/index/safety-overview-gpt-6-astra/