OpenAI Releases GPT-6 Astra for Coding and Computer Use
| Source: InfoQ AI/ML
Tags: OpenAI, GPT-6, Astra, computer use, AI agents, cybersecurity, coding AI, multimodal
OpenAI's GPT-6 Astra launches to ChatGPT and the API with 72.6% on OSWorld 2.0 computer-use tasks, 74.1% on DeepSWE coding, and 1M-token context — the first OpenAI model at the critical cybersecurity capability level, able to discover zero-day vulnerabilities and complete long-running agentic tasks across browsers, CRMs, and dev environments.
Details
GPT-6 Astra is OpenAI's most capable model to date, designed around agentic task completion rather than conversational response. It is rolling out to ChatGPT Plus, Pro, Business, and Enterprise users, plus the OpenAI API, Microsoft Azure, and AWS Bedrock — initially limited to select organizations. Benchmark results show clear progress across the board. Computer use hits 72.6% on OSWorld 2.0, up from 65.7% for GPT-5.6 Sol. Coding benchmarks reach 57.9% on Terminal-Bench 4.0 and 74.1% on DeepSWE v1.1. Long-context MRCR accuracy is 96.3% in the 512K-to-1M token range. Hallucination rates drop sharply — 4.2% for Astra versus 12.2% for GPT-5.6 Sol — a meaningful improvement for production factual tasks. A new Codex persistent memory feature lets the agent maintain notes across context windows instead of losing state during compaction. Earlier requirements, test results, and tool outputs remain searchable, enabling more reliable long-running coding workflows. The cybersecurity finding is the most consequential. Astra is the first OpenAI model classified at the critical cybersecurity capability level under the Preparedness Framework. In unguarded testing, it discovered two previously unknown vulnerabilities and built exploits against hardened browsers and operating systems. Production access restricts offensive use; defensive capabilities will be available through the Daybreak program. One concern: OpenAI found Astra's written reasoning harder to monitor than its predecessor — a transparency regression as capability grows that OpenAI acknowledges as an active research problem.