From answering questions to operating software
OpenAI introduced GPT‑6 Astra on September 3. The company positions it as a single frontier model for browser and desktop operation, software engineering, professional artifacts and scientific work. Access starts with a limited set of organizations and is scheduled to expand to ChatGPT Plus, Pro, Business and Enterprise, the OpenAI API, Microsoft Azure and AWS Bedrock. The API model is gpt-6-astra, priced at $10 per million input tokens and $50 per million output tokens for Standard processing.
OpenAI reports 72.6% on OSWorld 2.0 and 57.9% on Terminal-Bench 4.0. It also presents retrieval results into the one-million-token range and an experimental Codex capability that can search context from before compaction. Together, these features show that frontier competition is shifting from chat quality toward operating interfaces, using tools and maintaining long-running task state.
Benchmark strength is not deployment assurance
Most scores are vendor-reported results from OpenAI's research environment or a specified harness. Different system prompts, permissions, networks and confirmation policies can change production outcomes. OpenAI also says Astra reaches the Critical cybersecurity threshold under its Preparedness Framework, which brings stronger safeguards and can pause or stop advanced requests.
Teams should test task completion, unintended actions, approval frequency, total cost and recoverability on their own workflows instead of adopting from a leaderboard alone. Least privilege, sandboxing, destination checks and human confirmation remain necessary wherever an agent can change external systems.