Frontier model · visual briefing

OpenAI · 3 September 2026

GPT-6 Astra Is Here

OpenAI’s newest frontier model is built to carry difficult work from the first instruction to a finished result.

CodingResearchComputer use
Capability report02 / 10

One model, many tools

Built for End-to-End Work

Astra can move between code, browsers, scientific analysis and professional documents inside one multi-step workflow.

OpenAI describes these as integrated capabilities—not separate model variants.

Context window · API specification
1.05M

A Much Larger Working Memory

The API lists a 1,050,000-token context window and up to 128,000 output tokens.

Knowledge cutoff: 30 Apr 2026
API pricing04 / 10
$10 / $50

The API Price

$10 per million input tokens and $50 per million output tokens. Cached input costs $1 per million.

Prompts above 272K input tokens trigger higher long-context rates.

Agent workflow · new controls

Agent upgrade

Steer It While It Works

Async tools let independent work continue while a tool runs. Mid-turn steering can add corrections without restarting the whole response.

Tool calling requires OpenAI’s Responses API.

Model comparison06 / 10

Closest new rival

Astra vs. Claude Fable 5.1

GPT-6 Astra1.05Mcontext tokens
Fable 5.11Mcontext tokens

Both list 128K maximum output and the same $10/$50 headline API price.

Selected agent tests · coding

Benchmark report

Astra Takes a Narrow Coding Lead

Terminal-Bench 4.057.9%Astra · Fable 55.8%
TB Science 0.164.6%Astra · Fable 52.6%

OpenAI comparison, maximum effort; harness and safeguards affect scores.

Reasoning benchmarks08 / 10

But not every test

Fable Leads Broad Reasoning

Humanity’s Last Exam65.0%Fable · Astra 57.2%
AA Intelligence Index65.7Fable · Astra 61.2

Results with tools where stated. No benchmark proves a universal winner.

Safety context · cybersecurity

Important safety context

Critical Cyber Capability

OpenAI places Astra at its Critical cybersecurity threshold and adds monitoring that can pause agent work for review.

OpenAI also disclosed reduced monitorability in some adversarial evaluations.

Practical verdict10 / 10

Our assessment

Choose for the Work—not the Hype

AstraStart here for tool-heavy coding, computer use and active steering.
Fable 5.1Strong for broad reasoning and cache-heavy research.
Test both on your real tasks

Full comparison and sources on LinuxPanda