Models

GPT-6 Astra Is Live: Pricing, Message Caps and the Numbers

GPT-6 Astra is on ChatGPT Pro and Enterprise and in the API at $10/$50 per million tokens. Message caps are half of Sol's, the benchmark table is mixed, and some numbers changed on launch day.

Muhammet Fatih BatmanSeptember 5, 20264 min read2 views
GPT-6 Astra Is Live: Pricing, Message Caps and the Numbers

$10 in, $50 out. That is the per-million-token API price of GPT-6 Astra, which OpenAI announced on September 3, and it works out to 2.5 times what GPT-5.6 Sol costs at its current promotional rate. The context window is 1 million tokens, maximum output 128,000, cached input $1. The model ships as gpt-6-astra in the API and is also listed on Microsoft Azure and Amazon Bedrock.

We covered Astra on September 2 as the first model to cross OpenAI's "Critical" cybersecurity threshold, based on the pre-launch post. This piece is about the launch itself: who gets it, how many messages, and which numbers hold up.

Who has access, and how many messages

According to The Decoder's September 5 roundup, Astra is now available on ChatGPT Pro, Enterprise and Business Premium; Plus and Business Standard follow "in the coming days." Free and Go tiers are excluded. Message caps sit at roughly half of Sol's: 5 to 45 on Plus (Sol: 10 to 100), 200 Astra Pro messages per week on the $200 Pro plan, 50 per week on the $100 Pro and Business Premium plans, 15 per month on Business Standard. Enterprise admins can keep the model switched off by default.

What the benchmarks show, and what they don't

OpenAI's own table puts the clearest gains in computer use and mathematics. OSWorld 2.0: 72.6% against Sol's 65.7%, with about 47% less time per task. FrontierMath Tier 4: 97.6%. ARC Prize's independent run on the ARC-AGI-3 semi-private set reports 99.9% with the provider adapter harness at a cost of $18,817, and 62.7% in the standard harness. On 96% of levels the model used fewer actions than the human baseline.

Read the other columns and the picture flattens. Artificial Analysis' intelligence index has Claude Fable 5.1 at 65.7, ahead of Astra at 61.2. On Humanity's Last Exam Astra scores 57.2% to Fable 5.1's 65.0%, a line the launch post did not highlight. Terminal-Bench 4.0 is a two-point gap (57.7 to 55.8). "Best across the board" depends on which board.

Numbers that moved during the day

Fortune compared archived snapshots of the launch post and documented several figures changing within hours. The hallucination rate was first published as 4.2% (Sol 12.2%), dropped to 2% (Sol 9.4%) at midday, then reverted. Fable 5.1's FrontierMath score went from 87.8% to 78% and then to 83%. The embargoed draft listed ARC-AGI-3 at 98.6%; the live post said 99.99%. OpenAI's statement: most evaluations carry a few points of noise depending on checkpoint, scaffold and run, and the numbers were corrected to reflect "our best estimate." The system card also arrived later than usual.

Two safety figures worth keeping

Per the system card, in the indirect prompt injection test (instructions hidden inside a document the model reads, 15 attempts per scenario) Astra failed 8.5% of the time, down from Sol's 27%; Claude Opus 5 sits at 4.8%. Against persistent multi-turn jailbreaks the defense rate is around 67%. The advanced cyber capabilities are gated: exploit development and similar tasks are reserved for vetted defenders in the Daybreak program and blocked in the general API.

On architecture, OpenAI confirmed Astra uses a looped transformer, applying the same forward pass several times before emitting output. Chief scientist Jakub Pachocki says the computation depth stays within a factor of two of GPT-4. Safety researchers' concern is that the loops create room for reasoning that never appears in the visible chain of thought.

The math for a team deciding this week

Our reading: the price matches Fable 5.1 and is 2.5 times Sol. For a pipeline processing 10,000 documents a day that is a 2.5x line item. The gains that could justify it live in computer-using agents and long context (96.3% versus 73.8% on the 1-million-token needle test), not in routine text work. Run a 100-sample A/B on your own workload in the API before the Plus rollout lands. And although the injection failure rate fell by two thirds, it is not zero: keep the human approval step in any agent that reads external documents.

Sources: OpenAI: GPT-6 Astra, ARC Prize: Astra results, The Decoder: plans and message caps, Fortune: shifting metrics, The Decoder: system card safety data, OpenRouter: pricing and context

Share This Article

Muhammet Fatih Batman

Written by

Muhammet Fatih Batman

Founder & Editor

Founder of YZ Uzman, with 20+ years of experience in web design and software development.

More news

Want to put this technology to work in your business?

Let's talk