Skip to content

GPT-6 Astra reviewed: what OpenAI's new model really does, what it costs and where it pays off

OpenAI calls GPT-6 Astra the start of the AGI era, the community sees an expensive specialist model. This review puts benchmarks, prices and weaknesses in perspective and shows what agents that work a computer on their own mean for visibility, tracking and ad budgets.

Elias Domig Elias Domig Managing Director 15 min read
GPT-6 Astra by OpenAI: spiral galaxy as the announcement key visual, with the wordmark GPT on the left and Astra on the right

OpenAI unveiled GPT-6 Astra on September 3, 2026 and chose big words to do it: the „most intelligent and aligned model in the world”, and according to Greg Brockman it is „not unreasonable” to speak of an AGI era. The community does a more sober calculation: 2.5 times more expensive than its predecessor, and almost level on general intelligence. Both views are right, because they measure different things. This review breaks down what Astra actually does, where the numbers impress, where they disappoint, and what a model that works a computer on its own concretely means for companies in the DACH region.

In short: a specialist for autonomous work, not a leap in thinking

If you take away one assessment, make it this: GPT-6 Astra is not a „smarter ChatGPT” but a model trained to finish tasks on its own. The benchmark gains sit almost entirely where the model operates a computer, researches, tests software or writes code across many steps. In broad thinking and knowledge, Astra stays at its predecessor’s level.

  • Agentic work: computer tasks in 40 instead of 75 minutes, 72.6 percent on OSWorld 2.0, 92.7 percent at operating screen interfaces.
  • General intelligence: 61.2 points on the Artificial Analysis Intelligence Index, with predecessor GPT-5.6 Sol at 60.9. Practically a tie.
  • Price: 10 / 50 US dollars per million tokens, exactly the level of Claude Fable 5 and roughly 2.5 times above GPT-5.6 Sol.
  • Cybersecurity: the first OpenAI rating of „Critical", with offensive capabilities gated behind an access programme.
  • Availability: paid ChatGPT plans and API (gpt-6-astra), no free access.

What is GPT-6 Astra?

GPT-6 Astra is the first model generation of the GPT-6 family. A „model” here is the trained AI system itself; „agentic” means it does not merely answer tasks but executes them in work steps of its own: it plans, clicks, types, checks intermediate results and corrects itself. That is exactly what Astra is built for.

The key facts at a glance:

PropertyGPT-6 Astra
ProviderOpenAI, unveiled September 3, 2026
VariantsAstra and Astra Pro; fast mode (double speed, double price)
Context windowapprox. 1,050,000 tokens, output up to 128,000 tokens
Modalitiestext and image as input, text as output
Reasoningadjustable thinking depth from low to max, asynchronous tool calls
Knowledge cutoffApril 2026
API price10 $ input / 50 $ output per million tokens; cached input 1 $; batch/flex 50 %
AvailabilityChatGPT Plus, Pro, Business, Enterprise; API gpt-6-astra; AWS Bedrock announced

A token is the billing unit of these models, roughly a word fragment; a million tokens correspond to several hundred pages of text. The context window of around one million tokens means Astra can keep a very large codebase, a complete document archive or a multi-hour work session in view in a single run. The knowledge cutoff of April 2026 is notable: everything newer, the model has to work out itself, via browsing. That exactly this browsing is among its strongest disciplines is no coincidence but the design decision of this generation.

The benchmarks: where Astra dominates and where it stands still

Benchmarks are standardised test tasks for comparing models. With Astra, a close look pays off, because the results diverge sharply: historic bests in the special disciplines, standstill in the broad ones.

Computer use: the actual statement

„Computer use” means the model sees a screen and operates it itself, like a person with mouse and keyboard. This is where the clearest progress sits:

BenchmarkWhat it measuresGPT-6 AstraGPT-5.6 Sol
OSWorld 2.0completing real computer tasks72.6 %65.7 %
ScreenSpot-Prohitting interface elements92.7 %76.9 %
Agents’ Last Examhard agent tasks59.353.6
BrowseCompautonomous web research91.5 %90.4 %
AutomationBenchautomating professional workflows41.4 %18.1 %

Add speed to that: for OSWorld tasks Astra needs around 40 minutes where its predecessor needed 75. The model works not only more accurately, it finishes tasks almost twice as fast. On AutomationBench, a benchmark for professional automation, Astra with 41.4 percent also leads Claude Fable 5.1 clearly (31.4 percent).

Mathematics and abstract reasoning: best scores with one outlier

On FrontierMath Tier 4, the hardest published mathematics tasks, Astra reaches 97.6 percent (Claude Fable 5.1: 87.8). On GPQA Diamond, a doctorate-level test, it scores 96.0 percent. The most striking value is ARC-AGI-3 at 99.9 percent; this benchmark measures abstract pattern reasoning, and Claude Opus 5 reaches 30.2 percent there for comparison. Such an extreme gap deserves an honest footnote, though: when a single model virtually solves a benchmark while the competition stays far below, the question always arises how strongly it was optimised for exactly this task.

Coding: solidly ahead, no runaway win

On Terminal-Bench 4.0 (working in the command line), Astra leads with 57.7 ahead of Claude Fable 5.1 (55.8) and Claude Opus 5 (52.3). On DeepSWE v1.1, a benchmark for realistic software tasks, it scores 74.1 versus 72.7 percent for its predecessor. Good values, but the field is tight in this discipline; anyone following the open-weight competition, such as Kimi K3 with rank 1 in the Code Arena, knows that coding leads currently change within weeks.

The counter-check: general intelligence stands still

Now the other side of the calculation. The Intelligence Index from Artificial Analysis, an independent composite across many reasoning benchmarks, puts Astra at 61.2 points; GPT-5.6 Sol sits at 60.9, Claude Fable 5.1 at 65.7. On Humanitys Last Exam, a very broad knowledge and reasoning test, Astra reaches 57.2 percent with tools, while Claude Fable 5.1 scores 65.0 percent there. And the BenchLM composite improves from 81.8 (Sol) to 81.9 points, at a doubled list price.

Remember: GPT-6 Astra is the clearest evidence yet that model development has shifted from „generally smarter” to „drastically better in special disciplines”. Anyone choosing models by a single intelligence score misses both: the real leaps and the real standstill.

Cybersecurity: „Critical” for the first time, and deliberately behind a gate

Astra is the first OpenAI model whose cyber capabilities were internally rated „Critical”. The numbers explain why: 100 percent on ExploitBench (Sol: 78.5; Claude Fable 5.1: 70), 42.4 percent on the harder ExploitGym (Sol: 30.3), and on the operations benchmark SRE-Bench Astra resolves 88 percent of incidents on the first attempt. During safety testing, the model found two previously unknown vulnerabilities, so-called zero-days.

OpenAI responds in two ways. First, the offensive capabilities are gated behind the „Daybreak” access programme that security researchers and companies have to unlock. Second, an additional misalignment monitor runs alongside that can pause or stop tasks; in practice, users report that this protective layer occasionally halts legitimate work as well. Remarkably honest is a third statement from OpenAIs own documentation: the „monitorability”, meaning how well the models reasoning steps can be followed from the outside, has decreased compared to its predecessor.

Against that stand two strong safety results that OpenAI documents with charts of its own. In the computer-use safety stress test, which measures how often an agent working a computer autonomously arrives at an undesired outcome, Astra sits at 2.4 percent; Claude Fable 5.1 sits at 9.5 and Claude Opus 5 at 11.5 percent.

OpenAI chart computer-use safety stress test: misaligned outcome rate, GPT-6 Astra 2.4 percent, Claude Fable 5.1 9.5 percent, Claude Opus 5 11.5 percent, lower is better
Misaligned outcome rate in the computer-use safety stress test, lower is better. Source: OpenAI.

And in the ExploitGym honeypot, a prepared trap that deliberately invites models to run an unauthorised exploit, Astra took the bait in 0.0 percent of cases, its predecessor GPT-5.6 Sol in 48.2 percent.

OpenAI chart ExploitGym honeypot: successful exploit rate in the honeypot trap, GPT-5.6 Sol 48.2 percent, GPT-6 Astra 0.0 percent, lower is better
Successful exploits in the honeypot trap, lower is better: Astra does not fall for it. Source: OpenAI.

Prices: 2.5 times more expensive, but cheaper per task?

On price, Astra slots in exactly at the level of Anthropics most expensive model:

ModelInput per M tokensOutput per M tokensContext
GPT-6 Astra (OpenAI)10 $50 $1M
Claude Fable 5 (Anthropic)10 $50 $1M
GPT-5.6 Sol (OpenAI, promotional price)4 $20 $1M
Claude Opus 4.8 (Anthropic)5 $25 $1M
Kimi K3 (Moonshot AI)3 $15 $1M
GLM 5.2 max (Z.ai)1.40 $4.40 $1M

Then come the side conditions, which in practice often matter more than the list price: cached input at 1 US dollar per million tokens, batch and flex processing at half price, the fast mode at double.

OpenAI counters the price criticism with a change of argument. Greg Brockman put it this way: pricing tokens makes little sense, what matters is the price per completed task. The logic is not wrong: if Astra completes a computer task in 40 instead of 75 minutes, falls into error loops less often and needs fewer retries, the total bill can drop despite higher token prices; trade press estimates put the cost per completed OSWorld task at roughly half the predecessors.

OpenAI backs this argument with a chart of its own that makes the relationship visible: on Terminal-Bench Science 0.1, Astra solves between 54 and 65 percent of tasks at an API budget of roughly 10 to 26 US dollars, while Claude Fable 5.1 needs considerably more budget for lower resolution rates.

OpenAI chart Terminal-Bench Science 0.1: resolution rate against API cost, GPT-6 Astra reaches 54 to 65 percent at 10 to 26 US dollars, above Claude Fable 5.1, GPT-5.6 Sol, Claude Fable 5 and Claude Opus 5
Resolution rate against API cost on Terminal-Bench Science 0.1: the GPT-6 Astra curve sits above the whole field. Source: OpenAI.

Important context: the chart comes from OpenAI itself and shows a benchmark that suits its own model. But it cleanly demonstrates the principle at stake: for agent tasks, the curve of cost and resolution rate decides, not the list price per token.

The community pushes back, also rightly: for everything that is not a long agent task, you simply pay 2.5 times more for practically the same thinking performance. Users of coding subscriptions also warn that the predecessor was already token-hungry and that Astra drains allowances correspondingly faster. Both calculations are correct; they just apply to different tasks. How to find out systematically which calculation applies to your own case is described in our guide on evaluating AI tools.

What this means in practice: observations from our work

At Dometrics we manage campaigns, shops and tracking setups across industries from hospitality to B2B manufacturing. From that perspective, five consequences of Astra are more interesting than any single benchmark score.

1. When agents do the research, machine readability decides visibility

Astras strongest discipline is autonomous browsing and task completion. That accelerates a shift already visible in search data: research and shortlisting is increasingly done by an AI system, no longer by a human with ten open tabs. A hotel guest has an assistant compare the trip, an industrial buyer has supplier shortlists compiled, a consumer has products pre-filtered. Whoever does not appear in these answers and selection processes does not exist for a growing part of the customer journey. The consequence is not new: structured data, clear entities, citable content, cleanly crawlable pages. What is new is the speed at which it becomes a precondition. That is exactly the field of GEO, which we describe in the service GEO & AI Search and in the guide Generative Engine Optimization.

2. Agent traffic puts tracking to the test

An agent that browses, fills in forms and submits inquiries leaves traces that confuse classic measurement: visits without origin information that land in GA4 as „Direct”, completed forms without human intent behind them, sessions with untypical behaviour. In our tracking audit guide we described that clicks from AI assistants are already a growing source of unattributable traffic; with more capable computer-use models, this share keeps growing. Practically, this means: the reconciliation between ad account and backend becomes even more important, lead validation belongs in every form setup, and your campaigns origin parameters must be in order so that human and machine traffic remain distinguishable.

3. Routine work in ad accounts becomes more automatable, control more important

The AutomationBench jump (41.4 versus 18.1 percent for the predecessor) concerns exactly the kind of work that campaign operations largely consist of: gathering data, filling reports, checking settings, working through lists. That models increasingly handle such processes themselves is good news for advertisers, but it shifts the human work to where it belongs anyway: hypotheses, budget logic, quality control. Anyone letting agents work in ad accounts needs the same discipline as with automated bidding strategies; an algorithm optimising on wrong numbers does not get better through more autonomy, it gets wrong faster.

4. Model choice per task finally becomes a question of economics

Astra at Fable-5 price, its predecessor at a promotional price, Kimi K3 and GLM 5.2 far below: the price range between frontier models now spans more than tenfold, with completely task-dependent value in return. A hypothetical calculation makes the logic tangible: suppose a recurring research task consumes 200,000 tokens of output. With GLM 5.2 the run then costs around 0.88 US dollars, with Astra around 10 US dollars. If the cheap model needs three attempts plus rework on average while Astra reliably finishes in one run, the calculation flips; if the cheap model completes the task just as reliably, the Astra deployment simply burns budget. The numbers are freely chosen assumptions, the principle is not: it is not the token price that decides but the price per reliably completed task, measured on your own cases.

5. Governance is no longer a corporate exercise

A model with a „Critical” cyber rating, decreased traceability of its reasoning steps and a protective layer that sometimes wrongly stops tasks does not belong unmanaged in business processes. For industries with confidentiality obligations, such as legal and finance, this applies twice over: clear rules on which data an agent may see, which systems it may operate and where a human signs off. This is not a question of company size; an SME whose agent works autonomously in the shop backend carries the same risk as a corporation, just without its control apparatus.

How we test new models before they are allowed into client projects

Finally, the part that transfers to any organisation: a sober test protocol instead of gut feeling. With us, every new model goes through the same procedure, and only afterwards comes the decision whether and for what it gets used.

  • Your own task catalogue instead of third-party benchmarks: ten to twenty real, recurring tasks from your own day-to-day, with a documented target result. External benchmarks pre-sort; only your own cases get to decide.
  • Measure cost per completed task, not per token: per run, note token consumption, number of attempts and correction effort. Only that yields the honest comparison price.
  • Plan for batch and cache from the start: with Astra, batch processing halves the price, cached context costs a tenth of the input rate. Recurring tasks with constant context benefit the most.
  • Actively test the limits: deliberately include tasks that overwhelm the model or trigger the safeguards, to know the abort behaviour before it surprises you in production.
  • Do not migrate, complement: a new model rarely replaces the entire stack. It takes over the tasks where it is measurably better or cheaper; the rest stays where it is.

For Astra, as things stand, that means: a candidate for long, multi-step agent tasks and automated workflows; not a candidate for standard text work, where the price gap to the open models remains unjustifiable.

TL;DR

  • Astra is a specialist: historic bests in computer use, agent work, mathematics and cybersecurity, but practically no improvement in general intelligence over GPT-5.6 Sol.
  • The price is a statement: 10 / 50 US dollars per million tokens, exactly Fable-5 level and 2.5 times the predecessor. The bill only works out for tasks where Astras reliability saves attempts and working time.
  • Cybersecurity as a turning point: the first „Critical" rating, two zero-days found, access programme and monitoring layer included. Governance becomes mandatory, in mid-sized companies too.
  • For visibility, machine readability now counts: the more agents research and pre-select, the more structured, citable content decides who still appears in the customer journey at all.
  • For tracking, heightened vigilance applies: agent traffic without origin and machine-completed forms make the reconciliation between ad account and backend more important than ever.
  • Model choice remains a per-task calculation: between GLM 5.2, Kimi K3, the Claude models and Astra lies more than a factor of ten in price; which model pays off is decided by your own test cases, not by the announcement.

Keep reading: Kimi K3 reviewed · Evaluating AI tools systematically · Generative Engine Optimization

The service: GEO & AI Search at Dometrics

Look up terms: GEO · SEO · Server-side tracking · Attribution · all terms in the glossary

Frequently asked questions

What is GPT-6 Astra?

GPT-6 Astra is the first model generation of the GPT-6 family from OpenAI, unveiled on September 3, 2026. The model processes around 1.05 million tokens of context, understands text and images and is specialised in agentic work, meaning tasks it carries out on a computer by itself: filling in forms, researching, operating and testing software.

What does GPT-6 Astra cost?

Via the API, Astra costs 10 US dollars per million input tokens and 50 US dollars per million output tokens. Cached input is 1 US dollar, batch and flex processing cost half the standard rate, and the fast mode costs double. That makes Astra roughly 2.5 times more expensive than its predecessor GPT-5.6 Sol at the current promotional price.

Is GPT-6 Astra better than GPT-5.6 Sol?

For agent tasks, clearly: on OSWorld 2.0, Astra completes computer tasks in around 40 instead of 75 minutes and reaches 72.6 instead of 65.7 percent. For general intelligence, barely: on the Artificial Analysis Intelligence Index both are almost level (61.2 versus 60.9 points). Astra is a specialist for autonomous work, not a big leap in thinking.

How does GPT-6 Astra compare to Claude Fable 5?

Mixed. In agent and automation benchmarks such as AutomationBench (41.4 versus 31.4 percent) and in mathematics (FrontierMath Tier 4: 97.6 versus 87.8 percent) Astra leads. On the broad knowledge benchmark Humanitys Last Exam, Claude Fable 5.1 leads with 65.0 versus 57.2 percent, as it does on the Artificial Analysis Intelligence Index (65.7 versus 61.2).

What context window does GPT-6 Astra have?

Around 1.05 million tokens of context with up to 128,000 tokens of output. That is enough for very large codebases, long document collections or multi-hour agent work sessions in a single run. The model knowledge extends to April 2026.

What does computer use mean for GPT-6 Astra?

The model can see, understand and operate a screen itself: it clicks, types, fills in forms, updates CRM entries, researches in the browser and tests software. On the ScreenSpot-Pro benchmark, which measures exactly this operating of interfaces, Astra reaches 92.7 percent compared to 76.9 percent for its predecessor.

Why does OpenAI rate GPT-6 Astra as critical for cybersecurity?

Astra is the first OpenAI model with the rating „Critical" for cyber capabilities. On the ExploitBench benchmark it reaches 100 percent, and during testing the model found two previously unknown security vulnerabilities. Its offensive capabilities are therefore gated behind the Daybreak access programme and monitored by additional safeguards.

Who is GPT-6 Astra worth it for?

For tasks that run long and have many steps: multi-stage research, software testing, data maintenance in systems, complex coding tasks. There, the faster and more reliable execution often offsets the higher token price. For simple questions and standard text work, cheaper models remain the more economical choice.

Where is GPT-6 Astra available?

The rollout started on September 3, 2026, initially for selected organisations, then for paying ChatGPT users (Plus, Pro, Business, Enterprise) and via the API under the model ID gpt-6-astra. There is no free access. Alongside Astra there is the Astra Pro variant; AWS Bedrock has been announced.

What does GPT-6 Astra mean for marketing and visibility?

The better models browse and complete tasks on their own, the more often an agent researches and compares instead of a human. What stays visible is content that is machine-readable, structured and citable. At the same time, this agent traffic often shows up in tracking without an origin, which changes how channels are measured and evaluated.

Sources

#GPT-6 Astra#OpenAI#LLM#AI Agents#Computer Use#Model Choice

Clarity over guesswork.

A free initial analysis shows where a marketing budget holds potential and where it does not. Data-based, without obligation, with an honest assessment instead of a sales pitch.

Start the free analysis

A conversation starts with a question.

No sales pitch, no commitment. Dometrics responds with an honest assessment and says so when there is no leverage to be found.

  • Reply within 24 hours
  • No obligation. The data stays with Dometrics
  • A clear assessment instead of a sales pitch
info@dometrics.at Stuckgasse 1/10, 1070 Wien
What is it about?

By sending, the person consents to processing in line with the privacy policy .