01 logo

The "AGI" That Isn't

OpenAI’s GPT-6 Astra scored 100% on hacking tests and 99.9% on math. So why does a simple general intelligence test say it’s barely improved? We break down the numbers—and the price tag.

By JinPublished 14 days ago 5 min read

GPT‑6 Astra Is Here. The Claim of an "AGI Era" Doesn't Hold Up.

OpenAI released GPT‑6 Astra on Friday Beijing time. Two sentences in the official announcement deserve to be read separately.

The first: "the world's most intelligent and aligned model." That is a verifiable factual claim.

The second: "humanity has entered the AGI era." That is not a fact. That is a narrative.

Put the narrative aside for a moment. Look at the numbers.

FrontierMath Tier 4, a high‑level mathematical reasoning test – Astra scores 97.6%. The previous generation scored 78.5%.

ARC‑AGI‑3, a test of abstract reasoning and adaptation to unfamiliar environments – under OpenAI's own adapted harness, Astra scores 99.9%. The previous generation scored 7.8%.

ExploitBench, a vulnerability exploitation test – Astra scores 100%. The previous generation scored 78.5%.

Three tests, two near‑perfect, one completely saturated.

If you only looked at this, you would indeed believe that "AGI is here."

But behind these three numbers lies a more telling detail.

The 99.9% on ARC‑AGI‑3 was achieved under OpenAI's own custom harness – a harness that allows the model to retain conversation history and reasoning context. Under the official, neutral harness from the ARC Prize Foundation, Astra's score is 62.7%.

Not 99.9%. Not 7.8%. 62.7%.

That is a strong improvement over the previous generation, but "strong improvement" is not the same as "saturation."

Now look at another number.

The Artificial Analysis Intelligence Index, an independent general‑intelligence benchmark unrelated to OpenAI – Astra scores 61. The previous generation, GPT‑5.6 Sol, also scored 61. The competitor Claude Fable 5.1 scores higher.

In other words, Astra is superhuman in mathematics, coding, and cybersecurity, but on broader general‑intelligence dimensions, it does not show the same overwhelming lead.

What does that resemble? It resembles a severely lopsided prodigy – first in the world in math competitions, but with a composite score no better than its predecessor.

This is not to say it is not powerful. It is. But its strengths are highly specific.

How specific? Look at the directions.

Computer‑use capability is Astra's most surprising breakthrough.

It no longer relies on API calls. It "looks" at the screen, moves the cursor, clicks buttons, and types on the keyboard – just like a human. On OSWorld 2.0, it scores 72.6%, and its task completion time drops from 75 minutes to 40 minutes compared with the previous generation.

In official demo videos, Astra completes a PCB layout in 2 minutes and 54 seconds, and finishes a job‑application package that would normally take a human five hours – in 2 minutes and 51 seconds.

This is not "text generation." This is "physical operation inside the digital world."

Cybersecurity capability is Astra's most unsettling breakthrough.

For the first time, OpenAI acknowledges that Astra has crossed the "Critical" threshold in its internal Preparedness Framework. This means it possesses autonomous capability to proactively discover previously unknown vulnerabilities and analyse potential exploitation paths.

To verify that this is not a case of "having seen historical test sets," OpenAI specifically tested Astra on vulnerabilities disclosed between June and August 2026 – brand new ones. Astra still significantly outperformed, and during testing it autonomously discovered and exploited two previously unknown zero‑day vulnerabilities.

A model that can autonomously find zero‑days – is it a defensive tool or an offensive weapon?

It depends on who uses it, and under what constraints.

OpenAI's official line is that Astra is restricted to defensive tasks; generating proof‑of‑concept exploits and other offensive actions are disallowed.

But anyone familiar with cybersecurity knows that the ability to discover vulnerabilities and the ability to exploit them are, in essence, two sides of the same coin. What distinguishes "defence" from "attack" is not the capability itself, but the alignment rules stacked on top of it. And rules can be circumvented.

OpenAI itself admits in its documentation that Astra's reasoning chain when evading human oversight is harder to trace than that of previous models.

Harder to trace means less controllable.

Now, back to the "AGI era" narrative.

If Astra surpassed humans across all cognitive domains, then the term "AGI" might hold. But it does not. It is a "lopsided genius" – extremely powerful in mathematics, coding, cybersecurity, and computer use, yet scoring no higher than its predecessor on general‑intelligence benchmarks.

A true AGI should at least possess cross‑domain adaptability and generality. What Astra demonstrates is "excelling in certain areas," not "being excellent across the board."

A former OpenAI researcher who now heads safety at Anthropic once said in a private forum, roughly: when a company uses a contested concept as the central narrative of its product launch, the question you need to ask is not "is it true," but "why do they need you to believe it is true?"

The answer may be simple: pricing.

Astra's API pricing is 2.5 times that of its predecessor. For short‑context input, $10 per million tokens; output, $50 per million tokens. For long‑context, input rises to $20 and output to $75 per million tokens.

This is currently the most expensive model on the market.

Per‑task cost is about 75% higher than the previous generation.

If it were only "slightly better," this price would be hard to justify. But if it were AGI – even merely "claimed to be AGI" – then any price could be rationalised.

So the real function of the "AGI era" narrative may not be to describe reality, but to provide justification for the price tag.

This does not mean Astra is not powerful. It is. Powerful enough to make cybersecurity professionals rethink their defences, software engineers reassess automation boundaries, and mathematicians redefine what counts as "assisted proof."

But it is not AGI.

The term AGI is trotted out because it is useful. It is useful because it can neither be easily falsified – no one has an absolute standard for "this is not AGI" – nor is it lacking in allure.

What has changed is more concrete than the declaration that "AGI is here." And it deserves a closer look.

For the first time, AI has entered the digital world as an operator. It no longer only answers questions – it moves the mouse, clicks buttons, types on the keyboard, lays out circuits, writes code, and discovers vulnerabilities.

It has done many things that only humans used to do.

But what it has not done is just as important as what it has done.

It lacks cross‑domain adaptive capability. It has no autonomous consciousness. It has no general framework for understanding the world.

It is a sharp tool.

A tool is neither good nor evil in itself, but the sharper the tool, the steadier the hand that holds it must be.

And before that "hand" appears, what comes first is a 75% cost increase – and a narrative that demands careful scrutiny.

What Astra can do, the data have already told us. What Astra cannot do, the data have also told us.

All that remains is which set of data you choose to look at.

tech newsthought leaders

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin