01 logo

Google Just Released Gemini 4 Argon. It Is Not the GPT-6 Killer Everyone Expected.

Argon dominates legal, finance, long-context, and multimodal tests. It still trails GPT-6 Astra and Claude Opus 5.5 in from-scratch coding. Here is the full breakdown.

By JinPublished about 17 hours ago • Updated about 17 hours ago • 4 min read

On October 1, Google released Gemini 4 Argon.

Not Pro. Not Ultra. Argon. Argon is element 18, a noble gas. The name suggests Google is starting an element-themed naming run. Helium, neon, krypton, and xenon are waiting.

The name is not the important part. The release signal is. Google spent two years cycling through defeat, recovery, and overconfidence. Palm and Gemini 1 lost ground. Gemini 2.5 Pro recovered. Gemini 3 inflated benchmark scores while everyday use lagged. By Gemini 3.8 flash, Google had rebuilt a more cautious posture. Gemini 4 Argon is the first product of that posture.

Argon is a specialized release.

Knowledge Work: Where Argon Wins

Google says Argon beats GPT-6 on many tests. The wins cluster in knowledge work, long context, and multimodal understanding.

On the Harvey legal agent benchmark, Argon scored 19.6%. GPT-6 Astra scored 5.4%. Claude Opus 5.5 scored 3.8%.

On Vals financial agent v2, Argon scored 65.4%. Claude Opus 5.5 scored 58.6%. GPT-6 Astra scored 53.5%.

On GraphWalks, which tests 256K to 1M token context, Argon scored 84.2%.

On LVBench, a long-video test, Argon scored 91.7%. GPT-6 Astra scored 87.5%. Claude Opus 5.5 scored 83.7%.

On CWE-bench v1, a vulnerability repair test, Argon tied for first at 68%.

On the Gray Swan indirect prompt injection benchmark, Argon’s attack success rate was 0.7%.

Its hallucination rate was 15%. GPT-6 Astra’s was 51%.

On DeepSWE v1.1, Argon scored 77.9%, a new state of the art. Claude Opus 5.5 scored 74.2%. GPT-6 Astra scored 74.1%.

These numbers put Argon in the top tier for enterprise knowledge work, law, finance, long-document analysis, and multimodal tasks. In some niches, it is the best available model.

Coding: The Split

Coding is where Argon splits from the leaders.

On FrontierSWE v2, Argon scored 55.0%. GPT-6 Astra scored 65.5%. Claude Opus 5.5 scored 62.3%.

On Terminal-Bench 4.0, Argon came last.

On Code Arena: WebDev, it ranked eighth.

On Text Arena, it ranked first. That suggests strong general conversation, but weaker from-scratch project construction.

Bloomberg reported that some Google employees doubt Argon’s coding performance in practice. Benchmark scores do not equal productivity.

On the AA Intelligence Index, Argon scored 53. GPT-6 Astra also scored 53. Claude Opus 5.5 and Sonnet 5.5 scored higher.

At launch pricing, Argon costs about $1.99 per Intelligence Index task. GPT-6 Astra Max costs $3.26. The cost advantage comes from lower token prices, not lower token use. Argon generates about 62,000 output tokens per task. GPT-6 Astra generates about 27,000. Argon uses 2.3 times more. At standard pricing, the cost advantage shrinks.

Training: RL, RSI, and 1M Output

Google DeepMind researcher Tianfu Fu said much of Gemini 4’s improvement comes from scaling RL training. The training stack also added new methods. RSI is one key part. Earlier rumors about this were correct.

Argon supports 1M context. It also supports up to 1M output tokens. Most models cap output at 128K. A longer output ceiling lets the model think longer in one pass and handle longer, more complex tasks.

The tradeoff is token burn. If the model tends toward long thinking, everyday costs rise. Argon’s average 62,000 output tokens per task shows this.

Pricing and Launch: Sweet Start, Double Later

API pricing:

Launch discount: $2 per million input tokens, $10 per million output tokens, $0.1 per million cached input tokens.

Standard price: $4 per million input tokens, $20 per million output tokens.

Launch pricing matches GPT-6.1 Sol and Claude Sonnet 5.5. Cached input matches GPT-6.1 Sol. After the discount ends, prices double and the value case weakens.

Argon is not fully open. It is available to select trusted testers first. Paid API users and Google AI Ultra subscribers will get access next. The Fairwind program controls risk. Ultra priority drives subscriptions.

Internal doubt remains. Some Google employees say benchmark results and everyday use diverge. If users see that gap, the cautious recovery story could turn into overconfidence again.

Benchmark Fog: How Not to Get Fooled

When you read Argon benchmarks, watch for four traps.

First, when models cluster at the same score in a domain, suspect a ceiling. GPQA tops out around 90 to 94. SWEpro around 55 to 60. HLE around 50 to 60. DeepSWE sits near 74% for Astra, Opus, and Gemini 3.8 flash. That benchmark has lost signal. Terminal-Bench 4 shows similar wear, though scores below 50% still say something.

Second, treat working-style benchmarks with care unless they are OpenAI’s original GDPval. Many have messy judge details and environment issues. A workflow change can add 50 points. Search benchmarks are the worst offenders.

Third, LLM judges are fragile. Unless the judge is a top model, the benchmark needs scrutiny. Verification is often harder than generation.

Fourth, long-context tests matter. MRCR 8 needle 512K to 1M was the best, but it is saturated. GraphWalk is useful. AA-LCR is poor, with bad questions and judge abuse.

Against those rules, Argon looks like this: working ability near the frontier; coding in the first tier but behind Astra and Opus 5.5; research assistance much improved from earlier Gemini models; multimodal ability strong enough that benchmark scores are almost unnecessary.

The Bottom Line

Gemini 4 Argon is a specialized release. It wins in knowledge work, long context, and multimodal understanding. Legal, finance, security, and long-document teams should test it. Developers who need from-scratch coding should not treat it as the default.

At launch pricing, it is cheap to try. At standard pricing, the value depends on whether it can finish tasks in fewer tokens. It does not save tokens today.

Argon is a knowledge-work specialist. Google is back in the top tier. It still trails on from-scratch coding.

tech newsfact or fictionthought leadersapps

About the Creator

Jin

Writer of reamstories

https://reamstories.com/jin

Enjoyed the story? Support the Creator.

Subscribe for free to receive all their stories in your feed. You could also become a paid subscriber, letting them know you appreciate their work.

Subscribe For Free

Reader insights

Comments

There are no comments for this story

Be the first to respond and start the conversation.

Sign in to comment
    Written by Jin