Models

Gemini 4 Argon: Google's New Frontier Model You Can't Use Yet

Google's first frontier model in seven months ties GPT-6 Astra on independent tests, hallucinates far less, and launches only to cyber defenders. Its long-term price matches Claude Opus 5.5.

Muhammet Fatih BatmanOctober 1, 20263 min read7 views
Gemini 4 Argon: Google's New Frontier Model You Can't Use Yet

Is Google back in the frontier race? On paper, yes. In practice, almost nobody can check.

Gemini 4 Argon, announced on September 30, is Google's first frontier model in more than seven months. It tops a long list of benchmarks in Google's own post. It also launches to a single audience: the "trusted cyber defenders" in Google's Fairwind program. Everyone else waits.

Who can use Gemini 4 Argon right now?

Only Fairwind partners. Google says broader access will come "as soon as possible" in phases, starting with paid API customers and Google AI Ultra subscribers, but gives no date. A pre-release process for governments is also under way. The restriction follows from what Google says the model can do: find, validate and patch critical software vulnerabilities on its own. A capability like that goes to defenders first.

Fairwind isn't new. It first showed up with Gemini 3.8 Flash Cyber in early September. Argon is the first flagship to ship through it.

Do the benchmarks hold up?

Partly. Google reports 77.9% on DeepSWE v1.1, a first-place 51.3% on AutomationBench, 91.7% on the LVBench long-video test, and leading scores on Vals finance and Harvey legal agent evaluations.

Independent numbers are cooler. Artificial Analysis, as reported by The Decoder, puts Argon at 53 on its Intelligence Index. That ties GPT-6 Astra and trails Claude Opus 5.5 at 58. It is still a 23-point jump over Gemini 3.1 Pro, which tells you how far behind Google had fallen.

The more interesting results are elsewhere:

  • It makes things up less. On Artificial Analysis's knowledge test, Argon hallucinates on 15% of the questions it can't answer. GPT-6 Astra: 51%.
  • It talks a lot. Argon averages 62,000 output tokens per task. GPT-6 Astra finishes in 27,000.

Is the price as low as it looks?

The introductory rate is $2 per million input tokens and $10 per million output, with a 95% discount on cached input. After the promotion, it doubles to $4 / $20, exactly what Anthropic charges for Claude Opus 5.5.

Now combine that with the token count. At the standard rate, 62,000 output tokens come to about $1.24 per task on output alone. A model that matches a rival's sticker price but uses more than twice the tokens to finish the job is not the cheaper model. Per-task cost is the number that matters, and nobody outside Fairwind can measure it on their own workload yet.

One spec deserves credit. Argon can return up to 1 million output tokens in a single response, up from 64,000. For migrating a large codebase or generating a long document, that removes a lot of stitching.

So what should a business do with this news?

Not much today, and that is fine. Prepare a small evaluation set drawn from your real tasks, so that when Argon reaches the paid API you can run it next to Opus 5.5 and GPT-6 Sol within a day. Measure tokens per task and the error rate on your own documents. And budget on the $4 / $20 price, not the promo.

A lower hallucination rate is the claim worth testing first. If it survives contact with your data, Argon's verbosity may be a price worth paying. If it doesn't, the headline benchmarks won't save it.

Sources: Google Blog, The Decoder, TechCrunch

Share This Article

Muhammet Fatih Batman

Written by

Muhammet Fatih Batman

Founder & Editor

Founder of YZ Uzman, with 20+ years of experience in web design and software development.

More news

Newsletter

Want to put this technology to work in your business?

Let's talk