Gemini 4 Argon is now the highest-scoring model on the independent Vals Index, and it is cheaper per task than both Claude models it displaced. But most enterprises cannot buy it, and the price Google is advertising is not the price to budget. Vals scored Argon at 68.90%, #1 of 41 models, at $15.68 per test. Access today is limited to cyber defenders in Google's Fairwind Program and trusted testers, with no general-availability date. If you are mid-way through a GPT-6.1 Sol or Claude Sonnet 5.5 migration, do not pause it. Add Argon to your eval set, budget it at the $4/$20 standard rate, and decide when you can actually call it.
That is the short version. The longer one matters, because Argon's launch packs three separate claims — a benchmark lead, a price, and an availability promise — and each one holds up differently.
What Did Google Actually Ship on September 30?
Google shipped a model that very few enterprises can use yet. The Google announcement says Argon is "rolling out" to trusted cyber defenders through the Fairwind Program, with broader access to developers, enterprises and consumers "as soon as possible," starting with paid API customers and Google AI Ultra subscribers. No date is attached to "as soon as possible." Google also says the model is part of the U.S. government's voluntary pre-release access process — one more step before a wide release.
The pitch is aimed squarely at enterprise work. Google positions Argon for complex software engineering, legal and finance knowledge work, and cyber defense, and reports 77.9% on DeepSWE v1.1, 51.3% on AutomationBench and a tie for first at 68% on CWE-bench v1. Those are vendor-reported numbers. Treat them as such.
The one spec that will change how people build is output length. Google says Argon can produce up to 1 million output tokens, up from a previous 64K limit. Vals, however, lists Argon's maximum output at 262k tokens in the configuration it tested. Those two figures do not agree, and you should not design a pipeline around the larger one until your own API calls return it.
Is the Vals Lead Real, or Noise?
The lead is real but narrow — about two points over the next model. Vals puts Argon at 68.90% ±0.97. Claude Sonnet 5.5 scores 67.04% and Claude Opus 5.5 scores 66.97%. GPT-6.1 Sol, released the day before, scores 61.15% and ranks eighth.
A two-point gap on an aggregate index is the kind of gap that disappears on your workload. We covered this when only 3 of 36 published model gaps survived a proper sample-size check. The Vals Index is a composite. Your agent runs one or two of its tasks, not all of them.
Where Argon's lead is wide is the more useful signal. On Vals' sub-benchmarks, Argon scores 77.86% on CyberBench (#2) and 65.40% on Finance Agent v2 (#1). Sonnet 5.5 scores 59.58% on CyberBench, ranked #32. An 18-point gap on security work is not noise. If your shortlist is for a security-engineering or SOC workload, Argon belongs in the bake-off the moment you can get a key.
Steel-man the other side: on coding, the Claude models are not behind. Sonnet 5.5 is #1 on Vibe Code Bench at 92.39%, against Argon's 91.91% at #2. For a coding-agent fleet, Argon is a peer, not an upgrade.
Which Price Should Go in the Budget?
The $4/$20 standard rate — and Vals already used it. Google's launch price is $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after the introductory period, with cached input at a 95% discount. Google has not said when the introductory period ends.
Here is the detail most launch coverage misses: the Vals page lists Argon at $4.00 input / $20.00 output. The $15.68-per-test figure is the standard-price figure. That is good news for planning. The number you can defend to finance already assumes the price doubles.
It also tells you something about token efficiency. Sonnet 5.5 is priced at $2/$10 and costs $21.34 per Vals test. Argon, at twice that list price, comes in about 27% cheaper per test. That implies Argon spends far fewer tokens to finish the same tasks, though Vals does not publish the token breakdown, so caching and the input/output mix may account for part of the gap. Opus 5.5, at $4/$20, costs $32.14 per test — roughly double Argon at the same list price.
This is the lesson we keep relearning: per-token price is the wrong comparison, and Sonnet 5.5 can cost more than Opus 5.5 at a matched score. Cost per completed task is the only number that survives contact with a production workload.
The cheaper option is still not Argon. GPT-6.1 Sol costs $3.24 per Vals test — about a fifth of Argon — for a score 7.75 points lower. Whether those points are worth nearly 5x the spend depends on the task, which is the same routing question we raised when GPT-6.1 Sol's $5.47 task scored 11 points below Astra's.
Who Can Actually Use Argon Today?
Almost nobody outside a security team. Fairwind launched on September 2, 2026 with more than 650 participating partners, and those participants must limit access to internal cybersecurity, incident response or penetration-testing staff, behind multi-factor authentication. According to an analysis of the program's terms, eligibility covers governments, critical-infrastructure operators and core technology platforms, and partners may not share, redistribute or resell access.
Read that carefully if your company is a Fairwind member. Membership does not mean your legal or finance team gets Argon. The program's operating standard confines the model to defenders. The enterprise knowledge-work use cases Google is advertising — the ones behind the Finance Agent and LegalBench scores — are exactly the ones Fairwind members are not cleared to run.
As of October 1, 2026, Argon does not appear on the public Gemini API pricing page, which still lists Gemini 3.1 Pro Preview as the latest Pro model. A model you cannot put on a purchase order is a roadmap item, not a procurement option.
What Should You Do With a Migration Already in Flight?
Keep moving. A frontier model with no GA date is not a reason to stop a migration that has a date. The pattern is familiar: a new leader appears, a team pauses "to see," and three months later it is evaluating the next leader instead of running the last one. Argon may be generally available next week or next quarter. Google has not said.
This Week:
- Add Argon to your eval harness as a placeholder row, with the $4/$20 price pre-filled, so the moment API access opens you can run it against the same task set as Gemini, Claude and GPT-6.1 Sol in one afternoon.
- Write down which of your workloads map to Argon's wide leads — security engineering (CyberBench) and finance agents (Finance Agent v2) — versus its narrow one (coding). Only the first list justifies priority access.
This Month:
- If you are a Fairwind member, route Argon to your security team only, and log it as a separately governed model in your AI inventory. The program terms restrict who may use it; your access controls should match.
- Ask your Google account team for the introductory-period end date in writing, and for any committed-use pricing above the $4/$20 list. Do not sign anything that assumes $2/$10 persists.
- Test the output ceiling yourself. Google says 1M output tokens; Vals tested at 262k. If a design depends on very long single generations, verify the limit your project actually gets.
Before Renewal:
- Keep your current model contract flexible enough to add a second provider. If Argon lands at GA with its Vals profile intact, the security and finance workloads are the ones you will want to move, and a single-vendor commit will cost you the option.
The Bottom Line
Gemini list prices have moved before — we covered it when Google tripled a Gemini price and still undercut OpenAI — so an introductory rate with no end date is a number to plan around, not to plan on. What is different this time is that an independent evaluator ran the numbers at the standard price, and Argon still came out ahead of both Claude models on cost per task. That is a real result.
But a benchmark lead is not a delivery date. Budget at $4/$20, eval it the day you can, and keep shipping what you can buy.
Continue Reading
- GPT-6.1 Sol's $5.47 Task Scores 11 Points Below Astra's $23.80
- Claude Sonnet 5.5 Costs More Than Opus 5.5 at the Same Score
- Only 3 of 36 Model Gaps Were Real. Size Your Eval Set.
- Claude vs GPT vs Gemini: Stop Comparing Per-Token Prices
- 55 Zero-Days in 2 Hours. Google's $32B Security Bet Went Live.
- GPT-6 Sol Halves the Token Price but Benchmarks It at Top Effort
