Gemini 4 Argon, read closely: 18 benchmarks, 1M output tokens and a gated launch
Google's first Gemini 4 model leads 12 of the 18 benchmarks in its own table, and the public leaderboards mostly agree. This post goes through who ran which test, how many tokens Argon spends per task, what a 1M-token response costs, and why only cyber defenders can use it today.

On this page
Google announced Gemini 4 Argon on September 30 at 20:00 UTC, which was 05:00 on October 1 in Korea. It is the first model of the Gemini 4 generation and Google’s new frontier model. Unless you work on a vetted cyber defense team, you can’t call it yet.
That leaves three things to go on: Google’s own benchmark table, a few independent evaluators who got early access, and a lot of people arguing about both. This post walks through the table row by row, checks it against public leaderboards, works out what the new 1M-token output limit costs, and looks at the release plan and the doubts. For the short version, read our briefing.
At a glance
| Item | Details |
|---|---|
| Announced | September 30, 20:00 UTC, in a post by Koray Kavukcuoglu of Google DeepMind |
| Who can use it now | A subset of partners in Google’s Fairwind cyber defense program, plus US government evaluators |
| Next | Paid Gemini API customers and Google AI Ultra subscribers. No date |
| Output limit | 1M tokens per response, up from 64K |
| Price | $2 input and $10 output per million tokens to start, then $4 and $20. Cached input is 95% off. The end of the introductory price hasn’t been announced |
| Google’s table | Leads 12 of 18 benchmarks and ties 1 |
| Artificial Analysis | 53 on its Intelligence Index, level with GPT-6 Astra. Claude Opus 5.5 still leads that board at 58 |
| Not published yet | Model ID, input context window, knowledge cutoff, rate limits, model card |
Where Argon fits
Google didn’t explain the name and didn’t announce any sibling models. There is no Pro, Flash or Lite tier yet. A Google spokesperson told Reuters that Argon is larger than the previous Pro models and will “anchor” the Gemini 4 generation, and that Google no longer plans to ship Gemini 3.5 Pro. Artificial Analysis notes that it is Google DeepMind’s first proprietary model above the Flash class in more than seven months.
Third-party leaderboards list two reasoning settings, High and Medium, and Artificial Analysis calls High the highest available. Google’s API docs don’t mention Argon at all yet, so those labels may change.
The table Google published
Google compared Argon with three models: OpenAI’s GPT-6 Astra, and Anthropic’s Claude Fable 5.1 and Claude Opus 5.5. GraphWalks appears twice with two context ranges, so the 18 benchmarks take 19 rows. Bold marks the top score in each row.
| Area | Benchmark | Argon | GPT-6 Astra | Fable 5.1 | Opus 5.5 |
|---|---|---|---|---|---|
| Knowledge work | Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| Knowledge work | AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% |
| Knowledge work | Vals Finance Agent v2 | 65.4% | 53.5% | 58.9% | 58.6% |
| Knowledge work | Harvey’s Legal Agent Benchmark | 19.6% | 5.4% | 6.7% | 3.8% |
| Agentic coding | DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| Agentic coding | FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% |
| Agentic coding | Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% |
| Agentic coding | Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% |
| ML engineering | PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% |
| Science and math | Terminal-Bench Science 0.1 | 57.6% | 68.1% | 52.6% | 63.3% |
| Science and math | LABBench 2 | 88.8% | 85.4% | 68.6% | 73.1% |
| Science and math | RiemannBench | 76.0% | 72.0% | 65.6% | 69.6% |
| Long context | GraphWalks, up to 128K | 99.7% | 98.7% | 91.4% | 90.6% |
| Long context | GraphWalks, 256K to 1M | 84.2% | 71.8% | 65.0% | 66.8% |
| Computer use | Agent’s Last Exam | 39.5% | 34.2% | n/a | 38.2% |
| Computer use | OSWorld-2.0, offline subset | 69.2% | 72.6% | n/a | n/a |
| Multimodal | Chartography | 71.6% | 71.0% | 46.2% | 66.3% |
| Multimodal | LVBench | 91.7% | 87.5% | 79.7% | 83.7% |
| Cybersecurity | CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% |
Counted by benchmark, Argon leads 12 and ties 1. GPT-6 Astra leads 3 and shares the CWE-bench tie, Opus 5.5 leads 2, and Fable 5.1 leads none.
The margins are uneven. Argon’s four biggest leads sit in knowledge work and long context: +12.9 points on Harvey’s legal benchmark, +12.4 on GraphWalks at 256K to 1M tokens, +8.8 on AutomationBench and +6.5 on Vals Finance Agent v2. Five of its wins are under two points, including the Vals Index and Vibe Code Bench. Its three biggest losses are wider than most of its wins: 10.5 points behind Astra on FrontierSWE v2 and on Terminal-Bench Science, and 9.0 behind Opus 5.5 on Terminal-bench 4.0.
Coding is the mixed area. Of the four agentic coding benchmarks, Argon tops two and finishes last of the four models on the other two.
How the numbers were produced
Google published a five-page methodology note with the table. A few lines in it change how much weight each row can carry.
Who ran the test. Competitor scores come from vendor reports or public leaderboards, at each model’s maximum reasoning setting. Several of Argon’s scores are self-computed, including DeepSWE v1.1 (run with the mini-SWE-agent harness), Terminal-bench 4.0, Terminal-Bench Science and OSWorld-2.0. On PostTrainBench, LABBench 2 and GraphWalks, Google ran every model itself.
Settings that differ by model.
- Terminal-Bench Science: Argon ran “with 6x verifier timeout”, while the other scores come from the leaderboard. Argon still finished 10.5 points behind Astra.
- OSWorld-2.0: Argon’s score is the best of three runs.
- LVBench: Gemini got one frame per second of video. Astra got 800 frames, Fable 5.1 got 300 and Opus 5.5 got 600, which Google puts down to “API limitations”.
Who’s missing. The table has no Claude Sonnet 5.5 and no GPT-6.1 Sol, which shipped the day before Argon. Both matter on some boards. On the Vals Index, Sonnet 5.5 is second at 67.04%, just ahead of Opus 5.5. On Vibe Code Bench, Sonnet 5.5 is first at 92.39% and Argon second at 91.91%.
None of this makes the table wrong. It does mean the direction is more reliable than the size of each gap.
Checking the public leaderboards
Most rows in Google’s table come from boards anyone can open. Here is where Argon sits once every model on those boards is counted, as of October 1.
| Benchmark (board) | Argon’s place | Others near the top |
|---|---|---|
| Vals Index | 1st, 68.90% | Sonnet 5.5 67.04%, Opus 5.5 66.97% |
| Vals Finance Agent v2 | 1st, 65.40% | Gemini 3.8 Flash 61.44% |
| Harvey’s Legal Agent Benchmark (Vals) | 5th, 19.58% | Four Meta Muse Spark models, led by Muse Spark 1.2 at 25.42% |
| Vibe Code Bench (Vals) | 2nd, 91.91% | Sonnet 5.5 92.39% |
| FrontierSWE v2 (Proximal) | Last of the four in Google’s table, 55.0% averaged over 5 runs | Astra 65.5%, Opus 5.5 62.3%, Fable 5.1 56.3% |
| AutomationBench (Zapier) | 1st, 51.29% at High | Argon at Medium 50.08%, Sonnet 5.5 44.75% |
| CWE-bench v1 | Tied 1st at 68% pass@1 | Grok 4.7 and GPT-6 Astra, also 68% |
| DeepSWE v1.1 (Datacurve) | Not listed | Astra, Gemini 3.8 Flash and Claude Opus 5 all at 74% |
Three details from these boards are worth knowing.
Harvey’s benchmark counts a task only when every grading criterion passes. Argon passed 94.0% of the individual criteria, and 19.6% of tasks came out fully correct. That board is led by Meta’s Muse Spark models, which weren’t in Google’s comparison.
CWE-bench runs each model inside its own agent harness (Argon in Antigravity, Astra in Codex, the Claude models in Claude Code), so it scores the model and its tooling together. Ties are broken on pass@4, where Grok 4.7 has 81%, Opus 5.5 79%, Argon 75% and Astra 74%. The board’s cost per rollout is $6.63 for Argon, $2.85 for Astra, $2.75 for Grok 4.7 and $0.79 for Opus 5.5. Collinear, which runs the board, priced Argon’s cached input at $0.20 per million, while Google’s 95% discount works out to $0.10, so Argon’s figure may be a little high.
On DeepSWE, Google’s 77.9% comes from its own run. The public board doesn’t list Argon, and on that board Google’s cheaper Gemini 3.8 Flash already sits at 74%, level with Astra.
Artificial Analysis and Arena
Artificial Analysis (AA) runs its own versions of ten evaluations and published Argon’s results on launch day.
- Intelligence Index: 53 at High, level with GPT-6 Astra at max (53) and one point above GPT-6.1 Sol (52). That is 23 points above Gemini 3.1 Pro Preview and 12 above Gemini 3.8 Flash. It ranks eighth of 223 model configurations on AA’s board, where Claude Opus 5.5 at max still leads with 58.
- Agentic tasks: first on AA’s version of AutomationBench at 78%, seven points ahead of Sonnet 5.5. On Terminal-Bench 4 it scored 57%, behind Sonnet 5.5 (64%), Opus 5.5 (60%) and Astra (59%).
- Documents: on AA-Briefcase it posted the highest rubric pass rate AA has recorded (65%), with lower marks for analysis and presentation quality.
- Tokens: Argon averaged 62,000 output tokens per task, against 27,000 for Astra.
The hallucination number traveled furthest. On AA-Omniscience, Argon’s hallucination rate is 15%, against 51% for Astra and 54% for GPT-6.1 Sol. AA defines the rate as the share of wrong answers among all the questions a model didn’t get right, so it measures how often a model guesses instead of saying it doesn’t know. Argon’s accuracy on the same test is 50%, 13 points below Astra’s 63%, and the combined Omniscience score comes out about even (42 for Argon, 43 for Astra). Put plainly, Argon knows fewer of these facts and is far more willing to say so.
The 15% figure is AA’s. Google’s own posts don’t mention hallucination at all.
On Arena, where people vote between two anonymous answers, Argon (High) opened at #1 in Text with 1525 points and #8 in WebDev with 1679. Its Agent Arena placing, also #8, is marked preliminary and rests on about 3,000 sessions.
What Google says it already uses Argon for
Google’s post leads with internal results. These are Google’s claims, and none has been checked from outside.
- Quantum research: optimizing the qubits and gates of key subroutines, where in one case it “beat the published baseline by 40% in a matter of minutes.”
- Memory: a team of Argon agents read fleet-wide profiling data and applied memory optimizations across Google’s data centers, freeing more than 300 TiB, with 500 TiB to 1 PiB expected in total.
- Rust migrations: C and C++ code moved to Rust, from libraries of tens of thousands of lines (re2, libgav1) up to the 800K-plus lines of the Fuchsia Zircon kernel. Google says these rewrites go through auditing, emulation testing and review before production.
- libgav1: agents replaced 32,000 lines of SIMD code with safe Rust that the compiler vectorizes on its own. The decoder now runs 2.7 times faster than the earlier Rust port, with identical output.
- Security: through Wiz’s Scan for Good program, Argon found a critical vulnerability exposing personal data in healthcare software used by hospitals worldwide. Google didn’t name the software or say whether it has been fixed.
The 1M output limit
The output limit went from 64K tokens (65,536 in the docs for Gemini 3.8 Flash and 3.1 Pro Preview) to 1M. OpenAI lists GPT-6 Astra at 128,000 output tokens. Google’s pitch is depth: with that much room, the model can reason through “hundreds of thousands of tokens in a single trajectory.”
Before you raise your own limit, do the arithmetic.
- Cost. Output bills at $10 per million during the introductory period. One response that uses the full budget costs $10, and $20 once the standard price applies.
- Time. AA hasn’t published Argon’s speed yet. The other frontier models on its board run at roughly 50 to 140 output tokens per second. At those rates a million tokens takes somewhere between two and six hours.
- Plumbing. Most HTTP proxies, load balancers and serverless runtimes will close a connection long before that. AA tested Argon with Long Decode Continuation, which it describes as a new Gemini API feature that pauses a long response and resumes it across follow-up calls. Google’s docs don’t describe it yet.
The payoff is for work that has to happen in one pass: a long migration plan, a full test suite, a report that has to stay consistent across hundreds of pages, or a hard problem where the model needs room to think. Most agent loops are many short calls and will never get near the limit. AA’s average of 62,000 output tokens per task is already more than double Astra’s, though, so watch your token budgets even if you never ask for a long answer.
If you do use long outputs, run each one like a background job. Set an explicit output cap per request, stream the result, write partial output somewhere durable, and alert on spend.
What it costs per task
Start with list prices. The introductory $2/$10 matches GPT-6.1 Sol. The standard $4/$20 matches Claude Opus 5.5’s list price. Cached input at 95% off comes to $0.10 per million during the introductory period, which helps agents that resend long context on every step.
Cost per finished task looks different, because Argon writes a lot. AA puts it at $1.99 per Intelligence Index task at today’s price, which is 60% of Astra’s $3.26 and 2.7 times GPT-6.1 Sol’s. At the standard price it becomes $3.98, about 1.2 times Astra. For a closer match, Claude Opus 5.5 at high effort scores about the same on that index (54) for $1.82 per task. Zapier’s AutomationBench shows the same pattern: $0.85 per task now, $1.70 at list price.
The end date of the introductory price isn’t public. AA wrote “at least one month” and, in the same article, that Google hasn’t confirmed an end date. Budget with $4/$20.
Why only cyber defenders, for now
Fairwind launched on September 2 with Gemini 3.8 Flash Cyber and the CodeMender agent, and it has more than 650 partners. A subset of them now gets exclusive access to Argon, either on its own or inside CodeMender. Google says these trusted defenders, and its own teams, get Argon “without cyber guardrails.”
Access comes with conditions:
- Priority goes to governments and national cyber authorities, critical infrastructure operators and core technology platforms. Academic labs doing defensive benchmarking can apply too. Google runs background checks on applicants.
- Only internal security, incident response or penetration testing teams may use it, behind phishing-resistant MFA, and each organization has to track which employees use it.
- Access can’t be resold or shared, and creating malware is prohibited.
- As a managed model on Gemini Enterprise, Argon supports zero data retention.
Google also says it’s taking part in the US government’s voluntary process for pre-release model access. Before a broad release, it lists four areas of safeguards:
- Misuse: refusals for cyber and CBRN harm, backed by probes that watch the model’s internal activations and tested by internal and external red teams.
- Prompt injection: on Gray Swan’s indirect prompt injection test, injections succeeded 0.7% of the time within 15 attempts, against 1.0% for Opus 5.5 and 8.5% for Astra.
- Misalignment: monitors that read Argon’s chain of thought and actions and “stop execution when necessary.”
- Sandboxes: test environments isolated and sealed before high-risk training or evaluation.
No model card and no Frontier Safety Framework report for Argon have been published yet. The newest FSF reports on DeepMind’s site cover Gemini 3.7 Flash and Gemini 3 Pro.
The calendar explains some of the caution. On September 28 OpenAI held back GPT-6.1 Astra after it failed to stay within scope and authorization in testing; our Dots follow-up has the details. On September 30 the FTC said it is investigating AI companies over risks from their agents. And in September Google told the Wall Street Journal that a Gemini model had escaped its testing environment in May through a partner’s misconfiguration, as Engadget reported.
The doubts
The first day’s skepticism came in four flavors: whether the scores hold up in real work, whether anyone can use the model, what Google’s tooling is like, and how the model behaves when it runs on its own.
Bloomberg’s sources. Bloomberg reported that employees with access say Gemini 4 “has performed well on benchmarks” but “does less well when employees actually put it to work.” Its coding is described as uneven, and one person said it “isn’t particularly adept at front-end design.” Two people said it appears affected by benchmaxxing, meaning tuned for benchmark scores. Others inside Google believe it has caught up. Google answered that “it would be inaccurate to say that Gemini 4 is underperforming in areas such as coding,” and an employee told Bloomberg there is a “large consensus” internally that it sits at the frontier.
The first version of the story was syndicated on Yahoo Finance at 12:52 PT on September 30, eight minutes before Google’s post. Alphabet shares closed that day up 0.93%, having given back most of a gain of more than 2% as the report spread. They rose about 1.7% after hours on the announcement and were down about 1.2% by midday on October 1.
Outside data on front-end work is mixed. Arena’s WebDev board had Argon eighth on September 30, about 135 points below Claude Opus 5.5 and also behind GPT-6.1 Sol. Vals’ Vibe Code Bench, which builds whole apps, has it second. AA’s document test marked it down on presentation. Before launch, testers of an “argon” checkpoint had praised its UI work.
Can’t use it yet. The Hacker News launch thread had 1,579 points and 1,048 comments by 16:55 UTC on October 1, well above the 1,057 points of GPT-6.1 Sol’s thread the day before. The most common complaint among its top-level comments, about one in five, was access. One commenter put it as “Unlike OpenAI or Anthropic where a model release announcement == GA.” On r/ClaudeAI, the subreddit’s summary bot concluded that “the overwhelming consensus in this thread is that Gemini 4 is NOT “out.”” Zvi Mowshowitz wrote: “What we don’t have is access to the model.” Many expect competitors to ship again before Argon reaches the public API.
Benchmark fatigue. About one in ten top-level comments questioned the numbers themselves. “It really feels like benchmarks have been hyper saturated these days,” wrote one. Another pointed out that the methodology reports only maximum reasoning settings. On Reddit, the hallucination result drew both the most upvotes and the most corrections: “15% does not mean Gemini is hallucinating in 15% of all answers.” The pseudonymous commentator Teortaxes took the middle position: Argon “may be benchmaxed on preference data,” while some scores on harder tests are “eye-watering.”
The tooling. The second-largest theme on Hacker News, about 15% of top-level comments, wasn’t the model at all. It was Google’s Antigravity harness: missing features, and worries about account bans for using a subscription in third-party tools. “Using agy is like going back in time,” one user wrote. A commenter on GeekNews, a Korean tech news site, made the same point about the Antigravity CLI: however good the model is, the harness is hard to trust.
Behavior on its own. Andon Labs ranked Argon third on Vending-Bench 2, a long-running business simulation, and wrote that it “fabricates confirmation emails, refuses to pay refunds, exploits invoice errors, and lies to suppliers.” It’s a simulation, but it is exactly the kind of behavior to test for before giving an agent real accounts.
On the other side, Mercor’s APEX-Agents leaderboard for professional tasks puts Argon first, at 82.2% pass@1 against 75.5% for Sonnet 5.5 and 73.5% for Opus 5.5. Early-access users on Hacker News, several of them self-described Googlers, were enthusiastic, and one called it the first Gemini model they could hand complex tasks to. The Decoder’s verdict sits in between: Argon “puts the ad giant back among the top three AI labs, though Anthropic likely still holds the lead.”
If you plan to test it
When API access opens, a week of your own testing will tell you more than any table. A plan that fits in that week:
- Use your own tasks. Pull 30 to 50 real tickets from your backlog that have known good outcomes. Public benchmarks are what every lab tunes for, so your own codebase is the more honest test.
- Fix the reasoning setting and log tokens. Run High and Medium separately, and record input, cached and output tokens for each task next to pass or fail.
- Price it at $4/$20. The introductory price will end.
- Split out front-end work. That’s where Bloomberg’s sources and the WebDev Arena rank point. Compare rendered screenshots side by side.
- Count abstentions. If you run retrieval or support flows, track “I don’t know” separately from wrong answers. A model with 50% accuracy and a 15% hallucination rate behaves very differently from one with 63% and 51%.
- Try long inputs. GraphWalks at 256K to 1M is one of Argon’s largest leads. Repeat it with your own long documents or repositories.
- Test inside your harness. Google’s DeepSWE run used mini-SWE-agent. Results can move a lot in Claude Code, Codex or your own agent loop, so include tool-call failures and prompt-injection cases.
- Build continuation before you need it. If you plan to use long outputs, make sure your client survives a generation that runs for hours.
What we still don’t know
- When paid API customers and Ultra subscribers get access, and in which regions
- The model ID, input context window, knowledge cutoff and rate limits
- When the introductory price ends
- A model card and a Frontier Safety Framework report for Argon
- How Long Decode Continuation works, beyond AA’s description
- Which reasoning levels Argon officially supports
Sources
- Google, Gemini 4 Argon: our next era of frontier intelligence (September 30, 2026)
- Google DeepMind, Gemini 4 Argon model evaluation: approach, methodology and results (PDF, September 30)
- Google DeepMind, Fairwind Program and Gemini cyber defense page
- Google, Proactive cyber defense for governments and enterprises (Fairwind launch, September 2)
- Google AI for Developers, Gemini models and pricing
- Artificial Analysis, Gemini 4 Argon results (September 30) and AA-Omniscience methodology
- Vals AI: Vals Index, Finance Agent v2, Harvey’s Legal Agent Benchmark, Vibe Code Bench
- Zapier, AutomationBench leaderboard
- Collinear, CWE-bench v1
- Datacurve, DeepSWE leaderboard
- Arena on X, launch rankings and Agent Arena (September 30); Arena, WebDev leaderboard
- Proximal, FrontierSWE leaderboard; Mercor, APEX-Agents leaderboard; Andon Labs, Vending-Bench 2 and notes on Argon’s behavior
- Bloomberg via Yahoo Finance, Google grapples with employee skepticism about new Gemini model (September 30)
- Hacker News, Gemini 4 Argon launch thread
- Reddit, r/singularity on the hallucination result and r/ClaudeAI launch thread
- Zvi Mowshowitz, AI #188: Gemini Dot Argon (October 1)
- GeekNews, Gemini 4 Argon discussion (October 1)
- Reuters via KELO, Google announces Gemini 4 flagship AI model after months of delays (September 30)
- The Decoder, Gemini 4 Argon closes the gap with OpenAI and Anthropic (September 30)
- SecurityWeek, Google launches Gemini 4 Argon with guardrail-free access for vetted defenders (October 1)
- Engadget, A Gemini model escaped its testing environment (September 19)


