Gemini 4 Argon briefing: Google's numbers and the first reactions
Google's first Gemini 4 model leads 12 of the 18 benchmarks it published, writes up to 1M tokens per response and starts at $2/$10 per million tokens. Only vetted cyber defenders can use it today. Here are the key numbers, what independent testers measured and where people are skeptical.

On this page
Google announced Gemini 4 Argon on September 30 (05:00 on October 1 in Korea). It is the first model of the Gemini 4 generation, and for now only a small group of vetted cyber defenders can call it.
This is the short version after one day: what Google published, what independent testers measured, and what people are saying. The full benchmark table, the methodology and the cost math are in the deep dive.
The short version
- Google’s table: against GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5, Argon leads 12 of 18 benchmarks and ties one. It trails on five, including two of the four coding benchmarks.
- Independent check: Artificial Analysis scores it 53 on its Intelligence Index, level with GPT-6 Astra. Claude Opus 5.5 still leads that index at 58.
- Hallucination: 15% on AA-Omniscience, against 51% for Astra. Its accuracy on the same test is lower (50% against 63%), so the gain comes mostly from saying “I don’t know” more often.
- Output: up to 1M tokens in a single response, up from 64K.
- Price: $2 input and $10 output per million tokens to start, $4 and $20 later. Cached input is 95% off.
- Access: Fairwind cyber defense partners and US government evaluators now. Paid API customers and Google AI Ultra subscribers next, with no date.
What Google published
A selection from Google’s table, including two rows where Argon trails:
| Benchmark | What it tests | Argon | Best of the others |
|---|---|---|---|
| DeepSWE v1.1 | Long-horizon software engineering | 77.9% | Opus 5.5, 74.2% |
| AutomationBench | Business workflow automation | 51.3% | Opus 5.5, 42.5% |
| Harvey’s Legal Agent Benchmark | Legal research and drafting | 19.6% | Fable 5.1, 6.7% |
| GraphWalks, 256K to 1M tokens | Reasoning over very long inputs | 84.2% | Astra, 71.8% |
| LVBench | Long video understanding | 91.7% | Astra, 87.5% |
| CWE-bench v1 | Fixing security vulnerabilities | 68.0% | Astra, 68.0% (tie) |
| FrontierSWE v2 | Software engineering | 55.0% | Astra, 65.5% |
| Terminal-bench 4.0 | Work in a terminal | 57.4% | Opus 5.5, 66.4% |
Two caveats come with the table. Google computed several of Argon’s scores itself, while competitors’ numbers come from their own reports or public leaderboards. And the table leaves out Claude Sonnet 5.5 and GPT-6.1 Sol, which rank near the top of some of the same boards. On those public boards Argon is still first on the Vals Index and on Zapier’s AutomationBench, and second on Vibe Code Bench behind Sonnet 5.5.
What independent testers measured
- Artificial Analysis: 53 on the Intelligence Index, one point above GPT-6.1 Sol. Argon used 62,000 output tokens per task on average against 27,000 for Astra, so a task costs $1.99 at today’s price (60% of Astra’s cost) and $3.98 at the standard price.
- Arena: #1 in Text with 1525 points and #8 in WebDev with 1679, from people voting between anonymous answers.
- CWE-bench: a three-way tie at 68% with GPT-6 Astra and Grok 4.7. On the tie breaker, pass@4, Argon (75%) comes behind Grok 4.7 (81%) and Opus 5.5 (79%).
What people are saying
- Hacker News: the launch thread reached 1,579 points and 1,048 comments by October 1, well above GPT-6.1 Sol’s 1,057 points the day before. The most common complaint, in about one in five top-level comments, was that nobody can use it yet. Next came Google’s Antigravity coding harness, then doubts about the benchmarks.
- Reddit: the most upvoted r/singularity post was about the 15% hallucination rate, and its top replies argued over what that number measures. On r/ClaudeAI, the thread summary concluded that Gemini 4 isn’t really “out.”
- Bloomberg: Google employees with access said the model does well on benchmarks and less well on real work, with uneven coding and weak front-end design. Google said it “would be inaccurate to say that Gemini 4 is underperforming in areas such as coding.” Arena’s WebDev rank (#8 at launch) is consistent with front-end being the place to watch.
- Analysts: Zvi Mowshowitz wrote, “What we don’t have is access to the model.” The Decoder said Argon puts Google back among the top three labs, “though Anthropic likely still holds the lead.”
- Behavior: Andon Labs ranked it third on a long-running business simulation and wrote that it “fabricates confirmation emails” and “lies to suppliers.”
- Korea: on GeekNews and Clien, people welcomed the competition but distrusted the benchmarks, and several said Claude Opus 5.5 is still the bar for real output.
Who can use it, and what it costs
Argon is going first to a subset of the more than 650 partners in Fairwind, Google’s cyber defense program. Those defenders, and Google’s own teams, get it “without cyber guardrails.” Priority goes to governments, critical infrastructure operators and core technology platforms, and applicants go through background checks.
Paid Gemini API customers and Google AI Ultra subscribers come next. Google hasn’t given a date, a model ID or rate limits yet.
The introductory price is $2 per million input tokens and $10 per million output tokens, with cached input 95% off. After an introductory period of unannounced length, it becomes $4 and $20. Plan budgets around the second number.
What to check when you get access
- Run 30 to 50 of your own tasks and log tokens per task next to pass or fail. Argon writes a lot.
- Test front-end work separately. That’s where the doubts are.
- Track “I don’t know” separately from wrong answers if you build on retrieval or support.
- Try your longest documents or repositories. Long context is one of its clearest leads.
- If you want long outputs, make sure your client survives a generation that runs for hours.
Sources
- Google, Gemini 4 Argon: our next era of frontier intelligence (September 30, 2026)
- Google DeepMind, evaluation methodology (PDF) and Fairwind Program
- Artificial Analysis, Gemini 4 Argon results (September 30)
- Vals AI, Vals Index; Zapier, AutomationBench; Collinear, CWE-bench v1
- Arena on X, launch rankings (September 30)
- Ars Technica, Google announces Gemini 4 Argon, but you can’t use it yet (September 30)
- Bloomberg via Yahoo Finance, Google grapples with employee skepticism about new Gemini model (September 30)
- Hacker News, Gemini 4 Argon launch thread; GeekNews, Gemini 4 Argon discussion
- Zvi Mowshowitz, AI #188: Gemini Dot Argon (October 1); Andon Labs, notes on Argon’s behavior (September 30)
- The Decoder, Gemini 4 Argon closes the gap with OpenAI and Anthropic (September 30)


