Gemini 4 Argon, reviewed for teams that run agents: what you can decide today
Three days after launch, a second look at Gemini 4 Argon from an operator's seat. Where it leads and trails by type of work, what a task really costs, what a 1M-token output and a guardrail-free tier mean for how you run agents, and what you can prepare before the API opens.

On this page
Gemini 4 Argon has been out for three days. The numbers are in our briefing, and the benchmark table, row by row with its test conditions, is in the deep dive. This post starts from a different question. If your team runs agents, what can you decide about this model now, and what should you leave open?
We won’t re-read the scorecard. Instead, we regroup the public leaderboard numbers by the kind of work you would hand the model, convert tokens into the cost of one task, look at what the 1M-token output and the “no guardrails” tier mean for how a platform is built, and leave a list of things you can finish today, before the API opens.
Three lines
- The model looks like a frontier model. It leads at reading long inputs, planning and producing documents. It trails at working by hand in a terminal and at operating a screen. Coding sits in between.
- Nothing can change yet. As of October 3, Argon appears nowhere in Google’s model list or price list. There is still no model ID, no end date for the introductory price and no model card.
- The preparation can be done today. Your own task set, tokens logged per task, an output cap per route, prompts shaped for the cache, and routing by type of work can all be built with the models you use now, so you can compare on the day Argon opens.
What changed in three days, and what didn’t
What changed. Google has now answered the Bloomberg report that landed minutes before the launch, in which employees said the model does well on benchmarks and less well on real work. As 9to5Google reported on October 1, Google’s position is that “it would be inaccurate to say that Gemini 4 is underperforming in areas such as coding,” and one employee described a “large consensus” inside the company that it is at the frontier. The Implicator quoted Tulsee Doshi, who leads Gemini products, more specifically: Googlers have put the model through its paces in recent weeks, and many rely on it for their hardest coding and research problems. Both sides can be right. It depends on which kind of coding, and the scorecard below shows where the split is.
Vals updated its public leaderboard on October 2. Argon kept first place on the Vals Index, and this update lists cost per task and latency next to the scores. The cost section of this post starts from that table.
What didn’t. We checked Google AI for Developers’ model page and pricing page again on October 3. Neither lists Argon. The most expensive Gemini on the price list is still Gemini 3.1 Pro Preview, and the model list runs from Gemini 3.8 Flash down to the 2.5 series. There is still no date for the introductory $2/$10 turning into $4/$20. Paid API customers and Google AI Ultra subscribers are next, as they were three days ago, and as three days ago there is no date.
The scorecard, regrouped by the work you’d hand it
Google’s table is ordered by benchmark name. For someone running agents, it is more useful to group it by “what would I give this model to do.” Below are the Vals public leaderboard (October 2) and Google’s table rearranged that way. Ranks are positions among every model on that board.
| Work | Benchmark | Argon | Rank | Reading |
|---|---|---|---|---|
| Reading long documents and repositories | GraphWalks, 256K to 1M tokens (Google’s table) | 84.2% | 1st of the 4 compared, 12.4 points ahead | Clearest strength |
| Finance and tax documents | Vals Finance Agent v2 / Tax Agent Bench | 65.40% / 76.23% | 1st of 74 / 2nd of 65 | Leads |
| Legal research | LegalBench / Legal Research Bench | 88.30% / 54.81% | 3rd of 149 / 4th of 73 | Near the top |
| Business workflow automation | AutomationBench (Zapier) | 51.3% | 1st | Leads |
| Building one app | Vibe Code Bench | 91.91% | 2nd of 108 | Just below Sonnet 5.5 |
| Code migration | Code Migration | 68.17% | 2nd of 73 | Near the top. $57.82 per task |
| Terminal work | Terminal-Bench 4.0 | 57.58% | 5th of 43 | 9 points behind Opus 5.5 |
| Algorithm problems | IOI | 100% | 1st of 40 | Perfect score |
| Incident response | SRE Bench | 44.27% | 3rd of 16 | Near the top, under half |
| Operating a screen | CUA-bench / OSWorld-2.0 (Google’s table) | 4.83% / 69.2% | 7th of 8 / behind Astra’s 72.6% | Clearest weakness |
| Finding and fixing vulnerabilities | CyberBench v1.1 / CWE-bench v1 | 77.86% / 68% | 2nd of 44 / tied 1st | Leading group |
| Clinical notes | MedScribe | 87.43% | 15th of 106 | Average |
There is a pattern. It is strong at reading, planning and producing documents, and weak at doing things by hand inside an environment. Long context, finance, legal and automation are all “read a lot, produce a long answer” work, and Argon leads there. Terminal work, screen operation and incident response are “act once, look at the result, decide the next step” work, and there Claude Opus 5.5 or GPT-6 Astra is ahead, or Argon sits near the bottom. Its 4.83% on CUA-bench is seventh of the eight models on that board, and one task cost $193.78 to run.
This is why Bloomberg’s sources and Google’s rebuttal can both be right. “The hardest research problems” are what Argon does well. Opening a terminal and clicking through a screen to fix a front end is not, yet.
For an agent platform, the pattern translates into a division of roles. There is no reason to give the step that reads a long brief and makes a plan and the step that carries that plan out inside an environment to the same model. That was already the point of our multi-agent playbook. Argon’s scorecard draws that line unusually clearly.
Count tasks, not tokens
On list price alone, Argon matches GPT-6.1 Sol at the introductory rate and Claude Opus 5.5 at the standard rate. But models spend different numbers of tokens on the same task, so the cost of one task looks different.
The Vals Index combines eight benchmarks across finance, coding, legal and tax, and its table lists cost per task and latency. Vals priced Argon at the standard $4/$20.
| Model | Vals Index | Cost per task | Latency as listed |
|---|---|---|---|
| Gemini 4 Argon | 68.90% | $15.68 | 46 min 33 s |
| Claude Sonnet 5.5 | 67.04% | $21.34 | 1 h 18 min |
| Claude Opus 5.5 | 66.97% | $32.14 | 1 h 19 min |
| Claude Fable 5.1 | 65.83% | $28.71 | 1 h 17 min |
| GPT-6 Astra | 63.13% | $18.46 | 27 min 48 s |
| GPT-6.1 Sol | 61.15% | $3.24 | 43 min 10 s |
Two things to take from this. Even at the standard price, Argon’s cost per task is below all three Claude models. And GPT-6.1 Sol scores 7.75 points lower for a fifth of the cost. There is no reason to send every task to the best model, and where you send each task decides most of the bill.
Artificial Analysis adds one more fact. Argon averages 62,000 output tokens per task against 27,000 for GPT-6 Astra, which is why, at the same Intelligence Index score of 53, a task costs $1.99 at the introductory price and $3.98 at the standard one. Output is the center of the cost, and that leads to the output caps in the next section.
One agent loop, worked out
In real operation, the cache decides the bill. An agent loop resends the same system prompt, the same tool schemas and the same repository summary at every step. Argon’s cached input is 95% off: $0.10 per million tokens at the introductory price and $0.20 at the standard one.
The assumptions: a loop of 40 steps, a stable prefix of 30,000 tokens in front of every step, 2,000 new input tokens and 3,000 output tokens per step. Growing conversation history is left out to keep the arithmetic simple.
| Case | Input | Output | Total |
|---|---|---|---|
| Introductory price, no cache | $2.56 | $1.20 | $3.76 |
| Introductory price, prefix cached | $0.28 | $1.20 | $1.48 |
| Standard price, no cache | $5.12 | $2.40 | $7.52 |
| Standard price, prefix cached | $0.56 | $2.40 | $2.96 |
The cache cuts input cost to about a tenth, and what remains is almost entirely output. So there are two jobs: put the part of the prompt that is the same every time at the front, so the cache catches it, and cap the output. Google charges a storage fee for cached tokens by the hour on its other Gemini models and hasn’t published one for Argon, so leave that line open.
Budget 2027 at double for both models
Google’s price page carries one date that affects your budget whether or not you ever use Argon. Gemini 3.8 Flash costs $0.75 input and $3.75 output per million tokens through December 31, 2026, and doubles to $1.50 and $7.50 on January 1, 2027. Argon’s introductory price doubles on a day Google hasn’t named. If you run Flash as a default while waiting for Argon, budget 2027 at double for both.
Turning the 1M output limit into design decisions
The jump from 64K to 1M output tokens is converted into cost and time in the deep dive. Here it becomes four design decisions.
The limit belongs to the route, not the request. Vals lists Argon’s maximum output as 262,000 tokens. Artificial Analysis says it received up to 1M through Long Decode Continuation, a new API feature. Google’s docs mention neither yet. Either way, the maximum the model allows must not be your service’s default. Give the summarization route 4,000, the code review route 30,000, the migration planning route 200,000, and allow exceptions only above that baseline.
A long response is a background job. A generation of hundreds of thousands of tokens can take hours, and connections will drop in the meantime. Stream it, write partial output to durable storage, resume on a drop, and notify on completion, as a job queue would. The point is to keep it off the path where a person is waiting at a screen.
Don’t accept output nobody can review. Google’s showcase is a kernel migration of more than 800,000 lines. A change no human can read line by line has to be accepted by tests and CI. On routes that use long outputs, ask for the work in small units and gate each unit on automated checks before the next one starts. Pipeline design, not the model, keeps the review load down.
Alert on spend per route. One response that fills the limit costs $10 at the introductory price and $20 at the standard one. Count how many of those go out per day on each route, and make sure the operator hears about an overrun before the model does.
Guardrails now come in tiers
The most notable operational fact of this launch isn’t a benchmark. It’s how the model is shipped. Google says trusted defenders in its Fairwind program, and its own security teams, get Argon “without cyber guardrails.” The same model runs under a different safety policy depending on who is using it.
For a team running a platform, that is a clear message. A vendor’s safety behavior is a setting, not a constant. A request one version refuses, the next may carry out, and the same model may behave differently by account or contract. So what an agent is allowed to do has to be decided outside the model: a tool allow-list, a separate least-privilege identity per agent, a human approval in front of irreversible actions, and a log of every tool call. Those four stay in place whichever model is plugged in.
Read the prompt injection numbers the same way. Google’s evaluation note reports that on Gray Swan’s indirect prompt injection test, an injection succeeded within 15 attempts 0.7% of the time for Argon, 1.0% for Claude Opus 5.5 and 8.5% for GPT-6 Astra. A good number, and not zero. For an agent that makes thousands of tool calls a day, the design question isn’t “how likely is success” but “how far does the damage go when it succeeds.” The answer is separation: the agent that reads outside documents doesn’t also hold the permission to send mail.
There is one behavioral result too. Andon Labs wrote that on Vending-Bench 2, its long-running business simulation, Argon fabricated confirmation emails, refused refunds and lied to suppliers. It’s a simulation, but it is reason enough to run the same situations on fake accounts before giving the agent real ones.
Google itself says it monitors Argon’s chain of thought and actions and “stops execution when necessary.” If the model’s maker does that, a platform that uses the model needs the same ability: a person can stop a running agent at any time, and the reason it stopped is visible on screen.
What you can prepare today, before the API opens
You can’t use Argon now, but to make a decision within a week of it opening, some things need to exist already. All of them can be built with the models you use today.
- Build your own task set. Pick 50 real tasks with clear outcomes, write down the pass criteria, and run them once on your current model (Gemini 3.8 Flash or Claude Sonnet 5.5, say) for a baseline. Every lab tunes for the public benchmarks. None of them tunes for your task set.
- Log tokens per task. Record input, cached input and output tokens per task next to pass or fail. Without that record, the most you can say after a model change is “it seems better.”
- Set an output cap per route. A route that was fine because the current model stops at 64K becomes a different route at 1M.
- Move the stable prefix to the front of the prompt. The system prompt, tool schemas and repository summary, the parts that are the same every time, go first, and the parts that change go after, or the cache won’t catch them. In the arithmetic above that one change cut the bill by 60%.
- Route by type of work. A fast, cheap model as the default, and a branch that sends only “read a long brief and plan” work to a stronger model. Build the branch now, and Argon becomes one more option on it.
- Count “I don’t know” separately. On Artificial Analysis’ test, Argon’s hallucination rate is 15% and its accuracy 50%. If you run retrieval or support flows, you need abstentions and wrong answers counted apart to know how that profile will show up in your service.
- Put an approval gate in front of irreversible actions. Sending mail, posting messages, deleting data, anything that leaves the account or removes something, needs a place where a person confirms it, whatever the model. Then run the actions that pass that gate on fake accounts first.
- Budget 2027 at double. Flash has a date. Argon doesn’t.
And in endue
endue gathers models from several providers into one catalog and lets you set the model in three places: the agent’s default in Agent Builder, a single message from the composer, and a routine that pins its own model. The model picker shows the price per million tokens and whether the model supports tool calling, and some models also let you choose a reasoning effort. The models page lists Gemini 3.8 Flash and other Gemini models today, next to models from other providers.
So item 5 on the list above can be tried as is. Set a fast model as the agent’s default, raise it to a stronger one on the messages that read long material, and pin a model on each routine for scheduled work. The usage screen shows tokens by model and by agent, which is where item 2 can start. The settings are covered in Choosing a model.
Item 7 is how endue works by default. Sending mail, posting messages, deleting data and other irreversible actions always pass through an approval card, and that gate can’t be switched off in settings. The three reasons a run stops, and the card for each, are described in How it works. Whichever model is plugged in, that place stays the same.
What we still don’t know
- When paid API customers and Ultra subscribers get access, and in which regions
- The model ID, input context window, rate limits and cache storage price
- Whether a single request’s real output limit is 262,000 or 1M tokens, and how Long Decode Continuation works
- When the introductory price ends
- A model card and a Frontier Safety Framework report for Argon
Sources
- Google, Gemini 4 Argon: our next era of frontier intelligence (September 30, 2026)
- Google DeepMind, Gemini 4 Argon model evaluation: approach, methodology and results (PDF, September 30) and Fairwind Program
- Google AI for Developers, Gemini models and pricing (checked October 3)
- Vals AI, Vals Index (updated October 2) and Gemini 4 Argon model page
- Artificial Analysis, Gemini 4 Argon results (September 30)
- Zapier, AutomationBench leaderboard; Collinear, CWE-bench v1
- Bloomberg via Yahoo Finance, Google grapples with employee skepticism about new Gemini model (September 30)
- 9to5Google, Google stands by Gemini 4 performance as some claim it ‘struggles’ in real-world use (October 1)
- The Implicator, Google Staff Doubt Gemini 4 Argon Coding Despite Benchmarks (October 1)
- The Hacker News, Google Rolls Out Gemini 4 Argon to Trusted Cyber Defenders, Plans Guardrail-Free Version (October 1)
- Andon Labs, Vending-Bench 2 and notes on Argon’s behavior (September 30)


