Gemini 4 Argon: Benchmarks, API Pricing and Release Status
Gemini 4 Argon benchmarks from Google's launch table, $2/$10 intro API pricing ($4/$20 after), a 1M output limit, and when it reaches RouterPlex.

Google announced Gemini 4 Argon on 30 September 2026. It is Google DeepMind's new frontier model, and Google's launch table puts it ahead of GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on 14 of 19 benchmark rows. The introductory API price is $2 per 1M input tokens and $10 per 1M output tokens. After the introductory period it becomes $4 / $20. Its output limit is 1M tokens per response.
Argon is not generally available yet, and it is not on RouterPlex yet. Google is releasing it first to trusted cyber defenders, then to paid API customers and Google AI Ultra subscribers. Create a RouterPlex account now to use the 60+ models already on the marketplace. Argon joins that same key once it is live on RouterPlex.

Sources: Google's Gemini 4 Argon announcement by Koray Kavukcuoglu (30 September 2026), and Google DeepMind's evaluation methodology. Both checked 30 September 2026. Every score on this page is Google's. The images are Google's launch graphics. RouterPlex has not re-run these evals.
Gemini 4 Argon at a glance #
| Gemini 4 Argon | |
|---|---|
| Announced | 30 September 2026 |
| Maker | Google DeepMind |
| Intro API price | $2 input / $10 output per 1M tokens |
| Price after intro period | $4 input / $20 output per 1M tokens |
| Cached input | 95% off input ($0.10 intro, $0.20 standard) |
| Output limit | 1M tokens per response (previous Gemini limit: 64K) |
| Access today | Trusted cyber defenders, Fairwind Program |
| Next in line | Paid API customers and Google AI Ultra subscribers |
| Model ID | Not published yet |
| Input context window | Not published yet |
| On RouterPlex | Not yet. Register for the 60+ live models |
Gemini 4 Argon benchmarks #
This is Google's full launch table. Blue cells are where Google marks Argon as the best score. Grey cells mark another model's lead.

The same numbers as text, so you can search and copy them:
| Category | Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|---|
| Knowledge work | Vals Index | 68.9% | 63.1% | 65.8% | 67.0% |
| Knowledge work | AutomationBench | 51.3% | 41.4% | 31.4% | 42.5% |
| Knowledge work | Vals Finance Agent v2 | 65.4% | 53.5% | 58.9% | 58.6% |
| Knowledge work | Harvey's Legal Agent Benchmark | 19.6% | 5.4% | 6.7% | 3.8% |
| Agentic coding | DeepSWE v1.1 | 77.9% | 74.1% | 67.4% | 74.2% |
| Agentic coding | FrontierSWE v2 | 55.0% | 65.5% | 56.3% | 62.3% |
| Agentic coding | Vibe Code Bench | 91.9% | 89.6% | 90.3% | 90.3% |
| Agentic coding | Terminal-bench 4.0 | 57.4% | 58.2% | 57.9% | 66.4% |
| ML engineering | PostTrainBench | 45.3% | 44.3% | 40.2% | 49.3% |
| Science and math | Terminal-Bench Science 0.1 | 57.6% | 68.1% | 52.6% | 63.3% |
| Science and math | LABBench 2 | 88.8% | 85.4% | 68.6% | 73.1% |
| Science and math | RiemannBench | 76.0% | 72.0% | 65.6% | 69.6% |
| Long context | GraphWalks, up to 128K (BFS F1) | 99.7% | 98.7% | 91.4% | 90.6% |
| Long context | GraphWalks, 256K to 1M (BFS F1) | 84.2% | 71.8% | 65.0% | 66.8% |
| Computer use | Agent's Last Exam (pass rate) | 39.5% | 34.2% | — | 38.2% |
| Computer use | OSWorld-2.0 (offline subset, partial score) | 69.2% | 72.6% | — | — |
| Multimodal | Chartography | 71.6% | 71.0% | 46.2% | 66.3% |
| Multimodal | LVBench | 91.7% | 87.5% | 79.7% | 83.7% |
| Cybersecurity | CWE-bench v1 | 68.0% | 68.0% | 58.0% | 67.0% |
Where Argon leads
- Knowledge work is the clearest win. Argon leads all four rows. The gap is widest on Harvey's Legal Agent Benchmark: 19.6%, against 6.7% for the next model. Google also calls Argon the leading model on the Vals Index, which weights finance, coding, legal and tax work by each sector's share of U.S. GDP.
- Long-horizon coding. DeepSWE v1.1 is 77.9%, which Google calls a new state of the art. That is 3.7 points above Opus 5.5 and 3.8 above GPT-6 Astra.
- Very long prompts. On GraphWalks between 256K and 1M tokens, Argon scores 84.2%. The next best is 71.8%.
- Video and charts. LVBench, a long-video test, is 91.7%. Google calls that state of the art.

Where Argon trails
Google's own table shows five rows where another model is ahead:
- Terminal-bench 4.0: Opus 5.5 leads, 66.4% to 57.4%.
- FrontierSWE v2: GPT-6 Astra leads, 65.5% to 55.0%. Opus 5.5 is also ahead at 62.3%.
- Terminal-Bench Science 0.1: Astra leads, 68.1% to 57.6%.
- OSWorld-2.0: Astra leads, 72.6% to 69.2%. Anthropic models are not listed on this row.
- PostTrainBench: Opus 5.5 leads, 49.3% to 45.3%.
If your work is terminal-heavy agent coding, the table does not make Argon the default choice. Test your own tasks first.
How Google ran the numbers
The methodology document is short, and it matters:
- Argon scores are pass@1 and ran on the Gemini API at the highest thinking setting.
- Most competitor scores are the vendors' own published numbers, at their maximum reasoning setting when available. They were not re-run under one harness.
- Some Argon scores are Google's own runs. DeepSWE used a mini-swe-agent harness. Terminal-Bench Science used a 6× verifier timeout. OSWorld-2.0 is the best of 3 runs.
- PostTrainBench, LABBench 2 and GraphWalks were run by Google for every model.
- LVBench frame counts differ by model because of API limits: 1 FPS for Gemini, 800 frames for Astra, 300 for Fable 5.1 and 600 for Opus 5.5.
That is normal for a launch table. It is also a good reason to re-test on your own prompts before you switch.
Cybersecurity: Google's headline use case #
Google trained Argon for cyber defense. It says Argon can find, validate and patch serious software vulnerabilities on its own. Trusted defenders get Argon without cyber guardrails through the Fairwind Program. That is also why it is the first group with access.

- CWE-bench v1 (fixing security vulnerabilities): 68.0%, tied first with GPT-6 Astra.
- Real-world vulnerability discovery (Google internal, 20 languages): 85.8%, against 71.0% for Gemini 3.8 Flash Cyber.
- Wiz penetration test benchmark (no source code access): 70.9%, against 58.2%.
- Wiz used Argon in its Scan for Good program. Google says it found a critical vulnerability in hospital software that earlier frontier models had missed.
On prompt injection, Google reports the lowest attack success rate on Gray Swan's indirect prompt injection benchmark: 0.7% at 15 attempts. Claude Opus 5.5 and Claude Fable 5.1 are both at 1.0%. Lower is better on this chart.

Gemini 4 Argon API pricing #
| Tier | Input / 1M | Cached input / 1M | Output / 1M |
|---|---|---|---|
| Introductory | $2.00 | $0.10 | $10.00 |
| After intro period | $4.00 | $0.20 | $20.00 |
Google gives the cache price as "95% off input." The dollar cells are that discount applied to each input rate. Google has not said how long the introductory period lasts. It has not published a long-context tier, a batch price or a cache-storage price for Argon.
What one request costs
A 40,000-token prompt that returns 2,000 tokens, with no cache:
| Model | Input / Output per 1M | 40K in / 2K out | Where |
|---|---|---|---|
| Gemini 4 Argon (intro) | $2 / $10 | $0.100 | Google, not live yet |
| Gemini 4 Argon (standard) | $4 / $20 | $0.200 | Google, not live yet |
gpt-6.1-sol | $2 / $10 | $0.100 | Live on RouterPlex |
gemini-3.1-pro | $2 / $12 | $0.104 | Live on RouterPlex |
| Claude Opus 5.5 | $4 / $20 | $0.200 | Anthropic, not on RouterPlex yet |
gpt-6-astra | $10 / $50 | $0.500 | Live on RouterPlex |
gemini-3.8-flash | $0.75 / $3.75 | $0.038 | Live on RouterPlex |
At the intro price, Argon costs the same as GPT-6.1 Sol per token and one fifth of GPT-6 Astra. After the intro period it matches Opus 5.5's list price.
The 1M output limit changes the bill
Earlier Gemini models stopped at 64K output tokens. Argon can write up to 1M tokens in one response. Google says long reasoning is the point. Output tokens are the expensive side, though. One response that uses the full limit costs $10 at the intro rate and $20 after, before any input. If you run agents on Argon, set a max output token cap and a hard spend limit on each key.
Availability and release date #
- 30 September 2026: announced. Rolling out to trusted cyber defenders through the Fairwind Program.
- Next: paid API customers and Google AI Ultra subscribers. Google says "rolling out soon" and gives no date.
- Later: wider access for developers, enterprises and consumers, after more safety work.
Google says it is taking part in the U.S. government's voluntary pre-release access process. It is also adding safeguards in four areas before a broad release: misuse (cyber and CBRN), prompt injection, misalignment monitoring of the model's chain of thought and actions, and hardened sandboxes.
What Google says Argon already does inside Google #
These are Google's own examples. They are not independent tests.
- C/C++ to Rust migrations, from libraries like re2 and libgav1 up to the 800K+ line Fuchsia Zircon kernel. The Rust port of the libgav1 video decoder runs 2.7× faster with identical output after Argon replaced 32K lines of SIMD code.
- Memory savings across Google data centers: over 300 TiB freed so far, with 500 TiB to 1 PiB expected in total.
- Quantum algorithm work: beat a published baseline by 40% in minutes.
Use the RouterPlex marketplace while you wait for Argon #
Argon is not on any public API yet. The models it was compared against, and Google's current Gemini line, are live on RouterPlex today:
- GPT-6 Astra, which leads Argon on FrontierSWE, Terminal-Bench Science and OSWorld-2.0 in Google's table. $10 / $50.
- GPT-6.1 Sol, at the same $2 / $10 as Argon's intro price.
- Gemini 3.8 Flash and Gemini 3.1 Pro, Google's current API models. $0.75 / $3.75 and $2 / $12.
- The full model catalog, sorted by price on the cheapest-first ledger.
One OpenAI-compatible key reaches all of them:
import osfrom openai import OpenAIclient = OpenAI(api_key=os.environ["ROUTERPLEX_API_KEY"],base_url="https://api.routerplex.com/v1",)response = client.chat.completions.create(model="gemini-3.8-flash", # switch the model ID when Argon is livemessages=[{"role": "user", "content": "Summarize the risks in this migration plan."}],)print(response.choices[0].message.content)
When Argon lands on RouterPlex, changing that one model string is the whole migration. Claude Code works too, with ANTHROPIC_BASE_URL=https://api.routerplex.com. See the Claude Code setup guide.
Register now, use Argon when it goes live #
Create a RouterPlex account to use the 60+ models on the marketplace today, with one key, vendor list prices, 0% markup and $0 top-up fees. Credit is prepaid, and every key has a hard budget that stops at $0. When Google opens paid API access and Gemini 4 Argon is live on RouterPlex, it will be on the same key at Google's list price. This page will list the model ID on the day it goes live.
Common questions
Frequently asked questions
What is Gemini 4 Argon?
Gemini 4 Argon is Google DeepMind's new frontier model, announced on 30 September 2026 by Koray Kavukcuoglu. Google built it for long, multi-step work: software engineering, finance and legal research, and cybersecurity defense. Its output limit is 1M tokens per response, up from 64K.
How much does the Gemini 4 Argon API cost?
Google's launch price is an introductory $2 per 1M input tokens and $10 per 1M output tokens. Cached input is 95% off input, which is $0.10 per 1M at the intro rate. After the introductory period, the price becomes $4 input and $20 output per 1M. Google has not said how long the introductory period lasts.
Can I use Gemini 4 Argon today?
Not generally. On launch day Argon is rolling out to trusted cyber defenders in Google's Fairwind Program. Google says paid API customers and Google AI Ultra subscribers come next, and gave no date.
Is Gemini 4 Argon on RouterPlex?
Not yet. No RouterPlex model ID exists for Argon, and Google has not opened paid API access. Register now to use the 60+ models already on the RouterPlex marketplace, including Gemini 3.8 Flash, Gemini 3.1 Pro, GPT-6 Astra and GPT-6.1 Sol. Argon joins the same key once it is live on the marketplace.
How does Gemini 4 Argon compare with GPT-6 Astra and Claude Opus 5.5?
In Google's launch table, Argon scores highest or ties on 14 of 19 rows, including DeepSWE v1.1 (77.9% vs 74.1% Astra and 74.2% Opus 5.5) and AutomationBench (51.3% vs 41.4% and 42.5%). It trails on FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1 and OSWorld-2.0. Google compiled the table. RouterPlex has not re-run it.
What is the Gemini 4 Argon context window?
Google's announcement gives a 1M-token output limit. It does not give an input context window, a model ID or a knowledge cutoff. Google's GraphWalks results test prompts between 256K and 1M tokens, so Argon accepts at least that much input in Google's own tests.
Run the smallest paid test.
Add $5, cap the key, and verify the result with your own workload. No subscription, and credit never expires — a first top-up of $25+ is matched with $25 extra.



