Research index
Model releases/

Gemini 4 Argon: Benchmarks, API Pricing and Release Status

Gemini 4 Argon benchmarks from Google's launch table, $2/$10 intro API pricing ($4/$20 after), a 1M output limit, and when it reaches RouterPlex.

Written byRouterPlex
Reading time10 min
Gemini 4 Argon: Benchmarks, API Pricing and Release Status
Gemini 4 Argon: Benchmarks, API Pricing and Release Status

Google announced Gemini 4 Argon on 30 September 2026. It is Google DeepMind's new frontier model, and Google's launch table puts it ahead of GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 on 14 of 19 benchmark rows. The introductory API price is $2 per 1M input tokens and $10 per 1M output tokens. After the introductory period it becomes $4 / $20. Its output limit is 1M tokens per response.

Argon is not generally available yet, and it is not on RouterPlex yet. Google is releasing it first to trusted cyber defenders, then to paid API customers and Google AI Ultra subscribers. Create a RouterPlex account now to use the 60+ models already on the marketplace. Argon joins that same key once it is live on RouterPlex.

Google's launch art for Gemini 4 Argon: the Gemini spark and the words Gemini 4 Argon beside a large glowing 4 on a blue gradient.
Google's launch art for Gemini 4 Argon: the Gemini spark and the words Gemini 4 Argon beside a large glowing 4 on a blue gradient.

Sources: Google's Gemini 4 Argon announcement by Koray Kavukcuoglu (30 September 2026), and Google DeepMind's evaluation methodology. Both checked 30 September 2026. Every score on this page is Google's. The images are Google's launch graphics. RouterPlex has not re-run these evals.

Gemini 4 Argon at a glance #

Gemini 4 Argon
Announced30 September 2026
MakerGoogle DeepMind
Intro API price$2 input / $10 output per 1M tokens
Price after intro period$4 input / $20 output per 1M tokens
Cached input95% off input ($0.10 intro, $0.20 standard)
Output limit1M tokens per response (previous Gemini limit: 64K)
Access todayTrusted cyber defenders, Fairwind Program
Next in linePaid API customers and Google AI Ultra subscribers
Model IDNot published yet
Input context windowNot published yet
On RouterPlexNot yet. Register for the 60+ live models

Gemini 4 Argon benchmarks #

This is Google's full launch table. Blue cells are where Google marks Argon as the best score. Grey cells mark another model's lead.

Google's Gemini 4 Argon benchmark table comparing Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 across knowledge work, agentic coding, ML engineering, science and math, long context, computer use, multimodal understanding and cybersecurity.
Google's Gemini 4 Argon benchmark table comparing Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 across knowledge work, agentic coding, ML engineering, science and math, long context, computer use, multimodal understanding and cybersecurity.

The same numbers as text, so you can search and copy them:

CategoryBenchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5
Knowledge workVals Index68.9%63.1%65.8%67.0%
Knowledge workAutomationBench51.3%41.4%31.4%42.5%
Knowledge workVals Finance Agent v265.4%53.5%58.9%58.6%
Knowledge workHarvey's Legal Agent Benchmark19.6%5.4%6.7%3.8%
Agentic codingDeepSWE v1.177.9%74.1%67.4%74.2%
Agentic codingFrontierSWE v255.0%65.5%56.3%62.3%
Agentic codingVibe Code Bench91.9%89.6%90.3%90.3%
Agentic codingTerminal-bench 4.057.4%58.2%57.9%66.4%
ML engineeringPostTrainBench45.3%44.3%40.2%49.3%
Science and mathTerminal-Bench Science 0.157.6%68.1%52.6%63.3%
Science and mathLABBench 288.8%85.4%68.6%73.1%
Science and mathRiemannBench76.0%72.0%65.6%69.6%
Long contextGraphWalks, up to 128K (BFS F1)99.7%98.7%91.4%90.6%
Long contextGraphWalks, 256K to 1M (BFS F1)84.2%71.8%65.0%66.8%
Computer useAgent's Last Exam (pass rate)39.5%34.2%—38.2%
Computer useOSWorld-2.0 (offline subset, partial score)69.2%72.6%——
MultimodalChartography71.6%71.0%46.2%66.3%
MultimodalLVBench91.7%87.5%79.7%83.7%
CybersecurityCWE-bench v168.0%68.0%58.0%67.0%

Where Argon leads

  • Knowledge work is the clearest win. Argon leads all four rows. The gap is widest on Harvey's Legal Agent Benchmark: 19.6%, against 6.7% for the next model. Google also calls Argon the leading model on the Vals Index, which weights finance, coding, legal and tax work by each sector's share of U.S. GDP.
  • Long-horizon coding. DeepSWE v1.1 is 77.9%, which Google calls a new state of the art. That is 3.7 points above Opus 5.5 and 3.8 above GPT-6 Astra.
  • Very long prompts. On GraphWalks between 256K and 1M tokens, Argon scores 84.2%. The next best is 71.8%.
  • Video and charts. LVBench, a long-video test, is 91.7%. Google calls that state of the art.
Google's DeepSWE v1.1 chart: Gemini 4 Argon 77.9%, GPT-6 Astra 74.1%, Claude Fable 5.1 67.4%, Claude Opus 5.5 74.2%.
Google's DeepSWE v1.1 chart: Gemini 4 Argon 77.9%, GPT-6 Astra 74.1%, Claude Fable 5.1 67.4%, Claude Opus 5.5 74.2%.

Where Argon trails

Google's own table shows five rows where another model is ahead:

  • Terminal-bench 4.0: Opus 5.5 leads, 66.4% to 57.4%.
  • FrontierSWE v2: GPT-6 Astra leads, 65.5% to 55.0%. Opus 5.5 is also ahead at 62.3%.
  • Terminal-Bench Science 0.1: Astra leads, 68.1% to 57.6%.
  • OSWorld-2.0: Astra leads, 72.6% to 69.2%. Anthropic models are not listed on this row.
  • PostTrainBench: Opus 5.5 leads, 49.3% to 45.3%.

If your work is terminal-heavy agent coding, the table does not make Argon the default choice. Test your own tasks first.

How Google ran the numbers

The methodology document is short, and it matters:

  • Argon scores are pass@1 and ran on the Gemini API at the highest thinking setting.
  • Most competitor scores are the vendors' own published numbers, at their maximum reasoning setting when available. They were not re-run under one harness.
  • Some Argon scores are Google's own runs. DeepSWE used a mini-swe-agent harness. Terminal-Bench Science used a 6× verifier timeout. OSWorld-2.0 is the best of 3 runs.
  • PostTrainBench, LABBench 2 and GraphWalks were run by Google for every model.
  • LVBench frame counts differ by model because of API limits: 1 FPS for Gemini, 800 frames for Astra, 300 for Fable 5.1 and 600 for Opus 5.5.

That is normal for a launch table. It is also a good reason to re-test on your own prompts before you switch.

Cybersecurity: Google's headline use case #

Google trained Argon for cyber defense. It says Argon can find, validate and patch serious software vulnerabilities on its own. Trusted defenders get Argon without cyber guardrails through the Fairwind Program. That is also why it is the first group with access.

Google's vulnerability-discovery charts: Gemini 4 Argon scores 85.8% on real-world vulnerability discovery across 20 languages versus 71.0% for Gemini 3.8 Flash Cyber, and 70.9% on the Wiz penetration test benchmark versus 58.2%.
Google's vulnerability-discovery charts: Gemini 4 Argon scores 85.8% on real-world vulnerability discovery across 20 languages versus 71.0% for Gemini 3.8 Flash Cyber, and 70.9% on the Wiz penetration test benchmark versus 58.2%.
  • CWE-bench v1 (fixing security vulnerabilities): 68.0%, tied first with GPT-6 Astra.
  • Real-world vulnerability discovery (Google internal, 20 languages): 85.8%, against 71.0% for Gemini 3.8 Flash Cyber.
  • Wiz penetration test benchmark (no source code access): 70.9%, against 58.2%.
  • Wiz used Argon in its Scan for Good program. Google says it found a critical vulnerability in hospital software that earlier frontier models had missed.

On prompt injection, Google reports the lowest attack success rate on Gray Swan's indirect prompt injection benchmark: 0.7% at 15 attempts. Claude Opus 5.5 and Claude Fable 5.1 are both at 1.0%. Lower is better on this chart.

Google's Gray Swan indirect prompt injection chart. Attack success rate at 15 attempts: Gemini 4 Argon 0.7%, Claude Opus 5.5 1.0%, Claude Fable 5.1 1.0%, GPT-6 Astra 8.5%, GPT-6 Sol 10.1%. Lower is better.
Google's Gray Swan indirect prompt injection chart. Attack success rate at 15 attempts: Gemini 4 Argon 0.7%, Claude Opus 5.5 1.0%, Claude Fable 5.1 1.0%, GPT-6 Astra 8.5%, GPT-6 Sol 10.1%. Lower is better.

Gemini 4 Argon API pricing #

TierInput / 1MCached input / 1MOutput / 1M
Introductory$2.00$0.10$10.00
After intro period$4.00$0.20$20.00

Google gives the cache price as "95% off input." The dollar cells are that discount applied to each input rate. Google has not said how long the introductory period lasts. It has not published a long-context tier, a batch price or a cache-storage price for Argon.

What one request costs

A 40,000-token prompt that returns 2,000 tokens, with no cache:

ModelInput / Output per 1M40K in / 2K outWhere
Gemini 4 Argon (intro)$2 / $10$0.100Google, not live yet
Gemini 4 Argon (standard)$4 / $20$0.200Google, not live yet
gpt-6.1-sol$2 / $10$0.100Live on RouterPlex
gemini-3.1-pro$2 / $12$0.104Live on RouterPlex
Claude Opus 5.5$4 / $20$0.200Anthropic, not on RouterPlex yet
gpt-6-astra$10 / $50$0.500Live on RouterPlex
gemini-3.8-flash$0.75 / $3.75$0.038Live on RouterPlex

At the intro price, Argon costs the same as GPT-6.1 Sol per token and one fifth of GPT-6 Astra. After the intro period it matches Opus 5.5's list price.

The 1M output limit changes the bill

Earlier Gemini models stopped at 64K output tokens. Argon can write up to 1M tokens in one response. Google says long reasoning is the point. Output tokens are the expensive side, though. One response that uses the full limit costs $10 at the intro rate and $20 after, before any input. If you run agents on Argon, set a max output token cap and a hard spend limit on each key.

Availability and release date #

  • 30 September 2026: announced. Rolling out to trusted cyber defenders through the Fairwind Program.
  • Next: paid API customers and Google AI Ultra subscribers. Google says "rolling out soon" and gives no date.
  • Later: wider access for developers, enterprises and consumers, after more safety work.

Google says it is taking part in the U.S. government's voluntary pre-release access process. It is also adding safeguards in four areas before a broad release: misuse (cyber and CBRN), prompt injection, misalignment monitoring of the model's chain of thought and actions, and hardened sandboxes.

What Google says Argon already does inside Google #

These are Google's own examples. They are not independent tests.

  • C/C++ to Rust migrations, from libraries like re2 and libgav1 up to the 800K+ line Fuchsia Zircon kernel. The Rust port of the libgav1 video decoder runs 2.7× faster with identical output after Argon replaced 32K lines of SIMD code.
  • Memory savings across Google data centers: over 300 TiB freed so far, with 500 TiB to 1 PiB expected in total.
  • Quantum algorithm work: beat a published baseline by 40% in minutes.

Use the RouterPlex marketplace while you wait for Argon #

Argon is not on any public API yet. The models it was compared against, and Google's current Gemini line, are live on RouterPlex today:

One OpenAI-compatible key reaches all of them:

python
import os
from openai import OpenAI
 
client = OpenAI(
api_key=os.environ["ROUTERPLEX_API_KEY"],
base_url="https://api.routerplex.com/v1",
)
 
response = client.chat.completions.create(
model="gemini-3.8-flash", # switch the model ID when Argon is live
messages=[{"role": "user", "content": "Summarize the risks in this migration plan."}],
)
print(response.choices[0].message.content)

When Argon lands on RouterPlex, changing that one model string is the whole migration. Claude Code works too, with ANTHROPIC_BASE_URL=https://api.routerplex.com. See the Claude Code setup guide.

Register now, use Argon when it goes live #

Create a RouterPlex account to use the 60+ models on the marketplace today, with one key, vendor list prices, 0% markup and $0 top-up fees. Credit is prepaid, and every key has a hard budget that stops at $0. When Google opens paid API access and Gemini 4 Argon is live on RouterPlex, it will be on the same key at Google's list price. This page will list the model ID on the day it goes live.

Common questions

Frequently asked questions

What is Gemini 4 Argon?

Gemini 4 Argon is Google DeepMind's new frontier model, announced on 30 September 2026 by Koray Kavukcuoglu. Google built it for long, multi-step work: software engineering, finance and legal research, and cybersecurity defense. Its output limit is 1M tokens per response, up from 64K.

How much does the Gemini 4 Argon API cost?

Google's launch price is an introductory $2 per 1M input tokens and $10 per 1M output tokens. Cached input is 95% off input, which is $0.10 per 1M at the intro rate. After the introductory period, the price becomes $4 input and $20 output per 1M. Google has not said how long the introductory period lasts.

Can I use Gemini 4 Argon today?

Not generally. On launch day Argon is rolling out to trusted cyber defenders in Google's Fairwind Program. Google says paid API customers and Google AI Ultra subscribers come next, and gave no date.

Is Gemini 4 Argon on RouterPlex?

Not yet. No RouterPlex model ID exists for Argon, and Google has not opened paid API access. Register now to use the 60+ models already on the RouterPlex marketplace, including Gemini 3.8 Flash, Gemini 3.1 Pro, GPT-6 Astra and GPT-6.1 Sol. Argon joins the same key once it is live on the marketplace.

How does Gemini 4 Argon compare with GPT-6 Astra and Claude Opus 5.5?

In Google's launch table, Argon scores highest or ties on 14 of 19 rows, including DeepSWE v1.1 (77.9% vs 74.1% Astra and 74.2% Opus 5.5) and AutomationBench (51.3% vs 41.4% and 42.5%). It trails on FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1 and OSWorld-2.0. Google compiled the table. RouterPlex has not re-run it.

What is the Gemini 4 Argon context window?

Google's announcement gives a 1M-token output limit. It does not give an input context window, a model ID or a knowledge cutoff. Google's GraphWalks results test prompts between 256K and 1M tokens, so Argon accepts at least that much input in Google's own tests.

Run the smallest paid test.

Add $5, cap the key, and verify the result with your own workload. No subscription, and credit never expires — a first top-up of $25+ is matched with $25 extra.