Skip to content
AlphaCorp AI
Wave of light particles flowing through faint circuit traces on a dark background
News22 min read

Gemini 4 Argon Launch: Benchmarks, Pricing and Everything You Need to Know

Ignas Vaitukaitis, Founder & CEO of AlphaCorp AI

AI Agent Engineer ·

Gemini 4 Argon Launch: Benchmarks, Pricing and Everything You Need to Know
On this page(18)
  1. Gemini 4 Argon Launch: What Google Announced on September 30, 2026
  2. Gemini 4 Argon Benchmarks vs GPT-6 Astra and Claude
  3. Pricing: $2 Input, $10 Output, and How That Compares
  4. Availability: Fairwind Program First, Developers and Consumers Later
  5. The 1M-Token Output Limit and Long-Context Performance
  6. How Google Is Already Using Argon Internally
  7. Cybersecurity Defense: CWE-bench, Vulnerability Discovery, and the No-Guardrails Release
  8. Safety Safeguards Before Broad Release
  9. What the Benchmarks Don’t Tell You Yet
  10. Gemini 4 Argon FAQ
  11. Is Gemini 4 Argon available now?
  12. How much does Gemini 4 Argon cost?
  13. Is Gemini 4 Argon better than GPT-6 Astra or Claude Opus 5.5?
  14. What is the Fairwind Program?
  15. What is Gemini 4 Argon’s output limit?
  16. When will Gemini 4 Argon come to the Gemini app?
  17. Does Gemini 4 Argon have a technical report?
  18. How to Prepare for Gemini 4 Argon Access

Google announced Gemini 4 Argon on September 30, 2026, at $2 per million input tokens and $10 output, available today only to vetted cyber defenders. The Gemini 4 Argon launch comes with 19 published benchmark rows against GPT-6 Astra and Claude, a 1 million token output limit, and a staged rollout with no public dates. Below you’ll find every number Google released, where the model trails its rivals, and what to do while you wait for API access.

  • 77.9% on DeepSWE v1.1 in 2026, ahead of Claude Opus 5.5 (74.2%) and GPT-6 Astra (74.1%), per Google’s September 30 launch post.
  • 13 of 19 benchmark rows led outright in Google’s 2026 comparison table, with losses on FrontierSWE v2, Terminal-bench 4.0, and OSWorld-2.0.
  • $2 input / $10 output per million tokens introductory in 2026, rising to $4 / $20 at standard pricing, per Google DeepMind.
  • 0.7% attack success rate on Gray Swan’s 2026 Indirect Prompt Injection benchmark, the lowest of 13 models tested.
  • 1 million output tokens per response in 2026, up from 64K in the prior Gemini generation, per Google DeepMind’s model page.

Gemini 4 Argon Launch: What Google Announced on September 30, 2026

On September 30, 2026, Google announced Gemini 4 Argon, its new frontier model, and began rolling it out to a small group of trusted cyber defenders rather than the general public. That one detail shapes the whole launch. The model exists, Google has benchmarked it against GPT-6 Astra and Anthropic’s Claude line, and almost nobody outside Google can call it yet.

Google positions Argon around three kinds of work. Koray Kavukcuoglu’s launch post, written as SVP of Google DeepMind and Google’s Chief AI Architect, describes a model built “to sustain deep reasoning across complex, long-horizon workflows” in:

  • Real-world software engineering: debugging, algorithm design, and large codebase migrations that run for hours instead of minutes.
  • Enterprise knowledge work: legal research and drafting, multi-step financial research, and tax work.
  • Cybersecurity defense: finding, validating, and patching software vulnerabilities on its own.

“It delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense.” Koray Kavukcuoglu, Google DeepMind, September 30, 2026

Argon succeeds the Gemini 3 series. The last models in that line, Gemini 3.8 Flash and Gemini 3.8 Flash Cyber, shipped on September 2, 2026, four weeks before Argon. “Argon” was the internal codename, and Google kept it for the public name.

The rollout is deliberately narrow. Argon is available today only through Google’s Fairwind Program, which covers a vetted set of Google Cloud customers, government agencies, and cybersecurity partners, plus Google’s own internal teams. Google says it is also taking part in the U.S. government’s voluntary pre-release access process for frontier models. Developers, enterprises, and consumers come later, with paid API customers and Google AI Ultra subscribers first in line.

Two other headline claims travel with the announcement: a 1 million token output limit, and an introductory API price of $2 per million input tokens. Both get their own treatment further down.

One thing to hold onto while reading any of this. Every number in the launch, from benchmark scores to pricing, comes from Google’s own materials. No outside lab has replicated them yet.

Gemini 4 Argon Benchmarks vs GPT-6 Astra and Claude

Gemini 4 Argon leads 13 of the 19 benchmark rows Google published on September 30, 2026, ties one with GPT-6 Astra, and trails on the remaining five. GPT-6 Astra wins three rows. Claude Opus 5.5 wins two. Claude Fable 5.1 wins none.

Here is the full table, transcribed from Google’s launch charts:

CategoryBenchmarkGemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5Leader
Knowledge workVals Index68.9%63.1%65.8%67.0%Argon
Knowledge workAutomationBench51.3%41.4%31.4%42.5%Argon
Knowledge workVals Finance Agent v265.4%53.5%58.9%58.6%Argon
Knowledge workHarvey’s Legal Agent Benchmark19.6%5.4%6.7%3.8%Argon
Agentic codingDeepSWE v1.177.9%74.1%67.4%74.2%Argon
Agentic codingFrontierSWE v255.0%65.5%56.3%62.3%GPT-6 Astra
Agentic codingVibe Code Bench91.9%89.6%90.3%90.3%Argon
Agentic codingTerminal-bench 4.057.4%58.2%57.9%66.4%Claude Opus 5.5
ML engineeringPostTrainBench45.3%44.3%40.2%49.3%Claude Opus 5.5
Science and mathTerminal-Bench Science 0.157.6%68.1%52.6%63.3%GPT-6 Astra
Science and mathLABBench 288.8%85.4%68.6%73.1%Argon
Science and mathRiemannBench76.0%72.0%65.6%69.6%Argon
Long contextGraphWalks, up to 128K (BFS F1)99.7%98.7%91.4%90.6%Argon
Long contextGraphWalks, 256K to 1M (BFS F1)84.2%71.8%65.0%66.8%Argon
Computer useAgent’s Last Exam (pass rate)39.5%34.2%not reported38.2%Argon
Computer useOSWorld-2.0 (offline subset)69.2%72.6%not reportednot reportedGPT-6 Astra
MultimodalChartography71.6%71.0%46.2%66.3%Argon
MultimodalLVBench91.7%87.5%79.7%83.7%Argon
CybersecurityCWE-bench v168.0%68.0%58.0%67.0%Tie (Argon, GPT-6 Astra)

The knowledge-work block is a clean sweep. Argon’s 19.6% on Harvey’s Legal Agent Benchmark is more than triple the next model, though 19.6% also tells you how far legal drafting agents still have to go. AutomationBench, the Zapier-derived test of whether correct data lands in the right CRM, inbox, and calendar systems, shows the same pattern: a clear lead at 51.3%, on a task most models still fail half the time.

Coding is where I’d slow down. Argon sets the record on DeepSWE v1.1 at 77.9%, and that benchmark deserves the attention. The DeepSWE paper from Datacurve researchers, published July 2026, built 113 original long-horizon engineering tasks across TypeScript, Go, Python, JavaScript, and Rust specifically because frontier agents had bunched together on older suites. A 3.7-point gap over Claude Opus 5.5 there is real.

But look two rows down. On FrontierSWE v2, Argon posts 55.0% against GPT-6 Astra’s 65.5%, a 10.5-point deficit. On Terminal-bench 4.0, Claude Opus 5.5 leads by 9 points. Argon is the best model on one coding benchmark and the worst frontier model on another. Which one matches your workload is a question only your own eval can answer.

Grouped bar chart comparing four frontier models on three agentic coding benchmarks, September 2026. On FrontierSWE v2: Gemini 4 Argon 55.0%, Claude Opus 5.5 62.3%, GPT-6 Astra 65.5%, Claude Fable 5.1 56.3%. On Terminal-bench 4.0: Gemini 4 Argon 57.4%, Claude Opus 5.5 66.4%, GPT-6 Astra 58.2%, Claude Fable 5.1 57.9%. On DeepSWE v1.1: Gemini 4 Argon 77.9%, Claude Opus 5.5 74.2%, GPT-6 Astra 74.1%, Claude Fable 5.1 67.4%. Argon leads DeepSWE v1.1 and finishes last on FrontierSWE v2 and Terminal-bench 4.0.
Argon’s 77.9% on DeepSWE v1.1 leads every rival, while its 55.0% on FrontierSWE v2 sits 10.5 points behind GPT-6 Astra. Source: Google DeepMind, 2026.

The other places Argon trails are worth naming plainly: PostTrainBench (ML engineering), Terminal-Bench Science 0.1, and OSWorld-2.0. Three of those five losses involve terminal or desktop control. Where Argon reads, reasons over long inputs, or interprets charts and video, it wins. Where it has to operate a machine step by step, the field is level or ahead.

Pricing: $2 Input, $10 Output, and How That Compares

Gemini 4 Argon costs $2 per million input tokens and $10 per million output tokens at its introductory API price, announced September 30, 2026, with cached input discounted 95% to $0.10 per million. After the introductory period, standard pricing doubles to $4 per million input tokens and $20 per million output tokens.

That standard rate lands exactly on Claude Opus 5.5. Anthropic’s Claude Opus 5.5 pricing from September 22, 2026 is $4 input and $20 output per million tokens, with cache reads at $0.20 per million and cache writes at $5 per million. So Argon’s introductory price is a half-off window against its closest rival, and once it closes the two models cost the same per token.

Model (2026)Input / 1M tokensOutput / 1M tokensCached input / 1M tokens
Gemini 4 Argon (introductory)$2.00$10.00$0.10
Gemini 4 Argon (standard)$4.00$20.0095% off input
Claude Opus 5.5$4.00$20.00$0.20 (read), $5.00 (write)
Gemini 3.8 Flash (introductory)$0.75$3.75not listed
Gemini 3.8 Flash (standard)$1.50$7.50not listed
Grouped bar chart of API list prices per million tokens in September 2026, input and output shown side by side. Gemini 4 Argon standard: $4.00 input, $20.00 output. Claude Opus 5.5: $4.00 input, $20.00 output. Gemini 4 Argon introductory: $2.00 input, $10.00 output. Gemini 3.8 Flash standard: $1.50 input, $7.50 output. Gemini 3.8 Flash introductory: $0.75 input, $3.75 output.
Argon costs $2.00 per million input tokens and $10.00 output during the introductory window, then doubles to the $4.00 and $20.00 that Claude Opus 5.5 already charges. Source: Google DeepMind and Anthropic, 2026.

The Gemini 3.8 Flash row is the one enterprise buyers should stare at. Argon’s introductory price is 2.7 times Flash’s introductory input rate and 2.7 times its output rate, and the standard rates keep that same ratio. Argon is a different tier of spend, and Google isn’t pretending otherwise.

Per-token prices understate the gap. Anyone who has run agent loops in production knows the input price is the number you look at and the output price is the number you pay, because reasoning-heavy models generate far more output than a chat transcript suggests. The Vals AI leaderboard lists Argon’s inference cost at about $15.68 per test in 2026, against $5.73 for Gemini 3.8 Flash. That is a 2.7x per-token difference turning into a 2.7x per-task difference, which means Argon isn’t burning proportionally more tokens than Flash, at least on that suite. With a 1 million token output ceiling, a runaway trajectory can still be expensive.

The 95% cache discount is the lever that changes the math. For workloads that re-read the same large context on every call (a codebase, a contract set, a document corpus), cached input at $0.10 per million makes the input side nearly free. Output stays at full price. If you’re sizing production agent workloads against these numbers, model the output tokens per completed task first and the input rate second.

Availability: Fairwind Program First, Developers and Consumers Later

Gemini 4 Argon is available on September 30, 2026 only to trusted cyber defenders in Google’s Fairwind Program and to Google’s own internal teams. Everyone else waits. Google has committed to a sequence for wider access but has published no dates for any stage of it.

The rollout runs in three tiers, in this order:

  1. Fairwind Program members and Google internal teams (live now): a vetted group of Google Cloud customers, government agencies, and cybersecurity partners using Argon to find and fix vulnerabilities in critical infrastructure and public services.
  2. Paid API customers and Google AI Ultra subscribers (next): named in Kavukcuoglu’s post as first in line once the controlled phase ends.
  3. Developers, enterprises, and consumers broadly (later): Google’s stated goal is “as soon as possible,” with guardrails iterated on early-tester feedback before that happens.

Running alongside all of this is the U.S. government’s voluntary pre-release access process for frontier models, which Google says it is actively engaged in. That process has no public timeline either.

The consumer path runs through AI Ultra. Google restructured its AI subscriptions at I/O on May 19, 2026, setting a $100 per month AI Ultra tier with 5x Pro usage limits and 20TB of storage, and a $200 per month top tier (cut from $250) with 20x Pro limits plus Gemini Spark and Project Genie access, per Google’s 2026 subscription announcement. The Argon launch names “Google AI Ultra subscribers” without saying whether both price points get it at the same time.

Anyone who has tried to plan a Q4 deployment around a model that is “coming soon” knows how this goes. Budget for the wait. If your team is on Google Cloud and does security work, the Fairwind Program is the only door open today, and Google describes its members as invited rather than self-nominated.

The 1M-Token Output Limit and Long-Context Performance

Gemini 4 Argon can generate up to 1 million output tokens in a single response, up from the 64K ceiling of the prior Gemini generation, an increase of more than 15 times. Google calls the limit industry-leading and ties it directly to the model’s ability to work through a hard problem in one trajectory instead of many short ones.

Why does an output limit matter more than it sounds? Because agentic work is mostly output. A model migrating a library, drafting a legal brief from a document set, or patching a vulnerability spends its tokens on reasoning, tool calls, and generated code. At 64K, long tasks had to be chopped into segments with state handed between them. At 1M, Google’s argument is that the model can “think deeply and generate hundreds of thousands of tokens” without a handoff.

Reading long inputs is the other half, and here Google publishes GraphWalks scores. GraphWalks is a synthetic benchmark, originally released by OpenAI, that asks a model to traverse a graph scattered across its context window, so success depends on hopping between distant positions instead of pulling out one buried fact. The Google DeepMind Gemini model page reports the 2026 results as breadth-first-search F1:

  • Up to 128K tokens: Argon 99.7%, GPT-6 Astra 98.7%, Claude Fable 5.1 91.4%, Claude Opus 5.5 90.6%.
  • 256K to 1M tokens: Argon 84.2%, GPT-6 Astra 71.8%, Claude Opus 5.5 66.8%, Claude Fable 5.1 65.0%.

Every model falls off past 256K. Argon falls least, and its 12.4-point lead over GPT-6 Astra in the long band is the widest margin in the whole benchmark table.

One caution from anyone who has shipped long-context systems: a high GraphWalks score does not make the full window cheap to use. Feeding 1M input tokens costs $2 per call at Argon’s introductory rate, and $4 at standard. Retrieval that narrows the context before the model sees it still pays for itself, which is why RAG pipelines remain the sensible default for corpora larger than a single call should carry.

How Google Is Already Using Argon Internally

Google is already running Gemini 4 Argon across its own engineering, and its September 30, 2026 announcement publishes four internal results with numbers attached. Thousands of Googlers have used the model, according to Kavukcuoglu, mostly on specialized coding, deeper research, and writing.

Internal use case (2026)Result Google reportsStatus
Quantum subroutine optimizationBeat the published baseline on spacetime resources (qubits × gates) by 40%, in minutesResearch result
Data-center memory efficiencyOver 300 TiB freed after rollout; 500 TiB to 1 PiB projected totalPartly rolled out
C/C++ to Rust migrationsFrom tens of thousands of lines (re2, libgav1) up to 800K+ lines for the Fuchsia OS Zircon kernelUnder audit before production
libgav1 SIMD rewrite32K lines of SIMD replaced with safe Rust; decoder runs 2.7x faster than the prior Rust port with identical outputCompleted

The memory result is the one with the clearest business shape. A team of Argon agents read fleet-wide profiling telemetry, picked out memory optimizations, and applied them across Google’s data centers on their own. The 300 TiB figure is what has landed so far.

The libgav1 case is the one I’d show an engineering lead. Argon agents started from an existing Rust port of Google’s open-source video decoder and ran many rounds of profile-guided experiments, reading the compiler’s output each time and rewriting hot code as safe Rust the compiler would vectorize on its own. Hand-written SIMD intrinsics are exactly the code most teams never get around to porting. The result runs 2.7x faster than the Rust port it replaced. Google’s own wording is that this brings it “closer to the optimized C++,” which means the C++ is still ahead.

Read the Zircon kernel line carefully. Google says the 800K-line migration is undergoing automated and manual auditing, emulation testing, and review before anything reaches production. The work is in progress. Nothing at that scale has shipped.

AlphaCorp AIonline
Let's talk

Curious what AI could do for your business?

No jargon and no hard sell. Just a friendly look at where AI fits, and where it doesn't.

View Services

Two things are absent from all four cases. Google reports no failure modes or rework rates, and every example runs on infrastructure and codebases Google controls. They show what Argon can do with unlimited compute and a home-field advantage. They don’t yet show what it does in yours.

Cybersecurity Defense: CWE-bench, Vulnerability Discovery, and the No-Guardrails Release

Gemini 4 Argon ties for first place on CWE-bench v1 at 68% Pass@1, finds vulnerabilities in real source code at 85.8% against Gemini 3.8 Flash Cyber’s 71.0%, and ships to trusted defenders with its cyber guardrails switched off. Cyber defense is the pillar Google chose to lead the launch with, and on September 30, 2026 it is the only one where the model is in outside hands.

Start with patching. CWE-bench v1 scores whether a model can remediate a security vulnerability, and the top of the September 2026 leaderboard is crowded:

CWE-bench v1, September 2026 (Pass@1, ties broken by Pass@4)Agent harnessScore
Grok 4.7opencode68%
Gemini 4 ArgonAntigravity68%
GPT-6 AstraCodex68%
Claude Opus 5.5Claude Code67%
Claude Fable 5.1Claude Code58%
Grok 4.6opencode57%
DeepSeek-V4.1-Flashopencode55%
Muse Spark 1.3opencode55%
Hy4 Previewopencode53%
GPT-6 SolCodex52%
Inklingopencode37%

Three models at 68%, a fourth at 67%. Anyone who has run a patching agent knows a one-point spread on a suite like this is noise. The harness column deserves more attention than the ranking. Each model ran inside a different agent scaffold, so every row compares a model-plus-tooling pair, and some of the gap between 68% and 58% may live in the scaffold. The generational jump is the cleaner signal: Gemini 3.8 Flash Cyber scored 47.2% Pass@1 on CWE-bench v0 at its September 2, 2026 launch, and Argon posts 68% on the successor v1 four weeks later.

Finding vulnerabilities is where Argon pulls away from its own predecessor:

  • Real-world vulnerability discovery (scanning source code across 20 programming languages): Argon 85.8% Pass@1, Gemini 3.8 Flash Cyber 71.0%.
  • Wiz Penetration Test Benchmark (exploiting live web systems with no source code): Argon 70.9%, Gemini 3.8 Flash Cyber 58.2%.
Grouped bar chart comparing Gemini 4 Argon with Gemini 3.8 Flash Cyber on three cybersecurity measures, Pass@1, September 2026. Real world vulnerability discovery across source code in 20 languages: Argon 85.8%, Flash Cyber 71.0%. Wiz Penetration Test Benchmark on live web systems with no source code: Argon 70.9%, Flash Cyber 58.2%. CWE-bench remediation: Argon 68% on version 1, Flash Cyber 47.2% on version 0. Argon leads on every measure.
Reading real source code across 20 languages, Argon reaches 85.8% Pass@1 against 71.0% for Gemini 3.8 Flash Cyber, the model it replaces four weeks later. Source: Google DeepMind, 2026.

The Wiz number is the one to watch because it is black-box. The model gets a running web system and has to map the attack surface, find the flaw, and produce proof-of-concept evidence, with no code to read. That is closer to what an attacker faces. Wiz is already running Argon inside Scan for Good, its free program for protecting critical public infrastructure, and Google reports the model found a critical vulnerability exposing personal data across healthcare software used by hospitals worldwide, a flaw earlier frontier models had missed. No CVE, vendor name, or disclosure timeline has been published.

Then the guardrails. Google will hand trusted defenders and its own internal teams a version of Argon “without cyber guardrails,” so the full exploit-finding capability goes to people patching systems instead of attacking them. It is a deliberate break from Google’s usual restricted posture, and it explains why the Fairwind gate exists at all. The 70.9% pen-test score belongs to the unguarded model. Whether the broadly released version, guardrails on, keeps the same discovery numbers is a question Google has not answered.

Safety Safeguards Before Broad Release

Google is holding Gemini 4 Argon back from general release while it strengthens four safeguards: misuse refusals under its Frontier Safety Framework, resistance to indirect prompt injection, chain-of-thought monitoring for misalignment, and sealed sandboxes for high-risk testing. The governing document is version 3.1 of Google DeepMind’s Frontier Safety Framework, published April 17, 2026, which defines Critical Capability Levels for cyber, CBRN, machine-learning R&D, and deceptive alignment.

Misuse defenses. Argon is trained to refuse requests that would help with cyber or chemical, biological, radiological, and nuclear attacks while still answering legitimate dual-use research. For this launch Google says it is also monitoring the model’s internal activations to spot misuse, and that internal and external red teams attacked the safeguards with both manual and automated methods. The results of that red-teaming are unpublished.

Prompt-injection resistance. Indirect prompt injection is when instructions hidden in a document, web page, or tool result take over the model’s behavior. Google reports Argon has the lowest attack success rate on Gray Swan AI’s Indirect Prompt Injection benchmark in 2026, measured at 15 attack attempts:

ModelAttack success rate (k=15, lower is better)
Gemini 4 Argon0.7%
Claude Opus 5.51.0%
Claude Fable 5.11.0%
Claude Opus 54.6%
Gemini 3.8 Flash5.5%
Gemini 3.8 Flash Cyber6.0%
GPT-6 Astra8.5%
GPT-6 Sol10.1%
Muse Spark 1.315.9%
GPT-5.6 Sol27.0%
GLM 5.331.5%
Grok 4.651.8%
Kimi K352.7%
Bar chart of indirect prompt injection attack success rates at 15 attempts in 2026, where lower is better, showing 8 of the 13 models Gray Swan tested. Kimi K3 52.7%, Grok 4.6 51.8%, GLM 5.3 31.5%, Muse Spark 1.3 15.9%, GPT-6 Astra 8.5%, Gemini 3.8 Flash 5.5%, Claude Opus 5.5 1.0%, and Gemini 4 Argon lowest at 0.7%, highlighted.
Argon’s 0.7% attack success rate is the lowest of the 13 models Gray Swan tested, against 5.5% for Gemini 3.8 Flash and 8.5% for GPT-6 Astra. Source: Gray Swan AI via Google DeepMind, 2026.

Gray Swan is a private red-teaming vendor, with an arena co-designed alongside the UK AI Security Institute and the US CAISI, so treat this as industry benchmarking. The spread still matters. If you build agents that read inboxes or browse the web, this table should worry you more than the coding one, and a drop from 5.5% on Gemini 3.8 Flash to 0.7% on Argon is the kind of gain that changes what you can safely automate. Zero would be better. An attacker also gets more than 15 tries.

Misalignment monitoring. A monitor reads Argon’s chain-of-thought and actions and halts execution when the model tries to finish a task in a way that goes beyond what the user asked. Google ran a similar monitor during training, routing alerts to a dedicated incident response team, and says it deliberately kept those findings out of the training loop so the model would never learn to hide its reasoning from the monitor.

“We strongly encourage the rest of the industry to preserve reasoning transparency in these pivotal moments of increased capabilities.” Koray Kavukcuoglu, Google DeepMind, September 30, 2026

Hardened sandboxes. Google now isolates and seals its sandboxed environments before any high-risk training or evaluation run, following its agent control roadmap, and says it will share those practices with partners.

What the Benchmarks Don’t Tell You Yet

Every Gemini 4 Argon figure public on September 30, 2026 comes from Google, has been replicated by no outside lab, and arrives without a technical or safety report. That doesn’t make the numbers wrong. It makes them claims.

Five gaps sit behind the scores:

  • No independent replication. Argon is confined to the Fairwind Program, so nobody outside it can rerun DeepSWE, the Vals Index, or CWE-bench and check.
  • No model report. Google’s Frontier Safety Framework page lists safety reports for Gemini 3.7 Flash and Gemini 3 Pro as its most recent, and none for Argon. Google published technical reports for Gemini 2.5 in July 2025 and Gemma 4 in July 2026, so one may follow.
  • Different suites at each lab. Anthropic’s September 22, 2026 Claude Opus 5.5 materials report Terminal-Bench 4.0 at 66.4%, FrontierCode v1.1 at 54.4%, and a GDPval-AA v2.1 Elo of 1846. OpenAI’s September 3, 2026 GPT-6 Astra post reports OSWorld 2.0 at 72.6% and 100% on its internal ExploitBench. FrontierCode, GDPval, and ExploitBench appear nowhere in Google’s 19 rows. Each lab publishes the suite where it looks strongest.
  • Harness effects. The CWE-bench leaderboard ran each model in a different agent scaffold, so it ranks model-and-tooling pairs.
  • Pre-announcement leaks. Argon circulated on arena-style platforms in the days before the launch. Those scores are unverified and should stay separate from Google’s published ones.

None of this is unusual for a frontier launch. It is the reason to run your own eval before you move a workload.

Gemini 4 Argon FAQ

These are the questions people are searching about the Gemini 4 Argon launch on September 30, 2026, answered in a few sentences each.

Built for production

What could a custom AI agent take off your plate?

We build production-grade AI systems that quietly handle the busywork, so your team can focus on the work that actually matters.

View Services

Is Gemini 4 Argon available now?

No, unless you are in Google’s Fairwind Program or work at Google. On September 30, 2026, Argon is limited to a vetted set of cyber defenders, government agencies, and Google Cloud security partners. Paid API customers and Google AI Ultra subscribers are next in line, and Google has published no date for that stage.

How much does Gemini 4 Argon cost?

The introductory API price is $2 per million input tokens and $10 per million output tokens, with cached input at 95% off, or $0.10 per million. Standard pricing after the introductory window is $4 input and $20 output per million tokens. Consumer access is expected through the $100 and $200 per month Google AI Ultra tiers set at I/O in May 2026.

Is Gemini 4 Argon better than GPT-6 Astra or Claude Opus 5.5?

On Google’s own 19-benchmark table from September 2026, Argon leads 13 rows outright and ties one. It trails GPT-6 Astra on FrontierSWE v2 (55.0% vs 65.5%) and trails Claude Opus 5.5 on Terminal-bench 4.0 (57.4% vs 66.4%). Those figures are Google-reported and unreplicated, and OpenAI’s GPT-6 Astra announcement from September 3, 2026 and Anthropic’s Opus 5.5 materials use partly different suites, so no single answer holds across every workload.

What is the Fairwind Program?

Fairwind is Google’s controlled-access program for cyber defense. It gives a trusted group of Google Cloud customers, government agencies, and cybersecurity partners early use of Argon, without cyber guardrails, to find and fix vulnerabilities in critical infrastructure and public services. Membership is by Google’s invitation.

What is Gemini 4 Argon’s output limit?

Argon can generate up to 1 million output tokens in a single response, up from 64K in the prior Gemini generation. Google reports GraphWalks long-context accuracy of 99.7% up to 128K tokens and 84.2% in the 256K to 1M range as of September 2026.

When will Gemini 4 Argon come to the Gemini app?

Google has not said. The announcement names paid API customers and Google AI Ultra subscribers as the first groups after the Fairwind phase, and “developers, enterprises, and consumers” after that, with the timing described only as “as soon as possible.” Treat any specific date circulating elsewhere as unconfirmed.

Does Gemini 4 Argon have a technical report?

Not as of September 30, 2026. Google’s Frontier Safety Framework page lists no Argon safety report, and no arXiv paper exists. Google published technical reports for Gemini 2.5 in 2025 and Gemma 4 in July 2026, so one is plausible but unannounced.

How to Prepare for Gemini 4 Argon Access

The best preparation for the Gemini 4 Argon launch is to build the evaluation now, before the model reaches your API key, so the day access opens you can measure instead of guess. Five moves fit that window:

  1. Assemble a task set from your own work, sized to the long-horizon jobs Argon is built for, and run it against Gemini 3.8 Flash today so you have a baseline.
  2. Model spend at both $2/$10 and $4/$20 per million tokens, using output tokens per completed task as the driver, since the introductory rate will end.
  3. Add cache-friendly context layout to your agents now so the $0.10 cached-input rate pays off from the first call.
  4. Set a hard output cap per trajectory. A 1M-token ceiling is headroom, and headroom without a budget is a bill.
  5. If you run security tooling on Google Cloud, ask your account team about Fairwind.

Watch for three signals next: an Argon technical report, the first outside replication of DeepSWE or CWE-bench, and the standard-pricing switch date.

Run the eval before the model arrives. If you’d rather have engineers who’ve shipped production agents build that harness and cost model with you, talk to AlphaCorp AI about your Argon readiness.

Share
Newsletter · Weekly

Stay Ahead of AI

One email per week with the AI engineering insights, agent builds, and tools that actually matter.

No spamUnsubscribe anytimeFree forever

In every issue
  1. 01One agent build, taken apart step by step
  2. 02The tools that earned a place in our stack this week
  3. 03What broke in production, and what we changed

Written by Ignas Vaitukaitis, founder of AlphaCorp AI.

Wireframe cubes of circuitry linked by glowing strands above a dark circuit-board floor

Ready to Ship
Your AI System?

Book a free call and let's talk about what AI can do for your business. No sales pitch, just a real conversation.