Picture a model that replaces 32,000 lines of hand-written SIMD code with safer, auto-vectorized Rust, then produces a memory-safe video decoder that runs 2.7 times faster than the version it replaced. Same video output. Memory-safe Rust. And the work was done by Argon agents inside Google’s engineering workflows.
That’s Gemini 4 Argon in practice. Not just a benchmark. It’s already being used inside Google’s internal engineering workflows.
Google announced Gemini 4 Argon on September 30, 2026. It’s the company’s new frontier model, developed at Google DeepMind and introduced by SVP Koray Kavukcuoglu. It sits above the Gemini 3.8 line and is specifically built for work that runs for hours rather than seconds: large-scale software engineering, enterprise knowledge work across legal and finance, and cybersecurity defense. Here’s what it is, where it performs, where it doesn’t, and what you can actually do with it right now.
The One Engineering Change That Defines Argon
Before anything else: the output token limit.
Previous Gemini models topped out at 64,000 output tokens. Gemini 4 Argon raises that ceiling to 1 million tokens in a single generation pass. That’s a 16x increase. And according to Google’s announcement, this isn’t just about length. It’s about what the extra headroom does to reasoning quality.
When a model can generate across hundreds of thousands of tokens in one trajectory, it has more room to sustain a long task without being forced to stop, summarize, and restart around an output limit. For anyone who has run a long agentic coding session and watched a model lose the plot halfway through a refactor because it hit an output ceiling, this is the constraint being removed.
For context: GPT-6 Astra caps output at 128,000 tokens. Claude Opus 5.5 raises its limit to 128,000 tokens as well. Argon at 1M is in its own tier on that specific dimension.
The caveat, noted clearly in DataCamp’s coverage of the announcement: Argon generates output at $10 per million tokens (introductory) and $20 per million tokens (standard). Long reasoning traces at that rate add up fast. Google has also not published whether reasoning tokens bill at the output rate. That’s an open question worth tracking before you build an agent fleet around it.
What Google Is Already Using Argon For
The internal deployments are the most specific evidence available about what Argon can actually do. Google published three examples in the announcement.
Codebase migration at scale. Argon agents are migrating C and C++ codebases to Rust inside Google, from core libraries like re2 and libgav1 through to the Fuchsia OS Zircon kernel at over 800,000 lines. Google notes these rewrites go through automated and manual auditing, emulation testing, and review before production. They’re not just running agents and shipping.
The libgav1 example is specific enough to be worth quoting from the source. Per Google’s announcement: Argon agents took an existing Rust port of the video decoder and replaced 32,000 lines of hand-written SIMD code by running rounds of profile-guided experiments and studying compiler output until safe Rust could be vectorized automatically. The result runs 2.7 times faster than the Rust port, with identical video output.
Memory optimization across Google’s data centers. According to Google’s announcement, a team of Argon agents analyzed fleet-wide profiling telemetry and autonomously applied memory optimizations across Google’s data centers. Google reports over 300 TiB of memory freed once rolled out, with an estimated 500 TiB to 1 PiB in total savings. Neither figure has external verification, but both are specific enough to be checkable as results become available.
Quantum algorithm optimization. Argon is helping Google’s quantum computing researchers optimize the spacetime resources (qubits multiplied by gates) of subroutines that bottleneck important applications. Google says it beat a published baseline by 40% in minutes in one case. Again, vendor-reported. Worth noting specifically because of its precision.
What Gemini 4 Argon Is Built For: The Three Domains
Google framed Argon’s announcement around three specific domains. Each has its own benchmark story, covered properly in the comparison blog in this cluster.
Real-world software engineering
Google reports a 77.9% score on DeepSWE v1.1, the highest in the comparison published in its announcement. DeepSWE measures long-horizon software engineering against real repositories, rewarding models that can hold a plan across many edits rather than producing one good diff. The result is particularly relevant to the long-horizon repository work Google describes.
On Vibe Code Bench, Argon scores 91.9%, and on FrontierCode, 61.2%. More detail in the benchmarks post.
The pattern of strengths: Argon is a planner and reader. It holds context across a long task and coordinates across many files. Where it still trails on software engineering is in terminal-driven execution, which is covered honestly in the comparison section below.
Enterprise knowledge work
This is where the margins over competitors are widest, per Google’s published benchmarks. The Vals Index, which weights finance, coding, legal, and tax work by U.S. GDP contribution, gives Argon a score of 68.9%. Vals Finance Agent v2, measuring multi-step financial research, gives it 65.4%. AutomationBench, Zapier’s benchmark for end-to-end business workflow execution, gives it 51.3%, ranking it first.
Harvey’s Legal Agent Benchmark shows Argon at 19.6%, versus 6.7% for Fable 5.1 and 5.4% for GPT-6 Astra. Every model is failing most of this benchmark. The result is a relative signal that Argon performed substantially better than the other models tested on this benchmark, not evidence that any model is ready to draft contracts without supervision.
Argon is also notably strong when knowledge work requires visual understanding: reading charts, pulling details from long videos, acting across a series of documents. On LVBench, which measures long video understanding, Google reports Argon at 91.7%.
Cybersecurity defense
Argon can autonomously find, validate, and patch software vulnerabilities. That’s the stated capability, and the deployment model reflects how seriously Google takes it: for trusted defenders in the Fairwind Program and for Google’s internal teams, Argon ships without cyber guardrails so those teams get its full defensive capability.
Wiz is already using Argon through its Scan for Good initiative. In an early demonstration, per Google’s announcement, Argon uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, one that previous frontier models had missed. On CWE-bench v1, which measures vulnerability remediation, Argon ties GPT-6 Astra for first place at 68%.
Google also notes Argon is the company’s most resilient model against indirect prompt injection, leading Gray Swan’s IPI benchmark after automated red teaming and adversarial training.
What Gemini 4 Argon Is Multimodal On
Argon supports multimodal workflows, including visual understanding and long-video analysis. On the multimodal side, the strongest documented capability is long video understanding (LVBench: 91.7%) and chart analysis (Chartography: 71.6%, per Google’s benchmarks). The 1M output token limit applies to text generation only.
Argon is not a standalone image or video generator. The multimodal capabilities are about understanding and reasoning over visual content, not producing it.
What Argon Can’t Do Well (Per Published Data)
Google published its losses alongside its wins, which makes the announcement more credible than a clean-sweep table would.
FrontierSWE v2: Argon scores 55.0%. GPT-6 Astra scores 65.5%, and Claude Opus 5.5 scores 62.3%. A 10.5-point deficit on the harder of the two software engineering suites. Per DataCamp’s analysis of the pattern: Argon’s published scores are higher on several knowledge-work and long-context evaluations, while Astra and Opus 5.5 score higher on FrontierSWE v2 and Terminal-Bench 4.0.
Terminal-Bench 4.0: Argon scores 57.4%. Claude Opus 5.5 scores 66.4%. For terminal-driven agent loops, Argon trails by roughly 9 points.
Terminal-Bench Science 0.1: Argon scores 57.6%. GPT-6 Astra scores 68.1%. Scientific tasks executed through a shell, 10.5 points behind.
OSWorld-2.0 (offline subset): Argon at 69.2%, Astra at 72.6%.
These are Google’s own published numbers, not from independent third parties. No external researcher has yet reproduced any Argon score.

Safety: What Google Is Doing Before Wide Release
Google is holding Argon back from general availability while hardening four safeguard areas, per the announcement.
Misuse. Argon refuses cyber and CBRN (chemical, biological, radiological, nuclear) requests while preserving dual-use scientific research, per Google’s Frontier Safety Framework. Internal activations are monitored to spot misuse, and internal and external red teams ran both manual and automated attacks before the announcement.
Indirect prompt injection. Argon leads Gray Swan’s IPI benchmark after automated red teaming and adversarial training. These are complex, multi-layer attacks that require constant vigilance, and Google says it deliberately kept monitoring findings out of training so the model’s reasoning would not be shaped to evade the monitors.
Misalignment. Monitors watch chain-of-thought and actions during execution and halt when Argon exceeds user intent. The same monitoring system applied during training, with an incident response team receiving alerts.
System hardening. Sandboxed environments are isolated and sealed before high-risk training or evaluation runs, following Google’s agent control roadmap.
The safety framing also matters for the rollout sequence. Argon goes to a cohort of trusted cyber defenders through the Fairwind Program before anyone else, partly because the cybersecurity capability specifically requires frontier safeguards to be hardened before the model is in broad reach.
Availability: Where Things Actually Stand
As of September 30, 2026, Argon is not publicly available.
Access is limited to trusted cyber defenders in the Fairwind Program and Google’s internal teams. Google says broader access starts with paid API customers and Google AI Ultra subscribers, with no announced timeline.
As of the announcement date, per DataCamp’s coverage: no public model ID exists, Argon is not listed in the OpenRouter or models.dev catalogs, and it does not appear in Vertex AI, Gemini CLI, Cursor, or GitHub Copilot model documentation.
Google is engaged in the U.S. government’s voluntary process for pre-release model access while it expands gradually.
Frequently Asked Questions
What is Gemini 4 Argon?
Gemini 4 Argon is Google’s frontier model, announced September 30, 2026 by Google DeepMind. Per Google’s announcement, it’s built to sustain deep reasoning across complex, long-horizon workflows in software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. The key structural change is an output token limit of 1M, up from 64K on the previous generation.
Who made Gemini 4 Argon?
Google DeepMind, with the announcement authored by Koray Kavukcuoglu, SVP of Google DeepMind and Chief AI Architect at Google.
What is Gemini 4 Argon used for?
Per Google’s announcement, it targets real-world software engineering (including large-scale codebase migrations and rewrites), enterprise knowledge work requiring multi-step reasoning across legal, financial, and business domains, and defensive cybersecurity including autonomous vulnerability finding, validation, and patching.
What is Gemini 4 Argon’s context window and output limit?
The output token limit is 1M, per Google’s announcement. The input context window figure is not separately published in the announcement. Google describes the 1M output limit as the primary structural change, arguing that the headroom changes reasoning quality rather than just enabling longer responses.
How does Gemini 4 Argon handle multimodal tasks?
Argon supports multimodal workflows including visual understanding and long-video analysis. On long video understanding (LVBench), Google reports 91.7%. On chart analysis (Chartography), 71.6%. The multimodal capabilities are for understanding and reasoning over content, not generating images or video.
What are Gemini 4 Argon’s limitations?
Per Google’s own published benchmarks, Argon trails GPT-6 Astra on FrontierSWE v2 (55.0% vs 65.5%) and Terminal-Bench Science 0.1 (57.6% vs 68.1%), and trails Claude Opus 5.5 on Terminal-Bench 4.0 (57.4% vs 66.4%). No external researcher had independently reproduced any benchmark score as of the announcement date.
Is Gemini 4 Argon safe?
Google describes hardened safeguards across misuse prevention, prompt injection defense, misalignment monitoring, and system hardening. It refuses cyber and CBRN harmful requests. For trusted cybersecurity defenders through the Fairwind Program, it ships without cyber guardrails specifically to support defensive work.
When will Gemini 4 Argon be widely available?
Google says broader release starts with paid API customers and Google AI Ultra subscribers, after the initial Fairwind Program rollout. No timeline has been published.
Final Thoughts
Gemini 4 Argon has documented capabilities and numbers. The libgav1 rewrite, the data center memory optimization, the vulnerability discovery at Wiz: these go beyond standalone benchmark scores. Google says Argon is already being used in internal engineering and research workflows, although the reported results remain vendor-reported and subject to the review processes Google describes.
What’s still pending is everything outside the Fairwind cohort: independent benchmark reproduction, public API access, pricing behavior in real workloads, and the answer to whether reasoning tokens bill at the output rate. No one outside Google’s trusted circles has touched Argon yet, and vendor-published benchmarks without external reproduction carry their usual asterisks.
The capabilities are documented, but public availability remains limited. If you’re planning an enterprise deployment, the key milestone to watch is the paid API launch, for which Google has not yet announced a specific date.
All benchmark figures, internal deployment details, and feature descriptions in this article are drawn from Google’s official Gemini 4 Argon announcement at blog.google, published September 30, 2026, authored by Koray Kavukcuoglu. Competitor benchmark figures and additional context are drawn from DataCamp’s coverage of the same announcement. No independent third-party reproduction of any Argon benchmark score was available at publication time.
Curated byย Lorphic
Digital intelligence. Clarity. Truth.