From code to cybersecurity, Gemini 4 Argon takes on longer workflows
Gemini 4 Argon targets long-running tasks across software development, finance, legal work, and cybersecurity with an output limit expanded to one million tokens. Google is initially restricting its new frontier model to selected testers before a broader release.
One million tokens for a single trajectory
A model capable of producing hundreds of thousands of tokens during a single task. That is one of the defining features of Gemini 4 Argon, Google DeepMind's new frontier model, which increases the output limit from 64,000 to one million tokens.
The additional headroom targets problems that require sustained reasoning and long sequences of actions. Google is positioning Argon across software engineering, professional finance and legal work, and defensive cybersecurity.
Access remains controlled for now. An initial cohort of cyber defenders and trusted testers is using the model through Google's Fairwind Program before access expands to developers, enterprises, and consumers. Code migrations are already underway inside Google
Argon is already part of several internal Google workflows. Agents are being used to migrate C/C++ codebases to Rust, ranging from libraries with tens of thousands of lines to the Fuchsia Zircon kernel at more than 800,000 lines.
For libgav1, Google's open-source video decoder, Argon agents worked from an existing Rust port and replaced 32,000 lines of SIMD code. After multiple rounds of profile-guided experimentation, Google reports a memory-safe decoder running 2.7x faster than the previous Rust port while producing identical video output.
Another internal experiment used agents to analyze data center telemetry and identify memory optimizations. Google says deployed changes have already freed more than 300 TiB of memory, with estimated total savings ranging from 500 TiB to 1 PiB if all identified optimizations are implemented. 77.9% on DeepSWE v1.1
Google reports a 77.9% score for Argon on DeepSWE v1.1, which evaluates long-horizon software engineering tasks. The model also reaches 51.3% on AutomationBench for end-to-end professional workflows and 91.7% on LVBench for long-video understanding.
Google also places Argon at the top of the Vals Index, which combines evaluations across finance, coding, legal, and tax work based on their contribution to US GDP. Results are also reported for Vals Finance Agent v2 and Harvey's Legal Agent Benchmark.
These figures come from the evaluations published alongside the model. They cover different protocols and should not be treated as a single measure of real-world performance. Cybersecurity explains part of the restricted rollout
Argon's cybersecurity capabilities are one reason for the phased release. The model was trained to identify, validate, and patch vulnerabilities, with approved defenders and Google's internal teams receiving access without some of the standard cyber guardrails.
On CWE-bench v1, Google reports a 68% score, tied for the highest result in its comparison. Wiz is also testing Argon through its Scan for Good initiative for protecting critical infrastructure.
Google says an early deployment with Wiz uncovered a critical vulnerability affecting healthcare software used by hospitals after previous frontier models had missed it. The announcement does not provide enough technical information to independently assess that case.
The Fairwind Program controls early access, prioritizing governments, critical infrastructure operators, technology platforms, and selected cybersecurity teams. A phased release before broader API access
Google is still strengthening safeguards against malicious cyber and CBRN use, indirect prompt injection, and agent behavior that deviates from user intent.
Argon includes systems designed to monitor its reasoning and actions and stop execution when certain conditions are detected. Google is also isolating and hardening the environments used for higher-risk training and evaluations.
When commercial access expands, Gemini 4 Argon is set to carry an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at a 95% discount.
Paid API customers and Google AI Ultra subscribers are expected to be among the first users included in the broader rollout. Google has not provided a specific date for that stage.