Google has announced Gemini 4 Argon, a new frontier AI model built for long-horizon reasoning across software engineering, finance, legal work, and cybersecurity. The model raises the maximum output limit from 64,000 tokens to 1 million tokens, enabling complex multi-step tasks to run in a single continuous process. It is currently rolling out to trusted cyber defenders through the Fairwind Program before a broader release.
Expanded Context and Core Performance
Gemini 4 Argon is designed to reason for longer periods and complete complex workflows without interruption. Google reports that the model achieves 77.9% on DeepSWE v1.1 for long software engineering tasks and 91.7% on LVBench for understanding long video content. These results place Argon among the leading models for real-world coding and multimodal enterprise work.
The model also posts top rankings on the Vals Index, which measures economic impact across finance, coding, legal, and tax work. It ranks first on AutomationBench with a score of 51.3% for end-to-end execution across core business functions. Those gains support use cases such as financial research, legal drafting, and agentic business automation.
Cybersecurity Defense Focus
Cybersecurity is a central focus of the Argon release. Google says trusted specialists will receive access without cyber guardrails so they can thoroughly penetration-test systems and use the model for defensive work. Security firm Wiz has already used Argon through its Scan for Good initiative to identify a critical data exposure in healthcare software used worldwide.
On CWE-bench v1, which evaluates vulnerability remediation, Argon ties for first place with a score of 68%. Google states that the model can autonomously find, verify, and patch vulnerabilities across 20 programming languages. The company is also working with U.S. government voluntary processes to manage pre-release access.
Inside Google and Availability
Google has already integrated Argon into internal engineering workflows. Agents are helping migrate C and C++ codebases to memory-safe Rust, including the Fuchsia Zircon kernel with more than 800,000 lines of code. Teams also used the model to free up over 300 TB of memory in data centers through autonomous telemetry analysis.
In quantum computing research, Argon optimized spacetime resources for subroutines that bottleneck important applications and beat a published baseline by 40% within minutes. The model also improved a Rust video decoder replacement, producing safe Rust that runs 2.7 times faster than the existing Rust port while matching output. These internal results illustrate how long-context reasoning can translate into operational savings.
Pricing, Safeguards, and Rollout
Introductory pricing is set at US$2 per 1 million input tokens and US$10 per 1 million output tokens. Cached input tokens are priced at 95% off the input token price. After the introductory period, pricing will apply at US$4 per 1 million input tokens and US$20 per 1 million output tokens.
Google is strengthening safeguards before broader release across four main areas: misuse defense, prompt injection resistance, misalignment monitoring, and system hardening. The company says Argon is its most resilient model yet against indirect prompt injection attacks, according to the Gray Swan Indirect Prompt Injection benchmark. Early access is expected to expand to paid API customers and Google AI Ultra subscribers first, followed by developers, enterprises, and consumers.
Gemini 4 Argon represents a significant step toward AI models that sustain deep reasoning across long, multi-step workflows. Its combination of expanded context, cybersecurity focus, and internal engineering impact highlights a practical shift beyond short-form model outputs. Google will continue gathering feedback from trusted testers as it gradually widens access to developers, enterprises, and consumers.