Gemini 4 Argon Debuts With 1M-Token Output, 77.9% on DeepSWE and $2 Pricing

News Summary
Google announced Gemini 4 Argon at 1:23 p.m. Pacific Time on September 30, 2026 (4:23 p.m. Eastern Time), presenting it as its most powerful AI model to date. Built for long-horizon work such as real-world software engineering, legal and financial analysis, and cyber defense, Argon is the company's first flagship release since the Gemini 3 series arrived in November 2025. Access is initially limited to a small group of trusted partners, with a wider rollout planned afterward.
What Google Announced
Google describes Gemini 4 Argon as the start of its "next era of frontier intelligence." The model is aimed at tasks that unfold over many steps and long periods, such as fixing a large codebase, drafting a detailed legal document, or researching a financial question from many sources. Google also highlights improved understanding of visual material, including long videos and complex charts.
The headline engineering change is the output limit. Argon can generate up to 1 million tokens in a single response, up from 64,000 in earlier Gemini models. That makes it possible to produce very long documents, large code changes, or extended reasoning traces without stopping midway.
Benchmark Results
According to reported figures, Argon leads or ties for the top score in 13 of 18 benchmarks tracked in coverage of the launch. That compares with three outright leads for OpenAI's GPT-6 Astra and two for Anthropic's Claude Opus 5.5.
Selected results include:
- DeepSWE v1.1, a software engineering test: Argon scored 77.9%, versus 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra.
- AutomationBench: Argon scored 51.3%, versus 42.5% for Claude Opus 5.5 and 41.4% for GPT-6 Astra.
- Harvey's Legal Agent benchmark: Argon scored 19.6%, versus 5.4% for GPT-6 Astra and 3.8% for Claude Opus 5.5.
- LVBench, a long-video understanding test: Argon scored 91.7%, versus 87.5% for GPT-6 Astra and 83.7% for Claude Opus 5.5.
Argon did not win everywhere. GPT-6 Astra stayed ahead on FrontierSWE v2 (65.5% versus 55.0%) and Terminal-Bench Science 0.1 (68.1% versus 57.6%). As with any vendor-reported benchmark, independent testing will show how the numbers hold up in everyday use.
Cyber Defense Focus
Google says Argon was trained with defensive security in mind and can autonomously find, validate, and patch serious software vulnerabilities. For that reason, the first access goes to a set of trusted cyber defenders through the Fairwind Program, which Google introduced earlier in September. The idea is to let security teams strengthen software before the model reaches a broad audience.
Pricing
Google set an introductory price of $2 per million input tokens and $10 per million output tokens. Cached input tokens cost $0.10 per million, a 95% discount on the standard input rate. For comparison, GPT-6 Astra is reported at $10 and $50 per million input and output tokens.
After the introductory period, Argon is expected to cost $4 per million input tokens and $20 per million output tokens, which matches the base pricing reported for Claude Opus 5.5.
Availability and What Comes Next
Argon is not yet broadly available. Google says it is taking part in a voluntary pre-release model access and testing process before widening the rollout. The broader release is expected to begin with paid API customers and Google AI Ultra subscribers.
For developers and students, the key points are the very large output window, the low introductory price, and the model's strength on long, multi-step tasks. Many will wait for general availability and independent evaluations before deciding whether to build on it.