Home / News / Gemini 3.8 Flash Costs 40% More Per Task Despite Flat Pricing
Gemini

Gemini 3.8 Flash Costs 40% More Per Task Despite Flat Pricing

Sep 3, 20265 min read
Gemini 3.8 Flash Costs 40% More Per Task Despite Flat Pricing

News Summary

Google introduced Gemini 3.8 Flash on September 2, 2026, positioning it as its most capable "workhorse" model to date, one built to tackle software engineering and multi-step reasoning tasks by deliberately doing more computational work per request. Alongside it, Google released a specialized variant, Gemini 3.8 Flash Cyber, aimed at automated vulnerability discovery and patching for trusted government and infrastructure partners. The launch, Google's fourth Flash-tier release in about four months, arrives as the company works to reassure developers and enterprise customers that its model release cadence remains steady even as rivals continue to ship competitive systems.

What "Works Harder" Actually Means

Google's framing centers on a deliberate design choice: rather than optimizing purely for speed or minimal token usage, Gemini 3.8 Flash is tuned to execute additional reasoning steps and call external tools iteratively when a task calls for it. At higher "effort" settings, the model will consume more tokens in pursuit of a more thorough or accurate answer, a tradeoff Google frames as improving real-world task completion rather than just raw benchmark scores. Developers can dial this behavior down by selecting lower effort levels for latency- or cost-sensitive workloads, or continue using the prior Gemini 3.7 Flash model, which Google says remains fully supported for efficiency-first use cases.

Pricing and the Hidden Cost Increase

On paper, Gemini 3.8 Flash launches at the same introductory per-token price as its predecessor: $0.75 per million input tokens and $3.75 per million output tokens, effective through December 31, 2026 Eastern Time, after which standard pricing rises to $1.50 and $7.50 respectively starting in January 2027. However, because the model tends to generate more output tokens and take more turns when operating as an autonomous agent, several outlets that tested the model estimate real-world task costs run roughly 40 percent higher than with the previous Flash model, driven largely by the extra reasoning consuming close to 30 percent more output tokens per task. In effect, the sticker price stayed flat, but the total bill for a given job can climb because the model is doing, and billing for, more work to get there.

Benchmark Performance

Google says Gemini 3.8 Flash outperforms several larger frontier models on the DeepSWE v1.1 software engineering benchmark while costing a fraction as much to run, and it posted a 54.9 percent score on HLE-Verified, a benchmark spanning STEM, humanities, and professional reasoning tasks. Independent evaluation from Artificial Analysis placed the model at 59 on its Intelligence Index when run in high-reasoning mode, a three-point gain over the prior Flash release, putting it in the same tier as other recently released mid-size frontier models from competing labs, though still trailing the top-ranked flagship models on that index. On cost-efficiency measures, Artificial Analysis calculated roughly $0.58 per intelligence-index task for Gemini 3.8 Flash, making it one of the least expensive options at its performance tier among models tracked by the firm.

The Cybersecurity Variant

Gemini 3.8 Flash Cyber, offered to trusted testers through Google's Fairwind Program for government and critical-infrastructure partners, is tuned specifically for vulnerability detection and automated patch generation. Google reports the model exceeds prior versions on the CyberGym benchmark for vulnerability discovery, with more than 70 percent success identifying real-world vulnerabilities across roughly 20 programming languages, and a 47.2 percent pass rate on the CWE-Bench patch-generation benchmark at a notably lower operating cost than comparable commercial tools. Google also cited feedback from its own Chrome Security team, which reported the Cyber model produced 2.6 times more correct patches than the leading commercial alternative it had been using.

Availability and Rollout

Gemini 3.8 Flash is rolling out across Google's developer and consumer surfaces, including Google AI Studio, Android Studio, and the Google Antigravity coding environment for developers, Gemini Enterprise for business customers, and the consumer Gemini app, Google Search's AI Mode, and Gemini in Google Sheets for Google AI Pro and Ultra subscribers. The Cyber variant remains limited to Fairwind Program participants for now, without a stated timeline for broader availability.

Why the Release Cadence Matters

Industry observers note that this marks Google's third or fourth Flash-tier model release within a matter of weeks, a pace some read as an effort to keep pace with rapid releases from other major AI labs and to demonstrate continuity in Google DeepMind's model roadmap following recent internal leadership changes. The company had previously signaled a larger Gemini 3.5 Pro release earlier in the year that has yet to materialize, and the steady cadence of smaller Flash updates appears intended, in part, to keep Gemini's developer and enterprise ecosystem engaged while that larger release remains pending.

GeminiGoogle AI