Google launched Gemini 3.8 Flash in September 2026, just weeks after releasing its predecessor, Gemini 3.7 Flash. The new model is designed to perform more reasoning steps on complex tasks and call tools iteratively, which Google describes as the model “working harder.”
While the per-token pricing remains unchanged from 3.7 Flash — $0.75 per million input tokens and $3.75 per million output tokens — Google warns that actual costs could be higher. The company notes that “the model might use more tokens to maximize performance, especially at higher effort levels.” Early analysis from Artificial Analysis found total costs running roughly 40% higher than 3.7 Flash, driven by a 30% increase in output tokens per task and more turns on agentic evaluations. Developers who want to limit token usage can continue using Gemini 3.7 Flash.
Google says 3.8 Flash delivers significant improvements for software engineering and autonomous AI agents. The model outperforms its predecessor and several competitors on the DeepSWE v1.1 software engineering benchmark, as well as the Vals Finance Agent V2 and Harvey’s Legal Agent benchmarks. Aigora.ai CEO John Ennis described the model’s coding quality as comparable to Anthropic’s Opus 5 “but at a fraction of the cost and super fast.”
The model ships with built-in safeguards against misuse in Chemical, Biological, Radiological, and Nuclear (CBRN) domains and cyber offense. Alongside the main release, Google introduced Gemini 3.8 Flash Cyber and a new Fairwind Program, restricted to governments and trusted partners. The program’s 650 members — including CrowdStrike and the Center for Internet Security — gain access to 3.8 Flash Cyber and Google’s CodeMender agent, which Google says can autonomously find and fix vulnerabilities in critical infrastructure and public services.
Gemini 3.8 Flash is available now for consumers with a Google AI Pro or Ultra subscription, as well as developers and enterprise users.
Source: The Verge