Google began rolling out its Gemini 3.8 Flash AI model on September 2. It follows 3.7 Flash, released three weeks earlier, and is the third Flash-series release in six weeks. The model improves on software engineering, agentic tasks, and multi-step reasoning in specialized fields while keeping the same introductory pricing as 3.7 Flash. What Google highlights this time is a design philosophy: the harder the task, the more reasoning steps the model takes and the more it calls tools repeatedly. In Google's words, 3.8 Flash "works harder."

A New Flash Generation Every Three Weeks

The Flash series is the Gemini line that prioritizes speed and per-token cost for production workloads. This summer the update cadence has been striking: 3.6 Flash, 3.7 Flash, and now 3.8 Flash have arrived roughly three weeks apart. Independent evaluator Artificial Analysis counts it as the fourth Flash model in under four months.

Google positions 3.8 Flash as its "most intelligent workhorse model" and says it shares its foundational intelligence with Gemini 3.8 Flash Cyber, a cybersecurity-focused variant announced the same day. According to Google, rigorous training in the demanding domain of cybersecurity contributed to gains in coding and reasoning across that shared core, meaning security-oriented R&D is flowing back into the general-purpose model. The Cyber variant and the Fairwind Program that provides it to trusted defenders are covered in a separate article, so this piece focuses on the general-availability 3.8 Flash.

What Improved on the Benchmarks

In the results Google published, the biggest gains show up on tasks that require long, autonomous operation.

On DeepSWE v1.1, which measures long-horizon software engineering, Google says 3.8 Flash outperforms most larger frontier models at autonomously solving complex engineering problems end to end, and does so at a fraction of the cost.

In specialized fields, it beat 3.7 Flash and other frontier models on Vals Finance Agent V2, a financial-analysis agent benchmark, and on Harvey's Legal Agent Benchmark. On HLE-Verified, a cross-disciplinary set of hard problems, it scored 54.9 percent, which Google cites as evidence of multi-step reasoning across STEM, the humanities, and professional domains.

All of those figures come from Google. In measurements published the same day by Artificial Analysis, the model scored 59 on the Intelligence Index at high reasoning effort, up 3 points from 3.7 Flash's 56. The firm attributes most of the gain to agentic evaluations, with the largest jump on τ³-Banking, a tool-use benchmark built around banking tasks, where the model gained 12 points over 3.7 Flash to reach 45 percent. At medium effort it scores 57 and at low effort 52; the low setting matches Gemini 3.6 Flash at high effort while costing 30 percent less per task and taking roughly a third of the time.

Working Harder Means More Tokens

Google is unusually direct about where the gains come from. On complex tasks, 3.8 Flash is built to show greater diligence, executing extra reasoning steps and calling tools iteratively. As a result, Google states plainly that the model may use more tokens to maximize performance, especially at higher effort levels.

Independent measurements bear this out. According to Artificial Analysis, even though per-token pricing is unchanged from 3.7 Flash, cost per task rose about 40 percent, from 0.40 USD (about 64 yen) to 0.58 USD (about 92 yen). The reasons are a 30 percent increase in average output tokens per task, to 48,000, and more turns on agentic evaluations. Time per task at high effort also grew from 2.2 minutes to 2.5 minutes, though output speed itself remains fast at roughly 300 tokens per second.

Even so, the firm places 3.8 Flash on the Pareto frontier for intelligence versus cost, calling it the cheapest model at its level of intelligence. Cost per task drops to 0.41 USD (about 65 yen) at medium effort and 0.24 USD (about 38 yen) at low effort, so there is room to dial effort down depending on the workload. Google likewise advises that where compute efficiency is the primary constraint, developers can use lower effort levels or continue to rely on Gemini 3.7 Flash, which remains fully supported.

Pricing Holds Through Year-End, Doubles in 2027

Pricing matches 3.7 Flash's introductory rate: 0.75 USD (about 120 yen) per million input tokens and 3.75 USD (about 600 yen) per million output tokens. That rate is limited to December 31, 2026. From January 1, 2027, it switches to 1.50 USD (about 240 yen) for input and 7.50 USD (about 1,200 yen) for output. The 90 percent discount on cached input tokens continues to apply.

※1 USD = 159 JPY (based on the September 2, 2026 close)

As noted above, 3.8 Flash tends to consume more tokens than 3.7 Flash on the same task. Even at identical unit prices, workloads that run agents for long stretches could see higher bills than with 3.7 Flash, and the unit prices themselves double at the start of next year. Before locking in a design based on introductory pricing, it is prudent to factor in both the 2027 rates and the increase in token consumption.

The context window is 1 million tokens, unchanged from 3.7 Flash. Input covers text, images, video, and audio, with text output. The knowledge cutoff is March 2026 for some domains, while in others it may be limited to January 2025.

Where It Is Available

3.8 Flash is available through every channel from day one. Developers can use it via the Gemini API in Google AI Studio and Android Studio, in the agent development environment Google Antigravity, and in the UI generation tool Stitch. Enterprises get it through Gemini Enterprise. Consumers who subscribe to Google AI Pro or Ultra can use 3.8 Flash in the Gemini app, AI Mode in Google Search, and Gemini in Google Sheets.

Alongside the release, Google published demos such as a 3D puzzle game assembled in Antigravity from a simple prompt, a DOS-style recreation of a maps service, and a topographic visualization built on real U.S. Geological Survey data. The showcases are aimed squarely at agent use cases where a single instruction keeps the model working for a long time.

On safety, the model ships with safeguards against misuse in the chemical, biological, radiological, and nuclear (CBRN) domains and in cyber offense. Google also says the Gemini 3.8 generation shows a significant leap in robustness against prompt injection, as measured by Gray Swan.

Summary

Gemini 3.8 Flash advances the production-oriented Flash line one more generation just three weeks after 3.7 Flash, with gains on long-horizon software engineering, agentic tasks, and specialized fields such as finance and law. The independently measured Intelligence Index rose 3 points to 59. Introductory pricing of 0.75 USD input and 3.75 USD output holds through the end of the year, but the "works harder" design means higher token consumption, and cost per task is up about 40 percent. With unit prices doubling in 2027, both factors deserve a place in any adoption decision.