Google has released Gemini 3.6 Flash, the workhorse model of the Gemini lineup, together with the lightweight 3.5 Flash-Lite. The 3.6 Flash model consumes 17 percent fewer output tokens than its predecessor while also costing less per token[1]. Google also introduced 3.5 Flash Cyber, a security-focused model, and gave an update on the delayed Gemini 3.5 Pro as well as the start of work on Gemini 4.

The efficiency tier now competes on quality too

The Flash series is positioned to hit the sweet spot between efficiency and quality within the Gemini family. The new 3.6 Flash stays in that lane but beats the previous 3.5 Flash on both performance and unit cost.

Token efficiency is the clearest number. According to the Artificial Analysis Index, 3.6 Flash uses 17 percent fewer output tokens than 3.5 Flash, and on benchmarks such as DeepSWE by Datacurve the reduction reaches as much as 65 percent[1]. In practice it finishes the same job with shorter output.

Pricing is set at 1.50 USD per 1M input tokens (about 240 yen) and 7.50 USD per 1M output tokens (about 1,200 yen)[1]. Output on 3.5 Flash ran at 9 USD per 1M tokens, so the unit price itself has come down[2]. Fewer tokens at a lower rate compounds, which matters once agents run continuously.

What the benchmarks show for 3.6 Flash

The notable part is that the efficiency gain did not cost quality. Google cites the following comparisons[1].

On DeepSWE the score is 49 percent versus 37 percent, with fewer unwanted code edits and fewer execution loops. On MLE Bench, which measures machine learning research tasks, the gap is wider at 63.9 percent versus 49.7 percent. OSWorld-Verified, which involves computer operation, comes in at 83.0 percent versus 78.4 percent, and the practice-oriented GDPval-AA v2 at 1421 versus 1349.

Computer use is now available as a built-in client side tool through the Gemini API and Gemini Enterprise[1]. Customers including Hebbia and Harvey report strong results on multimodal work such as document parsing, chart and data analysis, and report drafting[1].

On safety, Google says the model ships with strengthened Frontier Safety safeguards covering Chemical, Biological, Radiological and Nuclear (CBRN) domains and cyber offense misuse, making it substantially more resistant to jailbreaks. At the same time it was trained to reduce unnecessary refusals for beneficial uses[1].

3.5 Flash-Lite runs at 350 tokens per second for high volume work

Released at the same time, 3.5 Flash-Lite targets workloads that need low latency and high throughput, such as agentic search and large-scale document processing.

Artificial Analysis measures it at 350 output tokens per second, the fastest in the 3.5 series[1]. Pricing is 0.3 USD per 1M input tokens (about 50 yen) and 2.5 USD per 1M output tokens (about 400 yen)[1].

Against the previous 3.1 Flash-Lite the margins are large: Terminal-Bench 2.1 at 54 percent versus 31 percent, the long-context GDM-MRCR v2 at 72.2 percent versus 60.1 percent, and GDPval-AA v2 at 1140 versus 642[1].

More interesting is that on several agentic and coding evaluations this lightweight model beats the higher-tier 3 Flash. SWE-Bench Pro comes in at 54.2 percent versus 49.6 percent, and OSWorld-Verified at 74.0 percent versus 65.1 percent[1]. A new generation of the lower tier overtaking the previous generation of the higher tier is a familiar pattern in this field.

Thinking levels are configurable, so developers can push toward low-latency, low-cost execution for high-volume tasks or engage deeper reasoning for multi-step subagent workloads[1].

3.5 Flash Cyber is limited to governments and trusted partners

The third model, 3.5 Flash Cyber, is built on 3.5 Flash and fine-tuned specifically for finding and fixing vulnerabilities. Multiple instances work together inside CodeMender, Google's code security agent, to produce a single combined report. Google says it reaches competitive frontier performance on the CyberGym vulnerability benchmark[1].

Access is restricted, however. Given the dual-use nature of the technology, the model will be offered through CodeMender only to governments and trusted partners as a limited-access pilot[1]. The stated backdrop is that AI can already find vulnerabilities faster than current systems can fix them, which makes a narrow rollout a reasonable call.

3.5 Pro is still delayed while attention moves to Gemini 4

Google also addressed the top-end Gemini 3.5 Pro. It is currently testing with partners and will be made broadly available as soon as it is ready[1], meaning the schedule slip continues.

In parallel, Google says it has already begun pre-training for Gemini 4 and describes it as its most ambitious run yet[1]. Acknowledging a delay at the top of the lineup while shipping volume in the efficiency tier is a fair snapshot of where the industry sits. For developers running agents in production, cost per task has become more pressing than raw capability.

3.6 Flash and 3.5 Flash-Lite are available starting today through the Gemini API via Google AI Studio and Android Studio. 3.6 Flash is also in Google Antigravity, and enterprises can access it through the Gemini Enterprise Agent Platform. Both are available to everyone via the Gemini app, and 3.5 Flash-Lite is rolling out in Google Search[1].

※1 USD = 162 JPY (based on the July 21, 2026 close)

Summary

Gemini 3.6 Flash cuts output tokens by 17 percent, lowers the unit price, and beats 3.5 Flash on major evaluations including DeepSWE and MLE Bench. The simultaneously released 3.5 Flash-Lite handles volume at 350 tokens per second and overtakes the higher-tier 3 Flash on some benchmarks. Security-focused 3.5 Flash Cyber is a limited pilot for governments and trusted partners. The top-end 3.5 Pro remains in testing, and Google's near-term battleground is shifting from raw capability toward cost per task.

Source: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/

Source: https://9to5google.com/2026/07/21/gemini-3-6-flash-launch/