Gemini 3.6 Flash & 3.5 Flash-Lite: Benchmarks, Pricing, Speed
Gemini 3.6 Flash scores 50 on the Intelligence Index at 304 tok/s. Flash-Lite hits 350 tok/s at $0.30/$2.50. Benchmarks, pricing, and comparisons vs GPT-5.6, Kimi K3.

Google shipped Gemini 3.6 Flash and 3.5 Flash-Lite today while everyone waits for 3.5 Pro. That model has now missed three deadlines. Instead of the flagship, Google released two Flash-tier updates that are cheaper and faster than what they replace.
Buried at the bottom of the blog post: Gemini 4 pre-training has started. That's the real news. The Flash releases are a bridge.
But the bridge is good. 3.6 Flash scores 50 on the Artificial Analysis Intelligence Index, matching GPT-5.6 Luna while running 2x faster and costing less per task. Flash-Lite at 350 tok/s and $0.30/$2.50 pricing is the cheapest model in its quality tier by a wide margin.
Here's the full breakdown.
3.6 Flash and 3.5 Flash-Lite Side by Side
| Gemini 3.6 Flash | Gemini 3.5 Flash-Lite | Gemini 3.5 Flash | GPT-5.6 Luna | |
|---|---|---|---|---|
| Intelligence Index | 50 | 36 | 50 | 51 |
| Speed | 304 tok/s | 350 tok/s | 156 tok/s | 150 tok/s |
| Input Price/1M | $1.50 | $0.30 | $1.50 | $1.00 |
| Output Price/1M | $7.50 | $2.50 | $9.00 | $6.00 |
| Cache Hit/1M | $0.15 | $0.03 | $0.15 | $0.10 |
| DeepSWE | 49% | — | 37% | 67% |
| SWE-Bench Pro | 58.7% | 54.2% | 55.1% | — |
| OSWorld-Verified | 83.0% | 74.0% | 78.4% | — |
| GDPval-AA v2 | 1421 | 1140 | 1349 | — |
| Context Window | 1M | 1M | 1M | 1.05M |
| Multimodal | Text, image, speech, video | Text, image, speech, video | Text, image, speech, video | Text, image |
3.6 Flash is the upgrade path from 3.5 Flash. Same Intelligence Index score (50), but 17% fewer output tokens per task and a $1.50 price drop on output (from $9 to $7.50). That efficiency gain compounds. Fewer tokens means faster responses and lower bills, even before the per-token discount.
Flash-Lite is a different animal. It's built for volume. At 350 tok/s, it's the fastest model in its intelligence class and costs a third of what Claude Haiku 4.5 charges. The 36 Intelligence Index score looks low until you realize it beats Gemini 3 Flash (46 Intelligence) on coding and agentic benchmarks while being cheaper and faster.
Where 3.6 Flash Improved Over 3.5 Flash
The headline number: time per task dropped from 2.7 minutes to 1.3 minutes, a 50%+ reduction measured by Artificial Analysis. That's driven by two things working together: 17% fewer output tokens per task and nearly double the output speed (304 tok/s vs 156 tok/s). On some benchmarks, the efficiency gain is much larger. Google says DeepSWE token usage dropped by 65%. That's not a typo.
Here's what the benchmark improvements look like:
| Benchmark | 3.6 Flash | 3.5 Flash | Change |
|---|---|---|---|
| DeepSWE | 49% | 37% | +12 points |
| MLE Bench | 63.9% | 49.7% | +14.2 points |
| OSWorld-Verified | 83.0% | 78.4% | +4.6 points |
| GDPval-AA v2 | 1421 | 1349 | +72 Elo |
| SWE-Bench Pro | 58.7% | 55.1% | +3.6 points |
| GDM-MRCR v2 (1M) | 54.0% | under 27% | 2x+ |
The coding gains are real. DeepSWE and MLE Bench both saw double-digit improvements. 3.6 Flash "delivers higher precision with fewer unwanted code edits and reduced execution loops," according to Google. In practice, that means less back-and-forth when using it as a coding agent.
Computer use is now a built-in client-side tool through the Gemini API. The 83% OSWorld-Verified score is the highest in Google's comparison table, ahead of GPT-5.6 Luna and Grok 4.5. If you're building browser automation or computer-use agents, 3.6 Flash is worth testing.
The long-context performance jump is the most dramatic. GDM-MRCR v2 at 1M tokens went from under 27% to 54%. That means 3.6 Flash can actually use its full 1M context window without quality collapsing. The previous generation couldn't say that.
Knowledge cutoff moved from January 2025 to March 2026. Fourteen months of fresher training data. That matters for anything involving recent APIs, libraries, or world events.
One caveat: early users are reporting that the knowledge cutoff doesn't always behave as expected. When asked about the best frontier model, 3.6 Flash confidently answered "Claude 3.5 Sonnet," a model from late 2024. A March 2026 cutoff should know better. This kind of confident-but-wrong factual recall was one of the issues that delayed Gemini 3.5 Pro, where internal checkpoints showed "frequent knowledge cutoff hallucinations" according to leaked test results. If your workflow depends on the model knowing recent facts, use Search Grounding or RAG rather than trusting the base weights.
3.5 Flash-Lite: The Speed and Cost Story
Flash-Lite isn't trying to compete with frontier models. It's built for the parts of your pipeline where you need fast, cheap, good-enough responses at scale.
Artificial Analysis measured time per task at 0.6 minutes, nearly half the 1.0 minutes of Gemini 3.1 Flash-Lite. Average output tokens per task dropped from 20K to 13K, meaning Flash-Lite is doing the same work with fewer tokens. The cost per task still went up ($0.04 to $0.09) because the per-token pricing increased, but the Intelligence gain is massive: +11 points on the Intelligence Index, with the biggest jumps in agentic evaluations (GDPval-AA went from 642 to 1140, TerminalBench from 31 to 53.6).
The comparison that matters here is against models in the same price bracket:
| Gemini 3.5 Flash-Lite | GPT-5.4 mini | Claude Haiku 4.5 | Gemini 3.1 Flash-Lite | |
|---|---|---|---|---|
| SWE-Bench Pro | 54.2% | ~54% | — | 38.3% |
| Terminal-Bench 2.1 | 54.0% | ~59% | — | 31.0% |
| OSWorld-Verified | 74.0% | — | — | 54.3% |
| GDPval-AA v2 | 1140 | — | — | 642 |
| Speed | 350 tok/s | — | — | — |
| Input Price/1M | $0.30 | $1.00 | $1.00 | — |
| Output Price/1M | $2.50 | $4.00 | $5.00 | — |
Flash-Lite is a third of the input cost of GPT-5.4 mini. Less than a third of Claude Haiku 4.5. And it lands within a few points on most benchmarks while leading on some.
What caught my attention: Flash-Lite beats Gemini 3 Flash on SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%). The "Lite" label is misleading. This model is genuinely better than last generation's main Flash at agentic and coding tasks.
Google designed configurable thinking levels for Flash-Lite. You can dial reasoning down to "minimal" for high-volume batch work, or push it to higher thinking for complex sub-agent tasks. That flexibility matters when you're routing different parts of a pipeline through the same model.
Gemini 3.5 Flash Cyber: Google's Security-Focused Model
This one's different from the other two. It's not a general-purpose model you can call through the API.
Gemini 3.5 Flash Cyber is fine-tuned from 3.5 Flash for one purpose: finding and fixing security vulnerabilities. It doesn't run standalone. It works inside CodeMender, Google's code security agent, where multiple Flash Cyber agents coordinate to scan code, validate findings, and produce a combined vulnerability report.
The architecture is multi-agent. Several Flash Cyber instances work in parallel on different aspects of a codebase, each handling discovery or verification. The combined report goes through validation before surfacing results. Google says this reaches "competitive frontier-level performance" on CyberGym at a cheaper cost than throwing larger models at the same problem.
Why this matters beyond security: Flash Cyber shows that fine-tuning a cheap, fast model for a specific domain and wrapping it in purpose-built agent infrastructure can match frontier performance. The model itself costs less per token than larger alternatives. The agent architecture (CodeMender) is doing the heavy lifting on coordination and accuracy.
The access model is deliberately restricted. Governments and trusted partners only, via a limited pilot. No public API. No timeline for broader access. Google is being cautious, and for good reason. A model optimized for finding vulnerabilities could also help exploit them. The restricted release is the right call.
For most developers, this doesn't change your workflow today. But it signals Google's playbook for specialized AI: take a cheap Flash model, fine-tune it for a domain, wrap it in an agent framework, and deploy it to targeted users. Expect similar specialized models for healthcare, legal, finance, and other verticals. If Kimi K3's unrestricted security capabilities make you nervous, Google's gated approach is the opposite end of that spectrum.
Gemini 3.5 Pro: When Is It Coming Out?
Google announced Gemini 3.5 Pro at I/O 2026 in May. It was supposed to ship in June. Then an intermediate deadline passed. Then July 17. It's missed all of them.
Today's blog post says "Gemini 3.5 Pro is currently testing with partners and we plan to make it broadly available as soon as it's ready." That's corporate for "it's not ready."
The backstory is worse than a simple delay. According to reporting from Bloomberg and The Verge, Google DeepMind scrapped the original Gemini 3.5 Pro entirely after internal testers found structural failures in recursive tool-calling and SVG generation. They did a full pretraining restart. The rebuilt model (internally called "Rev25") reportedly still has weak coding performance and frequent knowledge cutoff hallucinations. Some older "Rev24" checkpoints perform better on coding than the newer ones, which suggests the regression isn't fully resolved.
The hallucination issue deserves emphasis. Internal testing found that 3.5 Pro confidently generates answers about events it should acknowledge uncertainty about. For anyone building retrieval-augmented pipelines or agent workflows that depend on the model knowing what it doesn't know, that failure mode is worse than simply getting lower benchmark scores.
I'll be honest: this is becoming a pattern. Google ships Flash variants on time and delays Pro. It happened with 2.5 Pro too. Multiple Google employees told Bloomberg they're frustrated, concerned the company is losing ground to Anthropic and OpenAI. Google's response: "We're shipping quickly across models while keeping them effective for customers."
The Flash updates are solid and production-ready. But if you were planning API strategy around Gemini 3.5 Pro benchmarks, don't. No confirmed benchmark numbers exist for the rebuilt model. Build against what you can test today.
And here's the bigger picture. While everyone waits for 3.5 Pro, Google dropped this at the bottom of today's blog post: they've "started our most ambitious pre-training run yet" for Gemini 4. No timeline. No specs. But it tells you where the real investment is going. 3.5 Pro will eventually ship, but the team is already looking past it. The question developers are starting to ask: does anyone still need 3.5 Pro if 3.6 Flash covers most production use cases and Gemini 4 is the next real leap?
Compare Gemini 3.6 Flash against 200+ models
See how it stacks up on quality, speed, and price. Filter by task type and budget.
How 3.6 Flash Compares to the Competition
3.6 Flash sits in an interesting spot. It's not trying to be the smartest model. It's trying to be the best value for production agent workloads.
| Model | Intelligence | Speed | Input/Output $/1M | Best At |
|---|---|---|---|---|
| Claude Fable 5 | 60 | 71 tok/s | $10/$50 | Peak intelligence |
| GPT-5.6 Sol | 59 | 85 tok/s | $5/$30 | Coding agents |
| Kimi K3 | 57 | 62 tok/s | $3/$15 | Open-weight frontier |
| Claude Sonnet 5 | 53 | 78 tok/s | $2/$10 | Quality/price balance |
| GPT-5.6 Luna | 51 | 150 tok/s | $1/$6 | Budget agents |
| Gemini 3.6 Flash | 50 | 304 tok/s | $1.50/$7.50 | Fast production agents |
| Grok 4.5 | 54 | ~70 tok/s | $2/$6 | Token efficiency |
| Gemini 3.5 Flash | 50 | 156 tok/s | $1.50/$9 | (replaced by 3.6 Flash) |
3.6 Flash vs GPT-5.6 Luna: Near-identical Intelligence (50 vs 51). 3.6 Flash is 2x faster (304 vs 150 tok/s) but costs 25% more on output ($7.50 vs $6.00). Luna has a wider ecosystem through OpenAI's tools and OpenRouter availability. If speed matters more than ecosystem, 3.6 Flash wins. If you need Luna's max/ultra reasoning modes or Codex integration, stay on Luna.
3.6 Flash vs Claude Sonnet 5: Sonnet 5 scores higher (53 vs 50) and costs less on output during its intro pricing ($10/M vs $7.50/M). After August 31, Sonnet 5 moves to $15/M output and 3.6 Flash becomes the clear value pick. Speed isn't close. 3.6 Flash runs 4x faster.
3.6 Flash vs Grok 4.5: Grok scores higher (54 vs 50) and uses dramatically fewer tokens per task. At $2/$6, Grok's blended cost is competitive. But Grok runs at ~70 tok/s vs 3.6 Flash's 304 tok/s. For latency-sensitive agent loops, 3.6 Flash is the better choice.
3.6 Flash vs 3.5 Flash: Just upgrade. Same Intelligence score, cheaper output ($7.50 vs $9.00), 2x faster (304 vs 156 tok/s), better on every benchmark. There's no reason to stay on 3.5 Flash.
Real-World Cost Comparison
Using our cost calculator assumptions: 500 tasks/month at 20,000 average tokens per task.
| Model | Estimated Monthly Cost |
|---|---|
| Claude Fable 5 | ~$300 |
| GPT-5.6 Sol | ~$175 |
| Kimi K3 | ~$94 |
| Claude Sonnet 5 (intro) | ~$60 |
| GPT-5.6 Luna | ~$35 |
| Gemini 3.6 Flash | ~$45 |
| Grok 4.5 | ~$40 |
| Gemini 3.5 Flash-Lite | ~$15 |
Flash-Lite at ~$15/month is remarkable for a model that beats Gemini 3 Flash on coding benchmarks. For high-volume pipelines where you need hundreds of thousands of API calls, the savings add up fast. A workload that costs $175/month on Sol could run on Flash-Lite for under $20, assuming the 36 Intelligence Index score is sufficient for the task.
Artificial Analysis measured the actual cost per task (not just per-token pricing). 3.6 Flash comes in at $0.50 per task, down 18% from 3.5 Flash's $0.59. That's competitive with Luna despite the higher output price, because 3.6 Flash uses 17% fewer tokens to complete the same work. Flash-Lite's cost per task is $0.09. Even though that doubled from 3.1 Flash-Lite ($0.04), it's still a fraction of any frontier model.
Best Setup for Agent Workloads
Google clearly built these models for multi-agent architectures. Here's what the pricing and performance data suggests:
Two-model setup within Gemini:
- Master agent: Gemini 3.6 Flash ($1.50/$7.50). Handles orchestration, complex reasoning, computer use.
- Sub-agents: Gemini 3.5 Flash-Lite ($0.30/$2.50). Handles data extraction, translation, summarization, bulk processing.
Estimated monthly cost for moderate usage: $30-60. That's less than half what a GPT-5.6 Sol/Luna pairing costs.
Mixed-provider setup:
- Hard problems: GPT-5.6 Sol or Claude Fable 5 for tasks that need 55+ Intelligence scores.
- Production agents: Gemini 3.6 Flash for speed-sensitive agent loops.
- Bulk sub-tasks: Gemini 3.5 Flash-Lite for anything high-volume.
This splits the workload by what each model does best. Sol handles the 10% of tasks that need frontier intelligence. 3.6 Flash handles the 60% that need speed and good-enough quality. Flash-Lite handles the 30% that just need to be fast and cheap.
Use our model selector to find the right fit for your workload.
What Developers Are Saying
The community response is split. Both sides raise points worth considering.
On the skeptical side: 3.6 Flash ranks below Sonnet 5, Grok 4.5, GPT-5.6 Luna, and GPT-5.6 Terra on the Intelligence Index. Some developers testing it on coding tasks, particularly web development and 3D work, found Grok 4.5 performing better despite being in a similar price range ($2/$6 vs $1.50/$7.50). One common complaint: for the same budget, you could run GLM-5.2 or Grok 4.5 and get higher intelligence scores.
The knowledge cutoff problem drew attention too. Despite claiming a March 2026 cutoff, users found 3.6 Flash recommending "Claude 3.5 Sonnet" as the best frontier model. That's a model from late 2024. A March 2026 cutoff should know about Fable 5, Sol, and Grok 4.5. The base training data seems to know recent facts unevenly.
A few nuances the critics miss. As one commenter pointed out: "This is Flash tier. The number of parameters isn't comparable to Grok 4.5 or Sonnet 5." That's a valid correction. Flash models are designed for speed and cost, not peak intelligence. Comparing a Flash model against full-sized frontier models on raw quality is like comparing a Honda Civic to a BMW on acceleration. They're built for different things.
Others who tested it came away positive on multimodal tasks. Document parsing, chart analysis, and data extraction are where 3.6 Flash seems to shine. Philipp Schmid from Hugging Face noted it's a good model "especially on anything multimodal related" and that benchmarks only give a first hint. One user running a 200K context + 200K history test reported a 5-10% subjective quality improvement over 3.5 Flash on generated output. Not dramatic, but noticeable.
The pricing angle has supporters too. GPT-5.6 Luna is getting called "an underrated model" for its combination of cheap, fast, and good. But 3.6 Flash runs 2x faster than Luna (304 vs 150 tok/s) and is competitive on price. For latency-sensitive agent workflows, that speed advantage matters more than a 1-point Intelligence gap.
There's a strategic read that captures the situation well: Google may be losing the race for raw intelligence, but they're winning on cost efficiency. Flash-Lite at $0.30/$2.50 undercuts every model in its class. For production workloads running thousands of agent calls per day, economics and speed often matter more than peak benchmark scores.
The honest take: if you need the smartest model, 3.6 Flash isn't it. If you need a fast, affordable model for production agent pipelines, it deserves a serious look. Benchmarks give you one perspective. Test it on your own workload before deciding.
Benchmark Gaps Worth Knowing
A few things the blog post doesn't cover that affect your evaluation.
No SWE-Bench Verified numbers. Google reports SWE-Bench Pro but not Verified, which has historically been Anthropic's strongest benchmark. Every lab leads with the benchmarks where they shine. That's expected. But it means you should test 3.6 Flash on your own coding tasks rather than assuming the published numbers tell the complete story.
Artificial Analysis has published full Intelligence Index results and time-per-task benchmarks, confirming the efficiency gains Google claims. Their independent testing measured 304 tok/s for 3.6 Flash and 350 tok/s for Flash-Lite, consistent with Google's numbers.
Officechai.com's analysis made a useful observation: none of the mid-tier models (3.6 Flash, Luna, Sonnet 5, Grok 4.5) sweeps the board across all benchmarks. The spread between them is narrow enough that price and speed will decide most purchasing decisions, not raw scores. That's healthy for the market. It also means the "best model" depends entirely on your workload.
Estimate your monthly AI spend
Plug in your usage patterns and see costs across every model. Free, no signup.
Should You Switch to 3.6 Flash?
Switch from 3.5 Flash immediately. Better on every metric. Cheaper. Faster. No downsides. This is a free upgrade.
Switch from GPT-5.6 Luna if speed is your priority and you don't rely on OpenAI-specific tooling (Codex, max/ultra reasoning). 3.6 Flash is 2x faster at comparable intelligence.
Stay on Luna/Sol/Fable 5 if you need the OpenAI or Anthropic ecosystem, higher reasoning effort modes, or frontier intelligence above 50 on the Index.
Use Flash-Lite if you're running high-volume batch processing, sub-agent tasks, or cost-sensitive pipelines where 36 Intelligence is sufficient. At $0.30/$2.50, it's hard to beat the economics.
Wait if you're holding out for Gemini 3.5 Pro. It'll ship eventually. But don't block your development on it.
FAQ
Is Gemini 3.6 Flash better than GPT-5.6 Luna?
They're nearly identical on the Intelligence Index (50 vs 51). 3.6 Flash is 2x faster (304 vs 150 tok/s) and has stronger computer use capabilities (83% vs Luna's lower OSWorld score). Luna costs less on output ($6 vs $7.50) and has access to OpenAI's max/ultra reasoning modes. Pick based on whether you need speed or ecosystem.
How fast is Gemini 3.5 Flash-Lite?
350 tokens per second, making it one of the fastest models available. For context, GPT-5.6 Luna runs at 150 tok/s and Claude Sonnet 5 at 78 tok/s. Flash-Lite is built for workloads where latency matters more than peak intelligence.
What happened to Gemini 3.5 Pro?
It missed its June deadline, an intermediate date, and a July 17 target. Google says it's "currently testing with partners" with no new public timeline. The 3.6 Flash and Flash-Lite releases are filling the gap while Pro development continues. Gemini 4 pre-training has also started.
Can I use Gemini 3.5 Flash Cyber?
Not yet through the public API. Flash Cyber is available only to governments and trusted partners through a limited-access pilot via CodeMender, Google's code security agent. There's no timeline for broader availability.
How does Gemini 3.6 Flash compare to Kimi K3?
K3 scores higher on the Intelligence Index (57 vs 50) and has open weights coming July 27. 3.6 Flash is nearly 5x faster (304 vs 62 tok/s) and costs half as much per output token ($7.50 vs $15). For speed-sensitive production workloads, 3.6 Flash wins. For peak coding quality and self-hosting, K3 is the better pick.
Should I wait for Gemini 3.5 Pro or use 3.6 Flash now?
Use 3.6 Flash now. The gap between Flash and Pro has been narrowing with each generation. 3.6 Flash at Intelligence 50 is already competitive with mid-tier frontier models. If 3.5 Pro ships at 55-58 (speculative), the improvement may not justify waiting months. Build against what you can test today.
Is Gemini 3.5 Flash-Lite good enough for coding?
Yes, for many tasks. It scores 54.2% on SWE-Bench Pro, which beats Gemini 3 Flash (49.6%). It won't replace Sol or Fable 5 for complex debugging. But for code explanation, simple refactoring, test generation, and agent sub-tasks, Flash-Lite handles the work at a fraction of the cost.
When is Gemini 3.5 Pro coming out?
No confirmed date. Google announced it at I/O 2026 in May with a June target. It missed June, then an intermediate deadline, then July 17. The blog post from July 21 says it's "currently testing with partners." Internal reports suggest the model was rebuilt from scratch after failures in recursive tool-calling and still has hallucination issues. Don't plan your API strategy around it.
Why is Gemini 3.5 Pro delayed?
Google DeepMind scrapped the original model after finding structural failures in recursive tool-calling and SVG generation tasks. A full pretraining restart followed. The rebuilt checkpoints reportedly still show weak coding performance and knowledge cutoff hallucinations. Bloomberg reported that Google employees are frustrated, concerned about losing ground to Anthropic and OpenAI. Rather than ship a model with known regression issues, Google chose to delay and ship Flash updates instead.
Is Gemini 3.6 Flash better than Gemini 3.5 Flash?
Yes, across the board. Same Intelligence Index score (50), but 3.6 Flash runs 2x faster (304 vs 156 tok/s), costs less on output ($7.50 vs $9.00/1M), uses 17% fewer tokens per task, and scores higher on every benchmark Google published. The knowledge cutoff also moved from January 2025 to March 2026. There's no reason to stay on 3.5 Flash.
What is Gemini 3.5 Flash Cyber used for?
Finding and fixing security vulnerabilities in code. It's a specialized model fine-tuned from 3.5 Flash that runs inside CodeMender, Google's code security agent. Multiple Flash Cyber agents work in parallel to scan codebases, validate findings, and produce vulnerability reports. It's only available to governments and trusted partners through a limited pilot. No public API access.
What is Gemini 4?
Google confirmed they've started "the most ambitious pre-training run yet" for Gemini 4. No specs, no timeline, no benchmarks. It's a signal that Google's next major model generation is underway, separate from the 3.5 Pro delays. Expect more details later in 2026.
Benchmark data from Artificial Analysis Intelligence Index. Model details from Google's official blog post. 3.5 Pro delay reporting from Bloomberg, The Verge, and ChatForest. Additional analysis from officechai.com, 9to5Google, and CNBC. Updated July 21, 2026. See the LLM Leaderboard for live rankings, Best LLM in 2026 for the full comparison, or the Benchmark Dashboard for side-by-side model analysis.

