Google has released Gemini 3.8 Flash, its latest and most capable Flash-tier model, on September 2, 2026. This marks the third Flash update in just six weeks, following Gemini 3.7 Flash. The new model focuses on stronger reasoning, coding, and agentic workflows while keeping the same speed and introductory pricing as its predecessor.
If you work with large language models for software engineering, autonomous agents, or multi-step tasks, here’s a clear overview of what Gemini 3.8 Flash offers, how it performs, what it costs, and how it stacks up against similar models.
What Is Gemini 3.8 Flash?
Gemini 3.8 Flash is Google’s “most intelligent workhorse model” in the Flash line. It builds directly on Gemini 3.7 Flash and delivers measurable gains in:
Long-horizon software engineering
Agentic workflows and tool use
Multi-step reasoning in specialized domains
Google describes the core design choice as the model “working harder” — it performs more reasoning steps and calls tools more iteratively on complex tasks. This often brings performance close to higher-cost frontier models while staying in the efficient Flash tier.
Key technical details:
Context window: 1,048,576 tokens
Maximum output tokens: 65,536
Modalities: Text, image, audio, and video input; text output
Thinking levels: Low, medium (default), and high — allowing control over quality, latency, and token use
Model ID: gemini-3.8-flash
A specialized variant, Gemini 3.8 Flash Cyber, targets cybersecurity tasks such as vulnerability detection and automated patching. Access is limited to trusted defenders through Google’s new Fairwind Program.
Gemini 3.8 Flash Pricing
Google kept the introductory pricing identical to Gemini 3.7 Flash:
Period
Input (per 1M tokens)
Output (per 1M tokens)
Cached Input
Through December 31, 2026
$0.75
$3.75
$0.075
From January 1, 2027
$1.50
$7.50
$0.15
Important notes:
The model may consume more tokens on harder tasks (especially at higher thinking levels) to maximize performance.
Context caching and batch/flex inference options remain available at discounted rates.
Free-tier access exists with limits; production use is pay-as-you-go.
This pricing positions Gemini 3.8 Flash as significantly more affordable than most frontier models while delivering competitive results on coding and agent benchmarks.
Performance and Benchmarks
Gemini 3.8 Flash shows clear improvements over Gemini 3.7 Flash and approaches or exceeds larger models on several key evaluations:
Benchmark
Gemini 3.8 Flash
Gemini 3.7 Flash
Notes
Terminal-bench 2.1
90.8%
81.6%
Strong gain in terminal/agentic tasks
DeepSWE v1.1 (long-horizon coding)
~73.7%
Lower
Outperforms most larger frontier models at a fraction of the cost
SWE-Bench Pro
61.6%
60.4%
Solid software engineering progress
SWE-Atlas
51.9%
48.0%
Multi-file and complex coding
τ³-bench Banking
38.1%
30.9%
Multi-step reasoning
CharXiv (Multimodal)
86.2%
84.5%
Document understanding
Humanity’s Last Exam
45.4%
45.7%
Essentially flat
Google reports that on DeepSWE v1.1 the model often matches or exceeds models such as Claude Opus 5 and GPT-5.6 Sol while remaining far cheaper. Gains are strongest in coding, tool use, and agentic loops; general knowledge benchmarks show more modest changes.
How Gemini 3.8 Flash Compares to Similar Models
Here’s a practical comparison of the current landscape (as of early September 2026):
Model
Strengths
Approx. Input / Output (per 1M)
Best For
Gemini 3.8 Flash
Coding, agents, cost-efficiency, speed
$0.75 / $3.75 (intro)
Production agents, software engineering at scale
Gemini 3.7 Flash
Solid previous workhorse
Same as 3.8
Lower token use if maximum accuracy is not critical
Claude Opus 5
Highest reasoning in many domains
~$5 / $25
Complex research, high-stakes reasoning
Claude Sonnet 5
Strong balance of quality and cost
~$2 / $10
Everyday coding and analysis
GPT-5.6 Sol / Terra
Broad capability, tool use
Higher than Flash
General frontier tasks
Gemini 3.1 Pro
Deeper reasoning (Pro tier)
Higher
When maximum intelligence is needed over speed/cost
Key takeaways:
Gemini 3.8 Flash closes much of the gap to frontier models on coding and agentic benchmarks while staying in the efficient pricing tier.
It is especially attractive for teams running many agent loops or long-horizon coding tasks.
For pure maximum intelligence on the hardest problems, larger Pro/Opus-class models may still hold an edge, but at significantly higher cost.
Token usage can be higher on complex tasks, so real-world cost depends on your thinking level and workload.
Availability
Gemini 3.8 Flash is generally available now across:
Gemini API and Google AI Studio
Google Antigravity (default for managed agents)
Gemini Enterprise
Consumer Gemini app (Google AI Pro and Ultra subscribers)
AI Mode in Search and Gemini in Google Sheets
The Cyber variant remains restricted.
Who Should Consider Gemini 3.8 Flash?
This release is particularly relevant if you:
Build or run autonomous agents and tool-using workflows
Need strong coding and multi-file software engineering performance
Want frontier-level results without frontier-level pricing
Already use Gemini models and want a straightforward upgrade path
Because Google continues rapid iteration on the Flash line, teams can adopt the latest improvements quickly while controlling costs through thinking levels and caching.
Final Thoughts
Gemini 3.8 Flash continues Google’s strategy of shipping frequent, high-value updates to the efficient Flash tier. By improving reasoning depth and agentic reliability while holding the same introductory price, it strengthens the case for using Gemini models in production coding and agent systems.
For developers and companies already working with Gemini (or evaluating cost-effective alternatives to Claude and GPT frontier models), this is a meaningful step forward worth testing on your specific workloads.
Pricing and availability details are based on Google’s official announcements as of September 2–3, 2026. Always check the latest Gemini API pricing page for the most current rates.
Schneller aufbauen mit ReadyTools
Entdecke ReadyTools: die ultimative Produktivitätssuite für Creator. Wunderschöne Linksy-Seiten, smarte Lara-KI, Projektmanagement, sicherer Cloud-Speicher und alles andere, was du brauchst – vereint an einem Ort. Starte noch heute deine 7-tägige kostenlose Testphase.
The August 30 Workspace update brings full Lara Agent availability with rollback, real-time chat improvements, a new notification system, and dozens of...
ReadyTools Workspace Chat adds page-level and workspace-wide conversations with voice messages, file sharing, and Lara AI participation - all encrypted...