Google Gemini Flash
by Google (Google DeepMind)
Google's latency-optimised workhorse model for high-volume agentic and multimodal work
Gemini Flash is the speed-and-cost tier of Google's Gemini family, sold per token through the Gemini Developer API, Google AI Studio and the Gemini Enterprise Agent Platform. It targets engineering teams running high-volume agentic, coding and multimodal workloads that cannot justify frontier-model output prices, and it removes the need to trade a 1M-token context window for cheap inference.
Gemini Flash is the mid-tier, latency-optimised branch of Google's Gemini model family, sold through the Gemini Developer API at ai.google.dev, Google AI Studio, Google Antigravity and the Gemini Enterprise Agent Platform, which is the 2026 renamed form of Vertex AI. The current generally available flagship of the line is Gemini 3.6 Flash, shipped on 21 July 2026 alongside Gemini 3.5 Flash-Lite and a restricted, security-tuned Gemini 3.5 Flash Cyber that is limited to governments and trusted partners. Google positions 3.6 Flash as a workhorse optimised for token efficiency in coding, knowledge work and multimodal tasks: it accepts text, images, video, audio and PDF across a 1M-token input window, returns up to 64k output tokens, and supports function calling, Google Search as a tool and computer use. Google reports that it consumes roughly 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index while scoring 49% on DeepSWE against 37% for its predecessor, 63.9% on MLE-Bench against 49.7%, and 83.0% on OSWorld-Verified for agentic computer use — figures that are vendor-measured and, as independent coverage pointed out at launch, not independently replicated. Pricing is purely consumption-based with no seats: $1.50 per million input tokens and $7.50 per million output tokens, down from the $9.00 output price of Gemini 3.5 Flash, with batch processing at half rate, context caching at $0.15 per million input tokens, and a priority processing tier carrying an 80% premium. A free tier exists for prototyping, but Google's API terms permit human review of unpaid-tier prompts, whereas paid-tier prompts and responses are excluded from product improvement. The line sits deliberately below Gemini Pro on raw capability and above Flash-Lite on cost.
Platform and applied-AI engineering leads running production agent or document workloads at millions of calls a month, who need frontier-adjacent tool use without frontier-tier output pricing.
A 1M-token multimodal model with computer use and function calling at $1.50/$7.50 per million tokens, halving again under batch mode.
At a Glance
- Category
- AI Models & APIs
- Pricing
- Usage-based, Freemium
- Target Market
- CTOs, CIOs, VP Engineering, Enterprise Developers, ML Platform Engineers, Data Scientists
- Deployment
- Cloud-only, API-based
- Founded
- 1998
- Headquarters
- Mountain View, California, United States
- Team Size
- 500+
Key Features
- ✓1M-token input window with 64k output
Accepts entire codebases, contract sets or long video transcripts in a single request without an external chunking and retrieval layer.
- ✓Native multimodal input
Handles text, images, video, audio and PDF in the same call, so document and media pipelines need no separate transcription or OCR stage.
- ✓Computer use and tool calling
Supports function calling, Google Search as a tool and computer use, scoring 83.0% on OSWorld-Verified for agentic browser and desktop control.
- ✓Token-efficiency tuning
Google reports roughly 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, which lowers cost per completed task rather than just per token.
- ✓Batch, caching and priority modes
Batch runs at half the standard rate, cached input costs $0.15 per million, and a priority tier adds 80% for expedited real-time processing.
- ✓Search grounding as a billable tool
Grounding with Google Search gives 5,000 free requests a month across Gemini 3.x models, then charges $14 per 1,000 queries.
- ✓Paid-tier data exclusion
Google's API terms state paid prompts, cached content and responses are not used to improve Google products, which is what most enterprise reviews gate on.
Capabilities
Use Cases
- •Production coding agents
Drive multi-step repository edits and test loops where DeepSWE-class task completion at 49% and low output cost decide whether the agent is affordable to run continuously.
- •Contract and filing review
Load an entire agreement set into the 1M-token window and extract obligations, dates and counterparties without building a retrieval pipeline first.
- •Browser and desktop automation
Use the computer-use capability to complete UI workflows in legacy systems that expose no API, with tool calls checked against a real interface.
- •High-volume support triage
Classify, summarise and route millions of tickets per month at $1.50 per million input tokens, dropping to $0.75 under batch mode.
- •Media and meeting intelligence
Process audio and video natively to produce summaries, action items and compliance flags without a separate speech-to-text vendor in the chain.
Ideal For
Best For
- ✓High-volume agentic workflows that chain many tool calls and need low per-step token cost
- ✓Long-document and long-transcript processing inside a single 1M-token input window
- ✓Multimodal extraction across text, images, video, audio and PDF in one model call
- ✓Coding and software-engineering agents where DeepSWE-style task completion matters more than peak reasoning
- ✓Batch pipelines that can tolerate asynchronous processing for a 50% discount
Not Ideal For
- ✗Teams that need maximum reasoning depth on hard research or planning problems — that is what Gemini 3.1 Pro and rival flagships are priced for, and Flash is explicitly the tier below
- ✗The cheapest possible high-throughput classification or extraction work, where Gemini 3.5 Flash-Lite at $0.30/$2.50 per million tokens is five times cheaper on input and three times cheaper on output
- ✗Air-gapped or on-premise deployments — the model is only reachable as a hosted API through Google's own surfaces
- ✗Free-tier prototyping with confidential data, because Google's terms allow human reviewers to read and annotate unpaid-tier prompts
Integrations
Deployment
Market Analysis
Pros
- ✓Output pricing fell from $9.00 to $7.50 per million tokens between Gemini 3.5 Flash and 3.6 Flash, an unusual generational price cut
- ✓Vendor benchmarks show large jumps on agentic work: DeepSWE 49% versus 37%, MLE-Bench 63.9% versus 49.7%, OSWorld-Verified 83.0%
- ✓Batch at 50% off and cached input at 90% off make high-volume economics tractable without negotiating a contract
- ✓Paid-tier prompts and responses are contractually excluded from Google product improvement, and Vertex/Agent Platform coverage includes SOC 2, ISO 27001/27017/27018, ISO 42001 and a self-serve HIPAA BAA
Cons
- ✗Every headline benchmark at launch was Google's own; independent coverage explicitly flagged that the numbers lack independent replication
- ✗Independent analysis framed the release as incremental — Google competing on price and speed while its flagship Gemini 3.5 Pro slipped well past the June 2026 target set at I/O, leaving roadmap uncertainty for teams standardising on the family
- ✗The Flash tier is no longer cheap in absolute terms: at $1.50/$7.50 it costs five times more on input than Gemini 2.5 Flash did, so 'upgrade to the latest Flash' is a budget decision
- ✗Free-tier prompts may be read and annotated by human reviewers and are used to improve Google products, which rules the free tier out for most real evaluation data
- ✗Practitioner discussion on Hacker News reports that raising reasoning effort can reduce deep-research accuracy, so getting good results requires per-workload configuration rather than a single default
- ✗The security-tuned Gemini 3.5 Flash Cyber sibling is gated to governments and trusted partners, so the security-specific capability is not generally purchasable
Pricing
Free tier
$0
- ✓Rate-limited access to Gemini 3.6 Flash, 3.5 Flash and Flash-Lite
- ✓Prompts may be read and annotated by human reviewers
- ✓Data used to improve Google products
Gemini 3.6 Flash (pay-as-you-go)
From $1.50/1M input tokens
- ✓$7.50 per 1M output tokens
- ✓Context caching at $0.15 per 1M input tokens plus $1.00 per 1M tokens per hour storage
- ✓Paid-tier prompts excluded from product improvement
Gemini 3.6 Flash (batch)
From $0.75/1M input tokens
- ✓$3.75 per 1M output tokens
- ✓50% discount for asynchronous processing
- ✓Same 1M-token context window
Gemini 3.5 Flash-Lite
From $0.30/1M input tokens
- ✓$2.50 per 1M output tokens
- ✓Positioned for highest-throughput execution
- ✓Batch rate $0.15/$1.25 per 1M tokens
Gemini Enterprise Agent Platform
Contact for pricing
- ✓Same per-token rates with Standard, Priority (+80%) and Flex/Batch (-50%) modes
- ✓Regional ML processing and data residency controls
- ✓Grounding with your own data at $2.50 per 1,000 prompts
Metered per token, never per seat, with no upfront commitment and no charge for 4xx or 5xx failures. Gemini 3.6 Flash lists at $1.50 input and $7.50 output per million tokens; batch halves both, cached input drops to $0.15, and a priority tier adds 80%. Note the Flash tier has grown materially more expensive across generations — Gemini 2.5 Flash was $0.30/$2.50 — so a version upgrade is a real cost increase, not just a capability one. Search grounding is billed separately at $14 per 1,000 queries after 5,000 free monthly requests, and enterprise buyers on the Gemini Enterprise Agent Platform pay the same token rates rather than a licence fee.
Security & Compliance
Connect
Sources
This page was written from 8 sources, 6 on domains other than ai.google.dev.
Stay Ahead of the Curve
Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.
SubscribeRelated Products
Scaled Cognition
APT, a large action model trained to take actions instead of predicting text
Engram
A learned memory layer that makes AI actually know your organization — at up to 100x fewer tokens.
TwelveLabs
Video intelligence API that makes every hour of enterprise footage searchable, analyzable and agent-ready.
Mistral OCR 4
Structure-aware document AI that returns bounding boxes, typed blocks, and per-word confidence scores.