G

Google Gemini Flash

by Google (Google DeepMind)

AI Models & APIsAI Agents & OrchestrationDeveloper Tools

Google's latency-optimised workhorse model for high-volume agentic and multimodal work

Usage-based · Freemium·Added Mar 14, 2026·Updated Aug 2, 2026
Share:
THE DAILY BRIEF
Google Gemini Flash

by Google (Google DeepMind)

AI Models & APIsAI Agents & OrchestrationDeveloper Tools

Google's latency-optimised workhorse model for high-volume agentic and multimodal work

Usage-based · Freemium

Gemini Flash is the speed-and-cost tier of Google's Gemini family, sold per token through the Gemini Developer API, Google AI Studio and the Gemini Enterprise Agent Platform. It targets engineering teams running high-volume agentic, coding and multimodal workloads that cannot justify frontier-model output prices, and it removes the need to trade a 1M-token context window for cheap inference.

At a Glance

Category
AI Models & APIs
Pricing
Usage-based, Freemium
Target Market
CTOs, CIOs, VP Engineering, Enterprise Developers, ML Platform Engineers, Data Scientists
Deployment
Cloud-only, API-based
Founded
1998
Headquarters
Mountain View, California, United States
Team Size
500+

Key Features

  • 1M-token input window with 64k output
  • Native multimodal input
  • Computer use and tool calling
  • Token-efficiency tuning
  • Batch, caching and priority modes
  • Search grounding as a billable tool
  • Paid-tier data exclusion

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Production coding agents
  • Contract and filing review
  • Browser and desktop automation
  • High-volume support triage
  • Media and meeting intelligence

Ideal For

Best For

  • High-volume agentic workflows that chain many tool calls and need low per-step token cost
  • Long-document and long-transcript processing inside a single 1M-token input window
  • Multimodal extraction across text, images, video, audio and PDF in one model call
  • Coding and software-engineering agents where DeepSWE-style task completion matters more than peak reasoning
  • Batch pipelines that can tolerate asynchronous processing for a 50% discount

Not Ideal For

  • Teams that need maximum reasoning depth on hard research or planning problems — that is what Gemini 3.1 Pro and rival flagships are priced for, and Flash is explicitly the tier below
  • The cheapest possible high-throughput classification or extraction work, where Gemini 3.5 Flash-Lite at $0.30/$2.50 per million tokens is five times cheaper on input and three times cheaper on output
  • Air-gapped or on-premise deployments — the model is only reachable as a hosted API through Google's own surfaces
  • Free-tier prototyping with confidential data, because Google's terms allow human reviewers to read and annotate unpaid-tier prompts

Market Analysis

Enterprise-gradeCost-optimisedHigh-throughputMultimodal

Pros

  • Output pricing fell from $9.00 to $7.50 per million tokens between Gemini 3.5 Flash and 3.6 Flash, an unusual generational price cut
  • Vendor benchmarks show large jumps on agentic work: DeepSWE 49% versus 37%, MLE-Bench 63.9% versus 49.7%, OSWorld-Verified 83.0%
  • Batch at 50% off and cached input at 90% off make high-volume economics tractable without negotiating a contract
  • Paid-tier prompts and responses are contractually excluded from Google product improvement, and Vertex/Agent Platform coverage includes SOC 2, ISO 27001/27017/27018, ISO 42001 and a self-serve HIPAA BAA

Cons

  • Every headline benchmark at launch was Google's own; independent coverage explicitly flagged that the numbers lack independent replication
  • Independent analysis framed the release as incremental — Google competing on price and speed while its flagship Gemini 3.5 Pro slipped well past the June 2026 target set at I/O, leaving roadmap uncertainty for teams standardising on the family
  • The Flash tier is no longer cheap in absolute terms: at $1.50/$7.50 it costs five times more on input than Gemini 2.5 Flash did, so 'upgrade to the latest Flash' is a budget decision
  • Free-tier prompts may be read and annotated by human reviewers and are used to improve Google products, which rules the free tier out for most real evaluation data
  • Practitioner discussion on Hacker News reports that raising reasoning effort can reduce deep-research accuracy, so getting good results requires per-workload configuration rather than a single default
  • The security-tuned Gemini 3.5 Flash Cyber sibling is gated to governments and trusted partners, so the security-specific capability is not generally purchasable

Pricing

Free tier

$0

  • Rate-limited access to Gemini 3.6 Flash, 3.5 Flash and Flash-Lite
  • Prompts may be read and annotated by human reviewers
  • Data used to improve Google products

Gemini 3.6 Flash (pay-as-you-go)

From $1.50/1M input tokens

  • $7.50 per 1M output tokens
  • Context caching at $0.15 per 1M input tokens plus $1.00 per 1M tokens per hour storage
  • Paid-tier prompts excluded from product improvement

Gemini 3.6 Flash (batch)

From $0.75/1M input tokens

  • $3.75 per 1M output tokens
  • 50% discount for asynchronous processing
  • Same 1M-token context window

Gemini 3.5 Flash-Lite

From $0.30/1M input tokens

  • $2.50 per 1M output tokens
  • Positioned for highest-throughput execution
  • Batch rate $0.15/$1.25 per 1M tokens

Gemini Enterprise Agent Platform

Contact for pricing

  • Same per-token rates with Standard, Priority (+80%) and Flex/Batch (-50%) modes
  • Regional ML processing and data residency controls
  • Grounding with your own data at $2.50 per 1,000 prompts

Metered per token, never per seat, with no upfront commitment and no charge for 4xx or 5xx failures. Gemini 3.6 Flash lists at $1.50 input and $7.50 output per million tokens; batch halves both, cached input drops to $0.15, and a priority tier adds 80%. Note the Flash tier has grown materially more expensive across generations — Gemini 2.5 Flash was $0.30/$2.50 — so a version upgrade is a real cost increase, not just a capability one. Search grounding is billed separately at $14 per 1,000 queries after 5,000 free monthly requests, and enterprise buyers on the Gemini Enterprise Agent Platform pay the same token rates rather than a licence fee.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

THE DAILY BRIEF

Enterprise AI insights for technology and business leaders, twice weekly.

beri.net

Subscribe at beri.net/subscribe for twice-weekly AI insights delivered to your inbox.

LinkedIn: linkedin.com/in/rberi  |  X: x.com/rajeshberi

© 2026 Rajesh Beri. All rights reserved.

Gemini Flash is the speed-and-cost tier of Google's Gemini family, sold per token through the Gemini Developer API, Google AI Studio and the Gemini Enterprise Agent Platform. It targets engineering teams running high-volume agentic, coding and multimodal workloads that cannot justify frontier-model output prices, and it removes the need to trade a 1M-token context window for cheap inference.

Gemini Flash is the mid-tier, latency-optimised branch of Google's Gemini model family, sold through the Gemini Developer API at ai.google.dev, Google AI Studio, Google Antigravity and the Gemini Enterprise Agent Platform, which is the 2026 renamed form of Vertex AI. The current generally available flagship of the line is Gemini 3.6 Flash, shipped on 21 July 2026 alongside Gemini 3.5 Flash-Lite and a restricted, security-tuned Gemini 3.5 Flash Cyber that is limited to governments and trusted partners. Google positions 3.6 Flash as a workhorse optimised for token efficiency in coding, knowledge work and multimodal tasks: it accepts text, images, video, audio and PDF across a 1M-token input window, returns up to 64k output tokens, and supports function calling, Google Search as a tool and computer use. Google reports that it consumes roughly 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index while scoring 49% on DeepSWE against 37% for its predecessor, 63.9% on MLE-Bench against 49.7%, and 83.0% on OSWorld-Verified for agentic computer use — figures that are vendor-measured and, as independent coverage pointed out at launch, not independently replicated. Pricing is purely consumption-based with no seats: $1.50 per million input tokens and $7.50 per million output tokens, down from the $9.00 output price of Gemini 3.5 Flash, with batch processing at half rate, context caching at $0.15 per million input tokens, and a priority processing tier carrying an 80% premium. A free tier exists for prototyping, but Google's API terms permit human review of unpaid-tier prompts, whereas paid-tier prompts and responses are excluded from product improvement. The line sits deliberately below Gemini Pro on raw capability and above Flash-Lite on cost.

Ideal Buyer

Platform and applied-AI engineering leads running production agent or document workloads at millions of calls a month, who need frontier-adjacent tool use without frontier-tier output pricing.

Key Benefit

A 1M-token multimodal model with computer use and function calling at $1.50/$7.50 per million tokens, halving again under batch mode.

At a Glance

Category
AI Models & APIs
Pricing
Usage-based, Freemium
Target Market
CTOs, CIOs, VP Engineering, Enterprise Developers, ML Platform Engineers, Data Scientists
Deployment
Cloud-only, API-based
Founded
1998
Headquarters
Mountain View, California, United States
Team Size
500+

Key Features

  • 1M-token input window with 64k output

    Accepts entire codebases, contract sets or long video transcripts in a single request without an external chunking and retrieval layer.

  • Native multimodal input

    Handles text, images, video, audio and PDF in the same call, so document and media pipelines need no separate transcription or OCR stage.

  • Computer use and tool calling

    Supports function calling, Google Search as a tool and computer use, scoring 83.0% on OSWorld-Verified for agentic browser and desktop control.

  • Token-efficiency tuning

    Google reports roughly 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index, which lowers cost per completed task rather than just per token.

  • Batch, caching and priority modes

    Batch runs at half the standard rate, cached input costs $0.15 per million, and a priority tier adds 80% for expedited real-time processing.

  • Search grounding as a billable tool

    Grounding with Google Search gives 5,000 free requests a month across Gemini 3.x models, then charges $14 per 1,000 queries.

  • Paid-tier data exclusion

    Google's API terms state paid prompts, cached content and responses are not used to improve Google products, which is what most enterprise reviews gate on.

Capabilities

text generation
image generation
video generation
code generation
workflow automation
api access
audio generation
fine tuning
agent orchestration

Use Cases

  • Production coding agents

    Drive multi-step repository edits and test loops where DeepSWE-class task completion at 49% and low output cost decide whether the agent is affordable to run continuously.

  • Contract and filing review

    Load an entire agreement set into the 1M-token window and extract obligations, dates and counterparties without building a retrieval pipeline first.

  • Browser and desktop automation

    Use the computer-use capability to complete UI workflows in legacy systems that expose no API, with tool calls checked against a real interface.

  • High-volume support triage

    Classify, summarise and route millions of tickets per month at $1.50 per million input tokens, dropping to $0.75 under batch mode.

  • Media and meeting intelligence

    Process audio and video natively to produce summaries, action items and compliance flags without a separate speech-to-text vendor in the chain.

Ideal For

Best For

  • High-volume agentic workflows that chain many tool calls and need low per-step token cost
  • Long-document and long-transcript processing inside a single 1M-token input window
  • Multimodal extraction across text, images, video, audio and PDF in one model call
  • Coding and software-engineering agents where DeepSWE-style task completion matters more than peak reasoning
  • Batch pipelines that can tolerate asynchronous processing for a 50% discount

Not Ideal For

  • Teams that need maximum reasoning depth on hard research or planning problems — that is what Gemini 3.1 Pro and rival flagships are priced for, and Flash is explicitly the tier below
  • The cheapest possible high-throughput classification or extraction work, where Gemini 3.5 Flash-Lite at $0.30/$2.50 per million tokens is five times cheaper on input and three times cheaper on output
  • Air-gapped or on-premise deployments — the model is only reachable as a hosted API through Google's own surfaces
  • Free-tier prototyping with confidential data, because Google's terms allow human reviewers to read and annotate unpaid-tier prompts

Integrations

SDK Available
SDK:PythonJavaScriptGoJavaREST

Deployment

On-Premise

Market Analysis

Enterprise-gradeCost-optimisedHigh-throughputMultimodal

Pros

  • Output pricing fell from $9.00 to $7.50 per million tokens between Gemini 3.5 Flash and 3.6 Flash, an unusual generational price cut
  • Vendor benchmarks show large jumps on agentic work: DeepSWE 49% versus 37%, MLE-Bench 63.9% versus 49.7%, OSWorld-Verified 83.0%
  • Batch at 50% off and cached input at 90% off make high-volume economics tractable without negotiating a contract
  • Paid-tier prompts and responses are contractually excluded from Google product improvement, and Vertex/Agent Platform coverage includes SOC 2, ISO 27001/27017/27018, ISO 42001 and a self-serve HIPAA BAA

Cons

  • Every headline benchmark at launch was Google's own; independent coverage explicitly flagged that the numbers lack independent replication
  • Independent analysis framed the release as incremental — Google competing on price and speed while its flagship Gemini 3.5 Pro slipped well past the June 2026 target set at I/O, leaving roadmap uncertainty for teams standardising on the family
  • The Flash tier is no longer cheap in absolute terms: at $1.50/$7.50 it costs five times more on input than Gemini 2.5 Flash did, so 'upgrade to the latest Flash' is a budget decision
  • Free-tier prompts may be read and annotated by human reviewers and are used to improve Google products, which rules the free tier out for most real evaluation data
  • Practitioner discussion on Hacker News reports that raising reasoning effort can reduce deep-research accuracy, so getting good results requires per-workload configuration rather than a single default
  • The security-tuned Gemini 3.5 Flash Cyber sibling is gated to governments and trusted partners, so the security-specific capability is not generally purchasable

Pricing

Free tier

$0

  • Rate-limited access to Gemini 3.6 Flash, 3.5 Flash and Flash-Lite
  • Prompts may be read and annotated by human reviewers
  • Data used to improve Google products

Gemini 3.6 Flash (pay-as-you-go)

From $1.50/1M input tokens

  • $7.50 per 1M output tokens
  • Context caching at $0.15 per 1M input tokens plus $1.00 per 1M tokens per hour storage
  • Paid-tier prompts excluded from product improvement

Gemini 3.6 Flash (batch)

From $0.75/1M input tokens

  • $3.75 per 1M output tokens
  • 50% discount for asynchronous processing
  • Same 1M-token context window

Gemini 3.5 Flash-Lite

From $0.30/1M input tokens

  • $2.50 per 1M output tokens
  • Positioned for highest-throughput execution
  • Batch rate $0.15/$1.25 per 1M tokens

Gemini Enterprise Agent Platform

Contact for pricing

  • Same per-token rates with Standard, Priority (+80%) and Flex/Batch (-50%) modes
  • Regional ML processing and data residency controls
  • Grounding with your own data at $2.50 per 1,000 prompts

Metered per token, never per seat, with no upfront commitment and no charge for 4xx or 5xx failures. Gemini 3.6 Flash lists at $1.50 input and $7.50 output per million tokens; batch halves both, cached input drops to $0.15, and a priority tier adds 80%. Note the Flash tier has grown materially more expensive across generations — Gemini 2.5 Flash was $0.30/$2.50 — so a version upgrade is a real cost increase, not just a capability one. Search grounding is billed separately at $14 per 1,000 queries after 5,000 free monthly requests, and enterprise buyers on the Gemini Enterprise Agent Platform pay the same token rates rather than a licence fee.

Security & Compliance

soc2
gdpr
hipaa
iso27001
sso
data residency

Connect

Sources

This page was written from 8 sources, 6 on domains other than ai.google.dev.

  1. 1.ai.google.devpricingvendor
  2. 2.deepmind.googleflash
  3. 3.ai.google.devtermsvendor
  4. 4.cloud.google.compricing
  5. 5.9to5google.comgemini 3 6 flash launch
  6. 6.unite.aigoogle ships three gemini flash models as its flagship slips
  7. 7.hn.algolia.comhn.algolia.com
  8. 8.aiprovidertrust.comgemini vertex
Newsletter

Stay Ahead of the Curve

Weekly enterprise AI insights for technology leaders. No spam, no vendor pitches—unsubscribe anytime.

Subscribe