Home  /  Models  /  Gemini 2.0 Flash
⚡ Quick Answer (TL;DR)

Gemini 2.0 Flash costs $0.100 per 1M input tokens and $0.400 per 1M output tokens. It features a 128k context window, a blended 3:1 production rate of $0.175/1M, and is optimized for real-time multimodal voice/video streaming, low-latency agent loops.

Gemini 2.0 Flash Token Pricing

Google Fast / Multimodal

Next-gen multimodal workhorse with sub-second latency and 1M token context window.

Simulate Gemini 2.0 Flash Bill →
Input Price
$0.100
per 1 Million tokens
Output Price
$0.400
per 1 Million tokens
Prompt Caching
$0.025
Saves up to 75%
Context Window
128k
Max output: 4k tokens

Verified Specifications & Benchmark Data

Developer / Provider Google
Model Family Gemini 2.0
Blended 3:1 Rate (Production Benchmark) $0.175 / 1M tokens
Batch API Discount (24hr SLA) 50% off standard rate
Multimodal Vision ✅ Supported (Images, Diagrams, OCR)
Function Calling / Structured Outputs ✅ Native Tool Calling
MMLU Benchmark Score 84.6%
Latency & Throughput Tier Ultra-Fast (~110 t/s)
Optimal Architecture & Use Cases Real-time multimodal voice/video streaming, low-latency agent loops

Frequently Asked Questions about Gemini 2.0 Flash

How much does Gemini 2.0 Flash cost per 1M tokens?

Gemini 2.0 Flash pricing is set at $0.100 per 1 million input tokens and $0.400 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a 50% off standard rate.

How much money does Gemini 2.0 Flash prompt caching save?

With prompt caching enabled, cached input tokens are discounted to $0.025/1M, saving 75% on repeated system prompts and document vectors.

What is the context window limit of Gemini 2.0 Flash?

Gemini 2.0 Flash has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.

Hosting & Cloud GPU Options for Gemini 2.0 Flash

Calculate whether serverless API or dedicated GPU hosting (RunPod / Together AI / Vultr) is more cost-effective for your volume.

Request Custom Audit →

Direct Matchups with Gemini 2.0 Flash

See how Gemini 2.0 Flash compares against other leading frontier and open-weights models in cost and latency.