Gemini 1.5 Flash by Google costs $0.075 per 1 Million input tokens and $0.300 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $0.131/1M tokens.
Gemini 1.5 Flash
Lightweight model with 1M context window engineered for high speed and cost efficiency.
Verified Specifications & Benchmark Data
| Developer / Provider | |
| Model Family | Gemini 1.5 |
| Blended 3:1 Rate (Production Benchmark) | $0.131 / 1M tokens |
| Batch API Discount (24hr SLA) | 50% off standard rate |
| Multimodal Vision | ✅ Supported (Images, Diagrams, OCR) |
| Function Calling / Structured Outputs | ✅ Native Tool Calling |
| MMLU Benchmark Score | 78.9% |
| Latency & Throughput Tier | Ultra-Fast (~100 t/s) |
| Optimal Architecture & Use Cases | High-frequency API calls, audio processing, bulk document summarization |
Frequently Asked Questions about Gemini 1.5 Flash
How much does Gemini 1.5 Flash cost per 1M tokens?
Gemini 1.5 Flash pricing is set at $0.075 per 1 million input tokens and $0.300 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a 50% off standard rate.
How much money does Gemini 1.5 Flash prompt caching save?
With prompt caching enabled, cached input tokens are discounted to $0.019/1M, saving 75% on repeated system prompts and document vectors.
What is the context window limit of Gemini 1.5 Flash?
Gemini 1.5 Flash has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.
Direct Matchups with Gemini 1.5 Flash
See how Gemini 1.5 Flash compares against other leading frontier and open-weights models in cost and latency.
Compare Google vs OpenAI pricing, context limits, and cost per 1M tokens.
Compare Google vs Groq (Meta) pricing, context limits, and cost per 1M tokens.
Compare Google vs Google pricing, context limits, and cost per 1M tokens.
Compare Google vs OpenAI pricing, context limits, and cost per 1M tokens.