Home  /  Models  /  Gemini 1.5 Pro
⚡ Quick Pricing Summary

Gemini 1.5 Pro by Google costs $1.25 per 1 Million input tokens and $5.00 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $2.188/1M tokens.

Google Large Context (2M)

Gemini 1.5 Pro

Industry-record 2 Million token context window model with exceptional needle-in-a-haystack recall.

Calculate Spend in App →
Input Token Price
$1.25
Per 1,000,000 tokens
Output Token Price
$5.00
Per 1,000,000 tokens
Prompt Caching Rate
$0.312
Save up to 80% on cached inputs
Max Context Window
128k
Max output: 4k tokens

Verified Specifications & Benchmark Data

Developer / Provider Google
Model Family Gemini 1.5
Blended 3:1 Rate (Production Benchmark) $2.188 / 1M tokens
Batch API Discount (24hr SLA) 50% off standard rate
Multimodal Vision ✅ Supported (Images, Diagrams, OCR)
Function Calling / Structured Outputs ✅ Native Tool Calling
MMLU Benchmark Score 85.9%
Latency & Throughput Tier Standard (~35 t/s)
Optimal Architecture & Use Cases Massive codebase analysis, hours of video comprehension, full book translation

Frequently Asked Questions about Gemini 1.5 Pro

How much does Gemini 1.5 Pro cost per 1M tokens?

Gemini 1.5 Pro pricing is set at $1.25 per 1 million input tokens and $5.00 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a 50% off standard rate.

How much money does Gemini 1.5 Pro prompt caching save?

With prompt caching enabled, cached input tokens are discounted to $0.312/1M, saving 75% on repeated system prompts and document vectors.

What is the context window limit of Gemini 1.5 Pro?

Gemini 1.5 Pro has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.

Direct Matchups with Gemini 1.5 Pro

See how Gemini 1.5 Pro compares against other leading frontier and open-weights models in cost and latency.