Gemini 1.5 Flash is 98% cheaper than Claude 3.5 Sonnet (v2) on blended 3:1 token pricing. Gemini 1.5 Flash costs $0.07/1M in and $0.30/1M out, while Claude 3.5 Sonnet (v2) costs $3.00/1M in and $15.00/1M out.
Gemini 1.5 Flash vs Claude 3.5 Sonnet (v2)
Compare verified input/output token rates, prompt caching discounts, latency tiers, and context limits between Google and Anthropic.
| Metric / Feature | Gemini 1.5 Flash (Google) | Claude 3.5 Sonnet (v2) (Anthropic) |
|---|---|---|
| Input Price ($ / 1M Tokens) | $0.07 | $3.00 |
| Output Price ($ / 1M Tokens) | $0.30 | $15.00 |
| Prompt Caching Input Rate | $0.019/1M | $0.300/1M |
| Blended 3:1 Rate (Production) | $0.131 / 1M | $6.000 / 1M |
| Max Context Window | 128k | 200k |
| Multimodal Vision | ✅ Supported | ✅ Supported |
| MMLU Benchmark | 78.9% | 88.3% |
| Throughput & Latency | Ultra-Fast (~100 t/s) | Fast (~50 t/s) |
| Best For | High-frequency API calls, audio processing, bulk document summarization | Autonomous coding, complex tool use, multi-step agent reasoning |
When to Choose Gemini 1.5 Flash
Choose Gemini 1.5 Flash if your workload requires high-frequency api calls, audio processing, bulk document summarization. Ideal for teams needing Google's infrastructure and ecosystem tooling.
When to Choose Claude 3.5 Sonnet (v2)
Choose Claude 3.5 Sonnet (v2) if your primary objective is autonomous coding, complex tool use, multi-step agent reasoning. Ideal for scaling high-throughput pipelines with Anthropic.
Want to compare Gemini 1.5 Flash against another model?
Select another target model to view immediate pricing & latency trade-offs.