Home  /  Models  /  Claude 3.5 Haiku
⚡ Quick Pricing Summary

Claude 3.5 Haiku by Anthropic costs $0.800 per 1 Million input tokens and $4.00 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $1.600/1M tokens.

Anthropic Fast / Lightweight

Claude 3.5 Haiku

High-speed intelligence matching previous-generation flagship performance at a fraction of latency.

Calculate Spend in App →
Input Token Price
$0.800
Per 1,000,000 tokens
Output Token Price
$4.00
Per 1,000,000 tokens
Prompt Caching Rate
$0.080
Save up to 80% on cached inputs
Max Context Window
128k
Max output: 4k tokens

Verified Specifications & Benchmark Data

Developer / Provider Anthropic
Model Family Claude 3.5
Blended 3:1 Rate (Production Benchmark) $1.600 / 1M tokens
Batch API Discount (24hr SLA) 50% off standard rate
Multimodal Vision ❌ Text Only
Function Calling / Structured Outputs ✅ Native Tool Calling
MMLU Benchmark Score 75.2%
Latency & Throughput Tier Ultra-Fast (~95 t/s)
Optimal Architecture & Use Cases Real-time agents, code completion, fast document scanning

Frequently Asked Questions about Claude 3.5 Haiku

How much does Claude 3.5 Haiku cost per 1M tokens?

Claude 3.5 Haiku pricing is set at $0.800 per 1 million input tokens and $4.00 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a 50% off standard rate.

How much money does Claude 3.5 Haiku prompt caching save?

With prompt caching enabled, cached input tokens are discounted to $0.080/1M, saving 90% on repeated system prompts and document vectors.

What is the context window limit of Claude 3.5 Haiku?

Claude 3.5 Haiku has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.

Direct Matchups with Claude 3.5 Haiku

See how Claude 3.5 Haiku compares against other leading frontier and open-weights models in cost and latency.