Home  /  Models  /  GPT-4 Turbo
⚡ Quick Pricing Summary

GPT-4 Turbo by OpenAI costs $10.00 per 1 Million input tokens and $30.00 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $15.000/1M tokens.

OpenAI Flagship (Legacy)

GPT-4 Turbo

Previous-gen flagship model with 128k context window and JSON mode.

Calculate Spend in App →
Input Token Price
$10.00
Per 1,000,000 tokens
Output Token Price
$30.00
Per 1,000,000 tokens
Prompt Caching Rate
Not Supported
Save up to 80% on cached inputs
Max Context Window
128k
Max output: 4k tokens

Verified Specifications & Benchmark Data

Developer / Provider OpenAI
Model Family GPT-4
Blended 3:1 Rate (Production Benchmark) $15.000 / 1M tokens
Batch API Discount (24hr SLA) 50% off standard rate
Multimodal Vision ✅ Supported (Images, Diagrams, OCR)
Function Calling / Structured Outputs ✅ Native Tool Calling
MMLU Benchmark Score 86.4%
Latency & Throughput Tier Standard (~25 t/s)
Optimal Architecture & Use Cases Legacy enterprise pipelines, large document analysis

Frequently Asked Questions about GPT-4 Turbo

How much does GPT-4 Turbo cost per 1M tokens?

GPT-4 Turbo pricing is set at $10.00 per 1 million input tokens and $30.00 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a 50% off standard rate.

How much money does GPT-4 Turbo prompt caching save?

Prompt caching is currently not natively offered for GPT-4 Turbo on direct serverless endpoints.

What is the context window limit of GPT-4 Turbo?

GPT-4 Turbo has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.

Direct Matchups with GPT-4 Turbo

See how GPT-4 Turbo compares against other leading frontier and open-weights models in cost and latency.