Home  /  Models  /  GPT-4o Mini
⚡ Quick Pricing Summary

GPT-4o Mini by OpenAI costs $0.150 per 1 Million input tokens and $0.600 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $0.263/1M tokens.

OpenAI Fast / Lightweight

GPT-4o Mini

Ultra-affordable small model replacing GPT-3.5 Turbo with significantly superior reasoning and vision.

Calculate Spend in App →
Input Token Price
$0.150
Per 1,000,000 tokens
Output Token Price
$0.600
Per 1,000,000 tokens
Prompt Caching Rate
$0.075
Save up to 80% on cached inputs
Max Context Window
128k
Max output: 16k tokens

Verified Specifications & Benchmark Data

Developer / Provider OpenAI
Model Family GPT-4
Blended 3:1 Rate (Production Benchmark) $0.263 / 1M tokens
Batch API Discount (24hr SLA) 50% off standard rate
Multimodal Vision ✅ Supported (Images, Diagrams, OCR)
Function Calling / Structured Outputs ✅ Native Tool Calling
MMLU Benchmark Score 82.0%
Latency & Throughput Tier Ultra-Fast (~85 t/s)
Optimal Architecture & Use Cases High-volume classification, customer support bots, lightweight summarization

Frequently Asked Questions about GPT-4o Mini

How much does GPT-4o Mini cost per 1M tokens?

GPT-4o Mini pricing is set at $0.150 per 1 million input tokens and $0.600 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a 50% off standard rate.

How much money does GPT-4o Mini prompt caching save?

With prompt caching enabled, cached input tokens are discounted to $0.075/1M, saving 50% on repeated system prompts and document vectors.

What is the context window limit of GPT-4o Mini?

GPT-4o Mini has a maximum context window of 128k (128,000 tokens), supporting up to 16,384 completion tokens per response.

Direct Matchups with GPT-4o Mini

See how GPT-4o Mini compares against other leading frontier and open-weights models in cost and latency.