Home  /  Models  /  DeepSeek R1
⚡ Quick Answer (TL;DR)

DeepSeek R1 costs $0.550 per 1M input tokens and $2.19 per 1M output tokens. It features a 131k context window, a blended 3:1 production rate of $0.960/1M, and is optimized for math proofs, complex algorithmic coding, autonomous agent reasoning.

DeepSeek R1 Token Pricing

DeepSeek Reasoning / Deep Thinking

Open-weights reasoning model utilizing large-scale reinforcement learning, matching OpenAI o1 at 95% lower cost.

Simulate DeepSeek R1 Bill →
Input Price
$0.550
per 1 Million tokens
Output Price
$2.19
per 1 Million tokens
Prompt Caching
$0.140
Saves up to 75%
Context Window
131k
Max output: 65k tokens

Verified Specifications & Benchmark Data

Developer / Provider DeepSeek
Model Family DeepSeek
Blended 3:1 Rate (Production Benchmark) $0.960 / 1M tokens
Batch API Discount (24hr SLA) None
Multimodal Vision ❌ Text Only
Function Calling / Structured Outputs ✅ Native Tool Calling
MMLU Benchmark Score 90.8%
Latency & Throughput Tier Reasoning (~40 t/s)
Optimal Architecture & Use Cases Math proofs, complex algorithmic coding, autonomous agent reasoning

Frequently Asked Questions about DeepSeek R1

How much does DeepSeek R1 cost per 1M tokens?

DeepSeek R1 pricing is set at $0.550 per 1 million input tokens and $2.19 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a None.

How much money does DeepSeek R1 prompt caching save?

With prompt caching enabled, cached input tokens are discounted to $0.140/1M, saving 75% on repeated system prompts and document vectors.

What is the context window limit of DeepSeek R1?

DeepSeek R1 has a maximum context window of 131k (131,072 tokens), supporting up to 65,536 completion tokens per response.

Hosting & Cloud GPU Options for DeepSeek R1

Calculate whether serverless API or dedicated GPU hosting (RunPod / Together AI / Vultr) is more cost-effective for your volume.

Request Custom Audit →

Direct Matchups with DeepSeek R1

See how DeepSeek R1 compares against other leading frontier and open-weights models in cost and latency.