Home  /  Models  /  Llama 3.1 405B (Together AI)
⚡ Quick Pricing Summary

Llama 3.1 405B (Together AI) by Together AI (Meta) costs $3.50 per 1 Million input tokens and $3.50 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $3.500/1M tokens.

Together AI (Meta) Giant Open Weights

Llama 3.1 405B (Together AI)

The world's largest open-weights frontier model with 405 Billion parameters.

Calculate Spend in App →
Input Token Price
$3.50
Per 1,000,000 tokens
Output Token Price
$3.50
Per 1,000,000 tokens
Prompt Caching Rate
Not Supported
Save up to 80% on cached inputs
Max Context Window
128k
Max output: 4k tokens

Verified Specifications & Benchmark Data

Developer / Provider Together AI (Meta)
Model Family Llama 3.1
Blended 3:1 Rate (Production Benchmark) $3.500 / 1M tokens
Batch API Discount (24hr SLA) None
Multimodal Vision ❌ Text Only
Function Calling / Structured Outputs ✅ Native Tool Calling
MMLU Benchmark Score 88.6%
Latency & Throughput Tier Standard (~30 t/s)
Optimal Architecture & Use Cases Synthetic data generation, model distillation, enterprise on-premise benchmarking

Frequently Asked Questions about Llama 3.1 405B (Together AI)

How much does Llama 3.1 405B (Together AI) cost per 1M tokens?

Llama 3.1 405B (Together AI) pricing is set at $3.50 per 1 million input tokens and $3.50 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a None.

How much money does Llama 3.1 405B (Together AI) prompt caching save?

Prompt caching is currently not natively offered for Llama 3.1 405B (Together AI) on direct serverless endpoints.

What is the context window limit of Llama 3.1 405B (Together AI)?

Llama 3.1 405B (Together AI) has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.

Direct Matchups with Llama 3.1 405B (Together AI)

See how Llama 3.1 405B (Together AI) compares against other leading frontier and open-weights models in cost and latency.