Llama 3.3 70B (Groq LPU) by Groq (Meta) costs $0.590 per 1 Million input tokens and $0.790 per 1 Million output tokens. It features a 128k context window and has a blended 3:1 production rate of $0.640/1M tokens.
Llama 3.3 70B (Groq LPU)
Meta's flagship 70B open model hosted on Groq LPUs delivering 280+ tokens per second.
Verified Specifications & Benchmark Data
| Developer / Provider | Groq (Meta) |
| Model Family | Llama 3.3 |
| Blended 3:1 Rate (Production Benchmark) | $0.640 / 1M tokens |
| Batch API Discount (24hr SLA) | None |
| Multimodal Vision | ❌ Text Only |
| Function Calling / Structured Outputs | ✅ Native Tool Calling |
| MMLU Benchmark Score | 86.0% |
| Latency & Throughput Tier | Instant (~280 t/s) |
| Optimal Architecture & Use Cases | Instantaneous interactive chatbots, real-time voice translation, low-latency tool agents |
Frequently Asked Questions about Llama 3.3 70B (Groq LPU)
How much does Llama 3.3 70B (Groq LPU) cost per 1M tokens?
Llama 3.3 70B (Groq LPU) pricing is set at $0.590 per 1 million input tokens and $0.790 per 1 million output tokens. For high-volume batch workloads, 24-hour batch queues provide a None.
How much money does Llama 3.3 70B (Groq LPU) prompt caching save?
Prompt caching is currently not natively offered for Llama 3.3 70B (Groq LPU) on direct serverless endpoints.
What is the context window limit of Llama 3.3 70B (Groq LPU)?
Llama 3.3 70B (Groq LPU) has a maximum context window of 128k (128,000 tokens), supporting up to 4,096 completion tokens per response.
Direct Matchups with Llama 3.3 70B (Groq LPU)
See how Llama 3.3 70B (Groq LPU) compares against other leading frontier and open-weights models in cost and latency.
Compare Groq (Meta) vs DeepSeek pricing, context limits, and cost per 1M tokens.
Compare Groq (Meta) vs Anthropic pricing, context limits, and cost per 1M tokens.
Compare Groq (Meta) vs OpenAI pricing, context limits, and cost per 1M tokens.
Compare Groq (Meta) vs OpenAI pricing, context limits, and cost per 1M tokens.