DeepSeek: R1 Distill Llama 70B
DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across
Everything else
| Max output | 8K |
|---|---|
| Input / 1M | $0.8 |
| Output / 1M | $0.8 |
| Cache read / 1M | — |
| Cache write / 1M | — |
| 1M in + 1M out | $1.60 |
| Takes | text |
| Returns | text |
| Tools | no |
| Reasoning | optional |
| Open weights | yes |
| Released | 2025-01-23 |
| Knowledge cutoff | 2024-07-31 |
| Retires | — |
| Catalog id | deepseek/deepseek-r1-distill-llama-70b |
Benchmarks
| source | epoch |
|---|---|
| gpqa | 0.557 |
| gpqaStderr | 0.0297 |
| mathLevel5 | 0.899 |
| mathLevel5Stderr | 0.0068 |
| aime | 0.514 |
| aimeStderr | 0.0605 |
Benchmark scores by Epoch AI, CC BY 4.0.
Using it
In 00, this model is picked per agent — and per tier, so the model that answers your customers need not be the one that writes your code. Get 00.