/ Models

Language models

DeepSeek: R1 Distill Llama 70B

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across

Maker
Deepseek
Context window
8K
Max output
8K
Input / 1M
$0.8
Output / 1M
$0.8
1M in + 1M out
$1.60

Everything else

Max output8K
Input / 1M$0.8
Output / 1M$0.8
Cache read / 1M
Cache write / 1M
1M in + 1M out$1.60
Takestext
Returnstext
Toolsno
Reasoningoptional
Open weightsyes
Released2025-01-23
Knowledge cutoff2024-07-31
Retires
Catalog iddeepseek/deepseek-r1-distill-llama-70b

Benchmarks

sourceepoch
gpqa0.557
gpqaStderr0.0297
mathLevel50.899
mathLevel5Stderr0.0068
aime0.514
aimeStderr0.0605

Benchmark scores by Epoch AI, CC BY 4.0.

Using it

In 00, this model is picked per agent — and per tier, so the model that answers your customers need not be the one that writes your code. Get 00.

← all models