/ Models

Language models

StepFun: Step 3.7 Flash

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...

Maker
Stepfun
Context window
262K
Max output
256K
Input / 1M
$0.2
Output / 1M
$1.15
1M in + 1M out
$1.35

Everything else

Max output256K
Input / 1M$0.2
Output / 1M$1.15
Cache read / 1M$0.04
Cache write / 1M
1M in + 1M out$1.35
Takestext, image, video
Returnstext
Toolsyes
Reasoningalways on · high, medium, low
Open weightsyes
Released2026-05-28
Knowledge cutoff
Retires
Catalog idstepfun/step-3.7-flash

Using it

In 00, this model is picked per agent — and per tier, so the model that answers your customers need not be the one that writes your code. Get 00.

← all models