Qwen: Qwen3.5-Flash
The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the...
Everything else
| Max output | 66K |
|---|---|
| Input / 1M | $0.065 |
| Output / 1M | $0.26 |
| Cache read / 1M | — |
| Cache write / 1M | — |
| 1M in + 1M out | $0.325 |
| Takes | text, image, video |
| Returns | text |
| Tools | yes |
| Reasoning | optional |
| Open weights | no |
| Released | 2026-02-25 |
| Knowledge cutoff | — |
| Retires | — |
| Catalog id | qwen/qwen3.5-flash-02-23 |
Benchmarks
| source | epoch |
|---|---|
| gpqa | 0.838 |
| gpqaStderr | 0.022 |
| frontierMath | 0.062 |
| frontierMathStderr | 0.0142 |
| aime | 0.856 |
| aimeStderr | 0.046 |
| simpleQA | 0.198 |
| simpleQAStderr | 0.0126 |
Benchmark scores by Epoch AI, CC BY 4.0.
Using it
In 00, this model is picked per agent — and per tier, so the model that answers your customers need not be the one that writes your code. Get 00.