/ Models

Language models

ByteDance: UI-TARS 7B

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement...

Maker
Bytedance
Context window
128K
Max output
2K
Input / 1M
$0.1
Output / 1M
$0.2
1M in + 1M out
$0.3

Everything else

Max output2K
Input / 1M$0.1
Output / 1M$0.2
Cache read / 1M$0.1
Cache write / 1M
1M in + 1M out$0.3
Takesimage, text
Returnstext
Toolsno
Reasoningno
Open weightsyes
Released2025-07-22
Knowledge cutoff2025-01-31
Retires
Catalog idbytedance/ui-tars-1.5-7b

Using it

In 00, this model is picked per agent — and per tier, so the model that answers your customers need not be the one that writes your code. Get 00.

← all models