/ Models

GPT Realtime 2.1

OpenAI · Live conversation · openai/gpt-realtime-2.1

Live conversation streaming

The most capable speech-to-speech: tools, interruption, tone. Also the most expensive per minute.

Install on 00

Get the 00 app

Free · macOS

Your agents, your models, your machine — this model installs with one click once the app is on board.

Download free →

What it costs

Per minute of audio
$0.058
Quoted as
$32 / $64 per million audio tokens in / out
Runs
on the supplier's API

audio tokens: $32/1M in, $64/1M out (~10 tokens/s of audio).

The per-minute figure converts audio tokens at the supplier's own rate (10/10 a second in and out). Prices are ballpark, pay-as-you-go, mid-tier — they move constantly. Last checked 2026-07-29.

How it performs

Word errors
Latency
500 ms
Speed
Languages
60
Streaming
yes
Speaker labels
Timestamps
Voice cloning
Licence

A blank cell means the row does not state it, not that the answer is no. Word-error rates come from one independent harness, so they compare with each other only.

Using it

Runs on the supplier's API, on your own key. You bring the supplier's key and pay their rates directly.

These are not AI-token prices, and these engines do not answer at the hosted endpoints — 00 speaks to them from your machine, with your credentials. The speech models that do run on AI tokens are one click away on the media page.

Pick it per agent in the app. Get 00.

← every speech model