Inception: Mercury 2.5
Inception · inception/mercury-2.5
Runs in the cloud, on our API — not on your own machine.
tools reasoning optional
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
What it costs
Final AI-token prices, per million tokens. Bring your own provider key in the app to pay provider rates directly.
What it can do
Effort: high, medium, low and none — medium by default. More effort means more thinking tokens, billed at the output rate.
Dates
Using it
In the 00 app — pick it per agent, and per tier, then build your own agents on it. Get 00.
As a drop-in API — the same model at an OpenAI-compatible endpoint. Two settings:
- Create a workspace API key in the Overblast console — usage bills that workspace's AI tokens, at the prices on this page.
- Point your base URL at
https://brain.deployd.network/ai/v1. Chat completions with streaming, embeddings,/images,/audio/transcriptionsand/videos.
curl https://brain.deployd.network/ai/v1/chat/completions \
-H "Authorization: Bearer ob_live_YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "inception/mercury-2.5",
"messages": [{ "role": "user", "content": "Hello" }]
}'
← all models · compare it with another · calling image, video and speech models
Subscribe to new models