/ Models

Media models

Everything 00 can generate with — pictures, video, music, speech, 3D — from every place it can run them. Filter by what you need it to make and where you want it to run.

190 models

Seedance 2.5 Reference to Video

Dreamina Seedance 2.5 generates video from up to 50 multimodal references images, video, audio, and style inputs, locking a character, set, and palette across a full 30-second take for production-grade consistency.

Does
image, text → video
Release date
2026-08
Price
$0.4730 / second

via fal commercial

Seedance 2.5 Image to Video

Dreamina Seedance 2.5 animates a single still into a native 30-second clip at up to 720p, extending one frame into continuous, coherent motion without the drift or stitching of shorter multi-clip workflows.

Does
image, text → video
Release date
2026-08
Price
$0.4730 / second

via fal commercial

Seedance 2.5 Text to Video

Dreamina Seedance 2.5 generates native 30-second single-shot video at up to 720p from a single text prompt, reasoning about the whole shot at once so motion, lighting, and subject identity stay coherent from first frame to last.

Does
text → video
Release date
2026-08
Price
$0.4730 / second

via fal commercial

Hi3D Image to Relief

Generate a 3D relief depth map with Hi3D from a single image.

Does
image, text → image-edit
Release date
2026-08
Price
$0.2 / generation

via fal commercial

Seedream 5.0 Pro Layerize

Splits a finished image into independent, editable transparent-PNG layers — background plus separate elements, from a text description, returning 2 to 17 layers per call for non-destructive reuse in design tools.

Does
image, text → image-edit
Release date
2026-08
Price
$0.03375 / generated

via fal commercial

Hi3D Multiview to 3D

Generate 3D models from multiple view images using Hi3D.

Does
image → 3d
Release date
2026-08
Price
$0.02 / credit

via fal commercial

Hi3D Image to 3D

Generate 3D models from a single image with Hi3D.

Does
image → 3d
Release date
2026-08
Price
$0.02 / credit

via fal commercial

Flux 3 Image to Video

FLUX.3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.

Does
image, text → video
Release date
2026-08
Price
see supplier

via fal commercial

Flux 3 First Last Frame to Video

FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates the video between a defined start and end frame, interpolating a smooth, coherent transition from the first image to the last.

Does
image, text → video
Release date
2026-08
Price
see supplier

via fal commercial

Flux 3 Image to Video

FLUX.3 is Black Forest Labs' frontier video model. This endpoint builds video from a sequence of keyframes, generating the motion between each anchor point for precise control over how a shot progresses.

Does
image, text → video
Release date
2026-08
Price
see supplier

via fal commercial

Flux 3 Image To Video Draft

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that animate a still image, with a reusable draft cache for full-quality enhancement.

Does
image, text → video
Release date
2026-08
Price
see supplier

via fal commercial

Flux 3 First Last Frame to Video Draft

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews between a start and an end frame, with a reusable draft cache for full-quality enhancement.

Does
image, text → video
Release date
2026-08
Price
see supplier

via fal commercial

Flux 3 Keyframes To Video Draft

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews pinned to your keyframe images, with a reusable draft cache for full-quality enhancement.

Does
image, text → video
Release date
2026-08
Price
see supplier

via fal commercial

Flux 3 Text to Video

FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene.

Does
text → video
Release date
2026-08
Price
see supplier

via fal commercial

Flux 3 Text To Video Draft

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews from a text prompt, with a reusable draft cache for full-quality enhancement.

Does
text → video
Release date
2026-08
Price
see supplier

via fal commercial

Flux 3 Extend Video

FLUX.3 is Black Forest Labs' frontier video model. This endpoint continues an existing clip beyond its final frame, generating additional footage that stays consistent with the original motion and scene.

Does
video, text → video-edit
Release date
2026-08
Price
see supplier

via fal commercial

Flux 3 Extend Video Draft

FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that continue an existing clip, with a reusable draft cache for full-quality enhancement.

Does
video, text → video-edit
Release date
2026-08
Price
see supplier

via fal commercial

Flux 3 Draft Enhance

FLUX.3 is Black Forest Labs' frontier audio/video model. Re-render a previously generated draft at full quality — same seed, same motion, no re-planning.

Does
video, text → video-edit
Release date
2026-08
Price
see supplier

via fal commercial

Heygen

Use Heygen's Latest Model for Filler Word Removal.

Does
video, text → video-edit
Release date
2026-08
Price
see supplier

via fal commercial

Grok Imagine Video 1.5 Reference to Video

Generate videos from images and audio references using xAI's Grok Imagine 1.5 Video model.

Does
image, text → video
Release date
2026-08
Price
$0.08 / sec

via fal commercial

Grok Imagine Video 1.5 Text to Video

Generate videos from prompts with audio using xAI's Grok Imagine 1.5 Video model.

Does
text → video
Release date
2026-08
Price
$0.08 / sec

via fal commercial

MiniMax H3 Image to Video

MiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.

Does
image, text → video
Release date
2026-07
Price
$0.08 / second

via fal commercial

MiniMax H3 Reference to Video

MiniMax H3 is a frontier video model. This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio.

Does
image, text → video
Release date
2026-07
Price
$0.08 / second

via fal commercial

MiniMax H3 Text to Video

MiniMax H3 is a frontier video model. This endpoint generates video from a text prompt alone, rendering at 2K in durations from 5 to 15 seconds across seven aspect ratios.

Does
text → video
Release date
2026-07
Price
$0.08 / second

via fal commercial

Pixelcut Product Photo

Pixelcut's Background Remover produces fast, high-quality cutouts built for e-commerce product imagery

Does
image, text → image-edit
Release date
2026-07
Price
see supplier

via fal commercial

MAI Image 2.5 Pro (Edit)

Apply precise, controllable edits to a reference image while preserving composition, typography, identity, and fine visual detail.

Does
image, text → image-edit
Release date
2026-07
Price
$7.50 / 1m

via fal commercial

MAI Image 2.5 Pro (Text to Image)

Generate high-fidelity, design-ready images with precise typography, strong prompt alignment, and rich visual detail using Microsoft's flagship MAI Image 2.5 Pro.

Does
text → image
Release date
2026-07
Price
$7.50 / 1m

via fal commercial

Ideogram Object Removal

Prompt-free object removal from an image and mask, erasing objects with their shadows and reflections and reconstructing the scene cleanly.

Does
image, text → image-edit
Release date
2026-07
Price
$0.03 / generation

via fal commercial

Qwen Audio 3.0 TTS (Flash)

Generate natural multilingual speech from text with fast voice and language control using Qwen Audio 3.0 TTS Flash.

Does
text → speech
Release date
2026-07
Price
see supplier

via fal commercial

Feynobg Background Remover

FeyNobg is a state of the art AI model for background removal from feyninc

Does
image, text → image-edit
Release date
2026-07
Price
see supplier

via fal commercial

Ltx 2.3 Quality

Remove character from your video using Ltx 2.3

Does
video, text → video-edit
Release date
2026-07
Price
$0.0024075 / megapixel

via fal commercial

Qwen Image 3 Image Editing

Generates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen's strength in complex text rendering and precise prompt adherence

Does
image, text → image-edit
Release date
2026-07
Price
$0.04 / generated

via fal commercial

Qwen Image 3 Text to Image

Edits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes

Does
text → image
Release date
2026-07
Price
$0.04 / generated

via fal commercial

V1.1 Text to Sound Effects

Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.

Does
text → music
Release date
2026-07
Price
$0.0018 / second

via fal commercial

V1.1 Video to Video Music

Generates perfectly synced music for any video. Return a licensed music soundtrack ready for commercial use (optional preservation of the original speech in video)

Does
video, text → video-edit
Release date
2026-07
Price
$0.009 / second

via fal commercial

V1.1 Video to Video Sound Effects

Adds synchronized, royalty-free, commercial-use-safe sound effects to a video. Returns the finished video with the generated audio mixed in.

Does
video, text → video-edit
Release date
2026-07
Price
$0.009 / second

via fal commercial

Ltx 2.3

LTX-2.3 Reframe converts your videos to any aspect ratio without destructive cropping. It intelligently recenters the original footage and generatively fills the newly exposed areas with content that seamlessly matches the scene, so the result looks like it was shot natively in the target format. Tu

Does
video, text → video-edit
Release date
2026-07
Price
$0.10 / second

via fal commercial

Bria Product Dimensions

Bria Product Dimensions turns one product photo and its measurements into a marketplace-ready dimension image with callout lines, labels, and weight or capacity readouts

Does
image, text → image-edit
Release date
2026-07
Price
see supplier

via fal commercial

Reve 2.1

Remix images from text prompts with strong prompt adherence, layout intelligence, and accurate text rendering using Reve 2.1

Does
image, text → image-edit
Release date
2026-07
Price
see supplier

via fal commercial

Reve 2.1

Edit images from text prompts with strong prompt adherence, layout intelligence, and accurate text rendering using Reve 2.1

Does
image, text → image-edit
Release date
2026-07
Price
see supplier

via fal commercial

Reve 2.1

Generate high-quality images from text prompts with strong prompt adherence, layout intelligence, and accurate text rendering using Reve 2.1.

Does
text → image
Release date
2026-07
Price
see supplier

via fal commercial

VEED Lipsync

Generate production-quality lipsync from any audio using VEED's most advanced model yet.

Does
video, text → video-edit
Release date
2026-07
Price
$0.07

via fal commercial

Krea 2 Text to Image Turbo Style

Generate high-fidelity images from text with Krea 2 using a style reference image. Apply a reference image to guide the visual style into new generations, with aspect ratio, creativity, and seed controls.

Does
text → image
Release date
2026-07
Price
see supplier

via fal commercial

Seedream

Image editing endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent image editing with multiple inputs.

Does
image, text → image-edit
Release date
2026-07
Price
see supplier

via fal commercial

Seedream 5.0 Pro Image Editing

Seedream 5.0 Pro is grounded, region-precise image editing model that changes one element while keeping the rest of the frame intact with layer separation, sketch completion, and up to 10 reference images.

Does
image, text → image-edit
Release date
2026-07
Price
$0.0045

via fal commercial

Seedream

Text to Image endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent text-to-image generation.

Does
text → image
Release date
2026-07
Price
see supplier

via fal commercial

Seedream 5.0 Pro Text to Image

ByteDance's Seedream 5.0 Pro is flagship text-to-image model, with deep-thinking prompt understanding, native text in 14 languages, and precise control over dense layouts and structured designs.

Does
text → image
Release date
2026-07
Price
$0.0675 / image

via fal commercial

Ltx 2.3 Quality

Extend high-quality video with audio from input video using LTX-2.3 with Lora

Does
video, text → video-edit
Release date
2026-07
Price
$0.0024075 / megapixel

via fal commercial

Ltx 2.3 Quality

Extend high-quality video with audio from input video using LTX-2.3

Does
video, text → video-edit
Release date
2026-07
Price
$0.0024075 / megapixel

via fal commercial

V4.0q [instant]

Generate high-quality images, posters, and logos with Ideogram's latest V4.0q — producing crisp visuals with accurate text rendering, fine detail, and full creative control for polished, ready-to-use designs FRACTION OF A SECOND.

Does
text → image
Release date
2026-07
Price
$0.0075 / megapixel

via fal commercial

V4.0q [fast]

Generate high-quality images, posters, and logos with Ideogram's latest V4.0q — producing crisp visuals with accurate text rendering, fine detail, and full creative control for polished, ready-to-use designs IN A SECOND.

Does
text → image
Release date
2026-07
Price
$0.00525 / megapixel

via fal commercial

TRELLIS.2 LoRA Inference

Run inference on LoRA adapters for TRELLIS.2 model

Does
image → 3d
Release date
2026-07
Price
see supplier

via fal commercial

Nano Banana Lite Edit

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

Does
image, text → image-edit
Release date
2026-06
Price
$0.3125

via fal commercial

Nano Banana 2 Lite

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

Does
text → image
Release date
2026-06
Price
$0.3125

via fal commercial

Nano Banana Lite

Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.

Does
text → image
Release date
2026-06
Price
$0.3125

via fal commercial

Gemini Omni Flash

Generates video with audio from combined multimodal references. Accepts text, images, audio, and video together as input to guide subject, motion, style, and sound in the output.

Does
image, text → video
Release date
2026-06
Price
$1.875 / 1

via fal commercial

Gemini Omni Flash

Animates a still image into video with audio. Extends a single frame into coherent motion, grounded in Gemini's physical understanding of how scenes and subjects behave.

Does
image, text → video
Release date
2026-06
Price
$1.875 / 1

via fal commercial

Gemini Omni Flash

Creates video with synchronized audio from text input. Grounded in Gemini's real-world knowledge, with improved physics understanding for more coherent motion and interaction.

Does
text → video
Release date
2026-06
Price
$21.875 / 1

via fal commercial

Gemini Omni Flash

Edits generated video across multiple conversational turns while preserving scene coherence. Applies iterative changes through natural-language instructions without regenerating the full sequence from scratch.

Does
video, text → video-edit
Release date
2026-06
Price
$1.875 / 1

via fal commercial

Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

Does
image, text → image
Release date
2026-06
Price
$1.5 / 1M out

language-model key commercial

Extract Object

Bria Extract Object uses text prompts to isolate a selected object from an image and return it as an RGBA PNG with a transparent background. Ideal for product, ecommerce, advertising, and creative editing workflows. Bria's Extract Object API leads in product shot extraction, outperforming SAM 3.1 wh

Does
image, text → image-edit
Release date
2026-06
Price
see supplier

via fal commercial

Ltx 2.3 Quality

Transform your 3D video render into realistic using first frame with Ltx 2.3

Does
video, text → video-edit
Release date
2026-06
Price
$0.0024075 / megapixel

via fal commercial

Seed Audio 1.0

Seed Audio 1.0 is a new audio model from Bytedance that can generate high-quality, natural sounding audio using text, reference audios or an image.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Ltx 2.3 Quality

Deblur high-quality video using LTX-2.3

Does
video, text → video-edit
Release date
2026-06
Price
$0.0024075 / megapixel

via fal commercial

Seedance 2.0 Mini

Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.

Does
image, text → video
Release date
2026-06
Price
$0.0721 / second

via fal commercial

Seedance 2.0 Mini Text to Video

Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.

Does
text → video
Release date
2026-06
Price
$0.0721 / second

via fal commercial

Krea 2 Turbo

Generate high-fidelity images from text in seconds with Krea 2 Turbo, the speed-optimized open-source version of Krea 2, preserving its aesthetic range for rapid ideation.

Does
text → image
Release date
2026-06
Price
see supplier

via fal commercial

Krea 2 Text to Image Turbo LoRA

Generate high-fidelity images from text with Krea 2 using a custom-trained LoRA. Apply your LoRA weights to carry a learned subject, character, or style into new generations, with aspect ratio, creativity, and seed controls.

Does
text → image
Release date
2026-06
Price
see supplier

via fal commercial

Happy Horse 1.1 Text to Video

Happy Horse 1.1 is Alibaba's #1-ranked video model. This text-to-video endpoint generates 1080p video with synchronized native audio and multilingual lip-sync from a text prompt alone.

Does
text → video
Release date
2026-06
Price
$0.14 / second

via fal commercial

Hyper3D - Rodin V2.5 - Image to 3D - Fast

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.

Does
image → 3d
Release date
2026-06
Price
$0.1 / generation

via fal commercial

Hyper3D - Rodin V2.5 - Text to 3D - Fast

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.

Does
text → 3d
Release date
2026-06
Price
$0.1 / generation

via fal commercial

Boogu Image

Text To Image Model using Boogu-Image

Does
text → image
Release date
2026-06
Price
$0.04 / megapixel

via fal commercial

Async Text to Speech Pro V1.0

Generate professional-quality voiceovers in seconds with Async TTS Pro model text-based control over pauses, emphasis, and timing. Voice ids can be found at https://async.com/developer/voice-library

Does
text → speech
Release date
2026-06
Price
$0.01 / 1

via fal commercial

Google: Nano Banana Pro (Gemini 3 Pro Image)

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

Does
image, text → image
Release date
2026-06
Price
$12 / 1M out

language-model key commercial

Google: Nano Banana 2 (Gemini 3.1 Flash Image)

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced...

Does
image, text → image
Release date
2026-06
Price
$3 / 1M out

language-model key commercial

Sensenova U1 Infographic

Generate Infographic Image with Sensenova U1

Does
text → image
Release date
2026-06
Price
$0.05 / generated

via fal commercial

Ltx 2.3 Quality

Text to Audio high-quality using LTX-2.3 with Lora

Does
text → music
Release date
2026-06
Price
$0.0024075 / megapixel

via fal commercial

Ltx 2.3 Quality

Text to Audio high-quality using LTX-2.3

Does
text → music
Release date
2026-06
Price
$0.0024075 / megapixel

via fal commercial

Kling Video V3 Turbo Pro Text to Video

Generate high quality 1080p videos using Kling's Turbo 3.0 model, with improved lipsync and multishot generation capabilities.

Does
text → video
Release date
2026-06
Price
$0.14

via fal commercial

Kling Video V3 Standard Turbo Text to Video

Kling 3.0 Turbo Standard is a fast, cost-efficient video generation model that turns text prompts directly into 720P video with native audio, optimized for rapid iteration and high-volume production

Does
text → video
Release date
2026-06
Price
$0.112

via fal commercial

Zonos2 Text to Speech

Zonos2 is a text-to-speech model that clones a voice from a short sample and speaks naturally across many languages.

Does
text → speech
Release date
2026-06
Price
see supplier

via fal commercial

Nemotron Asr Multilingual

Nemotron-ASR-Streaming is a multi lingual, streaming Automatic Speech Recognition (ASR) engineered to deliver high-quality multi lingual transcription across both low-latency streaming and high-throughput batch workloads.

Does
audio → transcribe
Release date
2026-06
Price
see supplier

via fal commercial

Luma Ray 3.2 Text to Video

Luma Ray 3.2 generates cinematic video from a text prompt, with control over resolution, duration, and seamless looping, plus reference images to lock in subject and style.

Does
text → video
Release date
2026-06
Price
$0.50

via fal commercial

Bernini-R Text to Video

Generate high-quality video from a text prompt with Bernini-R, ByteDance's unified video generation and editing model.

Does
text → video
Release date
2026-06
Price
$0.08 / second

via fal commercial

Bytedance Seed Speech Text to Speech

Seed Speech developed by ByteDance, is a family of large-scale text-to-speech models capable of synthesizing speech that is virtually indistinguishable from human speech.

Does
text → speech
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Base Text to Audio

Stable Audio 3 Medium Base is the foundational 1.4 billion parameter text-to-audio checkpoint generating stereo music up to 6 minutes, intended as the unmodified base for custom fine-tuning workflows.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3

Stable Audio 3 Small Music Base is the foundational 459 million parameter checkpoint generating full music compositions up to 2 minutes from text prompts, intended as the unmodified base for fine-tuning.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Base Text to Audio

Stable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as the unmodified base for fine-tuning.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3

Stable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for safe commercial use.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small Music Text to Audio

Stable Audio 3 Small Music is a 459 million parameter latent diffusion model that generates full stereo music compositions up to 2 minutes from text prompts, lightweight enough for on-device deployment.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Text to Audio

Stable Audio 3 Small SFX is a 459 million parameter latent diffusion model that generates high-quality sound effects from text prompts, designed for on-device deployment on mobile phones and consumer laptops.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Sonilo V1.1 Text to Music

Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.

Does
text → music
Release date
2026-06
Price
$0.0025 / second

via fal commercial

Triposplat

TripoSplat is an open-source model from TripoAI / VAST AI Research that converts a single 2D image into high-quality 3D Gaussians using a novel learned density-control approach

Does
image → 3d
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Base Audio Outpainting

Stable Audio 3 Medium Base audio outpainting is the foundational 1.4 billion parameter checkpoint that extends existing stereo audio with causal continuation guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Base Audio Inpainting

Stable Audio 3 Medium Base audio inpainting is the foundational 1.4 billion parameter checkpoint for editing or filling selected stereo audio segments guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Base Audio to Audio

Stable Audio 3 Medium Base audio-to-audio is the foundational 1.4 billion parameter checkpoint that transforms input audio into new stereo variations up to 6 minutes guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Base Audio Outpainting

Stable Audio 3 Small SFX Base audio outpainting is the foundational 459 million parameter checkpoint that extends sound-effect tracks via causal continuation guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small Music Base Audio to Audio

Stable Audio 3 Small Music Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input music into new variations up to 2 minutes guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Base Audio to Audio

Stable Audio 3 Small SFX Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input audio into new sound-effect variations guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Audio Outpainting

Stable Audio 3 Medium audio outpainting is a 1.4 billion parameter latent diffusion model that extends existing stereo audio beyond its original endpoint via causal continuation guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Audio Inpainting

Stable Audio 3 Medium audio inpainting is a 1.4 billion parameter latent diffusion model that fills in or reworks selected segments of a stereo track guided by text prompts, supporting single- and multi-segment editing.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Audio to Audio

Stable Audio 3 Medium audio-to-audio is a 1.4 billion parameter latent diffusion model that transforms an input audio clip into new stereo variations up to 6 minutes guided by a text prompt.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small Music Audio Outpainting

Stable Audio 3 Small Music audio outpainting is a 459 million parameter latent diffusion model that extends music compositions beyond their original endpoint via causal continuation.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Audio Outpainting

Stable Audio 3 Small SFX audio outpainting is a 459 million parameter latent diffusion model that extends sound-effect tracks beyond their original endpoint via causal continuation.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small Music Audio Inpainting

Stable Audio 3 Small Music audio inpainting is a 459 million parameter latent diffusion model that fills in or reworks selected segments of a music track guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Audio Inpainting

Stable Audio 3 Small SFX audio inpainting is a 459 million parameter latent diffusion model that fills in or reworks selected segments of a sound-effect track guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3

Stable Audio 3 Small Music audio-to-audio is a 459 million parameter latent diffusion model that transforms input music into new variations up to 2 minutes guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Ltx 2.3 Quality

Generate high-quality video with audio from text using LTX-2.3 and custom LoRA

Does
text → video
Release date
2026-06
Price
$0.0027075 / megapixel

via fal commercial

Ltx 2.3 Quality

Generate high-quality video with audio from text using LTX-2.3

Does
text → video
Release date
2026-06
Price
$0.0024075 / megapixel

via fal commercial

Hyper3D - Rodin V2.5 - Image to 3D

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.

Does
image → 3d
Release date
2026-05
Price
$0.4 / generation

via fal commercial

Hyper3D - Rodin V2.5 - Text to 3D

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.

Does
text → 3d
Release date
2026-05
Price
$0.4 / generation

via fal commercial

Lyria 3 Pro

Lyria 3 Pro is the latest music model from Google

Does
text → music
Release date
2026-05
Price
see supplier

via fal commercial

Mirelo SFX1.6

Generate ambient sounds for any text prompt. Now you can turn any SFX into a natural loop for ambient soundscapes.

Does
text → music
Release date
2026-05
Price
see supplier

via fal commercial

Pixal3d

Pixal3D turns a single image into a high-fidelity 3D model with detailed geometry and realistic textures.

Does
image → 3d
Release date
2026-05
Price
$0.3

via fal commercial

Meshy 6 - Multi Image To 3D

Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.

Does
image → 3d
Release date
2026-04
Price
$0.8 / untextured

via fal commercial

Cohere Transcribe

Cohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation

Does
audio → transcribe
Release date
2026-04
Price
see supplier

via fal commercial

OpenAI: GPT-5.4 Image 2

[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...

Does
image, text, file → image
Release date
2026-04
Price
$15 / 1M out

language-model key commercial

Gemini 3.1 Flash Tts

Newest audio model from Google introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.

Does
text → speech
Release date
2026-04
Price
see supplier

via fal commercial

Tripo H3.1 Multiview to 3D

Generate 3D models from multiple view images using Tripo H3.1.

Does
image → 3d
Release date
2026-04
Price
$0.20

via fal commercial

Tripo H3.1 Image to 3D

Generate high-quality 3D models from a single image using Tripo H3.1.

Does
image → 3d
Release date
2026-04
Price
$0.20

via fal commercial

Tripo P1 Image to 3D

Generate 3D models from a single image using Tripo P1.

Does
image → 3d
Release date
2026-04
Price
$0.40

via fal commercial

Tripo H3.1 Text to 3D

Generate 3D models from text descriptions using Tripo H3.1.

Does
text → 3d
Release date
2026-04
Price
$0.10

via fal commercial

Tripo P1 Text to 3D

Generate 3D models from text descriptions using Tripo P1.

Does
text → 3d
Release date
2026-04
Price
see supplier

via fal commercial

Minimax Music 2.5

MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.

Does
text → music
Release date
2026-04
Price
see supplier

via fal commercial

ReconViaGen 0.5

Generate 3D models from one or more images using ReconViaGen 0.5

Does
image → 3d
Release date
2026-04
Price
$0.25

via fal commercial

Google: Lyria 3 Clip Preview

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...

Does
text, image → speech
Release date
2026-03
Price
free

language-model key free

Google: Lyria 3 Pro Preview

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...

Does
text, image → speech
Release date
2026-03
Price
free

language-model key free

xAI Text to Speech

Generate speech with expressive and realistic voices from xAI

Does
text → speech
Release date
2026-03
Price
$0.015 / 1000

via fal commercial

Inworld TTS-1.5 Max

Text to Speech Endpoint for Inworld's TTS-1.5 Max.

Does
text → speech
Release date
2026-03
Price
$0.01 / 1000

via fal commercial

Trellis 2

Generate 3D models from your images using Trellis 2. A native 3D generative model enabling versatile and high-quality 3D asset creation.

Does
image → 3d
Release date
2026-03
Price
see supplier

via fal commercial

Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines...

Does
image, text → image
Release date
2026-02
Price
$3 / 1M out

language-model key commercial

Meshy 6

Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.

Does
image → 3d
Release date
2026-02
Price
see supplier

via fal commercial

Meshy 6

Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.

Does
text → 3d
Release date
2026-02
Price
see supplier

via fal commercial

MiniMax Speech 2.8 [HD]

Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

Does
text → speech
Release date
2026-02
Price
see supplier

via fal commercial

MiniMax Speech 2.8 [Turbo]

Generate speech from text prompts and different voices using the MiniMax Speech-2.8 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.

Does
text → speech
Release date
2026-02
Price
see supplier

via fal commercial

Hunyuan 3d

Create detailed, fully-textured 3D models with text

Does
text → 3d
Release date
2026-01
Price
$0.225 / generation

via fal commercial

Hunyuan 3D Pro Text to 3D

Generate 3D models from text prompts with Hunyuan 3D Pro

Does
text → 3d
Release date
2026-01
Price
$0.375 / generation

via fal commercial

Qwen 3 TTS - Voice Design [1.7B]

Create custom voices using Qwen3-TTS Voice Design model and later use Clone Voice model to create your own voices!

Does
text → speech
Release date
2026-01
Price
see supplier

via fal commercial

Qwen 3 TTS - Text to Speech [1.7B]

Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model

Does
text → speech
Release date
2026-01
Price
see supplier

via fal commercial

Qwen 3 TTS - Text to Speech [0.6B]

Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model

Does
text → speech
Release date
2026-01
Price
see supplier

via fal commercial

OpenAI: GPT Audio

The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...

Does
text, audio → speech
Release date
2026-01
Price
$10 / 1M out

language-model key commercial

OpenAI: GPT Audio Mini

A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...

Does
text, audio → speech
Release date
2026-01
Price
$2.4 / 1M out

language-model key commercial

ElevenLabs Speech to Text - Scribe V2

Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!

Does
audio → transcribe
Release date
2026-01
Price
$0.008 / input

via fal commercial

Hunyuan Motion [0.46B]

Generate 3D human motions via text-to-generation interface of Hunyuan Motion!

Does
text → 3d
Release date
2025-12
Price
see supplier

via fal commercial

Hunyuan Motion [1B]

Generate 3D human motions via text-to-generation interface of Hunyuan Motion!

Does
text → 3d
Release date
2025-12
Price
see supplier

via fal commercial

Vibevoice

Generate long speech snippets fast using Microsoft's powerful TTS.

Does
text → speech
Release date
2025-12
Price
see supplier

via fal commercial

Hunyuan3d V3

Turn simple sketches into detailed, fully-textured 3D models. Instantly convert your concept designs into formats ready for Unity, Unreal, and Blender.

Does
text → 3d
Release date
2025-12
Price
$0.375 / generation

via fal commercial

Maya

Maya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise voice design.

Does
text → speech
Release date
2025-12
Price
see supplier

via fal commercial

Google: Nano Banana Pro (Gemini 3 Pro Image Preview)

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

Does
image, text → image
Release date
2025-11
Price
$12 / 1M out

language-model key commercial

OpenAI: GPT-5 Image Mini

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...

Does
file, image, text → image
Release date
2025-10
Price
$2 / 1M out

language-model key commercial

OpenAI: GPT-5 Image

[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...

Does
image, text, file → image
Release date
2025-10
Price
$10 / 1M out

language-model key commercial

Google: Nano Banana (Gemini 2.5 Flash Image)

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,...

Does
image, text → image
Release date
2025-10
Price
$2.5 / 1M out

language-model key commercial

Meshy 6 Preview

Meshy-6-Preview is the latest model from Meshy. It generates realistic and production ready 3D models.

Does
text → 3d
Release date
2025-10
Price
see supplier

via fal commercial

Pipecat's Smart Turn model

An open source, community-driven and native audio turn detection model by Pipecat AI.

Does
audio → transcribe
Release date
2025-04
Price
see supplier

via fal commercial

Speech-to-Text

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

Does
audio → transcribe
Release date
2025-04
Price
see supplier

via fal commercial

Speech-to-Text

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

Does
audio → transcribe
Release date
2025-04
Price
see supplier

via fal commercial

ElevenLabs Speech to Text

Generate text from speech using ElevenLabs advanced speech-to-text model.

Does
audio → transcribe
Release date
2025-02
Price
see supplier

via fal commercial

Wizper (Whisper v3 -- fal.ai edition)

[Experimental] Whisper v3 Large -- but optimized by our inference wizards. Same WER, double the performance!

Does
audio → transcribe
Release date
2024-04
Price
see supplier

via fal commercial

Wizper (Whisper v3 -- fal.ai edition)

[Experimental] Whisper v3 Large -- but optimized by our inference wizards. Same WER, double the performance!

Does
audio → transcribe
Release date
2024-04
Price
see supplier

via fal commercial

product shot with rendered copy — text drawn by the model
editorial portrait
isometric illustration

Nano Banana Pro (Gemini Image)

Google's image model — the reference for rendering copy/text inside images verbatim, plus strong instruction editing. The go-to for social assets with words on them.

Does
text, image → image, image-edit
Price
see supplier

cloud API commercial

action photography
cinematic street scene
macro detail

FLUX.1 dev

High-quality general image generation — photoreal scenes, cheap and fast.

Does
text → image
Price
see supplier

cloud API commercial

Veo 3.1 Fast

Google's video model — 8 s clips WITH generated sound and dialogue, vertical or landscape.

Does
text, image → video
Price
see supplier

cloud API commercial

LTX Video (cloud)

Fast, inexpensive silent clips — the default cloud backend of video_generate.

Does
text, image → video
Price
see supplier

cloud API commercial

Lyria 2

Google's music model — rich 30 s instrumentals from a mood prompt.

Does
text → music
Price
see supplier

cloud API commercial

storefront with rendered signage
food photography
flat-design poster

Nano Banana 2 (Gemini Image)

Google's newest fast image model — Nano Banana quality at higher speed, generation and editing.

Does
text, image → image, image-edit
Price
see supplier

cloud API commercial

detailed typography
photoreal scene
technical illustration

GPT Image 2

OpenAI's latest image model — extremely detailed output with fine typography, plus fine-grained edits.

Does
text, image → image, image-edit
Price
see supplier

cloud API commercial

Seedance 2.0

ByteDance's most advanced video model — cinematic output with native audio, real-world physics, and reference-to-video from up to 9 images, 3 videos, and 3 audio clips.

Does
text, image, images, audio → video
Price
see supplier

cloud API commercial

Kling Video v3 Pro

Kuaishou's top-tier video model — cinematic visuals, fluid motion, native audio, custom element support.

Does
text, image → video
Price
see supplier

cloud API commercial

Happy Horse 1.1

Alibaba's #1-ranked video model — 1080p output with synchronized native audio and multilingual lip-sync.

Does
text, image → video
Price
see supplier

cloud API commercial

Lyria 3 Pro

Google's latest music model — richer arrangements and cleaner mixes than Lyria 2.

Does
text → music
Price
see supplier

cloud API commercial

text-to-image
text-to-image
text-to-image

Grok Imagine (images)

xAI's image model — bold, stylized generation and editing from the Grok family.

Does
text, image → image, image-edit
Price
see supplier

cloud API commercial

Grok Imagine Video 1.5

xAI's video model — text/image to video with audio, plus reference-to-video, edit and extend endpoints.

Does
text, image, video → video
Price
see supplier

cloud API commercial

Hunyuan 3D v3.1 Pro

Tencent's 3D generation model — turn a prompt or a single image into a textured 3D mesh (GLB).

Does
text, image → 3d
Price
see supplier

cloud API commercial

4× upscale of a 448px sample

Topaz Upscale

Professional-grade upscaling for images and video — sharper detail, denoise, real resolution gains.

Does
image, video → upscale
Price
see supplier

cloud API commercial

ElevenLabs Eleven v3

State-of-the-art expressive text-to-speech — natural delivery, emotion, and multilingual voices.

Does
text → speech
Price
see supplier

cloud API commercial

barista pouring latte art · photoreal
autumn park · flat illustration
"OPEN LATE" shop sign · text rendering

Z-Image Turbo

Best open image model that fits this class of Mac — photorealism, sharp text, strong prompt following. 6B, 4-bit, runs on the Apple GPU via MLX.

Does
text → image
Price
free

on your Mac free Apache 2.0

luthier portrait · warm studio light
seaside village · isometric diorama
ceramic teapot · product shot

FLUX.2 klein 4B

Fast 4-step image generation AND instruction editing ("make it snow") — the only open commercial-use editing model that fits 16 GB.

Does
text, image, images → image, image-edit
Price
free

on your Mac free Apache 2.0

lighthouse in a storm · cinematic
"ILHAS DOS AÇORES" poster · text rendering
dew on a spider web · macro

FLUX.2 klein 9B

The bigger klein — noticeably stronger detail and prompt following than the 4B, still 4-step fast. Non-commercial license (unlike the 4B).

Does
text, image, images → image, image-edit
Price
free

on your Mac free FLUX klein 9B

Taipei night market · documentary
karst mountains · ink wash
coffee machine · cutaway illustration

ERNIE-Image Turbo

Baidu's 8B image model — strong photorealism and composition, Apache-licensed. Runs 4-bit on the Apple GPU via MLX (downloads the full official weights, quantized at load).

Does
text → image
Price
free

on your Mac free Apache 2.0

Juggernaut XL v9

The most popular SDXL finetune — polished photoreal people and scenes, the gateway to the civitai ecosystem. Runs on the managed ComfyUI.

Does
text → image
Price
free

on your Mac free CreativeML OpenRAIL-M

Wan 2.2 (5B)

Alibaba's open video model — the one local video generator that genuinely fits 16 GB (4-bit GGUF). Silent clips, ~3 s at reduced resolution.

Does
text, image → video
Price
free

on your Mac free Apache 2.0

Stable Audio 3 Small

Near-instant instrumental music and sound effects (44.1 kHz stereo) — pure MLX, tiny RAM footprint. No vocals.

Does
text, audio → music, sfx
Price
free

on your Mac free Stability Community

ACE-Step 1.5

Full songs WITH vocals and lyrics in 50+ languages — the best open local music model (between Suno v4.5 and v5). Also does covers and repaints.

Does
text, lyrics → song, music
Price
free

on your Mac free MIT

LTX-2.3 (video + audio)

The only local model anywhere that generates video WITH synchronized sound. 4-bit MLX port. Its pipeline pulls the full 56 GB weight set and wants real memory headroom — a 32 GB+ Mac.

Does
text, image → video
Price
free

on your Mac free LTX Community

cyclist at golden hour · 35mm film
"MAKE IT LOCAL" poster · swiss type
rainy Lisbon tram stop · watercolour

Qwen-Image (20B)

Top-tier image generation + the best open image editor. Needs a 32 GB+ Mac.

Does
text, image → image, image-edit
Price
free

on your Mac free Apache 2.0

Wan 2.7 (cloud only)

Alibaba's current video model, including prompt-based video editing — but API-only. No open weights: the newest downloadable Wan is 2.2.

Does
text, image, video → video, video-edit
Price
free

on your Mac free closed

InfiniteTalk (14B)

Dubs an existing video: sparse-frame video-to-video that re-lips footage to new audio while keeping the original motion. Unlimited length via chunked continuation. Built on Wan 2.1 I2V.

Does
video, audio, image → video, video-edit
Price
free

on your Mac free Apache 2.0

EchoMimicV3 Flash (1.3B)

Animates a portrait from audio at 768×768 in 8 steps — the best quality-per-VRAM in this category, and clearer licensing than LatentSync (whose weights are openrail++, not Apache).

Does
image, audio → video, video-edit
Price
free

on your Mac free Apache 2.0

MuseTalk 1.5

Real-time lip sync — single-step latent inpainting rather than a diffusion sampler, so it is fast and small. Capped at 256×256, and the authors note some jitter and lip-colour drift.

Does
video, audio → video, video-edit
Price
free

on your Mac free MIT (code + weights)

rustic galette · editorial food
solar-punk rooftop farm · concept art
border collie · studio portrait

FLUX.2 dev (32B)

Frontier-class open image model. Needs a 64 GB+ Mac.

Does
text, images → image, image-edit
Price
free

on your Mac free FLUX dev (non-commercial)

Examples are ours — generated with the model and hosted by us. Where a card has no example we have not made one yet; nothing on this page is loaded from a vendor's servers, so browsing it tells them nothing about you.

Speech 64

Prices are per MINUTE of audio, whatever unit the supplier quotes in — per-character text-to-speech is converted at 900 characters a minute so the columns can be compared at all. Word-error rates come from one independent harness, so they are comparable with each other.

Text to speech

No voice clips yet — we only publish samples we generated and host ourselves, and these have not been made.

ModelMaker$/min LatencyLanguagesDoes
Polly Standard AWS $0.0036 300 ms 29 streaming
Grok Voice TTS xAI $0.0038 30 streaming voice cloning
Octave 2 Hume $0.0068 11 streaming voice cloning
GPT-4o Mini TTS OpenAI $0.013 300 ms 50 streaming
TTS-1 OpenAI $0.013 400 ms 50 streaming
Aura-1 Deepgram $0.013 130 ms 1 streaming
OpenAudio S2 Pro Fish Audio $0.013 13 voice cloning
Neural2 Google Cloud $0.014 40
Neural Microsoft Azure $0.014 300 ms 140 streaming
Polly Neural AWS $0.014 400 ms 40 streaming
Neural HD Microsoft Azure $0.020 300 ms 140 streaming
TTS-1 HD OpenAI $0.027 50
Aura-2 Deepgram $0.027 150 ms 15 streaming
Chirp 3 HD Google Cloud $0.027 500 ms 31 streaming
Polly Generative AWS $0.027 12 streaming
Sonic 3 Cartesia $0.032 90 ms 42 streaming voice cloning
Mist v2 Rime $0.035 100 ms 2 streaming
Speech 2.6 MiniMax $0.041 40 streaming voice cloning
Flash v2.5 ElevenLabs $0.045 75 ms 32 streaming voice cloning
Multilingual v2 ElevenLabs $0.090 280 ms 32 streaming voice cloning
Eleven v3 ElevenLabs $0.090 70 voice cloning
Kokoro 82M local Kokoro · local 120 ms 8 offline
Piper local Piper · local 80 ms 30 offline
macOS say (built-in) local Apple · local 50 ms 40 offline
Chatterbox Turbo local Resemble · local 200 ms 23 voice cloning offline

Transcription

ModelMaker$/min Word errors LatencyLanguagesDoes
Wizper (Whisper v3) fal.ai $0.0005 4.7% 99
Whisper Large v3 Turbo Groq $0.0007 4.6% 99 timestamps
Parakeet TDT 0.6B v3 Together AI $0.0015 4.5% 25 timestamps
Soniox v5 Soniox $0.0017 3.8% 60 streaming speaker labels timestamps
Grok Speech-to-Text xAI $0.0017 4% 30 streaming
Universal-Streaming AssemblyAI $0.0025 300 ms 99 streaming
GPT-4o Mini Transcribe OpenAI $0.0030 4.5% 99 streaming
Voxtral Mini Transcribe 2 Mistral $0.0030 3.6% 30 speaker labels timestamps
Nova 2 Pro (transcribe) AWS Bedrock $0.0031 4.9% 100
Universal-3.5 Pro AssemblyAI $0.0035 3.1% 99 speaker labels timestamps
Scribe v2 ElevenLabs $0.0037 2.2% 90 speaker labels timestamps
Voxtral Small Mistral $0.0040 2.8% 30 timestamps
Melia Speechmatics $0.0040 4.9% 55 streaming speaker labels
Pulse Pro Smallest.ai $0.0040 2.4% 30 streaming
Nova-3 Deepgram $0.0043 5.2% 36 streaming speaker labels timestamps
GPT Transcribe OpenAI $0.0045 3.3% 99 streaming timestamps
Nova-3 Multilingual Deepgram $0.0052 36 streaming speaker labels timestamps
GPT-4o Transcribe OpenAI $0.0060 4% 99 streaming
GPT-4o Transcribe (diarize) OpenAI $0.0060 99 speaker labels timestamps
Whisper v2 (hosted) OpenAI $0.0060 4.1% 99 timestamps
Voxtral Realtime Mistral $0.0060 200 ms 30 streaming
MAI-Transcribe-1 Microsoft Azure $0.0060 2.6% 100 speaker labels timestamps
Amazon Transcribe AWS $0.0060 4.1% 100 streaming speaker labels timestamps
Scribe v2 Realtime ElevenLabs $0.0065 150 ms 90 streaming timestamps
Flux (conversational) Deepgram $0.0077 260 ms 1 streaming
Solaria-3 Gladia $0.010 3.2% 100 streaming speaker labels
Chirp 3 Google Cloud $0.016 125 streaming speaker labels timestamps
GPT Live Transcribe OpenAI $0.017 300 ms 99 streaming timestamps
Parakeet TDT 0.6B v3 (MLX) local NVIDIA · local 6.3% 25 timestamps offline
Whisper Large v3 Turbo (whisper.cpp) local OpenAI · local 7.4% 99 timestamps offline
Whisper Base (whisper.cpp) local OpenAI · local 99 timestamps offline
Voxtral Mini 4B (MLX, 4-bit) local Mistral · local 30 offline

Speech to speech (live conversation)

ModelMaker$/min LatencyLanguagesDoes
GPT Realtime Mini OpenAI $0.018 450 ms 60 streaming
Gemini Live Google $0.024 600 ms 70 streaming
Grok Voice Agent xAI $0.050 500 ms 30 streaming
GPT Realtime 2.1 OpenAI $0.058 500 ms 60 streaming
Line Cartesia $0.060 90 ms 42 streaming voice cloning
Voice Agent API Deepgram $0.075 36 streaming
Speech Engine ElevenLabs $0.080 90 streaming voice cloning

Speech prices checked 2026-07-29. Ballpark, pay-as-you-go, mid-tier — they move constantly. fal's registry read 2026-08-10 (1437 models scanned, newest per category kept).