/ Models

Media models

Everything 00 can generate with — pictures, video, music, speech, 3D.

286 models Use in your app
product shot with rendered copy — text drawn by the model
editorial portrait
isometric illustration

Nano Banana Pro (Gemini Image)

Google's image model — the reference for rendering copy/text inside images verbatim, plus strong instruction editing. The go-to for social assets with words on them.

Does
text, image → image, image-edit
Price
≈ $0.31 / image

AI tokens — built in commercial

action photography
cinematic street scene
macro detail

FLUX.1 dev

High-quality general image generation — photoreal scenes, cheap and fast.

Does
text → image
Price
see supplier

cloud API commercial

Veo 3.1 Fast

Google's video model — 8 s clips WITH generated sound and dialogue, vertical or landscape.

Does
text, image → video
Price
see supplier

cloud API commercial

LTX Video (cloud)

Fast, inexpensive silent clips — the default cloud backend of video_generate.

Does
text, image → video
Price
see supplier

cloud API commercial

Lyria 2

Google's music model — rich 30 s instrumentals from a mood prompt.

Does
text → music
Price
see supplier

cloud API commercial

storefront with rendered signage
food photography
flat-design poster

Nano Banana 2 (Gemini Image)

Google's newest fast image model — Nano Banana quality at higher speed, generation and editing.

Does
text, image → image, image-edit
Price
≈ $0.155 / image

AI tokens — built in commercial

detailed typography
photoreal scene
technical illustration

GPT Image 2

OpenAI's latest image model — extremely detailed output with fine typography, plus fine-grained edits.

Does
text, image → image, image-edit
Price
≈ $0.077 / image

AI tokens — built in commercial

Seedance 2.0

ByteDance's most advanced video model — cinematic output with native audio, real-world physics, and reference-to-video from up to 9 images, 3 videos, and 3 audio clips.

Does
text, image, images, audio → video
Price
see supplier

cloud API commercial

Kling Video v3 Pro

Kuaishou's top-tier video model — cinematic visuals, fluid motion, native audio, custom element support.

Does
text, image → video
Price
see supplier

cloud API commercial

Lyria 3 Pro

Google's latest music model — richer arrangements and cleaner mixes than Lyria 2.

Does
text → music
Price
see supplier

cloud API commercial

text-to-image
text-to-image
text-to-image

Grok Imagine (images)

xAI's image model — bold, stylized generation and editing from the Grok family.

Does
text, image → image, image-edit
Price
see supplier

cloud API commercial

Grok Imagine Video 1.5

xAI's video model — text/image to video with audio, plus reference-to-video, edit and extend endpoints.

Does
text, image, video → video
Price
see supplier

cloud API commercial

4× upscale of a 448px sample

Topaz Upscale

Professional-grade upscaling for images and video — sharper detail, denoise, real resolution gains.

Does
image, video → upscale
Price
see supplier

cloud API commercial

ElevenLabs Eleven v3

State-of-the-art expressive text-to-speech — natural delivery, emotion, and multilingual voices.

Does
text → speech
Price
see supplier

cloud API commercial

barista pouring latte art · photoreal
autumn park · flat illustration
"OPEN LATE" shop sign · text rendering

Z-Image Turbo

Best open image model that fits this class of Mac — photorealism, sharp text, strong prompt following. 6B, 4-bit, runs on the Apple GPU via MLX.

Does
text → image
Price
free

on your Mac free Apache 2.0

luthier portrait · warm studio light
seaside village · isometric diorama
ceramic teapot · product shot

FLUX.2 klein 4B

Fast 4-step image generation AND instruction editing ("make it snow") — the only open commercial-use editing model that fits 16 GB.

Does
text, image, images → image, image-edit
Price
free

on your Mac free Apache 2.0

lighthouse in a storm · cinematic
"ILHAS DOS AÇORES" poster · text rendering
dew on a spider web · macro

FLUX.2 klein 9B

The bigger klein — noticeably stronger detail and prompt following than the 4B, still 4-step fast. Non-commercial license (unlike the 4B).

Does
text, image, images → image, image-edit
Price
free

on your Mac free FLUX klein 9B

Taipei night market · documentary
karst mountains · ink wash
coffee machine · cutaway illustration

ERNIE-Image Turbo

Baidu's 8B image model — strong photorealism and composition, Apache-licensed. Runs 4-bit on the Apple GPU via MLX (downloads the full official weights, quantized at load).

Does
text → image
Price
free

on your Mac free Apache 2.0

stained-glass whale · flooded cathedral
island map · ink on parchment
astronaut and a flower · storybook

FLUX.1 schnell

Black Forest Labs' fast FLUX — four steps to a finished image, and the most permissive licence of any model this good. Apache 2.0: yours to use commercially, no strings. Runs on the managed ComfyUI.

Does
text → image
Price
free

on your Mac free Apache 2.0

typewriter · text rendered legibly
fisherman portrait · window light
dragonfly · macro

FLUX.1 dev

The bigger FLUX — slower than schnell and noticeably better at hands, text in images and fine detail. Its licence is NON-COMMERCIAL: fine for personal work, not for anything you sell.

Does
text → image
Price
free

on your Mac free FLUX.1-dev

lighthouse at dusk · cinematic
red bicycle · product light

Juggernaut XL v9

The most popular SDXL finetune — polished photoreal people and scenes, the gateway to the civitai ecosystem. Runs on the managed ComfyUI.

Does
text → image
Price
free

on your Mac free CreativeML OpenRAIL-M

before — the photograph it was given
after — “make the bicycle deep blue”
after — “turn it into a pencil sketch”
after — “add a wicker basket”

FLUX.1 Kontext dev

Instruction image EDITING — give it a photo and a sentence ("make it snow", "remove the car", "turn this into a pencil sketch") and it changes that and leaves the rest of the picture alone. The first editor here that does not need a Mac: it runs on a Linux GPU server. Non-commercial licence.

Does
image → image, image-edit
Price
free

on your Mac free FLUX.1-dev

Wan 2.2 (5B)

Alibaba's open video model — the one local video generator that genuinely fits 16 GB (4-bit GGUF). Silent clips, ~3 s at reduced resolution.

Does
text, image → video
Price
free

on your Mac free Apache 2.0

Wan 2.1 VACE 14B

Edit a clip you already have: restyle it, change what is in it, keep its motion ("the same dance, but the dancer is made of glass"). One model, Apache 2.0, on a Linux GPU server — the first thing here that edits video at all.

Does
video → video-edit
Price
free

on your Mac free Apache 2.0

Stable Audio 3 Small

Near-instant instrumental music and sound effects (44.1 kHz stereo) — pure MLX, tiny RAM footprint. No vocals.

Does
text, audio → music, sfx
Price
free

on your Mac free Stability Community

Stable Audio 3 Small-SFX

Sound effects and short ambiance from a sentence ("footsteps on gravel", "distant thunder") — 44.1 kHz stereo, up to two minutes. The first sound model here that does not need a Mac: it runs on a Linux GPU server.

Does
text → sfx
Price
free

on your Mac free Stability Community

ACE-Step 1.5

Full songs WITH vocals and lyrics in 50+ languages — the best open local music model (between Suno v4.5 and v5). Also does covers and repaints.

Does
text, lyrics → song, music
Price
free

on your Mac free MIT

LTX-2.3 (video + audio)

The only local model anywhere that generates video WITH synchronized sound. 4-bit MLX port. Its pipeline pulls the full 56 GB weight set and wants real memory headroom — a 32 GB+ Mac.

Does
text, image → video
Price
free

on your Mac free LTX Community

cyclist at golden hour · 35mm film
"MAKE IT LOCAL" poster · swiss type
rainy Lisbon tram stop · watercolour

Qwen-Image (20B)

Top-tier image generation + the best open image editor. Needs a 32 GB+ Mac.

Does
text, image → image, image-edit
Price
free

on your Mac free Apache 2.0

Wan 2.7 (cloud only)

Alibaba's current video model, including prompt-based video editing — but API-only. No open weights: the newest downloadable Wan is 2.2.

Does
text, image, video → video, video-edit
Price
free

on your Mac free closed

rustic galette · editorial food
solar-punk rooftop farm · concept art
border collie · studio portrait

FLUX.2 dev (32B)

Frontier-class open image model. Needs a 64 GB+ Mac.

Does
text, images → image, image-edit
Price
free

on your Mac free FLUX.1-dev

Seedream

Seedream 5.0 Flash is a fast image generation and editing model, built for workflows where speed and budget matter.

Does
image, text → image-edit
Release date
2026-09
Price
see supplier

via fal commercial

Seedream

Seedream 5.0 Flash is a fast image generation and editing model, built for workflows where speed and budget matter.

Does
image, text → image-edit
Release date
2026-09
Price
see supplier

via fal commercial

Seedream

Seedream 5.0 Flash is a fast image generation and editing model, built for workflows where speed and budget matter.

Does
text → image
Release date
2026-09
Price
see supplier

via fal commercial

Recraft V4.1 Flash Text to Image

Recraft V4.1 Flash generates raster images from text prompts, including photography, illustrations, and mixed-media compositions, with controls for image size, color palette, and background color.

Does
text → image
Release date
2026-09
Price
$0.007 / image

via fal commercial

inclusionAI: Ming Image 0.1 Design Layer

Ming Image 0.1 Design Layer is an image-to-image model from inclusionAI that decomposes a flattened design image into separate RGBA layers, such as a background layer and foreground elements, and...

Does
text, image → image, image-edit
Release date
2026-09
Price
≈ $0 / image

AI tokens — built in commercial

Recraft: Recraft V4.1 Flash

Recraft V4.1 Flash is a text-to-image model from Recraft, the speed and cost tier of the V4.1 family. It generates ~1K raster images in about 1.5 seconds end to end,...

Does
text → image
Release date
2026-09
Price
$0.014 / image

AI tokens — built in commercial

inclusionAI: Ming Image 0.1 Design

Ming Image 0.1 Design is a text-to-image model from inclusionAI aimed at graphic-design output, with an emphasis on legible text rendering inside the generated image. It generates from a prompt...

Does
text → image
Release date
2026-09
Price
≈ $0 / image

AI tokens — built in commercial

Meshy 7.1 Image to 3D

Meshy 7.1 generates 3D models from a single image, with standard, low-poly, and Smart Topology modes, optional textures and PBR maps, and geometry resolution up to 4K.

Does
image → 3d
Release date
2026-09
Price
$0.12 / call

via fal commercial

Meshy 7.1 Multi Image to 3D

Meshy 7.1 generates textured 3D models from one to four views of the same object, with polygon count, topology, symmetry, and optional PBR texture controls.

Does
image → 3d
Release date
2026-09
Price
$0.12 / call

via fal commercial

Meshy 7.1 Text to 3D

Meshy 7.1 generates 3D models from text prompts, with untextured preview and textured full modes, standard, low-poly, and Smart Topology options, and geometry resolution up to 4K.

Does
text → 3d
Release date
2026-09
Price
$0.12 / call

via fal commercial

H3 Max 16-bit Pixel

Generates 768p video with audio in a 16-bit pixel-art style from text prompts or an optional first-frame image. Supports durations of 5–15 seconds.

Does
text → video
Release date
2026-09
Price
$0.08 / second

via fal commercial

H3 Max Hand Drawn

Generates 768p video with audio in a hand-drawn animation style from text prompts or an optional first-frame image. Supports durations of 5–15 seconds.

Does
text → video
Release date
2026-09
Price
$0.08 / second

via fal commercial

H3 Max Low Poly

Generates 768p video with audio in a retro low-poly 3D style from text prompts or an optional first-frame image. Supports durations of 5–15 seconds.

Does
text → video
Release date
2026-09
Price
$0.08 / second

via fal commercial

H3 Max Retro Toon 70s

Generates 768p video with audio in a retro 1970s hand-painted animation style from text prompts or an optional first-frame image. Supports durations of 5–15 seconds.

Does
text → video
Release date
2026-09
Price
$0.08 / second

via fal commercial

H3 Max VHS

Generates 768p VHS-style video with audio from text prompts or an optional first-frame image. Supports 5–15 second clips and adjustable tape damage, from subtle analog noise to strong tracking distortion.

Does
text → video
Release date
2026-09
Price
$0.08 / second

via fal commercial

Tripo P2 Image to 3D

Tripo P2 generates 3D models from a single image, with optional PBR textures, adjustable face counts, and triangle or quad mesh topology.

Does
image → 3d
Release date
2026-09
Price
$1.00

via fal commercial

Tripo P2 Text to 3D

Tripo P2 generates 3D models from a text prompt, with optional PBR textures, adjustable face counts, and triangle or quad mesh topology.

Does
text → 3d
Release date
2026-09
Price
$1.00

via fal commercial

Seedance 2.5 US Image to Video

US-hosted ByteDance Seedance 2.5 animates still images with synchronized audio and optional end-frame control. Generate videos up to 30 seconds at 480p or 720p.

Does
image, text → video
Release date
2026-09
Price
$0.5676 / second

via fal commercial

Seedance 2.5 US Reference to Video

US-hosted ByteDance Seedance 2.5 generates video with native audio from up to 30 images, 10 videos, and 10 audio references. Supports reference-guided generation, video editing, and extension at 480p or 720p.

Does
image, text → video
Release date
2026-09
Price
$0.5676 / second

via fal commercial

Seedance 2.5 US Text to Video

US-hosted ByteDance Seedance 2.5 generates cinematic video from text with synchronized audio, up to 30-second duration, and 480p or 720p output.

Does
text → video
Release date
2026-09
Price
$0.5676 / second

via fal commercial

Lyria 3.5

Lyria 3.5 is Google DeepMind's latest music generation model, and you can generate almost any type of music with it

Does
text → music
Release date
2026-09
Price
see supplier

via fal commercial

H3 Max Lip Sync Image to Video

H3 Max Lip Sync generates a video from an image and supplied audio, synchronizing mouth movements to the soundtrack. It supports optional transcription guidance and output resolutions from 480p to 2K.

Does
image, text → video
Release date
2026-09
Price
$0.05 / second

via fal commercial

Bria Product Holding (FIBO-Edit-1.5)

Bria Product Holding edits a person photo to show the subject holding or carrying a product, using one to three product reference images and optional text instructions. Built on FIBO-Edit-1.5, it preserves the source aspect ratio by default and supports a selectable output aspect ratio.

Does
image, text → image-edit
Release date
2026-09
Price
$0.04

via fal commercial

Bria Virtual Try-On (FIBO-Edit-1.5)

Bria Virtual Try-On edits a person photo to show the subject wearing garments or accessories from one to three reference images, guided by optional text instructions. Built on FIBO-Edit-1.5, it supports multi-garment changes and preserves the source aspect ratio by default.

Does
image, text → image-edit
Release date
2026-09
Price
$0.04

via fal commercial

Seedance 2.0 US Image to Video

US hosted version of ByteDance's most advanced image-to-video model. Animate still images into cinematic video with synchronized audio, start and end frame control, and motion prompts.

Does
image, text → video
Release date
2026-09
Price
$0.37 / second

via fal commercial

Seedance 2.0 US Reference to Video

US hosted version of ByteDance's most advanced reference-to-video model. Generate video from up to 9 images, 3 videos, and 3 audio clips with native audio and cinematic camera control.

Does
image, text → video
Release date
2026-09
Price
$0.37 / second

via fal commercial

Seedance 2.0 US Text to Video

US hosted version of ByteDance's most advanced text-to-video model. Cinematic output with native audio, multi-shot editing, real-world physics, and director-level camera control.

Does
text → video
Release date
2026-09
Price
$0.37 / second

via fal commercial

Boreal

Create product, UGC, and presenter videos with synchronized native audio from text, with optional image and audio inputs.

Does
text → video
Release date
2026-09
Price
$0.01 / second

via fal commercial

Elevenlabs Music v2

Generate high quality, realistic music with fine controls using Elevenlabs Music v2!

Does
text → music
Release date
2026-09
Price
$0.6 / output

via fal commercial

Elevenlabs Music v2.5

Generate high quality, realistic music with fine controls using Elevenlabs Music v2.5!

Does
text → music
Release date
2026-09
Price
$0.6 / output

via fal commercial

ID-V2V

Restyle a video’s scene, lighting, and visual style from edited keyframes while preserving the source subjects’ identity, expressions, gaze, and motion. Developed by Eyeline Labs and Netflix researchers.

Does
video, text → video-edit
Release date
2026-09
Price
see supplier

via fal commercial

ID-V2V Relight

Change a video’s lighting using a relit reference frame while preserving the scene, subjects, and original performance. ID-V2V Relight propagates the new illumination across the video.

Does
video, text → video-edit
Release date
2026-09
Price
see supplier

via fal commercial

H3 Max Camera Controls

H3 Max Multi Angle turns a single image into a video with precise, keyframe-based control over the camera's orbit, elevation, and distance in 3D space

Does
image, text → video
Release date
2026-09
Price
$0.025 / second

via fal commercial

Flux 3 FAST Edit Video

FLUX.3 Edit Video [FAST] is Black Forest Labs' frontier video model. This endpoint edits an existing video from natural-language instructions, applying targeted changes while preserving the rest of the scene.

Does
video, text → video-edit
Release date
2026-09
Price
$0.03 / second

via fal commercial

Black Forest Labs: FLUX Video Edit

FLUX Video Edit [fast] takes a source video and an edit prompt and returns a precisely edited video. Add, remove, or replace objects and characters, rebuild the setting, edit on-screen...

Does
text → video
Release date
2026-09
Price
$6 / s

AI tokens — built in commercial

OpenAI: GPT Image 2.5 Flare

GPT Image 2.5 Flare is an image generation and editing model from OpenAI, positioned as the speed-oriented tier of the GPT Image 2.5 series. It is suited to high-volume everyday...

Does
text, image → image, image-edit
Release date
2026-09
Price
≈ $0.077 / image

AI tokens — built in commercial

OpenAI: GPT Image 2.5 Sunburst

GPT Image 2.5 Sunburst is an image generation and editing model from OpenAI, positioned as the precision-oriented tier of the GPT Image 2.5 series. It is suited to detailed creative...

Does
text, image → image, image-edit
Release date
2026-09
Price
≈ $0.077 / image

AI tokens — built in commercial

Elevenlabs - Forced Alignment

Align the transcript and your audio recording using Elevenlab's forced alignment feature!

Does
audio → transcribe
Release date
2026-09
Price
$0.22 / hour

via fal commercial

GPT Image 2.5 Flare Edit

Precise image editing that changes only what's asked, keeping subject, composition, and background intact, with reference subjects staying recognizable across styles and successive edits.

Does
image, text → image-edit
Release date
2026-09
Price
$5.00

via fal commercial

GPT Image 2.5 Flare Text to Image

OpenAI's default image model for most applications. Fast, high-quality generation with natural lighting, rich textures, and support for complex layouts including transparent backgrounds.

Does
text → image
Release date
2026-09
Price
$5.00

via fal commercial

GPT Image 2.5 Sunburst Edit

Editing built for the tightest control, edits scoped precisely to the instruction, with subject and composition preserved across many rounds of revision.

Does
image, text → image-edit
Release date
2026-09
Price
$5.00

via fal commercial

GPT Image 2.5 Sunburst Text to Image

OpenAI's precision-focused image model, built for premium visual work, extra fidelity on intricate detail, in exchange for longer generation times.

Does
text → image
Release date
2026-09
Price
$5.00

via fal commercial

Microsoft AI: MAI-Image-2.6

MAI-Image-2.6 is an image generation and editing model from Microsoft AI, the precision tier of the MAI-Image-2.6 family alongside the faster [MAI-Image-2.6 Flash](/microsoft/mai-image-2.6-flash). It is suited for design-ready visuals and...

Does
text, image → image, image-edit
Release date
2026-09
Price
≈ $0.098 / image

AI tokens — built in commercial

Microsoft AI: MAI-Image-2.6 Flash

MAI-Image-2.6 Flash is the lower-latency, lower-cost member of the [MAI-Image-2.6](/microsoft/mai-image-2.6) family from Microsoft AI, built for latency-sensitive, high-throughput production image generation and editing at comparable quality to the precision tier....

Does
text, image → image, image-edit
Release date
2026-09
Price
≈ $0.049 / image

AI tokens — built in commercial

H3 Max Director

Direct continuous, realtime video streams with live prompts while preserving characters, settings, and story continuity.

Does
text → video
Release date
2026-09
Price
$0.04 / second

via fal commercial

H3 Max Turbo Image to Video

fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

Does
image, text → video
Release date
2026-09
Price
$0.0125 / second

via fal commercial

H3 Max Turbo Text to Video

fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

Does
text → video
Release date
2026-09
Price
$0.0125 / second

via fal commercial

MiniMax: H3 Max

MiniMax H3 Max is a video-generation model from MiniMax, jointly released with fal.ai. Derived through additional training from MiniMax H3, it is designed for faster text-to-video and image-to-video generation with...

Does
text, image → video
Release date
2026-09
Price
$0.1 / s (480p)

AI tokens — built in commercial

Kling Video 4K Video to Video Edit

Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling

Does
video, text → video-edit
Release date
2026-08
Price
$0.42

via fal commercial

Kling Video 4K Video to Video

Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling

Does
video, text → video-edit
Release date
2026-08
Price
$0.42

via fal commercial

H3 Max Reference to Video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

Does
image, text → video
Release date
2026-08
Price
$0.05 / second

via fal commercial

Video

Remove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.

Does
video, text → video-edit
Release date
2026-08
Price
see supplier

via fal commercial

Gemini Omni Flash 1.1 Edit

Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint edits video through natural-language instruction, applying the requested change while preserving the parts of the scene you want kept, and carrying character and scene consistency across successive edits.

Does
video, text → video-edit
Release date
2026-08
Price
$0.03 / second

via fal commercial

Gemini Omni Flash 1.1 Image to Video

Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint animates a still image into video with synchronized audio, extending a single frame into coherent motion that reflects the logic of the real world.

Does
image, text → video
Release date
2026-08
Price
$0.03 / second

via fal commercial

Gemini Omni Flash 1.1 Reference to Video

Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint generates video from combined multimodal references, images, videos and text together. Reasoning across all inputs to produce a single coherent result, with characters retaining their face, clothing, and voice throughout

Does
image, text → video
Release date
2026-08
Price
$0.03 / second

via fal commercial

Gemini Omni Flash 1.1 Text to Video

Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint generates video with synchronized native audio from a text prompt, grounded in Gemini's real-world knowledge and physics understanding, with cinematic camera control expressed in natural language.

Does
text → video
Release date
2026-08
Price
$0.03 / second

via fal commercial

Alibaba: Wan 3.0 Prime

Wan 3.0 Prime is a fast-mode variant of Wan 3.0 from Alibaba. It supports text-to-video and first-frame image-to-video generation.

Does
text, image → video
Release date
2026-08
Price
$0.136 / s (480p)

AI tokens — built in commercial

Fibo Edit 1.5 Image Editing

Commercially safe, multi-reference image editing model. Follows natural language instructions alone or with up to 4 reference images, purpose-built for complex object and character combinations, virtual try-on, background replacement, style transfer, and more.

Does
image, text → image-edit
Release date
2026-08
Price
see supplier

via fal commercial

Meta Muse Image Edit

Meta's Muse Image model does precise edits that change only what you ask, stay coherent across turns, and compose from multiple reference images.

Does
image, text → image-edit
Release date
2026-08
Price
see supplier

via fal commercial

Meta Muse Image Text to Image

Meta's Muse Image model has faithful instruction-following and exceptional visual fidelity, with fine details like text, plots, and QR codes rendered accurately.

Does
text → image
Release date
2026-08
Price
see supplier

via fal commercial

H3 Max Image to Video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

Does
image, text → video
Release date
2026-08
Price
$0.025 / second

via fal commercial

MiniMax H3 Max Text to Video

fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality

Does
text → video
Release date
2026-08
Price
$0.025 / second

via fal commercial

Recraft V4 Styles Pro Text to Image

Generates raster images that hold a consistent style, from either a saved style ID or reference images attached directly.

Does
text → image
Release date
2026-08
Price
see supplier

via fal commercial

Recraft V4 Styles Pro Text to Vector

Generates vector images that hold a consistent style, from either a saved style ID or reference images attached directly.

Does
text → image
Release date
2026-08
Price
see supplier

via fal commercial

Recraft V4 Styles Text to Image

Generates raster images that hold a consistent style, from either a saved style ID or reference images attached directly.

Does
text → image
Release date
2026-08
Price
see supplier

via fal commercial

Recraft V4 Styles Text to Vector

Generates vector images that hold a consistent style, from either a saved style ID or reference images attached directly.

Does
text → image
Release date
2026-08
Price
see supplier

via fal commercial

Meta: Muse Image

Muse Image is an agentic image generation model from Meta that generates and edits images from text and reference images. Unlike single-pass image models, it reasons before it renders, breaking...

Does
text, image → image, image-edit
Release date
2026-08
Price
see supplier

AI tokens — built in commercial

Recraft: Recraft V4 Styles

Recraft V4 Styles is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's rendering...

Does
text, image → image, image-edit
Release date
2026-08
Price
$0.07 / image

AI tokens — built in commercial

Recraft: Recraft V4 Styles Pro

Recraft V4 Styles Pro is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's...

Does
text, image → image, image-edit
Release date
2026-08
Price
$0.2 / image

AI tokens — built in commercial

Recraft: Recraft V4 Styles Pro Vector

Recraft V4 Styles Pro Vector is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the...

Does
text, image → image, image-edit
Release date
2026-08
Price
$0.24 / image

AI tokens — built in commercial

Recraft: Recraft V4 Styles Vector

Recraft V4 Styles Vector is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's...

Does
text, image → image, image-edit
Release date
2026-08
Price
$0.1 / image

AI tokens — built in commercial

Wan 3.0 Prime

Wan 3.0 Prime Image-to-Video turns still images into dynamic, cinematic sequences with rapid turnaround, natural motion, and excellent visual continuity. It preserves the identity, composition, and atmosphere of the source image while introducing expressive movement, camera dynamics, and richly deta

Does
image, text → video
Release date
2026-08
Price
$0.068

via fal commercial

Wan 3.0 Prime

Wan 3.0 Prime Reference-to-Video combines reference images, videos, and audio into a unified video with fast generation and strong multimodal coherence. It follows character identity, visual style, movement, and sound cues across references to create controlled, consistent, and production-ready resu

Does
video, text → video-edit
Release date
2026-08
Price
$0.068

via fal commercial

Wan 3.0 Prime

Wan 3.0 Prime Text-to-Video transforms written prompts into polished videos with accelerated generation, fluid motion, strong scene fidelity, and coherent visual storytelling. Built for fast creative iteration, it brings complex ideas to life while preserving visual detail and cinematic consistency

Does
text → video
Release date
2026-08
Price
$0.068

via fal commercial

Wan 3.0

Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

Does
image, text → video
Release date
2026-08
Price
$0.05

via fal commercial

Wan 3.0

Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

Does
image, text → video
Release date
2026-08
Price
$0.05

via fal commercial

Wan Text to Video

Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.

Does
text → video
Release date
2026-08
Price
$0.05

via fal commercial

Fibo Gen 1.5 Text to Image

Text-to-image model with high-fidelity outputs, accurate typography, and style preset, strong in photorealism, textures, and beyond. JSON-structured prompts give enterprise and agentic workflows production-ready control. Trained on licensed data.

Does
text → image
Release date
2026-08
Price
see supplier

via fal commercial

Google Virtual Try On

Generate realistic virtual try-on images from a person image and a clothing product image.

Does
image, text → image-edit
Release date
2026-08
Price
$0.075 / generated

via fal commercial

V7 Text to 3D

Turns text into a fully textured, PBR-ready 3D mesh with complete geometry, in game-ready Smart Topology at a target polygon count

Does
text → 3d
Release date
2026-08
Price
$0.12 / call

via fal commercial

Alibaba: Wan 3.0

Wan 3.0 is a video generation model from Alibaba for text-to-video, image-to-video, and reference-guided video generation. It produces 480p, 720p, or 1080p video with durations from 2 to 30 seconds.

Does
text, image → video
Release date
2026-08
Price
$0.1 / s (480p)

AI tokens — built in commercial

HeyGen: Avatar IV

HeyGen: Avatar IV is an image-to-video model that animates a single photo into an expressive, lip-synced talking-head video. Rather than only matching mouth shapes to words, it interprets the vocal...

Does
text, image → video
Release date
2026-08
Price
$0.1 / s

AI tokens — built in commercial

Flux Video Upscale

Upscale videos to 1080p, 2K, or 4K via API. FLUX 3 powered super-resolution with a precise mode and a creative detail-enhancement mode.

Does
video, text → video-edit
Release date
2026-08
Price
$0.14 / second

via fal commercial

Hi3D Image to 3D

Generate 3D models from a single image with Hi3D V3.0.

Does
image → 3d
Release date
2026-08
Price
$0.02 / credit

via fal commercial

Hi3D Multiview to 3D

Generate 3D models from multiple view images using Hi3D V3.0.

Does
image → 3d
Release date
2026-08
Price
$0.02 / credit

via fal commercial

Black Forest Labs: FLUX Video Upscale

FLUX Video Upscale is a video upscaling model from Black Forest Labs. It enlarges a single source video by 1.5× to 3× while preserving its duration, with an optional prompt...

Does
text → video
Release date
2026-08
Price
$15 / s

AI tokens — built in commercial

Topaz Adjust Image

Professional color and lighting correction powered by Topaz Labs. Adjust V2 fixes exposure, White Balance corrects color casts, Colorize adds color to black-and-white photos. Best for one-click photo correction.

Does
image, text → image-edit
Release date
2026-08
Price
$0.08 / started

via fal commercial

Topaz Colorize Video

Professional video colorization powered by Topaz Labs. Brings natural color to black-and-white footage, upscaled to at least 1080p. Best for archival and historical clips.

Does
video, text → video-edit
Release date
2026-08
Price
$0.10

via fal commercial

Topaz Deblur Video

Professional motion deblur powered by Topaz Labs. Themis 2 restores clarity to fast-moving, motion-blurred footage at source resolution. Best for sports and action footage.

Does
video, text → video-edit
Release date
2026-08
Price
$0.10

via fal commercial

Topaz Denoise Image

Professional photo denoising powered by Topaz Labs. Normal, Strong and Extreme presets clean noise at source resolution; Denoise Max adds generative detail recovery. Best for high-ISO and night photography.

Does
image, text → image-edit
Release date
2026-08
Price
$0.08

via fal commercial

Topaz Denoise Video

Professional video denoising powered by Topaz Labs. Nyx models remove noise at source resolution, with Nyx Fast as a lighter, cheaper pass. Best for low-light and high-ISO footage.

Does
video, text → video-edit
Release date
2026-08
Price
$0.10

via fal commercial

Topaz Interpolate Video

Professional frame interpolation powered by Topaz Labs. Apollo, Chronos and Aion retime footage up to 120 fps, from smooth motion to extreme slow motion. Best for fluid 60fps output and slow-motion effects.

Does
video, text → video-edit
Release date
2026-08
Price
$0.30

via fal commercial

Topaz Restore Image

Professional image restoration powered by Topaz Labs. Recover 3 generatively rebuilds natural detail; Dust-Scratch V2 cleans film dust and scratches. Best for old, damaged or degraded photos.

Does
image, text → image-edit
Release date
2026-08
Price
$0.08 / started

via fal commercial

Topaz Sdr To Hdr Video

Professional SDR-to-HDR conversion powered by Topaz Labs. Hyperion 2.5 redistributes luminance and color while preserving detail in text, faces and motion. Best for giving flat SDR footage a true HDR look.

Does
video, text → video-edit
Release date
2026-08
Price
$2.40

via fal commercial

Topaz Sharpen Image

Professional photo sharpening powered by Topaz Labs. Models tuned per blur type (lens, motion, portrait, wildlife), plus Super Focus for generative recovery of severely blurred shots. Best for out-of-focus and motion-blurred photos.

Does
image, text → image-edit
Release date
2026-08
Price
$0.08

via fal commercial

MiniMax Music 3

MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long

Does
text → music
Release date
2026-08
Price
see supplier

via fal commercial

ByteDance Seed: Seedream 5.0 Lite

Seedream 5.0 Lite is an image generation model from ByteDance Seed. It is suited for professional visual creation that benefits from web-connected retrieval, complex-prompt comprehension, visual references, and broad knowledge...

Does
text, image → image, image-edit
Release date
2026-08
Price
$0.07 / image

AI tokens — built in commercial

ByteDance Seed: Seedream 5.0 Pro

Seedream 5.0 Pro is an image generation and editing model from ByteDance Seed. It is suited for commercial visual-production workflows that require precise editing control, lifelike scenes, and natural rendering.

Does
text, image → image, image-edit
Release date
2026-08
Price
$0.09 / image

AI tokens — built in commercial

ByteDance: Seedance 2.0 Mini

Seedance 2.0 Mini is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video with image, video, and audio inputs. It...

Does
text, image → video
Release date
2026-08
Price
$0 / s

AI tokens — built in commercial

V7 Image to 3D

Turns a single image into a fully textured, PBR-ready 3D mesh with complete geometry, in game-ready Smart Topology at a target polygon count

Does
image → 3d
Release date
2026-08
Price
$0.12 / call

via fal commercial

V7 Multi Image to 3D

econstructs a high-fidelity textured 3D model from multiple angle views of one object, with game-ready topology and polygon control

Does
image → 3d
Release date
2026-08
Price
$0.12 / call

via fal commercial

Grok Imagine Image 2.0

Generate images from text using xAi's Grok Imagine 2.0 model.

Does
text → image
Release date
2026-08
Price
$0.04

via fal commercial

xAI: Grok Imagine Image 2.0

Grok Imagine Image 2.0 is an image generation and editing model from xAI. It is suited for creating images from text prompts and editing images from references, with low and...

Does
text, image → image, image-edit
Release date
2026-08
Price
$0.08 / image

AI tokens — built in commercial

ByteDance: Seedance 2.5

Seedance 2.5 is a video generation model from ByteDance. It is suited for long-form storytelling, multimodal reference-based generation, video editing, and video extension. It supports first-frame and first-and-last-frame control, up...

Does
text, image → video
Release date
2026-08
Price
$0 / s

AI tokens — built in commercial

Hi3D Image to 3D

Generate 3D models from a single image with Hi3D.

Does
image → 3d
Release date
2026-08
Price
$0.02 / credit

via fal commercial

Hi3D Multiview to 3D

Generate 3D models from multiple view images using Hi3D.

Does
image → 3d
Release date
2026-08
Price
$0.02 / credit

via fal commercial

Qwen: Qwen Image 3

Qwen Image 3 is a unified image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with a richer world...

Does
text, image → image, image-edit
Release date
2026-08
Price
$0.06 / image

AI tokens — built in commercial

Qwen: Qwen Image 3 Pro

Qwen Image 3 Pro is an image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with richer world knowledge...

Does
text, image → image, image-edit
Release date
2026-08
Price
$0.08 / image

AI tokens — built in commercial

Black Forest Labs: FLUX.3 Video

FLUX.3 Video is a video generation model from Black Forest Labs. It supports text-to-video, image-guided generation with opening and closing keyframes, and video continuation workflows, making it suited for controlled...

Does
text, image → video
Release date
2026-08
Price
$34 / s

AI tokens — built in commercial

MAI Image 2.5 Pro (Text to Image)

Generate high-fidelity, design-ready images with precise typography, strong prompt alignment, and rich visual detail using Microsoft's flagship MAI Image 2.5 Pro.

Does
text → image
Release date
2026-07
Price
$7.50 / 1m

via fal commercial

Qwen Audio 3.0 TTS (Flash)

Generate natural multilingual speech from text with fast voice and language control using Qwen Audio 3.0 TTS Flash.

Does
text → speech
Release date
2026-07
Price
see supplier

via fal commercial

MiniMax: H3

MiniMax H3 is a lightweight, open-weights video generation model from MiniMax. It is designed for precise multimodal editing and controlled content generation, including instruction-guided edits, text and brand rendering, and...

Does
text, image → video
Release date
2026-07
Price
$0.08 / s

AI tokens — built in commercial

Runway: Aleph 2.0

Runway Aleph 2.0 is an in-context video editing model from Runway. It applies text instructions and keyframe-guided edits across existing footage while preserving details that are not meant to change....

Does
text → video
Release date
2026-07
Price
$56 / s

AI tokens — built in commercial

Runway: Gen-4.5

Runway Gen-4.5 is a video generation model from Runway for text-to-video and image-to-video workflows. It is designed for cinematic scene creation with strong motion quality, visual fidelity, and prompt adherence....

Does
text, image → video
Release date
2026-07
Price
$24 / s

AI tokens — built in commercial

Microsoft AI: MAI-Image-2.5 Pro

Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

Does
text, image → image, image-edit
Release date
2026-07
Price
≈ $0.279 / image

AI tokens — built in commercial

Qwen Image 3 Text to Image

Generates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen's strength in complex text rendering and precise prompt adherence

Does
text → image
Release date
2026-07
Price
$0.04 / generated

via fal commercial

V1.1 Text to Sound Effects

Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.

Does
text → music
Release date
2026-07
Price
$0.0018 / second

via fal commercial

Krea: Krea 2 Large

Krea 2 Large is Krea's high-capability image generation model, more than twice the size of Krea 2 Medium. Its lighter post-training gives images a rawer, more textured, and flexible character,...

Does
text, image → image, image-edit
Release date
2026-07
Price
see supplier

AI tokens — built in commercial

Krea: Krea 2 Medium

Krea 2 Medium is Krea's balanced, cost-efficient image generation model and a practical starting point for a broad range of use cases. Its extensive post-training supports stable, consistent generations, with...

Does
text, image → image, image-edit
Release date
2026-07
Price
see supplier

AI tokens — built in commercial

Krea: Krea 2 Medium Turbo

Krea 2 Medium Turbo is a distilled, speed-focused variant of Krea 2 Medium from Krea. It is designed for rapid iteration and graphic design exploration where fast generation is the...

Does
text, image → image, image-edit
Release date
2026-07
Price
see supplier

AI tokens — built in commercial

SpaceXAI: Grok Imagine Video 1.5

Grok Imagine Video 1.5 is a video generation model from SpaceXAI. It creates videos from text prompts, with an optional starting image to guide the scene. It can direct subject...

Does
text, image → video
Release date
2026-07
Price
$2 / s

AI tokens — built in commercial

Krea 2 Text to Image Turbo Style

Generate high-fidelity images from text with Krea 2 using a style reference image. Apply a reference image to guide the visual style into new generations, with aspect ratio, creativity, and seed controls.

Does
text → image
Release date
2026-07
Price
see supplier

via fal commercial

TRELLIS.2 LoRA Inference

Run inference on LoRA adapters for TRELLIS.2 model

Does
image → 3d
Release date
2026-07
Price
see supplier

via fal commercial

Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

Does
text, image → image, image-edit
Release date
2026-06
Price
≈ $0.077 / image

AI tokens — built in commercial

Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

Does
image, text → image
Release date
2026-06
Price
$3 / 1M out

language-model key commercial

Seed Audio 1.0

Seed Audio 1.0 is a new audio model from Bytedance that can generate high-quality, natural sounding audio using text, reference audios or an image.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Alibaba: HappyHorse 1.0

HappyHorse 1.0 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up...

Does
text, image → video
Release date
2026-06
Price
$0.198 / s (720p)

AI tokens — built in commercial

Alibaba: HappyHorse 1.1

HappyHorse 1.1 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up...

Does
text, image → video
Release date
2026-06
Price
$0.198 / s (720p)

AI tokens — built in commercial

OpenAI: GPT Image 1

OpenAI's GPT Image 1 generates and edits images via the dedicated Images API. Features accurate text rendering, transparent backgrounds, and up to 16 reference images for edits.

Does
text, image → image, image-edit
Release date
2026-06
Price
≈ $0.103 / image

AI tokens — built in commercial

OpenAI: GPT Image 1 Mini

A cost-efficient variant of GPT Image 1 for high-quality image generation at reduced latency and cost via OpenAI's dedicated Images API.

Does
text, image → image, image-edit
Release date
2026-06
Price
≈ $0.021 / image

AI tokens — built in commercial

OpenAI: GPT Image 2

OpenAI's latest image generation model. Supports high-fidelity image generation and editing via the dedicated Images API.

Does
text, image → image, image-edit
Release date
2026-06
Price
≈ $0.077 / image

AI tokens — built in commercial

Hyper3D - Rodin V2.5 - Image to 3D - Fast

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.

Does
image → 3d
Release date
2026-06
Price
$0.1 / generation

via fal commercial

Hyper3D - Rodin V2.5 - Text to 3D - Fast

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.

Does
text → 3d
Release date
2026-06
Price
$0.1 / generation

via fal commercial

Async Text to Speech Pro V1.0

Generate professional-quality voiceovers in seconds with Async TTS Pro model text-based control over pauses, emphasis, and timing. Voice ids can be found at https://async.com/developer/voice-library

Does
text → speech
Release date
2026-06
Price
$0.01 / 1

via fal commercial

Google: Nano Banana Pro (Gemini 3 Pro Image)

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

Does
image, text → image
Release date
2026-06
Price
$24 / 1M out

language-model key commercial

Google: Nano Banana 2 (Gemini 3.1 Flash Image)

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced...

Does
image, text → image
Release date
2026-06
Price
$6 / 1M out

language-model key commercial

Ltx 2.3 Quality

Text to Audio high-quality using LTX-2.3

Does
text → music
Release date
2026-06
Price
$0.0024075 / megapixel

via fal commercial

Ltx 2.3 Quality

Text to Audio high-quality using LTX-2.3 with Lora

Does
text → music
Release date
2026-06
Price
$0.0024075 / megapixel

via fal commercial

Zonos2 Text to Speech

Zonos2 is a text-to-speech model that clones a voice from a short sample and speaks naturally across many languages.

Does
text → speech
Release date
2026-06
Price
see supplier

via fal commercial

Nemotron Asr Multilingual

Nemotron-ASR-Streaming is a multi lingual, streaming Automatic Speech Recognition (ASR) engineered to deliver high-quality multi lingual transcription across both low-latency streaming and high-throughput batch workloads.

Does
audio → transcribe
Release date
2026-06
Price
see supplier

via fal commercial

Bytedance Seed Speech Text to Speech

Seed Speech developed by ByteDance, is a family of large-scale text-to-speech models capable of synthesizing speech that is virtually indistinguishable from human speech.

Does
text → speech
Release date
2026-06
Price
see supplier

via fal commercial

Sourceful: Riverflow V2.5 Fast

Riverflow V2.5 Fast is the speed-optimized variant of Sourceful's Riverflow 2.5 lineup, best for production deployments and latency-critical workflows. The Riverflow 2.5 series is a unified text-to-image and image-to-image family...

Does
text, image → image, image-edit
Release date
2026-06
Price
$0.038 / image

AI tokens — built in commercial

Sourceful: Riverflow V2.5 Pro

Riverflow V2.5 Pro is the most powerful variant of Sourceful's Riverflow 2.5 lineup, best for top-tier control and quality-sensitive outputs. The Riverflow 2.5 series is a unified text-to-image and image-to-image...

Does
text, image → image, image-edit
Release date
2026-06
Price
$0.26 / image

AI tokens — built in commercial

Stable Audio 3 Medium Audio Inpainting

Stable Audio 3 Medium audio inpainting is a 1.4 billion parameter latent diffusion model that fills in or reworks selected segments of a stereo track guided by text prompts, supporting single- and multi-segment editing.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Audio Outpainting

Stable Audio 3 Medium audio outpainting is a 1.4 billion parameter latent diffusion model that extends existing stereo audio beyond its original endpoint via causal continuation guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Audio to Audio

Stable Audio 3 Medium audio-to-audio is a 1.4 billion parameter latent diffusion model that transforms an input audio clip into new stereo variations up to 6 minutes guided by a text prompt.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Base Audio Inpainting

Stable Audio 3 Medium Base audio inpainting is the foundational 1.4 billion parameter checkpoint for editing or filling selected stereo audio segments guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Base Audio Outpainting

Stable Audio 3 Medium Base audio outpainting is the foundational 1.4 billion parameter checkpoint that extends existing stereo audio with causal continuation guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Base Audio to Audio

Stable Audio 3 Medium Base audio-to-audio is the foundational 1.4 billion parameter checkpoint that transforms input audio into new stereo variations up to 6 minutes guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Medium Base Text to Audio

Stable Audio 3 Medium Base is the foundational 1.4 billion parameter text-to-audio checkpoint generating stereo music up to 6 minutes, intended as the unmodified base for custom fine-tuning workflows.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3

Stable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for safe commercial use.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small Music Audio Outpainting

Stable Audio 3 Small Music audio outpainting is a 459 million parameter latent diffusion model that extends music compositions beyond their original endpoint via causal continuation.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small Music Base Audio Inpainting

Stable Audio 3 Small Music Base audio inpainting is the foundational 459 million parameter checkpoint for editing or filling selected music segments guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small Music Base Audio Outpainting

Stable Audio 3 Small Music Base audio outpainting is the foundational 459 million parameter checkpoint that extends music tracks via causal continuation guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small Music Base Audio to Audio

Stable Audio 3 Small Music Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input music into new variations up to 2 minutes guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3

Stable Audio 3 Small Music Base is the foundational 459 million parameter checkpoint generating full music compositions up to 2 minutes from text prompts, intended as the unmodified base for fine-tuning.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small Music Text to Audio

Stable Audio 3 Small Music is a 459 million parameter latent diffusion model that generates full stereo music compositions up to 2 minutes from text prompts, lightweight enough for on-device deployment.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Audio Outpainting

Stable Audio 3 Small SFX audio outpainting is a 459 million parameter latent diffusion model that extends sound-effect tracks beyond their original endpoint via causal continuation.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Base Audio Inpainting

Stable Audio 3 Small SFX Base audio inpainting is the foundational 459 million parameter checkpoint for editing or filling selected sound-effect segments guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Base Audio Outpainting

Stable Audio 3 Small SFX Base audio outpainting is the foundational 459 million parameter checkpoint that extends sound-effect tracks via causal continuation guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Base Audio to Audio

Stable Audio 3 Small SFX Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input audio into new sound-effect variations guided by text prompts.

Does
audio → audio-edit
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Base Text to Audio

Stable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as the unmodified base for fine-tuning.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Stable Audio 3 Small SFX Text to Audio

Stable Audio 3 Small SFX is a 459 million parameter latent diffusion model that generates high-quality sound effects from text prompts, designed for on-device deployment on mobile phones and consumer laptops.

Does
text → music
Release date
2026-06
Price
see supplier

via fal commercial

Triposplat

TripoSplat is an open-source model from TripoAI / VAST AI Research that converts a single 2D image into high-quality 3D Gaussians using a novel learned density-control approach

Does
image → 3d
Release date
2026-06
Price
see supplier

via fal commercial

Microsoft AI: MAI-Image-2.5

Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.

Does
text, image → image, image-edit
Release date
2026-06
Price
≈ $0.121 / image

AI tokens — built in commercial

Hyper3D - Rodin V2.5 - Image to 3D

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.

Does
image → 3d
Release date
2026-05
Price
$0.4 / generation

via fal commercial

Hyper3D - Rodin V2.5 - Text to 3D

Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.

Does
text → 3d
Release date
2026-05
Price
$0.4 / generation

via fal commercial

SpaceXAI: Grok Imagine Image Quality

Grok Imagine Image Quality is SpaceXAI's fast, high-fidelity image generation and editing model. It accepts text prompts and optional reference images, producing photorealistic outputs at 1K or 2K across a...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.1 / image

AI tokens — built in commercial

SpaceXAI: Grok Imagine Video

Grok Imagine Video is SpaceXAI's fast, text-, image-, and reference-conditioned video generation model. It produces short videos (1–15 seconds, 24 fps) at 480p or 720p across seven aspect ratios -...

Does
text, image → video
Release date
2026-05
Price
$0.4 / s

AI tokens — built in commercial

Pixal3d

Pixal3D turns a single image into a high-fidelity 3D model with detailed geometry and realistic textures.

Does
image → 3d
Release date
2026-05
Price
$0.3

via fal commercial

Recraft: Recraft V4 Pro Vector

Recraft V4 Pro Vector is the vector (SVG) variant of Recraft V4 Pro. It supports text and image inputs and produces vector image output across multiple aspect ratios at the...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.6 / image

AI tokens — built in commercial

Recraft: Recraft V4 Vector

Recraft V4 Vector is the vector (SVG) variant of Recraft V4. It supports text and image inputs and produces vector image output across multiple aspect ratios. Compared to the raster...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.16 / image

AI tokens — built in commercial

Recraft: Recraft V4.1

Recraft V4.1 is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios, with...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.07 / image

AI tokens — built in commercial

Recraft: Recraft V4.1 Pro

Recraft V4.1 Pro is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.42 / image

AI tokens — built in commercial

Recraft: Recraft V4.1 Pro Vector

Recraft V4.1 Pro Vector is the vector (SVG) variant of Recraft V4.1 Pro, tuned for high aesthetics. It supports text and image inputs and produces higher-resolution SVG image output across...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.6 / image

AI tokens — built in commercial

Recraft: Recraft V4.1 Utility

Recraft V4.1 Utility is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios, with typical generation...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.07 / image

AI tokens — built in commercial

Recraft: Recraft V4.1 Utility Pro

Recraft V4.1 Utility Pro is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios — double...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.42 / image

AI tokens — built in commercial

Recraft: Recraft V4.1 Vector

Recraft V4.1 Vector is the vector (SVG) variant of Recraft V4.1, tuned for high aesthetics. It supports text and image inputs and produces SVG image output across multiple aspect ratios,...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.16 / image

AI tokens — built in commercial

Recraft: Recraft V3

Recraft V3 is an image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios. Supports the following `image_config` parameters:...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.08 / image

AI tokens — built in commercial

Recraft: Recraft V4

Recraft V4 is an image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios. It delivers stronger compositional judgment,...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.08 / image

AI tokens — built in commercial

Recraft: Recraft V4 Pro

Recraft V4 Pro is an image generation model from Recraft. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios, double the resolution of...

Does
text, image → image, image-edit
Release date
2026-05
Price
$0.5 / image

AI tokens — built in commercial

Kling: Video v3.0 Pro

Kling v3.0 Pro is Kuaishou's premium video generation model, offering higher visual quality than the Standard tier. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for precise...

Does
text, image → video
Release date
2026-04
Price
$0.224 / s

AI tokens — built in commercial

Kling: Video v3.0 Standard

Kling v3.0 Standard is a video generation model from Kuaishou. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for guided scene composition. Clips range from 3 to...

Does
text, image → video
Release date
2026-04
Price
$0.168 / s

AI tokens — built in commercial

Google: Veo 3.1 Fast

Google's mid-tier video generation model balancing speed and quality. Veo 3.1 Fast generates high-quality video from text or image prompts with native synchronized audio, offering faster turnaround than Veo 3.1...

Does
text, image → video
Release date
2026-04
Price
$0.16 / s (720p)

AI tokens — built in commercial

Google: Veo 3.1 Lite

Google's most cost-effective video generation model, designed for high-volume applications and rapid iteration. Veo 3.1 Lite generates 720p and 1080p video from text or image prompts with native synchronized audio...

Does
text, image → video
Release date
2026-04
Price
$0.06 / s (720p)

AI tokens — built in commercial

Cohere Transcribe

Cohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation

Does
audio → transcribe
Release date
2026-04
Price
see supplier

via fal commercial

OpenAI: GPT-5.4 Image 2

GPT-5.4 Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...

Does
image, text, file → image
Release date
2026-04
Price
$30 / 1M out

language-model key commercial

Kling: Video O1

Kling Video O1 is a video generation model from Kuaishou. It supports text and image inputs with video output, enabling text-to-video and image-to-video workflows. It is suited for cinematic content...

Does
text, image → video
Release date
2026-04
Price
$0.224 / s

AI tokens — built in commercial

MiniMax: Hailuo 2.3

Hailuo 2.3 is a video generation model from MiniMax. It accepts text prompts and reference images as input and generates video output, supporting both text-to-video and image-to-video workflows. It is...

Does
text, image → video
Release date
2026-04
Price
$0.163 / s

AI tokens — built in commercial

Gemini 3.1 Flash Tts

Newest audio model from Google introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.

Does
text → speech
Release date
2026-04
Price
see supplier

via fal commercial

Tripo H3.1 Text to 3D

Generate 3D models from text descriptions using Tripo H3.1.

Does
text → 3d
Release date
2026-04
Price
$0.10

via fal commercial

Tripo P1 Text to 3D

Generate 3D models from text descriptions using Tripo P1.

Does
text → 3d
Release date
2026-04
Price
see supplier

via fal commercial

Alibaba: Wan 2.7

Wan 2.7 is a video generation model from Alibaba. It supports text-to-video, image-to-video with first and last frame control, and reference-to-video, where multiple reference images guide the style and content...

Does
text, image → video
Release date
2026-04
Price
$0.2 / s

AI tokens — built in commercial

ByteDance: Seedance 2.0

Seedance 2.0 is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It is particularly strong at preserving character consistency,...

Does
text, image → video
Release date
2026-04
Price
$0 / s

AI tokens — built in commercial

ByteDance: Seedance 2.0 Fast

Seedance 2.0 Fast is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It prioritizes generation speed and lower cost...

Does
text, image → video
Release date
2026-04
Price
$0 / s

AI tokens — built in commercial

Google: Lyria 3 Clip Preview

30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...

Does
text, image → speech
Release date
2026-03
Price
free

language-model key free

Google: Lyria 3 Pro Preview

Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...

Does
text, image → speech
Release date
2026-03
Price
free

language-model key free

Alibaba: Wan 2.6

Alibaba's most advanced video generation model, supporting over 10 visual creation capabilities in a unified system. Wan 2.6 generates 1080p video at 24fps from text, images, reference videos, or audio,...

Does
text, image → video
Release date
2026-03
Price
$0.08 / s (480p)

AI tokens — built in commercial

ByteDance: Seedance 1.5 Pro

ByteDance's next-generation audio-visual generation model with a 4.5B parameter Dual-Branch Diffusion Transformer architecture. Seedance 1.5 Pro generates video and audio simultaneously in a single unified pass — eliminating the timing...

Does
text, image → video
Release date
2026-03
Price
$0 / s

AI tokens — built in commercial

Google: Veo 3.1

Google's state-of-the-art video generation model, built for maximum visual fidelity in final production cuts. Veo 3.1 generates high-quality 1080p video from text or image prompts with native synchronized audio —...

Does
text, image → video
Release date
2026-03
Price
$0.4 / s

AI tokens — built in commercial

OpenAI: Sora 2 Pro

OpenAI's flagship video generation model, delivering production-quality video with physics-accurate motion, synchronized audio, and world-state persistence across shots. Sora 2 Pro follows intricate multi-shot instructions while maintaining consistent spatial relationships...

Does
text → video
Release date
2026-03
Price
$0.6 / s (720p)

AI tokens — built in commercial

xAI Text to Speech

Generate speech with expressive and realistic voices from xAI

Does
text → speech
Release date
2026-03
Price
$0.015 / 1000

via fal commercial

Inworld TTS-1.5 Max

Text to Speech Endpoint for Inworld's TTS-1.5 Max.

Does
text → speech
Release date
2026-03
Price
$0.01 / 1000

via fal commercial

Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines...

Does
text, image → image, image-edit
Release date
2026-02
Price
≈ $0.155 / image

AI tokens — built in commercial

Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines...

Does
image, text → image
Release date
2026-02
Price
$6 / 1M out

language-model key commercial

Meshy 6

Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.

Does
text → 3d
Release date
2026-02
Price
see supplier

via fal commercial

MiniMax Speech 2.8 [HD]

Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.

Does
text → speech
Release date
2026-02
Price
see supplier

via fal commercial

MiniMax Speech 2.8 [Turbo]

Generate speech from text prompts and different voices using the MiniMax Speech-2.8 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.

Does
text → speech
Release date
2026-02
Price
see supplier

via fal commercial

Sourceful: Riverflow V2 Fast

Riverflow V2 Fast is the fastest variant of Sourceful's Riverflow 2.0 lineup, best for production deployments and latency-critical workflows. The Riverflow 2.0 series represents SOTA performance on image generation and...

Does
text, image → image, image-edit
Release date
2026-02
Price
$0.04 / image

AI tokens — built in commercial

Sourceful: Riverflow V2 Pro

Riverflow V2 Pro is the most powerful variant of Sourceful's Riverflow 2.0 lineup, best for top-tier control and perfect text rendering. The Riverflow 2.0 series represents SOTA performance on image...

Does
text, image → image, image-edit
Release date
2026-02
Price
$0.3 / image

AI tokens — built in commercial

Hunyuan 3d

Create detailed, fully-textured 3D models with text

Does
text → 3d
Release date
2026-01
Price
$0.225 / generation

via fal commercial

Hunyuan 3D Pro Text to 3D

Generate 3D models from text prompts with Hunyuan 3D Pro

Does
text → 3d
Release date
2026-01
Price
$0.375 / generation

via fal commercial

Qwen 3 TTS - Text to Speech [0.6B]

Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model

Does
text → speech
Release date
2026-01
Price
see supplier

via fal commercial

Qwen 3 TTS - Text to Speech [1.7B]

Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model

Does
text → speech
Release date
2026-01
Price
see supplier

via fal commercial

Qwen 3 TTS - Voice Design [1.7B]

Create custom voices using Qwen3-TTS Voice Design model and later use Clone Voice model to create your own voices!

Does
text → speech
Release date
2026-01
Price
see supplier

via fal commercial

OpenAI: GPT Audio

The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...

Does
text, audio → speech
Release date
2026-01
Price
$20 / 1M out

language-model key commercial

OpenAI: GPT Audio Mini

A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...

Does
text, audio → speech
Release date
2026-01
Price
$4.8 / 1M out

language-model key commercial

ElevenLabs Speech to Text - Scribe V2

Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!

Does
audio → transcribe
Release date
2026-01
Price
$0.008 / input

via fal commercial

Black Forest Labs: FLUX.2 Klein 4B

FLUX.2 [klein] 4B is the fastest and most cost-effective model in the FLUX.2 family, optimized for high-throughput use cases while maintaining excellent image quality. Pricing is based on the output...

Does
text, image → image, image-edit
Release date
2026-01
Price
$0.028 / image

AI tokens — built in commercial

Hunyuan Motion [1B]

Generate 3D human motions via text-to-generation interface of Hunyuan Motion!

Does
text → 3d
Release date
2025-12
Price
see supplier

via fal commercial

Hunyuan Motion [0.46B]

Generate 3D human motions via text-to-generation interface of Hunyuan Motion!

Does
text → 3d
Release date
2025-12
Price
see supplier

via fal commercial

ByteDance Seed: Seedream 4.5

Seedream 4.5 is the latest in-house image generation model developed by ByteDance. Compared with Seedream 4.0, it delivers comprehensive improvements, especially in editing consistency, including better preservation of subject details,...

Does
text, image → image, image-edit
Release date
2025-12
Price
$0.08 / image

AI tokens — built in commercial

Vibevoice

Generate long speech snippets fast using Microsoft's powerful TTS.

Does
text → speech
Release date
2025-12
Price
see supplier

via fal commercial

Hunyuan3d V3

Turn simple sketches into detailed, fully-textured 3D models. Instantly convert your concept designs into formats ready for Unity, Unreal, and Blender.

Does
text → 3d
Release date
2025-12
Price
$0.375 / generation

via fal commercial

Black Forest Labs: FLUX.2 Max

FLUX.2 [max] is the new top-tier image model from Black Forest Labs, pushing image quality, prompt understanding, and editing consistency to the highest level yet. Pricing is as follows, [per...

Does
text, image → image, image-edit
Release date
2025-12
Price
$0.14 / image

AI tokens — built in commercial

Maya

Maya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise voice design.

Does
text → speech
Release date
2025-12
Price
see supplier

via fal commercial

Black Forest Labs: FLUX.2 Flex

FLUX.2 [flex] excels at rendering complex text, typography, and fine details, and supports multi-reference editing in the same unified architecture. Pricing is as follows, [per the docs](https://bfl.ai/pricing?category=flux.2): We charge $0.06...

Does
text, image → image, image-edit
Release date
2025-11
Price
$0.12 / image

AI tokens — built in commercial

Black Forest Labs: FLUX.2 Pro

A high-end image generation and editing model focused on frontier-level visual quality and reliability. It delivers strong prompt adherence, stable lighting, sharp textures, and consistent character/style reproduction across multi-reference inputs....

Does
text, image → image, image-edit
Release date
2025-11
Price
$0.06 / image

AI tokens — built in commercial

Google: Nano Banana Pro (Gemini 3 Pro Image Preview)

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

Does
text, image → image, image-edit
Release date
2025-11
Price
≈ $0.155 / image

AI tokens — built in commercial

Google: Nano Banana Pro (Gemini 3 Pro Image Preview)

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

Does
image, text → image
Release date
2025-11
Price
$24 / 1M out

language-model key commercial

OpenAI: GPT-5 Image Mini

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by GPT-5 Mini, with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...

Does
text, image → image, image-edit
Release date
2025-10
Price
≈ $0.021 / image

AI tokens — built in commercial

OpenAI: GPT-5 Image Mini

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by GPT-5 Mini, with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...

Does
file, image, text → image
Release date
2025-10
Price
$4 / 1M out

language-model key commercial

OpenAI: GPT-5 Image

GPT-5 Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...

Does
text, image → image, image-edit
Release date
2025-10
Price
≈ $0.103 / image

AI tokens — built in commercial

OpenAI: GPT-5 Image

GPT-5 Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...

Does
image, text, file → image
Release date
2025-10
Price
$20 / 1M out

language-model key commercial

Google: Nano Banana (Gemini 2.5 Flash Image)

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,...

Does
text, image → image, image-edit
Release date
2025-10
Price
≈ $0.139 / image

AI tokens — built in commercial

Google: Nano Banana (Gemini 2.5 Flash Image)

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,...

Does
image, text → image
Release date
2025-10
Price
$5 / 1M out

language-model key commercial

Meshy 6 Preview

Meshy-6-Preview is the latest model from Meshy. It generates realistic and production ready 3D models.

Does
text → 3d
Release date
2025-10
Price
see supplier

via fal commercial

Pipecat's Smart Turn model

An open source, community-driven and native audio turn detection model by Pipecat AI.

Does
audio → transcribe
Release date
2025-04
Price
see supplier

via fal commercial

Speech-to-Text

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

Does
audio → transcribe
Release date
2025-04
Price
see supplier

via fal commercial

Speech-To-text

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

Does
audio → transcribe
Release date
2025-04
Price
see supplier

via fal commercial

Speech-to-Text

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

Does
audio → transcribe
Release date
2025-04
Price
see supplier

via fal commercial

Speech-to-Text

Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.

Does
audio → transcribe
Release date
2025-04
Price
see supplier

via fal commercial

ElevenLabs Speech to Text

Generate text from speech using ElevenLabs advanced speech-to-text model.

Does
audio → transcribe
Release date
2025-02
Price
see supplier

via fal commercial

Wizper (Whisper v3 -- fal.ai edition)

[Experimental] Whisper v3 Large -- but optimized by our inference wizards. Same WER, double the performance!

Does
audio → transcribe
Release date
2024-04
Price
see supplier

via fal commercial

Happy Horse 1.1

Alibaba's #1-ranked video model — 1080p output with synchronized native audio and multilingual lip-sync.

Does
text, image → video
Price
see supplier

cloud API commercial

Hunyuan 3D v3.1 Pro

Tencent's 3D generation model — turn a prompt or a single image into a textured 3D mesh (GLB).

Does
text, image → 3d
Price
see supplier

cloud API commercial

Z-Image Turbo (ComfyUI)

Tongyi's fast 6B model — eight steps to a photoreal or illustrated image, and it renders text in English and Chinese well. The same model as Z-Image Turbo above, on ComfyUI so it runs on a Linux GPU server as well as a Mac. Takes LoRAs. Apache 2.0.

Does
text → image
Price
free

on your Mac free Apache 2.0

Z-Image

The full, undistilled Z-Image — slower than Turbo, with real guidance: a negative prompt that works, more variety between seeds, and wider range of styles. It's the base that LoRAs are trained on. Apache 2.0. Runs on the managed ComfyUI.

Does
text → image
Price
free

on your Mac free Apache 2.0

Chroma1-HD

An open 8.9B FLUX-family model from the community (lodestones), trained on an unfiltered dataset — it has no built-in content refusals, so what it makes is up to you and your prompt. Strong photography, illustration and anime, a real negative prompt, and LoRAs. Apache 2.0. Runs on the managed ComfyUI.

Does
text → image
Price
free

on your Mac free Apache 2.0

Chroma1-HD (4-bit)

Chroma1-HD quantised to 4 bits so it fits a 16 GB Mac — the same open, unfiltered 8.9B model, with a small loss of fine detail. Real negative prompt, takes LoRAs. Apache 2.0. Runs on the managed ComfyUI.

Does
text → image
Price
free

on your Mac free Apache 2.0

Chroma1-Flash (4-bit)

The fast Chroma: distilled to about 8 steps with guidance baked in, and quantised to 4 bits for 16 GB machines. Same unfiltered model family; no negative prompt. Takes LoRAs. Apache 2.0. Runs on the managed ComfyUI.

Does
text → image
Price
free

on your Mac free Apache 2.0

InfiniteTalk (14B)

Dubs an existing video: sparse-frame video-to-video that re-lips footage to new audio while keeping the original motion. Unlimited length via chunked continuation. Built on Wan 2.1 I2V.

Does
video, audio, image → video, video-edit
Price
free

on your Mac free Apache 2.0

EchoMimicV3 Flash (1.3B)

Animates a portrait from audio at 768×768 in 8 steps — the best quality-per-VRAM in this category, and clearer licensing than LatentSync (whose weights are openrail++, not Apache).

Does
image, audio → video, video-edit
Price
free

on your Mac free Apache 2.0

MuseTalk 1.5

Real-time lip sync — single-step latent inpainting rather than a diffusion sampler, so it is fast and small. Capped at 256×256, and the authors note some jitter and lip-colour drift.

Does
video, audio → video, video-edit
Price
free

on your Mac free MIT (code + weights)

Examples are ours: generated with the model, hosted by us. Nothing on this page loads from a vendor's server, so browsing it tells them nothing about you.

Call these on your AI tokens — how to: pictures · video · transcription · speech

Pictures on your AI tokens

Create a workspace API key in the Overblast console and point your base URL at https://brain.deployd.network/ai/v1. Usage bills that workspace's AI tokens, at the prices shown here.

curl https://brain.deployd.network/ai/v1/images \
  -H "Authorization: Bearer ob_live_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemini-3-pro-image",
    "prompt": "a red bicycle against a wet brick wall at night"
  }'

# → { "data": [ { "b64_json": "iVBORw0KGgo…" } ] }

Video on your AI tokens

A job, not a request: start it, poll it, download it. Polling is free — the tokens are charged to the render, and each render in flight needs $5.00 of AI tokens available.

Create a workspace API key in the Overblast console and point your base URL at https://brain.deployd.network/ai/v1. Usage bills that workspace's AI tokens, at the prices shown here.

# 1 · start the render. A render in flight needs $5.00 of AI tokens available.
curl https://brain.deployd.network/ai/v1/videos \
  -H "Authorization: Bearer ob_live_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "black-forest-labs/flux-video-edit",
    "prompt": "a paper plane crossing a sunlit kitchen, slow motion"
  }'

# → 202 { "id": "vid_123", "status": "queued" }

# 2 · poll until it says completed — the cost rides the job, not the poll
curl https://brain.deployd.network/ai/v1/videos/vid_123 \
  -H "Authorization: Bearer ob_live_YOUR_KEY"

# 3 · download it
curl https://brain.deployd.network/ai/v1/videos/vid_123/content \
  -H "Authorization: Bearer ob_live_YOUR_KEY" \
  -o out.mp4

Transcription on your AI tokens

Any model that takes audio transcribes here — 47 do, each at its own token prices on the Language tab.

Create a workspace API key in the Overblast console and point your base URL at https://brain.deployd.network/ai/v1. Usage bills that workspace's AI tokens, at the prices shown here.

curl https://brain.deployd.network/ai/v1/audio/transcriptions \
  -H "Authorization: Bearer ob_live_YOUR_KEY" \
  -F file=@audio.mp3 \
  -F model=google/gemini-2.5-flash

# → { "text": "…what was said…" }

Speech on your AI tokens

A chat call with the audio modality: the reply carries a base64 clip beside the text. Voices and formats are the model's own.

Create a workspace API key in the Overblast console and point your base URL at https://brain.deployd.network/ai/v1. Usage bills that workspace's AI tokens, at the prices shown here.

curl https://brain.deployd.network/ai/v1/chat/completions \
  -H "Authorization: Bearer ob_live_YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-audio",
    "modalities": ["text", "audio"],
    "audio": { "voice": "alloy", "format": "mp3" },
    "messages": [{ "role": "user", "content": "Read this aloud: your order is on its way." }]
  }'

# → base64 audio in choices[0].message.audio.data

In your own app

Anything that speaks the OpenAI completions API runs every model here: one key, one base URL, one model id.

  1. Pick your tool.
  2. Get your key. Not needed for 00. For the rest: create a workspace API key in the Overblast console — usage bills that workspace's AI tokens, at the prices on this site.
  3. Turn it on.

    Pick the model per agent — Settings → your agent → Model, and per tier, so different agents run different models.

    The media models are built in on your AI tokens: ask for a picture, a clip, a transcript or speech. No key, no provider account, no setup.

    Download free →

    Two variables and a model id.

    export OPENAI_BASE_URL="https://brain.deployd.network/ai/v1"
    export OPENAI_API_KEY="ob_live_YOUR_KEY"
    
    codex -m deepseek/deepseek-v4.1-flash

    The SDK takes the same two settings.

    from openai import OpenAI
    
    client = OpenAI(base_url="https://brain.deployd.network/ai/v1", api_key="ob_live_YOUR_KEY")
    client.chat.completions.create(model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "Hello"}])

    The SDK takes the same two settings.

    import OpenAI from "openai";
    
    const client = new OpenAI({ baseURL: "https://brain.deployd.network/ai/v1", apiKey: "ob_live_YOUR_KEY" });
    await client.chat.completions.create({ model: "deepseek/deepseek-v4.1-flash", messages: [{ role: "user", content: "Hello" }] });

    Paste these into its model-provider settings.

    Base URL   https://brain.deployd.network/ai/v1
    API key    ob_live_YOUR_KEY
    Model      deepseek/deepseek-v4.1-flash

Speech 65

Two routes to sound. Through the endpoints above: any audio-taking model transcribes, any audio-returning one speaks, at their token prices on the Language tab. In the app: the engines below.

These run in the app on your own provider key or your own machine — provider rates and local costs, not AI-token prices, and not callable at the endpoints above. Prices are per minute of audio; per-character quotes convert at 900 characters a minute. Word-error rates come from one independent harness.

Text to speech how to use →

No voice clips yet — we only publish samples we made and host ourselves.

ModelMaker$/min LatencyLanguagesDoes
Kokoro 82M local Kokoro · local free 120 ms 8 offline
Piper local Piper · local free 80 ms 30 offline
macOS say (built-in) local Apple · local free 50 ms 40 offline
Chatterbox Turbo local Resemble · local free 200 ms 23 voice cloning offline
Polly Standard AWS $0.0036 300 ms 29 streaming
Grok Voice TTS xAI $0.0038 — 30 streaming voice cloning
Octave 2 Hume $0.0068 — 11 streaming voice cloning
GPT-4o Mini TTS OpenAI $0.013 300 ms 50 streaming
TTS-1 OpenAI $0.013 400 ms 50 streaming
Aura-1 Deepgram $0.013 130 ms 1 streaming
OpenAudio S2 Pro Fish Audio $0.013 — 13 voice cloning
Neural2 Google Cloud $0.014 — 40
Neural Microsoft Azure $0.014 300 ms 140 streaming
Polly Neural AWS $0.014 400 ms 40 streaming
Neural HD Microsoft Azure $0.020 300 ms 140 streaming
TTS-1 HD OpenAI $0.027 — 50
Aura-2 Deepgram $0.027 150 ms 15 streaming
Chirp 3 HD Google Cloud $0.027 500 ms 31 streaming
Polly Generative AWS $0.027 — 12 streaming
Sonic 3 Cartesia $0.032 90 ms 42 streaming voice cloning
Mist v2 Rime $0.035 100 ms 2 streaming
Speech 2.6 MiniMax $0.041 — 40 streaming voice cloning
Flash v2.5 ElevenLabs $0.045 75 ms 32 streaming voice cloning
Multilingual v2 ElevenLabs $0.090 280 ms 32 streaming voice cloning
Eleven v3 ElevenLabs $0.090 — 70 voice cloning

Transcription how to use →

ModelMaker$/min Word errors LatencyLanguagesDoes
Parakeet TDT 0.6B v3 (MLX) local NVIDIA · local free 6.3% — 25 timestamps offline
Whisper Large v3 Turbo (whisper.cpp) local OpenAI · local free 7.4% — 99 timestamps offline
Whisper Base (whisper.cpp) local OpenAI · local free — — 99 timestamps offline
Voxtral Mini 4B (MLX, 4-bit) local Mistral · local free — — 30 offline
Wizper (Whisper v3) fal.ai $0.0005 4.7% — 99
Whisper Large v3 Turbo Groq $0.0007 4.6% — 99 timestamps
Parakeet TDT 0.6B v3 Together AI $0.0015 4.5% — 25 timestamps
Soniox v5 Soniox $0.0017 3.8% — 60 streaming speaker labels timestamps
Grok Speech-to-Text xAI $0.0017 4% — 30 streaming
Universal-Streaming AssemblyAI $0.0025 — 300 ms 99 streaming
GPT-4o Mini Transcribe OpenAI $0.0030 4.5% — 99 streaming
Voxtral Mini Transcribe 2 Mistral $0.0030 3.6% — 30 speaker labels timestamps
Nova 2 Pro (transcribe) AWS Bedrock $0.0031 4.9% — 100
Universal-3.5 Pro AssemblyAI $0.0035 3.1% — 99 speaker labels timestamps
Scribe v2 ElevenLabs $0.0037 2.2% — 90 speaker labels timestamps
Voxtral Small Mistral $0.0040 2.8% — 30 timestamps
Melia Speechmatics $0.0040 4.9% — 55 streaming speaker labels
Pulse Pro Smallest.ai $0.0040 2.4% — 30 streaming
Nova-3 Deepgram $0.0043 5.2% — 36 streaming speaker labels timestamps
GPT Transcribe OpenAI $0.0045 3.3% — 99 streaming timestamps
Gemini 3.5 Transcribe Google $0.0050 — — 85 speaker labels timestamps
Nova-3 Multilingual Deepgram $0.0052 — — 36 streaming speaker labels timestamps
GPT-4o Transcribe OpenAI $0.0060 4% — 99 streaming
GPT-4o Transcribe (diarize) OpenAI $0.0060 — — 99 speaker labels timestamps
Whisper v2 (hosted) OpenAI $0.0060 4.1% — 99 timestamps
Voxtral Realtime Mistral $0.0060 — 200 ms 30 streaming
MAI-Transcribe-1 Microsoft Azure $0.0060 2.6% — 100 speaker labels timestamps
Amazon Transcribe AWS $0.0060 4.1% — 100 streaming speaker labels timestamps
Scribe v2 Realtime ElevenLabs $0.0065 — 150 ms 90 streaming timestamps
Flux (conversational) Deepgram $0.0077 — 260 ms 1 streaming
Solaria-3 Gladia $0.010 3.2% — 100 streaming speaker labels
Chirp 3 Google Cloud $0.016 — — 125 streaming speaker labels timestamps
GPT Live Transcribe OpenAI $0.017 — 300 ms 99 streaming timestamps

Speech to speech (live conversation)

ModelMaker$/min LatencyLanguagesDoes
GPT Realtime Mini OpenAI $0.018 450 ms 60 streaming
Gemini Live Google $0.024 600 ms 70 streaming
Grok Voice Agent xAI $0.050 500 ms 30 streaming
GPT Realtime 2.1 OpenAI $0.058 500 ms 60 streaming
Line Cartesia $0.060 90 ms 42 streaming voice cloning
Voice Agent API Deepgram $0.075 — 36 streaming
Speech Engine ElevenLabs $0.080 — 90 streaming voice cloning

Speech prices checked 2026-07-29. Ballpark mid-tier rates; they move constantly. fal's registry read 2026-09-23 (1496 models scanned, newest per category kept).