Everything 00 can generate with — pictures, video, music, speech, 3D — from every place
it can run them. Filter by what you need it to make and where you want it to run.
Dreamina Seedance 2.5 generates video from up to 50 multimodal references images, video, audio, and style inputs, locking a character, set, and palette across a full 30-second take for production-grade consistency.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- $0.4730 / second
via fal
commercial
Dreamina Seedance 2.5 animates a single still into a native 30-second clip at up to 720p, extending one frame into continuous, coherent motion without the drift or stitching of shorter multi-clip workflows.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- $0.4730 / second
via fal
commercial
Dreamina Seedance 2.5 generates native 30-second single-shot video at up to 720p from a single text prompt, reasoning about the whole shot at once so motion, lighting, and subject identity stay coherent from first frame to last.
- Does
- text → video
- Release date
- 2026-08
- Price
- $0.4730 / second
via fal
commercial
Generate a 3D relief depth map with Hi3D from a single image.
- Does
- image, text → image-edit
- Release date
- 2026-08
- Price
- $0.2 / generation
via fal
commercial
Splits a finished image into independent, editable transparent-PNG layers — background plus separate elements, from a text description, returning 2 to 17 layers per call for non-destructive reuse in design tools.
- Does
- image, text → image-edit
- Release date
- 2026-08
- Price
- $0.03375 / generated
via fal
commercial
Generate 3D models from multiple view images using Hi3D.
- Does
- image → 3d
- Release date
- 2026-08
- Price
- $0.02 / credit
via fal
commercial
Generate 3D models from a single image with Hi3D.
- Does
- image → 3d
- Release date
- 2026-08
- Price
- $0.02 / credit
via fal
commercial
FLUX.3 is Black Forest Labs' frontier video model. This endpoint animates a single still image into video, extending one frame into coherent, natural motion.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates the video between a defined start and end frame, interpolating a smooth, coherent transition from the first image to the last.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
FLUX.3 is Black Forest Labs' frontier video model. This endpoint builds video from a sequence of keyframes, generating the motion between each anchor point for precise control over how a shot progresses.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that animate a still image, with a reusable draft cache for full-quality enhancement.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews between a start and an end frame, with a reusable draft cache for full-quality enhancement.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews pinned to your keyframe images, with a reusable draft cache for full-quality enhancement.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
FLUX.3 is Black Forest Labs' frontier video model. This endpoint generates video directly from a text prompt, translating a written description into motion, composition, and scene.
- Does
- text → video
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews from a text prompt, with a reusable draft cache for full-quality enhancement.
- Does
- text → video
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
FLUX.3 is Black Forest Labs' frontier video model. This endpoint continues an existing clip beyond its final frame, generating additional footage that stays consistent with the original motion and scene.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
FLUX.3 is Black Forest Labs' frontier audio/video model. Generate fast, low-cost draft previews that continue an existing clip, with a reusable draft cache for full-quality enhancement.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
FLUX.3 is Black Forest Labs' frontier audio/video model. Re-render a previously generated draft at full quality — same seed, same motion, no re-planning.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
Use Heygen's Latest Model for Filler Word Removal.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
Generate videos from images and audio references using xAI's Grok Imagine 1.5 Video model.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- $0.08 / sec
via fal
commercial
Generate videos from prompts with audio using xAI's Grok Imagine 1.5 Video model.
- Does
- text → video
- Release date
- 2026-08
- Price
- $0.08 / sec
via fal
commercial
MiniMax H3 is a frontier video model. This endpoint animates a supplied image into 2K video, using it as the opening frame or pairs a first and last frame to control a transition between two images with the aspect ratio following the input.
- Does
- image, text → video
- Release date
- 2026-07
- Price
- $0.08 / second
via fal
commercial
MiniMax H3 is a frontier video model. This endpoint generates 2K video from multimodal references up to 9 images for subject and style, 3 video clips for motion, and 3 audio clips each cited in the prompt by order, keeping subjects consistent while following the referenced motion and audio.
- Does
- image, text → video
- Release date
- 2026-07
- Price
- $0.08 / second
via fal
commercial
MiniMax H3 is a frontier video model. This endpoint generates video from a text prompt alone, rendering at 2K in durations from 5 to 15 seconds across seven aspect ratios.
- Does
- text → video
- Release date
- 2026-07
- Price
- $0.08 / second
via fal
commercial
Pixelcut's Background Remover produces fast, high-quality cutouts built for e-commerce product imagery
- Does
- image, text → image-edit
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
Apply precise, controllable edits to a reference image while preserving composition, typography, identity, and fine visual detail.
- Does
- image, text → image-edit
- Release date
- 2026-07
- Price
- $7.50 / 1m
via fal
commercial
Generate high-fidelity, design-ready images with precise typography, strong prompt alignment, and rich visual detail using Microsoft's flagship MAI Image 2.5 Pro.
- Does
- text → image
- Release date
- 2026-07
- Price
- $7.50 / 1m
via fal
commercial
Prompt-free object removal from an image and mask, erasing objects with their shadows and reflections and reconstructing the scene cleanly.
- Does
- image, text → image-edit
- Release date
- 2026-07
- Price
- $0.03 / generation
via fal
commercial
Generate natural multilingual speech from text with fast voice and language control using Qwen Audio 3.0 TTS Flash.
- Does
- text → speech
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
FeyNobg is a state of the art AI model for background removal from feyninc
- Does
- image, text → image-edit
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
Remove character from your video using Ltx 2.3
- Does
- video, text → video-edit
- Release date
- 2026-07
- Price
- $0.0024075 / megapixel
via fal
commercial
Generates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen's strength in complex text rendering and precise prompt adherence
- Does
- image, text → image-edit
- Release date
- 2026-07
- Price
- $0.04 / generated
via fal
commercial
Edits images from one to three reference images and a natural-language instruction, preserving key details such as facial features and identity while applying the requested changes
- Does
- text → image
- Release date
- 2026-07
- Price
- $0.04 / generated
via fal
commercial
Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.
- Does
- text → music
- Release date
- 2026-07
- Price
- $0.0018 / second
via fal
commercial
Generates perfectly synced music for any video. Return a licensed music soundtrack ready for commercial use (optional preservation of the original speech in video)
- Does
- video, text → video-edit
- Release date
- 2026-07
- Price
- $0.009 / second
via fal
commercial
Adds synchronized, royalty-free, commercial-use-safe sound effects to a video. Returns the finished video with the generated audio mixed in.
- Does
- video, text → video-edit
- Release date
- 2026-07
- Price
- $0.009 / second
via fal
commercial
LTX-2.3 Reframe converts your videos to any aspect ratio without destructive cropping. It intelligently recenters the original footage and generatively fills the newly exposed areas with content that seamlessly matches the scene, so the result looks like it was shot natively in the target format. Tu
- Does
- video, text → video-edit
- Release date
- 2026-07
- Price
- $0.10 / second
via fal
commercial
Bria Product Dimensions turns one product photo and its measurements into a marketplace-ready dimension image with callout lines, labels, and weight or capacity readouts
- Does
- image, text → image-edit
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
Remix images from text prompts with strong prompt adherence, layout intelligence, and accurate text rendering using Reve 2.1
- Does
- image, text → image-edit
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
Edit images from text prompts with strong prompt adherence, layout intelligence, and accurate text rendering using Reve 2.1
- Does
- image, text → image-edit
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
Generate high-quality images from text prompts with strong prompt adherence, layout intelligence, and accurate text rendering using Reve 2.1.
- Does
- text → image
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
Generate production-quality lipsync from any audio using VEED's most advanced model yet.
- Does
- video, text → video-edit
- Release date
- 2026-07
- Price
- $0.07
via fal
commercial
Generate high-fidelity images from text with Krea 2 using a style reference image. Apply a reference image to guide the visual style into new generations, with aspect ratio, creativity, and seed controls.
- Does
- text → image
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
Image editing endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent image editing with multiple inputs.
- Does
- image, text → image-edit
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
Seedream 5.0 Pro is grounded, region-precise image editing model that changes one element while keeping the rest of the frame intact with layer separation, sketch completion, and up to 10 reference images.
- Does
- image, text → image-edit
- Release date
- 2026-07
- Price
- $0.0045
via fal
commercial
Text to Image endpoint for the fast Lite version of Seedream 5.0, supporting high quality intelligent text-to-image generation.
- Does
- text → image
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
ByteDance's Seedream 5.0 Pro is flagship text-to-image model, with deep-thinking prompt understanding, native text in 14 languages, and precise control over dense layouts and structured designs.
- Does
- text → image
- Release date
- 2026-07
- Price
- $0.0675 / image
via fal
commercial
Extend high-quality video with audio from input video using LTX-2.3 with Lora
- Does
- video, text → video-edit
- Release date
- 2026-07
- Price
- $0.0024075 / megapixel
via fal
commercial
Extend high-quality video with audio from input video using LTX-2.3
- Does
- video, text → video-edit
- Release date
- 2026-07
- Price
- $0.0024075 / megapixel
via fal
commercial
Generate high-quality images, posters, and logos with Ideogram's latest V4.0q — producing crisp visuals with accurate text rendering, fine detail, and full creative control for polished, ready-to-use designs FRACTION OF A SECOND.
- Does
- text → image
- Release date
- 2026-07
- Price
- $0.0075 / megapixel
via fal
commercial
Generate high-quality images, posters, and logos with Ideogram's latest V4.0q — producing crisp visuals with accurate text rendering, fine detail, and full creative control for polished, ready-to-use designs IN A SECOND.
- Does
- text → image
- Release date
- 2026-07
- Price
- $0.00525 / megapixel
via fal
commercial
Run inference on LoRA adapters for TRELLIS.2 model
- Does
- image → 3d
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.
- Does
- image, text → image-edit
- Release date
- 2026-06
- Price
- $0.3125
via fal
commercial
Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.
- Does
- text → image
- Release date
- 2026-06
- Price
- $0.3125
via fal
commercial
Nano banana lite is the efficiency-focused model in the image generation family. Sub-2 second latency with cost-effective generation and editing, fast multi-turn local edits, and 14 supported aspect ratios.
- Does
- text → image
- Release date
- 2026-06
- Price
- $0.3125
via fal
commercial
Generates video with audio from combined multimodal references. Accepts text, images, audio, and video together as input to guide subject, motion, style, and sound in the output.
- Does
- image, text → video
- Release date
- 2026-06
- Price
- $1.875 / 1
via fal
commercial
Animates a still image into video with audio. Extends a single frame into coherent motion, grounded in Gemini's physical understanding of how scenes and subjects behave.
- Does
- image, text → video
- Release date
- 2026-06
- Price
- $1.875 / 1
via fal
commercial
Creates video with synchronized audio from text input. Grounded in Gemini's real-world knowledge, with improved physics understanding for more coherent motion and interaction.
- Does
- text → video
- Release date
- 2026-06
- Price
- $21.875 / 1
via fal
commercial
Edits generated video across multiple conversational turns while preserving scene coherence. Applies iterative changes through natural-language instructions without regenerating the full sequence from scratch.
- Does
- video, text → video-edit
- Release date
- 2026-06
- Price
- $1.875 / 1
via fal
commercial
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...
- Does
- image, text → image
- Release date
- 2026-06
- Price
- $1.5 / 1M out
language-model key
commercial
Bria Extract Object uses text prompts to isolate a selected object from an image and return it as an RGBA PNG with a transparent background. Ideal for product, ecommerce, advertising, and creative editing workflows. Bria's Extract Object API leads in product shot extraction, outperforming SAM 3.1 wh
- Does
- image, text → image-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Transform your 3D video render into realistic using first frame with Ltx 2.3
- Does
- video, text → video-edit
- Release date
- 2026-06
- Price
- $0.0024075 / megapixel
via fal
commercial
Seed Audio 1.0 is a new audio model from Bytedance that can generate high-quality, natural sounding audio using text, reference audios or an image.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Deblur high-quality video using LTX-2.3
- Does
- video, text → video-edit
- Release date
- 2026-06
- Price
- $0.0024075 / megapixel
via fal
commercial
Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.
- Does
- image, text → video
- Release date
- 2026-06
- Price
- $0.0721 / second
via fal
commercial
Seedance 2.0 Mini is a faster version of Seedance 2.0 that brings great performance and high generation speed at a lower cost.
- Does
- text → video
- Release date
- 2026-06
- Price
- $0.0721 / second
via fal
commercial
Generate high-fidelity images from text in seconds with Krea 2 Turbo, the speed-optimized open-source version of Krea 2, preserving its aesthetic range for rapid ideation.
- Does
- text → image
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Generate high-fidelity images from text with Krea 2 using a custom-trained LoRA. Apply your LoRA weights to carry a learned subject, character, or style into new generations, with aspect ratio, creativity, and seed controls.
- Does
- text → image
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Happy Horse 1.1 is Alibaba's #1-ranked video model. This text-to-video endpoint generates 1080p video with synchronized native audio and multilingual lip-sync from a text prompt alone.
- Does
- text → video
- Release date
- 2026-06
- Price
- $0.14 / second
via fal
commercial
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.
- Does
- image → 3d
- Release date
- 2026-06
- Price
- $0.1 / generation
via fal
commercial
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.
- Does
- text → 3d
- Release date
- 2026-06
- Price
- $0.1 / generation
via fal
commercial
Text To Image Model using Boogu-Image
- Does
- text → image
- Release date
- 2026-06
- Price
- $0.04 / megapixel
via fal
commercial
Generate professional-quality voiceovers in seconds with Async TTS Pro model text-based control over pauses, emphasis, and timing. Voice ids can be found at https://async.com/developer/voice-library
- Does
- text → speech
- Release date
- 2026-06
- Price
- $0.01 / 1
via fal
commercial
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...
- Does
- image, text → image
- Release date
- 2026-06
- Price
- $12 / 1M out
language-model key
commercial
Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced...
- Does
- image, text → image
- Release date
- 2026-06
- Price
- $3 / 1M out
language-model key
commercial
Generate Infographic Image with Sensenova U1
- Does
- text → image
- Release date
- 2026-06
- Price
- $0.05 / generated
via fal
commercial
Text to Audio high-quality using LTX-2.3 with Lora
- Does
- text → music
- Release date
- 2026-06
- Price
- $0.0024075 / megapixel
via fal
commercial
Text to Audio high-quality using LTX-2.3
- Does
- text → music
- Release date
- 2026-06
- Price
- $0.0024075 / megapixel
via fal
commercial
Generate high quality 1080p videos using Kling's Turbo 3.0 model, with improved lipsync and multishot generation capabilities.
- Does
- text → video
- Release date
- 2026-06
- Price
- $0.14
via fal
commercial
Kling 3.0 Turbo Standard is a fast, cost-efficient video generation model that turns text prompts directly into 720P video with native audio, optimized for rapid iteration and high-volume production
- Does
- text → video
- Release date
- 2026-06
- Price
- $0.112
via fal
commercial
Zonos2 is a text-to-speech model that clones a voice from a short sample and speaks naturally across many languages.
- Does
- text → speech
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Nemotron-ASR-Streaming is a multi lingual, streaming Automatic Speech Recognition (ASR) engineered to deliver high-quality multi lingual transcription across both low-latency streaming and high-throughput batch workloads.
- Does
- audio → transcribe
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Luma Ray 3.2 generates cinematic video from a text prompt, with control over resolution, duration, and seamless looping, plus reference images to lock in subject and style.
- Does
- text → video
- Release date
- 2026-06
- Price
- $0.50
via fal
commercial
Generate high-quality video from a text prompt with Bernini-R, ByteDance's unified video generation and editing model.
- Does
- text → video
- Release date
- 2026-06
- Price
- $0.08 / second
via fal
commercial
Seed Speech developed by ByteDance, is a family of large-scale text-to-speech models capable of synthesizing speech that is virtually indistinguishable from human speech.
- Does
- text → speech
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium Base is the foundational 1.4 billion parameter text-to-audio checkpoint generating stereo music up to 6 minutes, intended as the unmodified base for custom fine-tuning workflows.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music Base is the foundational 459 million parameter checkpoint generating full music compositions up to 2 minutes from text prompts, intended as the unmodified base for fine-tuning.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as the unmodified base for fine-tuning.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for safe commercial use.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music is a 459 million parameter latent diffusion model that generates full stereo music compositions up to 2 minutes from text prompts, lightweight enough for on-device deployment.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX is a 459 million parameter latent diffusion model that generates high-quality sound effects from text prompts, designed for on-device deployment on mobile phones and consumer laptops.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Generates licensed, commercial-use-safe music from a single text prompt, with full control over style, mood, instrumentation, and exact duration.
- Does
- text → music
- Release date
- 2026-06
- Price
- $0.0025 / second
via fal
commercial
TripoSplat is an open-source model from TripoAI / VAST AI Research that converts a single 2D image into high-quality 3D Gaussians using a novel learned density-control approach
- Does
- image → 3d
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium Base audio outpainting is the foundational 1.4 billion parameter checkpoint that extends existing stereo audio with causal continuation guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium Base audio inpainting is the foundational 1.4 billion parameter checkpoint for editing or filling selected stereo audio segments guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium Base audio-to-audio is the foundational 1.4 billion parameter checkpoint that transforms input audio into new stereo variations up to 6 minutes guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX Base audio outpainting is the foundational 459 million parameter checkpoint that extends sound-effect tracks via causal continuation guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input music into new variations up to 2 minutes guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input audio into new sound-effect variations guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium audio outpainting is a 1.4 billion parameter latent diffusion model that extends existing stereo audio beyond its original endpoint via causal continuation guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium audio inpainting is a 1.4 billion parameter latent diffusion model that fills in or reworks selected segments of a stereo track guided by text prompts, supporting single- and multi-segment editing.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium audio-to-audio is a 1.4 billion parameter latent diffusion model that transforms an input audio clip into new stereo variations up to 6 minutes guided by a text prompt.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music audio outpainting is a 459 million parameter latent diffusion model that extends music compositions beyond their original endpoint via causal continuation.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX audio outpainting is a 459 million parameter latent diffusion model that extends sound-effect tracks beyond their original endpoint via causal continuation.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music audio inpainting is a 459 million parameter latent diffusion model that fills in or reworks selected segments of a music track guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX audio inpainting is a 459 million parameter latent diffusion model that fills in or reworks selected segments of a sound-effect track guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music audio-to-audio is a 459 million parameter latent diffusion model that transforms input music into new variations up to 2 minutes guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Generate high-quality video with audio from text using LTX-2.3 and custom LoRA
- Does
- text → video
- Release date
- 2026-06
- Price
- $0.0027075 / megapixel
via fal
commercial
Generate high-quality video with audio from text using LTX-2.3
- Does
- text → video
- Release date
- 2026-06
- Price
- $0.0024075 / megapixel
via fal
commercial
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.
- Does
- image → 3d
- Release date
- 2026-05
- Price
- $0.4 / generation
via fal
commercial
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.
- Does
- text → 3d
- Release date
- 2026-05
- Price
- $0.4 / generation
via fal
commercial
Lyria 3 Pro is the latest music model from Google
- Does
- text → music
- Release date
- 2026-05
- Price
- see supplier
via fal
commercial
Generate ambient sounds for any text prompt. Now you can turn any SFX into a natural loop for ambient soundscapes.
- Does
- text → music
- Release date
- 2026-05
- Price
- see supplier
via fal
commercial
Pixal3D turns a single image into a high-fidelity 3D model with detailed geometry and realistic textures.
- Does
- image → 3d
- Release date
- 2026-05
- Price
- $0.3
via fal
commercial
Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.
- Does
- image → 3d
- Release date
- 2026-04
- Price
- $0.8 / untextured
via fal
commercial
Cohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation
- Does
- audio → transcribe
- Release date
- 2026-04
- Price
- see supplier
via fal
commercial
[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...
- Does
- image, text, file → image
- Release date
- 2026-04
- Price
- $15 / 1M out
language-model key
commercial
Newest audio model from Google introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.
- Does
- text → speech
- Release date
- 2026-04
- Price
- see supplier
via fal
commercial
Generate 3D models from multiple view images using Tripo H3.1.
- Does
- image → 3d
- Release date
- 2026-04
- Price
- $0.20
via fal
commercial
Generate high-quality 3D models from a single image using Tripo H3.1.
- Does
- image → 3d
- Release date
- 2026-04
- Price
- $0.20
via fal
commercial
Generate 3D models from a single image using Tripo P1.
- Does
- image → 3d
- Release date
- 2026-04
- Price
- $0.40
via fal
commercial
Generate 3D models from text descriptions using Tripo H3.1.
- Does
- text → 3d
- Release date
- 2026-04
- Price
- $0.10
via fal
commercial
Generate 3D models from text descriptions using Tripo P1.
- Does
- text → 3d
- Release date
- 2026-04
- Price
- see supplier
via fal
commercial
MiniMax Music 2.5 creates complete tracks with singing, backing music, and detailed arrangements from lyrics and a style description.
- Does
- text → music
- Release date
- 2026-04
- Price
- see supplier
via fal
commercial
Generate 3D models from one or more images using ReconViaGen 0.5
- Does
- image → 3d
- Release date
- 2026-04
- Price
- $0.25
via fal
commercial
30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...
- Does
- text, image → speech
- Release date
- 2026-03
- Price
- free
language-model key
free
Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...
- Does
- text, image → speech
- Release date
- 2026-03
- Price
- free
language-model key
free
Generate speech with expressive and realistic voices from xAI
- Does
- text → speech
- Release date
- 2026-03
- Price
- $0.015 / 1000
via fal
commercial
Text to Speech Endpoint for Inworld's TTS-1.5 Max.
- Does
- text → speech
- Release date
- 2026-03
- Price
- $0.01 / 1000
via fal
commercial
Generate 3D models from your images using Trellis 2. A native 3D generative model enabling versatile and high-quality 3D asset creation.
- Does
- image → 3d
- Release date
- 2026-03
- Price
- see supplier
via fal
commercial
Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines...
- Does
- image, text → image
- Release date
- 2026-02
- Price
- $3 / 1M out
language-model key
commercial
Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.
- Does
- image → 3d
- Release date
- 2026-02
- Price
- see supplier
via fal
commercial
Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.
- Does
- text → 3d
- Release date
- 2026-02
- Price
- see supplier
via fal
commercial
Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
- Does
- text → speech
- Release date
- 2026-02
- Price
- see supplier
via fal
commercial
Generate speech from text prompts and different voices using the MiniMax Speech-2.8 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.
- Does
- text → speech
- Release date
- 2026-02
- Price
- see supplier
via fal
commercial
Create detailed, fully-textured 3D models with text
- Does
- text → 3d
- Release date
- 2026-01
- Price
- $0.225 / generation
via fal
commercial
Generate 3D models from text prompts with Hunyuan 3D Pro
- Does
- text → 3d
- Release date
- 2026-01
- Price
- $0.375 / generation
via fal
commercial
Create custom voices using Qwen3-TTS Voice Design model and later use Clone Voice model to create your own voices!
- Does
- text → speech
- Release date
- 2026-01
- Price
- see supplier
via fal
commercial
Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model
- Does
- text → speech
- Release date
- 2026-01
- Price
- see supplier
via fal
commercial
Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model
- Does
- text → speech
- Release date
- 2026-01
- Price
- see supplier
via fal
commercial
The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...
- Does
- text, audio → speech
- Release date
- 2026-01
- Price
- $10 / 1M out
language-model key
commercial
A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...
- Does
- text, audio → speech
- Release date
- 2026-01
- Price
- $2.4 / 1M out
language-model key
commercial
Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!
- Does
- audio → transcribe
- Release date
- 2026-01
- Price
- $0.008 / input
via fal
commercial
Generate 3D human motions via text-to-generation interface of Hunyuan Motion!
- Does
- text → 3d
- Release date
- 2025-12
- Price
- see supplier
via fal
commercial
Generate 3D human motions via text-to-generation interface of Hunyuan Motion!
- Does
- text → 3d
- Release date
- 2025-12
- Price
- see supplier
via fal
commercial
Generate long speech snippets fast using Microsoft's powerful TTS.
- Does
- text → speech
- Release date
- 2025-12
- Price
- see supplier
via fal
commercial
Turn simple sketches into detailed, fully-textured 3D models. Instantly convert your concept designs into formats ready for Unity, Unreal, and Blender.
- Does
- text → 3d
- Release date
- 2025-12
- Price
- $0.375 / generation
via fal
commercial
Maya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise voice design.
- Does
- text → speech
- Release date
- 2025-12
- Price
- see supplier
via fal
commercial
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...
- Does
- image, text → image
- Release date
- 2025-11
- Price
- $12 / 1M out
language-model key
commercial
GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...
- Does
- file, image, text → image
- Release date
- 2025-10
- Price
- $2 / 1M out
language-model key
commercial
[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...
- Does
- image, text, file → image
- Release date
- 2025-10
- Price
- $10 / 1M out
language-model key
commercial
Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,...
- Does
- image, text → image
- Release date
- 2025-10
- Price
- $2.5 / 1M out
language-model key
commercial
Meshy-6-Preview is the latest model from Meshy. It generates realistic and production ready 3D models.
- Does
- text → 3d
- Release date
- 2025-10
- Price
- see supplier
via fal
commercial
An open source, community-driven and native audio turn detection model by Pipecat AI.
- Does
- audio → transcribe
- Release date
- 2025-04
- Price
- see supplier
via fal
commercial
Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.
- Does
- audio → transcribe
- Release date
- 2025-04
- Price
- see supplier
via fal
commercial
Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.
- Does
- audio → transcribe
- Release date
- 2025-04
- Price
- see supplier
via fal
commercial
Generate text from speech using ElevenLabs advanced speech-to-text model.
- Does
- audio → transcribe
- Release date
- 2025-02
- Price
- see supplier
via fal
commercial
[Experimental] Whisper v3 Large -- but optimized by our inference wizards. Same WER, double the performance!
- Does
- audio → transcribe
- Release date
- 2024-04
- Price
- see supplier
via fal
commercial
[Experimental] Whisper v3 Large -- but optimized by our inference wizards. Same WER, double the performance!
- Does
- audio → transcribe
- Release date
- 2024-04
- Price
- see supplier
via fal
commercial
Google's image model — the reference for rendering copy/text inside images verbatim, plus strong instruction editing. The go-to for social assets with words on them.
- Does
- text, image → image, image-edit
- Price
- see supplier
cloud API
commercial
High-quality general image generation — photoreal scenes, cheap and fast.
- Does
- text → image
- Price
- see supplier
cloud API
commercial
Google's video model — 8 s clips WITH generated sound and dialogue, vertical or landscape.
- Does
- text, image → video
- Price
- see supplier
cloud API
commercial
Fast, inexpensive silent clips — the default cloud backend of video_generate.
- Does
- text, image → video
- Price
- see supplier
cloud API
commercial
Google's music model — rich 30 s instrumentals from a mood prompt.
- Does
- text → music
- Price
- see supplier
cloud API
commercial
Google's newest fast image model — Nano Banana quality at higher speed, generation and editing.
- Does
- text, image → image, image-edit
- Price
- see supplier
cloud API
commercial
OpenAI's latest image model — extremely detailed output with fine typography, plus fine-grained edits.
- Does
- text, image → image, image-edit
- Price
- see supplier
cloud API
commercial
ByteDance's most advanced video model — cinematic output with native audio, real-world physics, and reference-to-video from up to 9 images, 3 videos, and 3 audio clips.
- Does
- text, image, images, audio → video
- Price
- see supplier
cloud API
commercial
Kuaishou's top-tier video model — cinematic visuals, fluid motion, native audio, custom element support.
- Does
- text, image → video
- Price
- see supplier
cloud API
commercial
Alibaba's #1-ranked video model — 1080p output with synchronized native audio and multilingual lip-sync.
- Does
- text, image → video
- Price
- see supplier
cloud API
commercial
Google's latest music model — richer arrangements and cleaner mixes than Lyria 2.
- Does
- text → music
- Price
- see supplier
cloud API
commercial
xAI's image model — bold, stylized generation and editing from the Grok family.
- Does
- text, image → image, image-edit
- Price
- see supplier
cloud API
commercial
xAI's video model — text/image to video with audio, plus reference-to-video, edit and extend endpoints.
- Does
- text, image, video → video
- Price
- see supplier
cloud API
commercial
Tencent's 3D generation model — turn a prompt or a single image into a textured 3D mesh (GLB).
- Does
- text, image → 3d
- Price
- see supplier
cloud API
commercial
Professional-grade upscaling for images and video — sharper detail, denoise, real resolution gains.
- Does
- image, video → upscale
- Price
- see supplier
cloud API
commercial
State-of-the-art expressive text-to-speech — natural delivery, emotion, and multilingual voices.
- Does
- text → speech
- Price
- see supplier
cloud API
commercial
Best open image model that fits this class of Mac — photorealism, sharp text, strong prompt following. 6B, 4-bit, runs on the Apple GPU via MLX.
- Does
- text → image
- Price
- free
on your Mac
free
Apache 2.0
Fast 4-step image generation AND instruction editing ("make it snow") — the only open commercial-use editing model that fits 16 GB.
- Does
- text, image, images → image, image-edit
- Price
- free
on your Mac
free
Apache 2.0
The bigger klein — noticeably stronger detail and prompt following than the 4B, still 4-step fast. Non-commercial license (unlike the 4B).
- Does
- text, image, images → image, image-edit
- Price
- free
on your Mac
free
FLUX klein 9B
Baidu's 8B image model — strong photorealism and composition, Apache-licensed. Runs 4-bit on the Apple GPU via MLX (downloads the full official weights, quantized at load).
- Does
- text → image
- Price
- free
on your Mac
free
Apache 2.0
The most popular SDXL finetune — polished photoreal people and scenes, the gateway to the civitai ecosystem. Runs on the managed ComfyUI.
- Does
- text → image
- Price
- free
on your Mac
free
CreativeML OpenRAIL-M
Alibaba's open video model — the one local video generator that genuinely fits 16 GB (4-bit GGUF). Silent clips, ~3 s at reduced resolution.
- Does
- text, image → video
- Price
- free
on your Mac
free
Apache 2.0
Near-instant instrumental music and sound effects (44.1 kHz stereo) — pure MLX, tiny RAM footprint. No vocals.
- Does
- text, audio → music, sfx
- Price
- free
on your Mac
free
Stability Community
Full songs WITH vocals and lyrics in 50+ languages — the best open local music model (between Suno v4.5 and v5). Also does covers and repaints.
- Does
- text, lyrics → song, music
- Price
- free
on your Mac
free
MIT
The only local model anywhere that generates video WITH synchronized sound. 4-bit MLX port. Its pipeline pulls the full 56 GB weight set and wants real memory headroom — a 32 GB+ Mac.
- Does
- text, image → video
- Price
- free
on your Mac
free
LTX Community
Top-tier image generation + the best open image editor. Needs a 32 GB+ Mac.
- Does
- text, image → image, image-edit
- Price
- free
on your Mac
free
Apache 2.0
Alibaba's current video model, including prompt-based video editing — but API-only. No open weights: the newest downloadable Wan is 2.2.
- Does
- text, image, video → video, video-edit
- Price
- free
on your Mac
free
closed
Dubs an existing video: sparse-frame video-to-video that re-lips footage to new audio while keeping the original motion. Unlimited length via chunked continuation. Built on Wan 2.1 I2V.
- Does
- video, audio, image → video, video-edit
- Price
- free
on your Mac
free
Apache 2.0
Animates a portrait from audio at 768×768 in 8 steps — the best quality-per-VRAM in this category, and clearer licensing than LatentSync (whose weights are openrail++, not Apache).
- Does
- image, audio → video, video-edit
- Price
- free
on your Mac
free
Apache 2.0
Real-time lip sync — single-step latent inpainting rather than a diffusion sampler, so it is fast and small. Capped at 256×256, and the authors note some jitter and lip-colour drift.
- Does
- video, audio → video, video-edit
- Price
- free
on your Mac
free
MIT (code + weights)
Frontier-class open image model. Needs a 64 GB+ Mac.
- Does
- text, images → image, image-edit
- Price
- free
on your Mac
free
FLUX dev (non-commercial)
Nothing matches that.
Examples are ours — generated with the model and hosted by us. Where a card has no
example we have not made one yet; nothing on this page is loaded from a vendor's servers, so browsing
it tells them nothing about you.
Prices are per MINUTE of audio, whatever unit the supplier quotes in — per-character
text-to-speech is converted at 900 characters a minute so the columns can be compared at all.
Word-error rates come from one independent harness, so they are comparable with each other.
No voice clips yet — we only publish samples we generated and host ourselves, and these have not been made.
Speech prices checked 2026-07-29. Ballpark,
pay-as-you-go, mid-tier — they move constantly. fal's registry read 2026-08-10 (1437 models scanned, newest per category kept).