Google's image model — the reference for rendering copy/text inside images verbatim, plus strong instruction editing. The go-to for social assets with words on them.
- Does
- text, image → image, image-edit
- Price
- ≈ $0.31 / image
AI tokens — built in
commercial
High-quality general image generation — photoreal scenes, cheap and fast.
- Does
- text → image
- Price
- see supplier
cloud API
commercial
Google's video model — 8 s clips WITH generated sound and dialogue, vertical or landscape.
- Does
- text, image → video
- Price
- see supplier
cloud API
commercial
Fast, inexpensive silent clips — the default cloud backend of video_generate.
- Does
- text, image → video
- Price
- see supplier
cloud API
commercial
Google's music model — rich 30 s instrumentals from a mood prompt.
- Does
- text → music
- Price
- see supplier
cloud API
commercial
Google's newest fast image model — Nano Banana quality at higher speed, generation and editing.
- Does
- text, image → image, image-edit
- Price
- ≈ $0.155 / image
AI tokens — built in
commercial
OpenAI's latest image model — extremely detailed output with fine typography, plus fine-grained edits.
- Does
- text, image → image, image-edit
- Price
- ≈ $0.077 / image
AI tokens — built in
commercial
ByteDance's most advanced video model — cinematic output with native audio, real-world physics, and reference-to-video from up to 9 images, 3 videos, and 3 audio clips.
- Does
- text, image, images, audio → video
- Price
- see supplier
cloud API
commercial
Kuaishou's top-tier video model — cinematic visuals, fluid motion, native audio, custom element support.
- Does
- text, image → video
- Price
- see supplier
cloud API
commercial
Google's latest music model — richer arrangements and cleaner mixes than Lyria 2.
- Does
- text → music
- Price
- see supplier
cloud API
commercial
xAI's image model — bold, stylized generation and editing from the Grok family.
- Does
- text, image → image, image-edit
- Price
- see supplier
cloud API
commercial
xAI's video model — text/image to video with audio, plus reference-to-video, edit and extend endpoints.
- Does
- text, image, video → video
- Price
- see supplier
cloud API
commercial
Professional-grade upscaling for images and video — sharper detail, denoise, real resolution gains.
- Does
- image, video → upscale
- Price
- see supplier
cloud API
commercial
State-of-the-art expressive text-to-speech — natural delivery, emotion, and multilingual voices.
- Does
- text → speech
- Price
- see supplier
cloud API
commercial
Best open image model that fits this class of Mac — photorealism, sharp text, strong prompt following. 6B, 4-bit, runs on the Apple GPU via MLX.
- Does
- text → image
- Price
- free
on your Mac
free
Apache 2.0
Fast 4-step image generation AND instruction editing ("make it snow") — the only open commercial-use editing model that fits 16 GB.
- Does
- text, image, images → image, image-edit
- Price
- free
on your Mac
free
Apache 2.0
The bigger klein — noticeably stronger detail and prompt following than the 4B, still 4-step fast. Non-commercial license (unlike the 4B).
- Does
- text, image, images → image, image-edit
- Price
- free
on your Mac
free
FLUX klein 9B
Baidu's 8B image model — strong photorealism and composition, Apache-licensed. Runs 4-bit on the Apple GPU via MLX (downloads the full official weights, quantized at load).
- Does
- text → image
- Price
- free
on your Mac
free
Apache 2.0
Black Forest Labs' fast FLUX — four steps to a finished image, and the most permissive licence of any model this good. Apache 2.0: yours to use commercially, no strings. Runs on the managed ComfyUI.
- Does
- text → image
- Price
- free
on your Mac
free
Apache 2.0
The bigger FLUX — slower than schnell and noticeably better at hands, text in images and fine detail. Its licence is NON-COMMERCIAL: fine for personal work, not for anything you sell.
- Does
- text → image
- Price
- free
on your Mac
free
FLUX.1-dev
The most popular SDXL finetune — polished photoreal people and scenes, the gateway to the civitai ecosystem. Runs on the managed ComfyUI.
- Does
- text → image
- Price
- free
on your Mac
free
CreativeML OpenRAIL-M
Instruction image EDITING — give it a photo and a sentence ("make it snow", "remove the car", "turn this into a pencil sketch") and it changes that and leaves the rest of the picture alone. The first editor here that does not need a Mac: it runs on a Linux GPU server. Non-commercial licence.
- Does
- image → image, image-edit
- Price
- free
on your Mac
free
FLUX.1-dev
Alibaba's open video model — the one local video generator that genuinely fits 16 GB (4-bit GGUF). Silent clips, ~3 s at reduced resolution.
- Does
- text, image → video
- Price
- free
on your Mac
free
Apache 2.0
Edit a clip you already have: restyle it, change what is in it, keep its motion ("the same dance, but the dancer is made of glass"). One model, Apache 2.0, on a Linux GPU server — the first thing here that edits video at all.
- Does
- video → video-edit
- Price
- free
on your Mac
free
Apache 2.0
Near-instant instrumental music and sound effects (44.1 kHz stereo) — pure MLX, tiny RAM footprint. No vocals.
- Does
- text, audio → music, sfx
- Price
- free
on your Mac
free
Stability Community
Sound effects and short ambiance from a sentence ("footsteps on gravel", "distant thunder") — 44.1 kHz stereo, up to two minutes. The first sound model here that does not need a Mac: it runs on a Linux GPU server.
- Does
- text → sfx
- Price
- free
on your Mac
free
Stability Community
Full songs WITH vocals and lyrics in 50+ languages — the best open local music model (between Suno v4.5 and v5). Also does covers and repaints.
- Does
- text, lyrics → song, music
- Price
- free
on your Mac
free
MIT
The only local model anywhere that generates video WITH synchronized sound. 4-bit MLX port. Its pipeline pulls the full 56 GB weight set and wants real memory headroom — a 32 GB+ Mac.
- Does
- text, image → video
- Price
- free
on your Mac
free
LTX Community
Top-tier image generation + the best open image editor. Needs a 32 GB+ Mac.
- Does
- text, image → image, image-edit
- Price
- free
on your Mac
free
Apache 2.0
Alibaba's current video model, including prompt-based video editing — but API-only. No open weights: the newest downloadable Wan is 2.2.
- Does
- text, image, video → video, video-edit
- Price
- free
on your Mac
free
closed
Frontier-class open image model. Needs a 64 GB+ Mac.
- Does
- text, images → image, image-edit
- Price
- free
on your Mac
free
FLUX.1-dev
Seedream 5.0 Flash is a fast image generation and editing model, built for workflows where speed and budget matter.
- Does
- image, text → image-edit
- Release date
- 2026-09
- Price
- see supplier
via fal
commercial
Seedream 5.0 Flash is a fast image generation and editing model, built for workflows where speed and budget matter.
- Does
- image, text → image-edit
- Release date
- 2026-09
- Price
- see supplier
via fal
commercial
Seedream 5.0 Flash is a fast image generation and editing model, built for workflows where speed and budget matter.
- Does
- text → image
- Release date
- 2026-09
- Price
- see supplier
via fal
commercial
Recraft V4.1 Flash generates raster images from text prompts, including photography, illustrations, and mixed-media compositions, with controls for image size, color palette, and background color.
- Does
- text → image
- Release date
- 2026-09
- Price
- $0.007 / image
via fal
commercial
Ming Image 0.1 Design Layer is an image-to-image model from inclusionAI that decomposes a flattened design image into separate RGBA layers, such as a background layer and foreground elements, and...
- Does
- text, image → image, image-edit
- Release date
- 2026-09
- Price
- ≈ $0 / image
AI tokens — built in
commercial
Recraft V4.1 Flash is a text-to-image model from Recraft, the speed and cost tier of the V4.1 family. It generates ~1K raster images in about 1.5 seconds end to end,...
- Does
- text → image
- Release date
- 2026-09
- Price
- $0.014 / image
AI tokens — built in
commercial
Ming Image 0.1 Design is a text-to-image model from inclusionAI aimed at graphic-design output, with an emphasis on legible text rendering inside the generated image. It generates from a prompt...
- Does
- text → image
- Release date
- 2026-09
- Price
- ≈ $0 / image
AI tokens — built in
commercial
Meshy 7.1 generates 3D models from a single image, with standard, low-poly, and Smart Topology modes, optional textures and PBR maps, and geometry resolution up to 4K.
- Does
- image → 3d
- Release date
- 2026-09
- Price
- $0.12 / call
via fal
commercial
Meshy 7.1 generates textured 3D models from one to four views of the same object, with polygon count, topology, symmetry, and optional PBR texture controls.
- Does
- image → 3d
- Release date
- 2026-09
- Price
- $0.12 / call
via fal
commercial
Meshy 7.1 generates 3D models from text prompts, with untextured preview and textured full modes, standard, low-poly, and Smart Topology options, and geometry resolution up to 4K.
- Does
- text → 3d
- Release date
- 2026-09
- Price
- $0.12 / call
via fal
commercial
Generates 768p video with audio in a 16-bit pixel-art style from text prompts or an optional first-frame image. Supports durations of 5–15 seconds.
- Does
- text → video
- Release date
- 2026-09
- Price
- $0.08 / second
via fal
commercial
Generates 768p video with audio in a hand-drawn animation style from text prompts or an optional first-frame image. Supports durations of 5–15 seconds.
- Does
- text → video
- Release date
- 2026-09
- Price
- $0.08 / second
via fal
commercial
Generates 768p video with audio in a retro low-poly 3D style from text prompts or an optional first-frame image. Supports durations of 5–15 seconds.
- Does
- text → video
- Release date
- 2026-09
- Price
- $0.08 / second
via fal
commercial
Generates 768p video with audio in a retro 1970s hand-painted animation style from text prompts or an optional first-frame image. Supports durations of 5–15 seconds.
- Does
- text → video
- Release date
- 2026-09
- Price
- $0.08 / second
via fal
commercial
Generates 768p VHS-style video with audio from text prompts or an optional first-frame image. Supports 5–15 second clips and adjustable tape damage, from subtle analog noise to strong tracking distortion.
- Does
- text → video
- Release date
- 2026-09
- Price
- $0.08 / second
via fal
commercial
Tripo P2 generates 3D models from a single image, with optional PBR textures, adjustable face counts, and triangle or quad mesh topology.
- Does
- image → 3d
- Release date
- 2026-09
- Price
- $1.00
via fal
commercial
Tripo P2 generates 3D models from a text prompt, with optional PBR textures, adjustable face counts, and triangle or quad mesh topology.
- Does
- text → 3d
- Release date
- 2026-09
- Price
- $1.00
via fal
commercial
US-hosted ByteDance Seedance 2.5 animates still images with synchronized audio and optional end-frame control. Generate videos up to 30 seconds at 480p or 720p.
- Does
- image, text → video
- Release date
- 2026-09
- Price
- $0.5676 / second
via fal
commercial
US-hosted ByteDance Seedance 2.5 generates video with native audio from up to 30 images, 10 videos, and 10 audio references. Supports reference-guided generation, video editing, and extension at 480p or 720p.
- Does
- image, text → video
- Release date
- 2026-09
- Price
- $0.5676 / second
via fal
commercial
US-hosted ByteDance Seedance 2.5 generates cinematic video from text with synchronized audio, up to 30-second duration, and 480p or 720p output.
- Does
- text → video
- Release date
- 2026-09
- Price
- $0.5676 / second
via fal
commercial
Lyria 3.5 is Google DeepMind's latest music generation model, and you can generate almost any type of music with it
- Does
- text → music
- Release date
- 2026-09
- Price
- see supplier
via fal
commercial
H3 Max Lip Sync generates a video from an image and supplied audio, synchronizing mouth movements to the soundtrack. It supports optional transcription guidance and output resolutions from 480p to 2K.
- Does
- image, text → video
- Release date
- 2026-09
- Price
- $0.05 / second
via fal
commercial
Bria Product Holding edits a person photo to show the subject holding or carrying a product, using one to three product reference images and optional text instructions. Built on FIBO-Edit-1.5, it preserves the source aspect ratio by default and supports a selectable output aspect ratio.
- Does
- image, text → image-edit
- Release date
- 2026-09
- Price
- $0.04
via fal
commercial
Bria Virtual Try-On edits a person photo to show the subject wearing garments or accessories from one to three reference images, guided by optional text instructions. Built on FIBO-Edit-1.5, it supports multi-garment changes and preserves the source aspect ratio by default.
- Does
- image, text → image-edit
- Release date
- 2026-09
- Price
- $0.04
via fal
commercial
US hosted version of ByteDance's most advanced image-to-video model. Animate still images into cinematic video with synchronized audio, start and end frame control, and motion prompts.
- Does
- image, text → video
- Release date
- 2026-09
- Price
- $0.37 / second
via fal
commercial
US hosted version of ByteDance's most advanced reference-to-video model. Generate video from up to 9 images, 3 videos, and 3 audio clips with native audio and cinematic camera control.
- Does
- image, text → video
- Release date
- 2026-09
- Price
- $0.37 / second
via fal
commercial
US hosted version of ByteDance's most advanced text-to-video model. Cinematic output with native audio, multi-shot editing, real-world physics, and director-level camera control.
- Does
- text → video
- Release date
- 2026-09
- Price
- $0.37 / second
via fal
commercial
Create product, UGC, and presenter videos with synchronized native audio from text, with optional image and audio inputs.
- Does
- text → video
- Release date
- 2026-09
- Price
- $0.01 / second
via fal
commercial
Generate high quality, realistic music with fine controls using Elevenlabs Music v2!
- Does
- text → music
- Release date
- 2026-09
- Price
- $0.6 / output
via fal
commercial
Generate high quality, realistic music with fine controls using Elevenlabs Music v2.5!
- Does
- text → music
- Release date
- 2026-09
- Price
- $0.6 / output
via fal
commercial
Restyle a video’s scene, lighting, and visual style from edited keyframes while preserving the source subjects’ identity, expressions, gaze, and motion. Developed by Eyeline Labs and Netflix researchers.
- Does
- video, text → video-edit
- Release date
- 2026-09
- Price
- see supplier
via fal
commercial
Change a video’s lighting using a relit reference frame while preserving the scene, subjects, and original performance. ID-V2V Relight propagates the new illumination across the video.
- Does
- video, text → video-edit
- Release date
- 2026-09
- Price
- see supplier
via fal
commercial
H3 Max Multi Angle turns a single image into a video with precise, keyframe-based control over the camera's orbit, elevation, and distance in 3D space
- Does
- image, text → video
- Release date
- 2026-09
- Price
- $0.025 / second
via fal
commercial
FLUX.3 Edit Video [FAST] is Black Forest Labs' frontier video model. This endpoint edits an existing video from natural-language instructions, applying targeted changes while preserving the rest of the scene.
- Does
- video, text → video-edit
- Release date
- 2026-09
- Price
- $0.03 / second
via fal
commercial
FLUX Video Edit [fast] takes a source video and an edit prompt and returns a precisely edited video. Add, remove, or replace objects and characters, rebuild the setting, edit on-screen...
- Does
- text → video
- Release date
- 2026-09
- Price
- $6 / s
AI tokens — built in
commercial
GPT Image 2.5 Flare is an image generation and editing model from OpenAI, positioned as the speed-oriented tier of the GPT Image 2.5 series. It is suited to high-volume everyday...
- Does
- text, image → image, image-edit
- Release date
- 2026-09
- Price
- ≈ $0.077 / image
AI tokens — built in
commercial
GPT Image 2.5 Sunburst is an image generation and editing model from OpenAI, positioned as the precision-oriented tier of the GPT Image 2.5 series. It is suited to detailed creative...
- Does
- text, image → image, image-edit
- Release date
- 2026-09
- Price
- ≈ $0.077 / image
AI tokens — built in
commercial
Align the transcript and your audio recording using Elevenlab's forced alignment feature!
- Does
- audio → transcribe
- Release date
- 2026-09
- Price
- $0.22 / hour
via fal
commercial
Precise image editing that changes only what's asked, keeping subject, composition, and background intact, with reference subjects staying recognizable across styles and successive edits.
- Does
- image, text → image-edit
- Release date
- 2026-09
- Price
- $5.00
via fal
commercial
OpenAI's default image model for most applications. Fast, high-quality generation with natural lighting, rich textures, and support for complex layouts including transparent backgrounds.
- Does
- text → image
- Release date
- 2026-09
- Price
- $5.00
via fal
commercial
Editing built for the tightest control, edits scoped precisely to the instruction, with subject and composition preserved across many rounds of revision.
- Does
- image, text → image-edit
- Release date
- 2026-09
- Price
- $5.00
via fal
commercial
OpenAI's precision-focused image model, built for premium visual work, extra fidelity on intricate detail, in exchange for longer generation times.
- Does
- text → image
- Release date
- 2026-09
- Price
- $5.00
via fal
commercial
Upscale any image 2x or 4x, up to 8192×8192, with Bria Increase Resolution. Preserves the original content — no regeneration, no altered details. Commercial-safe
- Does
- image, text → image-edit
- Release date
- 2026-09
- Price
- $0.04 / image
via fal
commercial
MAI-Image-2.6 is an image generation and editing model from Microsoft AI, the precision tier of the MAI-Image-2.6 family alongside the faster [MAI-Image-2.6 Flash](/microsoft/mai-image-2.6-flash). It is suited for design-ready visuals and...
- Does
- text, image → image, image-edit
- Release date
- 2026-09
- Price
- ≈ $0.098 / image
AI tokens — built in
commercial
MAI-Image-2.6 Flash is the lower-latency, lower-cost member of the [MAI-Image-2.6](/microsoft/mai-image-2.6) family from Microsoft AI, built for latency-sensitive, high-throughput production image generation and editing at comparable quality to the precision tier....
- Does
- text, image → image, image-edit
- Release date
- 2026-09
- Price
- ≈ $0.049 / image
AI tokens — built in
commercial
Direct continuous, realtime video streams with live prompts while preserving characters, settings, and story continuity.
- Does
- text → video
- Release date
- 2026-09
- Price
- $0.04 / second
via fal
commercial
fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
- Does
- image, text → video
- Release date
- 2026-09
- Price
- $0.0125 / second
via fal
commercial
fal's H3 Max Turbo is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
- Does
- text → video
- Release date
- 2026-09
- Price
- $0.0125 / second
via fal
commercial
MiniMax H3 Max is a video-generation model from MiniMax, jointly released with fal.ai. Derived through additional training from MiniMax H3, it is designed for faster text-to-video and image-to-video generation with...
- Does
- text, image → video
- Release date
- 2026-09
- Price
- $0.1 / s (480p)
AI tokens — built in
commercial
Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- $0.42
via fal
commercial
Kling's Native 4K is a video generation model that directly outputs professional-grade 4K video in one step, eliminating the need for post-production upscaling
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- $0.42
via fal
commercial
fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
- Does
- image, text → video
- Release date
- 2026-08
- Price
- $0.05 / second
via fal
commercial
Remove background from videos filmed using chromakey, with automatic green spill suppression for clean, professional edges.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint edits video through natural-language instruction, applying the requested change while preserving the parts of the scene you want kept, and carrying character and scene consistency across successive edits.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- $0.03 / second
via fal
commercial
Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint animates a still image into video with synchronized audio, extending a single frame into coherent motion that reflects the logic of the real world.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- $0.03 / second
via fal
commercial
Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint generates video from combined multimodal references, images, videos and text together. Reasoning across all inputs to produce a single coherent result, with characters retaining their face, clothing, and voice throughout
- Does
- image, text → video
- Release date
- 2026-08
- Price
- $0.03 / second
via fal
commercial
Gemini Omni Flash 1.1 is Google's multimodal video model. This endpoint generates video with synchronized native audio from a text prompt, grounded in Gemini's real-world knowledge and physics understanding, with cinematic camera control expressed in natural language.
- Does
- text → video
- Release date
- 2026-08
- Price
- $0.03 / second
via fal
commercial
Wan 3.0 Prime is a fast-mode variant of Wan 3.0 from Alibaba. It supports text-to-video and first-frame image-to-video generation.
- Does
- text, image → video
- Release date
- 2026-08
- Price
- $0.136 / s (480p)
AI tokens — built in
commercial
Commercially safe, multi-reference image editing model. Follows natural language instructions alone or with up to 4 reference images, purpose-built for complex object and character combinations, virtual try-on, background replacement, style transfer, and more.
- Does
- image, text → image-edit
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
Meta's Muse Image model does precise edits that change only what you ask, stay coherent across turns, and compose from multiple reference images.
- Does
- image, text → image-edit
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
Meta's Muse Image model has faithful instruction-following and exceptional visual fidelity, with fine details like text, plots, and QR codes rendered accurately.
- Does
- text → image
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
- Does
- image, text → video
- Release date
- 2026-08
- Price
- $0.025 / second
via fal
commercial
fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality
- Does
- text → video
- Release date
- 2026-08
- Price
- $0.025 / second
via fal
commercial
Generates raster images that hold a consistent style, from either a saved style ID or reference images attached directly.
- Does
- text → image
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
Generates vector images that hold a consistent style, from either a saved style ID or reference images attached directly.
- Does
- text → image
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
Generates raster images that hold a consistent style, from either a saved style ID or reference images attached directly.
- Does
- text → image
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
Generates vector images that hold a consistent style, from either a saved style ID or reference images attached directly.
- Does
- text → image
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
Muse Image is an agentic image generation model from Meta that generates and edits images from text and reference images. Unlike single-pass image models, it reasons before it renders, breaking...
- Does
- text, image → image, image-edit
- Release date
- 2026-08
- Price
- see supplier
AI tokens — built in
commercial
Recraft V4 Styles is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's rendering...
- Does
- text, image → image, image-edit
- Release date
- 2026-08
- Price
- $0.07 / image
AI tokens — built in
commercial
Recraft V4 Styles Pro is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's...
- Does
- text, image → image, image-edit
- Release date
- 2026-08
- Price
- $0.2 / image
AI tokens — built in
commercial
Recraft V4 Styles Pro Vector is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the...
- Does
- text, image → image, image-edit
- Release date
- 2026-08
- Price
- $0.24 / image
AI tokens — built in
commercial
Recraft V4 Styles Vector is a style-consistent image generation model from Recraft. Every request requires at least one style reference image and generates a new image that reproduces the reference's...
- Does
- text, image → image, image-edit
- Release date
- 2026-08
- Price
- $0.1 / image
AI tokens — built in
commercial
Wan 3.0 Prime Image-to-Video turns still images into dynamic, cinematic sequences with rapid turnaround, natural motion, and excellent visual continuity. It preserves the identity, composition, and atmosphere of the source image while introducing expressive movement, camera dynamics, and richly deta
- Does
- image, text → video
- Release date
- 2026-08
- Price
- $0.068
via fal
commercial
Wan 3.0 Prime Reference-to-Video combines reference images, videos, and audio into a unified video with fast generation and strong multimodal coherence. It follows character identity, visual style, movement, and sound cues across references to create controlled, consistent, and production-ready resu
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- $0.068
via fal
commercial
Wan 3.0 Prime Text-to-Video transforms written prompts into polished videos with accelerated generation, fluid motion, strong scene fidelity, and coherent visual storytelling. Built for fast creative iteration, it brings complex ideas to life while preserving visual detail and cinematic consistency
- Does
- text → video
- Release date
- 2026-08
- Price
- $0.068
via fal
commercial
Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- $0.05
via fal
commercial
Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
- Does
- image, text → video
- Release date
- 2026-08
- Price
- $0.05
via fal
commercial
Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual coherence.
- Does
- text → video
- Release date
- 2026-08
- Price
- $0.05
via fal
commercial
Text-to-image model with high-fidelity outputs, accurate typography, and style preset, strong in photorealism, textures, and beyond. JSON-structured prompts give enterprise and agentic workflows production-ready control. Trained on licensed data.
- Does
- text → image
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
Generate realistic virtual try-on images from a person image and a clothing product image.
- Does
- image, text → image-edit
- Release date
- 2026-08
- Price
- $0.075 / generated
via fal
commercial
Turns text into a fully textured, PBR-ready 3D mesh with complete geometry, in game-ready Smart Topology at a target polygon count
- Does
- text → 3d
- Release date
- 2026-08
- Price
- $0.12 / call
via fal
commercial
Wan 3.0 is a video generation model from Alibaba for text-to-video, image-to-video, and reference-guided video generation. It produces 480p, 720p, or 1080p video with durations from 2 to 30 seconds.
- Does
- text, image → video
- Release date
- 2026-08
- Price
- $0.1 / s (480p)
AI tokens — built in
commercial
HeyGen: Avatar IV is an image-to-video model that animates a single photo into an expressive, lip-synced talking-head video. Rather than only matching mouth shapes to words, it interprets the vocal...
- Does
- text, image → video
- Release date
- 2026-08
- Price
- $0.1 / s
AI tokens — built in
commercial
Upscale videos to 1080p, 2K, or 4K via API. FLUX 3 powered super-resolution with a precise mode and a creative detail-enhancement mode.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- $0.14 / second
via fal
commercial
Generate 3D models from a single image with Hi3D V3.0.
- Does
- image → 3d
- Release date
- 2026-08
- Price
- $0.02 / credit
via fal
commercial
Generate 3D models from multiple view images using Hi3D V3.0.
- Does
- image → 3d
- Release date
- 2026-08
- Price
- $0.02 / credit
via fal
commercial
FLUX Video Upscale is a video upscaling model from Black Forest Labs. It enlarges a single source video by 1.5× to 3× while preserving its duration, with an optional prompt...
- Does
- text → video
- Release date
- 2026-08
- Price
- $15 / s
AI tokens — built in
commercial
Professional color and lighting correction powered by Topaz Labs. Adjust V2 fixes exposure, White Balance corrects color casts, Colorize adds color to black-and-white photos. Best for one-click photo correction.
- Does
- image, text → image-edit
- Release date
- 2026-08
- Price
- $0.08 / started
via fal
commercial
Professional video colorization powered by Topaz Labs. Brings natural color to black-and-white footage, upscaled to at least 1080p. Best for archival and historical clips.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- $0.10
via fal
commercial
Professional motion deblur powered by Topaz Labs. Themis 2 restores clarity to fast-moving, motion-blurred footage at source resolution. Best for sports and action footage.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- $0.10
via fal
commercial
Professional photo denoising powered by Topaz Labs. Normal, Strong and Extreme presets clean noise at source resolution; Denoise Max adds generative detail recovery. Best for high-ISO and night photography.
- Does
- image, text → image-edit
- Release date
- 2026-08
- Price
- $0.08
via fal
commercial
Professional video denoising powered by Topaz Labs. Nyx models remove noise at source resolution, with Nyx Fast as a lighter, cheaper pass. Best for low-light and high-ISO footage.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- $0.10
via fal
commercial
Professional frame interpolation powered by Topaz Labs. Apollo, Chronos and Aion retime footage up to 120 fps, from smooth motion to extreme slow motion. Best for fluid 60fps output and slow-motion effects.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- $0.30
via fal
commercial
Professional image restoration powered by Topaz Labs. Recover 3 generatively rebuilds natural detail; Dust-Scratch V2 cleans film dust and scratches. Best for old, damaged or degraded photos.
- Does
- image, text → image-edit
- Release date
- 2026-08
- Price
- $0.08 / started
via fal
commercial
Professional SDR-to-HDR conversion powered by Topaz Labs. Hyperion 2.5 redistributes luminance and color while preserving detail in text, faces and motion. Best for giving flat SDR footage a true HDR look.
- Does
- video, text → video-edit
- Release date
- 2026-08
- Price
- $2.40
via fal
commercial
Professional photo sharpening powered by Topaz Labs. Models tuned per blur type (lens, motion, portrait, wildlife), plus Super Focus for generative recovery of severely blurred shots. Best for out-of-focus and motion-blurred photos.
- Does
- image, text → image-edit
- Release date
- 2026-08
- Price
- $0.08
via fal
commercial
MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long
- Does
- text → music
- Release date
- 2026-08
- Price
- see supplier
via fal
commercial
Seedream 5.0 Lite is an image generation model from ByteDance Seed. It is suited for professional visual creation that benefits from web-connected retrieval, complex-prompt comprehension, visual references, and broad knowledge...
- Does
- text, image → image, image-edit
- Release date
- 2026-08
- Price
- $0.07 / image
AI tokens — built in
commercial
Seedream 5.0 Pro is an image generation and editing model from ByteDance Seed. It is suited for commercial visual-production workflows that require precise editing control, lifelike scenes, and natural rendering.
- Does
- text, image → image, image-edit
- Release date
- 2026-08
- Price
- $0.09 / image
AI tokens — built in
commercial
Seedance 2.0 Mini is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video with image, video, and audio inputs. It...
- Does
- text, image → video
- Release date
- 2026-08
- Price
- $0 / s
AI tokens — built in
commercial
Turns a single image into a fully textured, PBR-ready 3D mesh with complete geometry, in game-ready Smart Topology at a target polygon count
- Does
- image → 3d
- Release date
- 2026-08
- Price
- $0.12 / call
via fal
commercial
econstructs a high-fidelity textured 3D model from multiple angle views of one object, with game-ready topology and polygon control
- Does
- image → 3d
- Release date
- 2026-08
- Price
- $0.12 / call
via fal
commercial
Generate images from text using xAi's Grok Imagine 2.0 model.
- Does
- text → image
- Release date
- 2026-08
- Price
- $0.04
via fal
commercial
Grok Imagine Image 2.0 is an image generation and editing model from xAI. It is suited for creating images from text prompts and editing images from references, with low and...
- Does
- text, image → image, image-edit
- Release date
- 2026-08
- Price
- $0.08 / image
AI tokens — built in
commercial
Seedance 2.5 is a video generation model from ByteDance. It is suited for long-form storytelling, multimodal reference-based generation, video editing, and video extension. It supports first-frame and first-and-last-frame control, up...
- Does
- text, image → video
- Release date
- 2026-08
- Price
- $0 / s
AI tokens — built in
commercial
Generate 3D models from a single image with Hi3D.
- Does
- image → 3d
- Release date
- 2026-08
- Price
- $0.02 / credit
via fal
commercial
Generate 3D models from multiple view images using Hi3D.
- Does
- image → 3d
- Release date
- 2026-08
- Price
- $0.02 / credit
via fal
commercial
Qwen Image 3 is a unified image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with a richer world...
- Does
- text, image → image, image-edit
- Release date
- 2026-08
- Price
- $0.06 / image
AI tokens — built in
commercial
Qwen Image 3 Pro is an image generation and editing model from Qwen. It supports precise rendering of text and details as small as 10px, along with richer world knowledge...
- Does
- text, image → image, image-edit
- Release date
- 2026-08
- Price
- $0.08 / image
AI tokens — built in
commercial
FLUX.3 Video is a video generation model from Black Forest Labs. It supports text-to-video, image-guided generation with opening and closing keyframes, and video continuation workflows, making it suited for controlled...
- Does
- text, image → video
- Release date
- 2026-08
- Price
- $34 / s
AI tokens — built in
commercial
Generate high-fidelity, design-ready images with precise typography, strong prompt alignment, and rich visual detail using Microsoft's flagship MAI Image 2.5 Pro.
- Does
- text → image
- Release date
- 2026-07
- Price
- $7.50 / 1m
via fal
commercial
Generate natural multilingual speech from text with fast voice and language control using Qwen Audio 3.0 TTS Flash.
- Does
- text → speech
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
MiniMax H3 is a lightweight, open-weights video generation model from MiniMax. It is designed for precise multimodal editing and controlled content generation, including instruction-guided edits, text and brand rendering, and...
- Does
- text, image → video
- Release date
- 2026-07
- Price
- $0.08 / s
AI tokens — built in
commercial
Runway Aleph 2.0 is an in-context video editing model from Runway. It applies text instructions and keyframe-guided edits across existing footage while preserving details that are not meant to change....
- Does
- text → video
- Release date
- 2026-07
- Price
- $56 / s
AI tokens — built in
commercial
Runway Gen-4.5 is a video generation model from Runway for text-to-video and image-to-video workflows. It is designed for cinematic scene creation with strong motion quality, visual fidelity, and prompt adherence....
- Does
- text, image → video
- Release date
- 2026-07
- Price
- $24 / s
AI tokens — built in
commercial
Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.
- Does
- text, image → image, image-edit
- Release date
- 2026-07
- Price
- ≈ $0.279 / image
AI tokens — built in
commercial
Generates images from a text prompt at resolutions up to 2048×2048, with automatic prompt rewriting and prompt-guided resolution selection, building on Qwen's strength in complex text rendering and precise prompt adherence
- Does
- text → image
- Release date
- 2026-07
- Price
- $0.04 / generated
via fal
commercial
Generates high-quality, commercial-use-safe sound effects from a text prompt, with full control over type, texture, intensity, and exact duration.
- Does
- text → music
- Release date
- 2026-07
- Price
- $0.0018 / second
via fal
commercial
Krea 2 Large is Krea's high-capability image generation model, more than twice the size of Krea 2 Medium. Its lighter post-training gives images a rawer, more textured, and flexible character,...
- Does
- text, image → image, image-edit
- Release date
- 2026-07
- Price
- see supplier
AI tokens — built in
commercial
Krea 2 Medium is Krea's balanced, cost-efficient image generation model and a practical starting point for a broad range of use cases. Its extensive post-training supports stable, consistent generations, with...
- Does
- text, image → image, image-edit
- Release date
- 2026-07
- Price
- see supplier
AI tokens — built in
commercial
Krea 2 Medium Turbo is a distilled, speed-focused variant of Krea 2 Medium from Krea. It is designed for rapid iteration and graphic design exploration where fast generation is the...
- Does
- text, image → image, image-edit
- Release date
- 2026-07
- Price
- see supplier
AI tokens — built in
commercial
Grok Imagine Video 1.5 is a video generation model from SpaceXAI. It creates videos from text prompts, with an optional starting image to guide the scene. It can direct subject...
- Does
- text, image → video
- Release date
- 2026-07
- Price
- $2 / s
AI tokens — built in
commercial
Generate high-fidelity images from text with Krea 2 using a style reference image. Apply a reference image to guide the visual style into new generations, with aspect ratio, creativity, and seed controls.
- Does
- text → image
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
Run inference on LoRA adapters for TRELLIS.2 model
- Does
- image → 3d
- Release date
- 2026-07
- Price
- see supplier
via fal
commercial
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...
- Does
- text, image → image, image-edit
- Release date
- 2026-06
- Price
- ≈ $0.077 / image
AI tokens — built in
commercial
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...
- Does
- image, text → image
- Release date
- 2026-06
- Price
- $3 / 1M out
language-model key
commercial
Seed Audio 1.0 is a new audio model from Bytedance that can generate high-quality, natural sounding audio using text, reference audios or an image.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
HappyHorse 1.0 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up...
- Does
- text, image → video
- Release date
- 2026-06
- Price
- $0.198 / s (720p)
AI tokens — built in
commercial
HappyHorse 1.1 is a video generation model from Alibaba. It generates short videos from a text prompt, a single starting image, or a set of reference images, with output up...
- Does
- text, image → video
- Release date
- 2026-06
- Price
- $0.198 / s (720p)
AI tokens — built in
commercial
OpenAI's GPT Image 1 generates and edits images via the dedicated Images API. Features accurate text rendering, transparent backgrounds, and up to 16 reference images for edits.
- Does
- text, image → image, image-edit
- Release date
- 2026-06
- Price
- ≈ $0.103 / image
AI tokens — built in
commercial
A cost-efficient variant of GPT Image 1 for high-quality image generation at reduced latency and cost via OpenAI's dedicated Images API.
- Does
- text, image → image, image-edit
- Release date
- 2026-06
- Price
- ≈ $0.021 / image
AI tokens — built in
commercial
OpenAI's latest image generation model. Supports high-fidelity image generation and editing via the dedicated Images API.
- Does
- text, image → image, image-edit
- Release date
- 2026-06
- Price
- ≈ $0.077 / image
AI tokens — built in
commercial
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.
- Does
- image → 3d
- Release date
- 2026-06
- Price
- $0.1 / generation
via fal
commercial
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images. Do fast prototyping using the fast model.
- Does
- text → 3d
- Release date
- 2026-06
- Price
- $0.1 / generation
via fal
commercial
Generate professional-quality voiceovers in seconds with Async TTS Pro model text-based control over pauses, emphasis, and timing. Voice ids can be found at https://async.com/developer/voice-library
- Does
- text → speech
- Release date
- 2026-06
- Price
- $0.01 / 1
via fal
commercial
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...
- Does
- image, text → image
- Release date
- 2026-06
- Price
- $24 / 1M out
language-model key
commercial
Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced...
- Does
- image, text → image
- Release date
- 2026-06
- Price
- $6 / 1M out
language-model key
commercial
Text to Audio high-quality using LTX-2.3
- Does
- text → music
- Release date
- 2026-06
- Price
- $0.0024075 / megapixel
via fal
commercial
Text to Audio high-quality using LTX-2.3 with Lora
- Does
- text → music
- Release date
- 2026-06
- Price
- $0.0024075 / megapixel
via fal
commercial
Zonos2 is a text-to-speech model that clones a voice from a short sample and speaks naturally across many languages.
- Does
- text → speech
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Nemotron-ASR-Streaming is a multi lingual, streaming Automatic Speech Recognition (ASR) engineered to deliver high-quality multi lingual transcription across both low-latency streaming and high-throughput batch workloads.
- Does
- audio → transcribe
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Seed Speech developed by ByteDance, is a family of large-scale text-to-speech models capable of synthesizing speech that is virtually indistinguishable from human speech.
- Does
- text → speech
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Riverflow V2.5 Fast is the speed-optimized variant of Sourceful's Riverflow 2.5 lineup, best for production deployments and latency-critical workflows. The Riverflow 2.5 series is a unified text-to-image and image-to-image family...
- Does
- text, image → image, image-edit
- Release date
- 2026-06
- Price
- $0.038 / image
AI tokens — built in
commercial
Riverflow V2.5 Pro is the most powerful variant of Sourceful's Riverflow 2.5 lineup, best for top-tier control and quality-sensitive outputs. The Riverflow 2.5 series is a unified text-to-image and image-to-image...
- Does
- text, image → image, image-edit
- Release date
- 2026-06
- Price
- $0.26 / image
AI tokens — built in
commercial
Stable Audio 3 Medium audio inpainting is a 1.4 billion parameter latent diffusion model that fills in or reworks selected segments of a stereo track guided by text prompts, supporting single- and multi-segment editing.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium audio outpainting is a 1.4 billion parameter latent diffusion model that extends existing stereo audio beyond its original endpoint via causal continuation guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium audio-to-audio is a 1.4 billion parameter latent diffusion model that transforms an input audio clip into new stereo variations up to 6 minutes guided by a text prompt.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium Base audio inpainting is the foundational 1.4 billion parameter checkpoint for editing or filling selected stereo audio segments guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium Base audio outpainting is the foundational 1.4 billion parameter checkpoint that extends existing stereo audio with causal continuation guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium Base audio-to-audio is the foundational 1.4 billion parameter checkpoint that transforms input audio into new stereo variations up to 6 minutes guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium Base is the foundational 1.4 billion parameter text-to-audio checkpoint generating stereo music up to 6 minutes, intended as the unmodified base for custom fine-tuning workflows.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for safe commercial use.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music audio outpainting is a 459 million parameter latent diffusion model that extends music compositions beyond their original endpoint via causal continuation.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music Base audio inpainting is the foundational 459 million parameter checkpoint for editing or filling selected music segments guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music Base audio outpainting is the foundational 459 million parameter checkpoint that extends music tracks via causal continuation guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input music into new variations up to 2 minutes guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music Base is the foundational 459 million parameter checkpoint generating full music compositions up to 2 minutes from text prompts, intended as the unmodified base for fine-tuning.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small Music is a 459 million parameter latent diffusion model that generates full stereo music compositions up to 2 minutes from text prompts, lightweight enough for on-device deployment.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX audio outpainting is a 459 million parameter latent diffusion model that extends sound-effect tracks beyond their original endpoint via causal continuation.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX Base audio inpainting is the foundational 459 million parameter checkpoint for editing or filling selected sound-effect segments guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX Base audio outpainting is the foundational 459 million parameter checkpoint that extends sound-effect tracks via causal continuation guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input audio into new sound-effect variations guided by text prompts.
- Does
- audio → audio-edit
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as the unmodified base for fine-tuning.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Stable Audio 3 Small SFX is a 459 million parameter latent diffusion model that generates high-quality sound effects from text prompts, designed for on-device deployment on mobile phones and consumer laptops.
- Does
- text → music
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
TripoSplat is an open-source model from TripoAI / VAST AI Research that converts a single 2D image into high-quality 3D Gaussians using a novel learned density-control approach
- Does
- image → 3d
- Release date
- 2026-06
- Price
- see supplier
via fal
commercial
Microsoft AI's MAI-Image-2.5 is a high-quality image generation model available via Azure AI Foundry. It produces photorealistic and artistic images from text prompts with support for various aspect ratios.
- Does
- text, image → image, image-edit
- Release date
- 2026-06
- Price
- ≈ $0.121 / image
AI tokens — built in
commercial
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.
- Does
- image → 3d
- Release date
- 2026-05
- Price
- $0.4 / generation
via fal
commercial
Rodin V2.5 by Hyper3D generates realistic and production ready 3D models from text or images.
- Does
- text → 3d
- Release date
- 2026-05
- Price
- $0.4 / generation
via fal
commercial
Grok Imagine Image Quality is SpaceXAI's fast, high-fidelity image generation and editing model. It accepts text prompts and optional reference images, producing photorealistic outputs at 1K or 2K across a...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.1 / image
AI tokens — built in
commercial
Grok Imagine Video is SpaceXAI's fast, text-, image-, and reference-conditioned video generation model. It produces short videos (1–15 seconds, 24 fps) at 480p or 720p across seven aspect ratios -...
- Does
- text, image → video
- Release date
- 2026-05
- Price
- $0.4 / s
AI tokens — built in
commercial
Pixal3D turns a single image into a high-fidelity 3D model with detailed geometry and realistic textures.
- Does
- image → 3d
- Release date
- 2026-05
- Price
- $0.3
via fal
commercial
Recraft V4 Pro Vector is the vector (SVG) variant of Recraft V4 Pro. It supports text and image inputs and produces vector image output across multiple aspect ratios at the...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.6 / image
AI tokens — built in
commercial
Recraft V4 Vector is the vector (SVG) variant of Recraft V4. It supports text and image inputs and produces vector image output across multiple aspect ratios. Compared to the raster...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.16 / image
AI tokens — built in
commercial
Recraft V4.1 is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios, with...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.07 / image
AI tokens — built in
commercial
Recraft V4.1 Pro is an image generation model from Recraft tuned for high aesthetics. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.42 / image
AI tokens — built in
commercial
Recraft V4.1 Pro Vector is the vector (SVG) variant of Recraft V4.1 Pro, tuned for high aesthetics. It supports text and image inputs and produces higher-resolution SVG image output across...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.6 / image
AI tokens — built in
commercial
Recraft V4.1 Utility is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios, with typical generation...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.07 / image
AI tokens — built in
commercial
Recraft V4.1 Utility Pro is a general-purpose image generation model from Recraft. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios — double...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.42 / image
AI tokens — built in
commercial
Recraft V4.1 Vector is the vector (SVG) variant of Recraft V4.1, tuned for high aesthetics. It supports text and image inputs and produces SVG image output across multiple aspect ratios,...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.16 / image
AI tokens — built in
commercial
Recraft V3 is an image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios. Supports the following `image_config` parameters:...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.08 / image
AI tokens — built in
commercial
Recraft V4 is an image generation model from Recraft. It supports text and image inputs with image output at ~1K resolution across multiple aspect ratios. It delivers stronger compositional judgment,...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.08 / image
AI tokens — built in
commercial
Recraft V4 Pro is an image generation model from Recraft. It supports text and image inputs with image output at ~2K resolution across multiple aspect ratios, double the resolution of...
- Does
- text, image → image, image-edit
- Release date
- 2026-05
- Price
- $0.5 / image
AI tokens — built in
commercial
Kling v3.0 Pro is Kuaishou's premium video generation model, offering higher visual quality than the Standard tier. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for precise...
- Does
- text, image → video
- Release date
- 2026-04
- Price
- $0.224 / s
AI tokens — built in
commercial
Kling v3.0 Standard is a video generation model from Kuaishou. It supports text-to-video and image-to-video workflows, with first-frame and last-frame control for guided scene composition. Clips range from 3 to...
- Does
- text, image → video
- Release date
- 2026-04
- Price
- $0.168 / s
AI tokens — built in
commercial
Google's mid-tier video generation model balancing speed and quality. Veo 3.1 Fast generates high-quality video from text or image prompts with native synchronized audio, offering faster turnaround than Veo 3.1...
- Does
- text, image → video
- Release date
- 2026-04
- Price
- $0.16 / s (720p)
AI tokens — built in
commercial
Google's most cost-effective video generation model, designed for high-volume applications and rapid iteration. Veo 3.1 Lite generates 720p and 1080p video from text or image prompts with native synchronized audio...
- Does
- text, image → video
- Release date
- 2026-04
- Price
- $0.06 / s (720p)
AI tokens — built in
commercial
Cohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation
- Does
- audio → transcribe
- Release date
- 2026-04
- Price
- see supplier
via fal
commercial
GPT-5.4 Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...
- Does
- image, text, file → image
- Release date
- 2026-04
- Price
- $30 / 1M out
language-model key
commercial
Kling Video O1 is a video generation model from Kuaishou. It supports text and image inputs with video output, enabling text-to-video and image-to-video workflows. It is suited for cinematic content...
- Does
- text, image → video
- Release date
- 2026-04
- Price
- $0.224 / s
AI tokens — built in
commercial
Hailuo 2.3 is a video generation model from MiniMax. It accepts text prompts and reference images as input and generates video output, supporting both text-to-video and image-to-video workflows. It is...
- Does
- text, image → video
- Release date
- 2026-04
- Price
- $0.163 / s
AI tokens — built in
commercial
Newest audio model from Google introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.
- Does
- text → speech
- Release date
- 2026-04
- Price
- see supplier
via fal
commercial
Generate 3D models from text descriptions using Tripo H3.1.
- Does
- text → 3d
- Release date
- 2026-04
- Price
- $0.10
via fal
commercial
Generate 3D models from text descriptions using Tripo P1.
- Does
- text → 3d
- Release date
- 2026-04
- Price
- see supplier
via fal
commercial
Wan 2.7 is a video generation model from Alibaba. It supports text-to-video, image-to-video with first and last frame control, and reference-to-video, where multiple reference images guide the style and content...
- Does
- text, image → video
- Release date
- 2026-04
- Price
- $0.2 / s
AI tokens — built in
commercial
Seedance 2.0 is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It is particularly strong at preserving character consistency,...
- Does
- text, image → video
- Release date
- 2026-04
- Price
- $0 / s
AI tokens — built in
commercial
Seedance 2.0 Fast is a video generation model from ByteDance. It supports text-to-video, image-to-video with first and last frame control, and multimodal reference-to-video. It prioritizes generation speed and lower cost...
- Does
- text, image → video
- Release date
- 2026-04
- Price
- $0 / s
AI tokens — built in
commercial
30 second duration clips are priced at $0.04 per clip. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate...
- Does
- text, image → speech
- Release date
- 2026-03
- Price
- free
language-model key
free
Full-length songs are priced at $0.08 per song. Lyria 3 is Google's family of music generation models, available through the Gemini API. With Lyria 3, you can generate high-quality, 48kHz...
- Does
- text, image → speech
- Release date
- 2026-03
- Price
- free
language-model key
free
Alibaba's most advanced video generation model, supporting over 10 visual creation capabilities in a unified system. Wan 2.6 generates 1080p video at 24fps from text, images, reference videos, or audio,...
- Does
- text, image → video
- Release date
- 2026-03
- Price
- $0.08 / s (480p)
AI tokens — built in
commercial
ByteDance's next-generation audio-visual generation model with a 4.5B parameter Dual-Branch Diffusion Transformer architecture. Seedance 1.5 Pro generates video and audio simultaneously in a single unified pass — eliminating the timing...
- Does
- text, image → video
- Release date
- 2026-03
- Price
- $0 / s
AI tokens — built in
commercial
Google's state-of-the-art video generation model, built for maximum visual fidelity in final production cuts. Veo 3.1 generates high-quality 1080p video from text or image prompts with native synchronized audio —...
- Does
- text, image → video
- Release date
- 2026-03
- Price
- $0.4 / s
AI tokens — built in
commercial
OpenAI's flagship video generation model, delivering production-quality video with physics-accurate motion, synchronized audio, and world-state persistence across shots. Sora 2 Pro follows intricate multi-shot instructions while maintaining consistent spatial relationships...
- Does
- text → video
- Release date
- 2026-03
- Price
- $0.6 / s (720p)
AI tokens — built in
commercial
Generate speech with expressive and realistic voices from xAI
- Does
- text → speech
- Release date
- 2026-03
- Price
- $0.015 / 1000
via fal
commercial
Text to Speech Endpoint for Inworld's TTS-1.5 Max.
- Does
- text → speech
- Release date
- 2026-03
- Price
- $0.01 / 1000
via fal
commercial
Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines...
- Does
- text, image → image, image-edit
- Release date
- 2026-02
- Price
- ≈ $0.155 / image
AI tokens — built in
commercial
Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines...
- Does
- image, text → image
- Release date
- 2026-02
- Price
- $6 / 1M out
language-model key
commercial
Meshy-6 is the latest model from Meshy. It generates realistic and production ready 3D models.
- Does
- text → 3d
- Release date
- 2026-02
- Price
- see supplier
via fal
commercial
Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to create high-quality text-to-speech.
- Does
- text → speech
- Release date
- 2026-02
- Price
- see supplier
via fal
commercial
Generate speech from text prompts and different voices using the MiniMax Speech-2.8 Turbo model, which leverages advanced AI techniques to create high-quality text-to-speech.
- Does
- text → speech
- Release date
- 2026-02
- Price
- see supplier
via fal
commercial
Riverflow V2 Fast is the fastest variant of Sourceful's Riverflow 2.0 lineup, best for production deployments and latency-critical workflows. The Riverflow 2.0 series represents SOTA performance on image generation and...
- Does
- text, image → image, image-edit
- Release date
- 2026-02
- Price
- $0.04 / image
AI tokens — built in
commercial
Riverflow V2 Pro is the most powerful variant of Sourceful's Riverflow 2.0 lineup, best for top-tier control and perfect text rendering. The Riverflow 2.0 series represents SOTA performance on image...
- Does
- text, image → image, image-edit
- Release date
- 2026-02
- Price
- $0.3 / image
AI tokens — built in
commercial
Create detailed, fully-textured 3D models with text
- Does
- text → 3d
- Release date
- 2026-01
- Price
- $0.225 / generation
via fal
commercial
Generate 3D models from text prompts with Hunyuan 3D Pro
- Does
- text → 3d
- Release date
- 2026-01
- Price
- $0.375 / generation
via fal
commercial
Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model
- Does
- text → speech
- Release date
- 2026-01
- Price
- see supplier
via fal
commercial
Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice model
- Does
- text → speech
- Release date
- 2026-01
- Price
- see supplier
via fal
commercial
Create custom voices using Qwen3-TTS Voice Design model and later use Clone Voice model to create your own voices!
- Does
- text → speech
- Release date
- 2026-01
- Price
- see supplier
via fal
commercial
The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...
- Does
- text, audio → speech
- Release date
- 2026-01
- Price
- $20 / 1M out
language-model key
commercial
A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...
- Does
- text, audio → speech
- Release date
- 2026-01
- Price
- $4.8 / 1M out
language-model key
commercial
Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!
- Does
- audio → transcribe
- Release date
- 2026-01
- Price
- $0.008 / input
via fal
commercial
FLUX.2 [klein] 4B is the fastest and most cost-effective model in the FLUX.2 family, optimized for high-throughput use cases while maintaining excellent image quality. Pricing is based on the output...
- Does
- text, image → image, image-edit
- Release date
- 2026-01
- Price
- $0.028 / image
AI tokens — built in
commercial
Generate 3D human motions via text-to-generation interface of Hunyuan Motion!
- Does
- text → 3d
- Release date
- 2025-12
- Price
- see supplier
via fal
commercial
Generate 3D human motions via text-to-generation interface of Hunyuan Motion!
- Does
- text → 3d
- Release date
- 2025-12
- Price
- see supplier
via fal
commercial
Seedream 4.5 is the latest in-house image generation model developed by ByteDance. Compared with Seedream 4.0, it delivers comprehensive improvements, especially in editing consistency, including better preservation of subject details,...
- Does
- text, image → image, image-edit
- Release date
- 2025-12
- Price
- $0.08 / image
AI tokens — built in
commercial
Generate long speech snippets fast using Microsoft's powerful TTS.
- Does
- text → speech
- Release date
- 2025-12
- Price
- see supplier
via fal
commercial
Turn simple sketches into detailed, fully-textured 3D models. Instantly convert your concept designs into formats ready for Unity, Unreal, and Blender.
- Does
- text → 3d
- Release date
- 2025-12
- Price
- $0.375 / generation
via fal
commercial
FLUX.2 [max] is the new top-tier image model from Black Forest Labs, pushing image quality, prompt understanding, and editing consistency to the highest level yet. Pricing is as follows, [per...
- Does
- text, image → image, image-edit
- Release date
- 2025-12
- Price
- $0.14 / image
AI tokens — built in
commercial
Maya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise voice design.
- Does
- text → speech
- Release date
- 2025-12
- Price
- see supplier
via fal
commercial
FLUX.2 [flex] excels at rendering complex text, typography, and fine details, and supports multi-reference editing in the same unified architecture. Pricing is as follows, [per the docs](https://bfl.ai/pricing?category=flux.2): We charge $0.06...
- Does
- text, image → image, image-edit
- Release date
- 2025-11
- Price
- $0.12 / image
AI tokens — built in
commercial
A high-end image generation and editing model focused on frontier-level visual quality and reliability. It delivers strong prompt adherence, stable lighting, sharp textures, and consistent character/style reproduction across multi-reference inputs....
- Does
- text, image → image, image-edit
- Release date
- 2025-11
- Price
- $0.06 / image
AI tokens — built in
commercial
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...
- Does
- text, image → image, image-edit
- Release date
- 2025-11
- Price
- ≈ $0.155 / image
AI tokens — built in
commercial
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...
- Does
- image, text → image
- Release date
- 2025-11
- Price
- $24 / 1M out
language-model key
commercial
GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by GPT-5 Mini, with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...
- Does
- text, image → image, image-edit
- Release date
- 2025-10
- Price
- ≈ $0.021 / image
AI tokens — built in
commercial
GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by GPT-5 Mini, with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...
- Does
- file, image, text → image
- Release date
- 2025-10
- Price
- $4 / 1M out
language-model key
commercial
GPT-5 Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...
- Does
- text, image → image, image-edit
- Release date
- 2025-10
- Price
- ≈ $0.103 / image
AI tokens — built in
commercial
GPT-5 Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...
- Does
- image, text, file → image
- Release date
- 2025-10
- Price
- $20 / 1M out
language-model key
commercial
Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,...
- Does
- text, image → image, image-edit
- Release date
- 2025-10
- Price
- ≈ $0.139 / image
AI tokens — built in
commercial
Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation,...
- Does
- image, text → image
- Release date
- 2025-10
- Price
- $5 / 1M out
language-model key
commercial
Meshy-6-Preview is the latest model from Meshy. It generates realistic and production ready 3D models.
- Does
- text → 3d
- Release date
- 2025-10
- Price
- see supplier
via fal
commercial
An open source, community-driven and native audio turn detection model by Pipecat AI.
- Does
- audio → transcribe
- Release date
- 2025-04
- Price
- see supplier
via fal
commercial
Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.
- Does
- audio → transcribe
- Release date
- 2025-04
- Price
- see supplier
via fal
commercial
Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.
- Does
- audio → transcribe
- Release date
- 2025-04
- Price
- see supplier
via fal
commercial
Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.
- Does
- audio → transcribe
- Release date
- 2025-04
- Price
- see supplier
via fal
commercial
Leverage the rapid processing capabilities of AI models to enable accurate and efficient real-time speech-to-text transcription.
- Does
- audio → transcribe
- Release date
- 2025-04
- Price
- see supplier
via fal
commercial
Generate text from speech using ElevenLabs advanced speech-to-text model.
- Does
- audio → transcribe
- Release date
- 2025-02
- Price
- see supplier
via fal
commercial
[Experimental] Whisper v3 Large -- but optimized by our inference wizards. Same WER, double the performance!
- Does
- audio → transcribe
- Release date
- 2024-04
- Price
- see supplier
via fal
commercial
Alibaba's #1-ranked video model — 1080p output with synchronized native audio and multilingual lip-sync.
- Does
- text, image → video
- Price
- see supplier
cloud API
commercial
Tencent's 3D generation model — turn a prompt or a single image into a textured 3D mesh (GLB).
- Does
- text, image → 3d
- Price
- see supplier
cloud API
commercial
Tongyi's fast 6B model — eight steps to a photoreal or illustrated image, and it renders text in English and Chinese well. The same model as Z-Image Turbo above, on ComfyUI so it runs on a Linux GPU server as well as a Mac. Takes LoRAs. Apache 2.0.
- Does
- text → image
- Price
- free
on your Mac
free
Apache 2.0
The full, undistilled Z-Image — slower than Turbo, with real guidance: a negative prompt that works, more variety between seeds, and wider range of styles. It's the base that LoRAs are trained on. Apache 2.0. Runs on the managed ComfyUI.
- Does
- text → image
- Price
- free
on your Mac
free
Apache 2.0
An open 8.9B FLUX-family model from the community (lodestones), trained on an unfiltered dataset — it has no built-in content refusals, so what it makes is up to you and your prompt. Strong photography, illustration and anime, a real negative prompt, and LoRAs. Apache 2.0. Runs on the managed ComfyUI.
- Does
- text → image
- Price
- free
on your Mac
free
Apache 2.0
Chroma1-HD quantised to 4 bits so it fits a 16 GB Mac — the same open, unfiltered 8.9B model, with a small loss of fine detail. Real negative prompt, takes LoRAs. Apache 2.0. Runs on the managed ComfyUI.
- Does
- text → image
- Price
- free
on your Mac
free
Apache 2.0
The fast Chroma: distilled to about 8 steps with guidance baked in, and quantised to 4 bits for 16 GB machines. Same unfiltered model family; no negative prompt. Takes LoRAs. Apache 2.0. Runs on the managed ComfyUI.
- Does
- text → image
- Price
- free
on your Mac
free
Apache 2.0
Dubs an existing video: sparse-frame video-to-video that re-lips footage to new audio while keeping the original motion. Unlimited length via chunked continuation. Built on Wan 2.1 I2V.
- Does
- video, audio, image → video, video-edit
- Price
- free
on your Mac
free
Apache 2.0
Animates a portrait from audio at 768×768 in 8 steps — the best quality-per-VRAM in this category, and clearer licensing than LatentSync (whose weights are openrail++, not Apache).
- Does
- image, audio → video, video-edit
- Price
- free
on your Mac
free
Apache 2.0
Real-time lip sync — single-step latent inpainting rather than a diffusion sampler, so it is fast and small. Capped at 256×256, and the authors note some jitter and lip-colour drift.
- Does
- video, audio → video, video-edit
- Price
- free
on your Mac
free
MIT (code + weights)
Nothing matches that.
Examples are ours: generated with the model, hosted by us. Nothing on this page loads
from a vendor's server, so browsing it tells them nothing about you.
Two routes to sound. Through the endpoints above: any audio-taking model transcribes,
any audio-returning one speaks, at their token prices on the Language tab. In the app:
the engines below.
These run in the app on your own provider key or your own machine — provider rates and
local costs, not AI-token prices, and not callable at the endpoints above. Prices are per minute of
audio; per-character quotes convert at 900 characters a minute. Word-error rates come from one
independent harness.
No voice clips yet — we only publish samples we made and host ourselves.
Speech prices checked 2026-07-29. Ballpark
mid-tier rates; they move constantly. fal's registry read 2026-09-23 (1496 models scanned, newest per category kept).