FLUX 3 jointly learns images, video, and audio in one architecture, generating 20-second clips with native sound. Put it to work inside VideoGen workflows.
By Black Forest Labs · Announced July 2026; video generally available August 2026 · Multimodal frontier model; strong on facial expression, event-synced sound, and lip-synced multilingual dialogue
Real prompts, real outputs, generated with FLUX 3 in VideoGen.

“A blacksmith hammers glowing steel on an anvil in a dim workshop, each strike lands with a sharp metallic ring and a burst of sparks, bellows breathe in the background”
Generated with FLUX 3 in VideoGen

“Close-up of a barista finishing latte art, she looks up with a warm smile and says 'One rosetta, just for you', espresso machine hisses and cafe chatter hums behind her”
Generated with FLUX 3 in VideoGen

“A neon sign reading 'OPEN ALL NIGHT' flickers to life above a rainy diner entrance, each letter buzzes as it ignites, reflections shimmer in puddles on the sidewalk”
Generated with FLUX 3 in VideoGen
| Specification | Details |
|---|---|
| Max duration | 20 seconds per generation |
| Resolutions | HD (720p) and Full HD (1080p via upscaling) |
| Native audio | Yes: dialogue, sound effects, and ambient sound |
| Modes | Text-to-video, image-to-video, keyframes, video continuation, multi-shot scenes in one generation |
| Input types | Text, image references, start and end frames, up to 4 seconds of video and audio |
| Languages | 14+ lip-synced languages, including English, Chinese, Spanish, Hindi, and Japanese |
| Draft mode | Fast low-cost previews; approved drafts re-render at full quality |
| Open weights | FLUX 3 Dev backbone announced for later release |
| Elo Benchmark | Unknown |
The fastest and most affordable tier, built for drafts and quick iteration.
The balanced default: good quality video clips at an everyday cost.
Premium quality video clips for the content you publish.
MAX always integrates the state of the art. FLUX 3 is part of that set.
Every quality tier routes each request to the best model based on our internal evals, compliance requirements, and your intent. Requests stay fully compliant at every tier: sensitive data is never routed to models hosted by foreign adversaries such as China.
Sign up for VideoGen to open the AI playground and editor.
Start the storyboard to video or prompt to video workflow, or use the AI video tool inside any project.
Choose the tier that fits your needs and describe what you want to create.
VideoGen builds the finished video with visuals, narration, and captions, ready to export.
Developers can generate video clips through the VideoGen API with the same quality tier routing as the app, so requests reach models like FLUX 3 without managing provider keys.
curl -X POST "https://api.videogen.io/v1/tools/generate-video-clip" \
-H "Authorization: Bearer $VIDEOGEN_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt": "A drone shot over a coastal town at sunrise", "durationSeconds": 8}'
Generate video with synchronized sound and high-fidelity images from the same multimodal backbone, keeping style consistent across your assets.
“VideoGen is awesome for scaling video production and improving turnaround time... The complex process of video editing, which could take days or months, now takes minutes!”
“VideoGen solves the biggest pain points of video production—complexity, cost, and time. In just a few clicks, anyone can create professional, copyright-free videos.”
“VideoGen is the most underrated tool for content creators who want to put out the highest quality content in the shortest amount of time.”
Every VideoGen workflow uses models like FLUX 3 automatically to build complete videos with visuals, narration, music, and captions.
Sign in to generate images, video clips, and more with the models you choose.