FLUX 3 Video just went live — 20-second HD clips, native audio, 14-language lip-sync, and one feature no other model has: multi-shot in one prompt.

The reels with quick cuts — restaurant to sidewalk to the punchline — used to be three clips stitched together. This week that stopped being true.
FLUX 3 Video went live on August 4 — text-to-video and image-to-video up to 20 seconds, HD and 1080p, with native audio and dialogue with lip-sync in 14 languages.
The unlock nobody else has: switching multiple camera angles inside one generation, coherent through the cuts. Every other model asks you to render clips separately and edit them together.
Draft mode runs at $0.06 per second — a 20-second preview costs $1.20, and the full-quality render inherits the composition and motion of the version you approved.
Black Forest Labs — the team behind the FLUX image models most creators already use — opened FLUX 3 Video to general API access on August 4. It is their first product outside image generation, and the launch shipped with three modes: text-to-video, image-to-video with keyframes, and video continuation from up to four seconds of existing footage.
The pitch is that one model handles the whole reel — picture, motion, dialogue, ambient sound — in a single pass. It is what Grok Imagine 1.5 promised at the still-to-reel end. FLUX 3 pushes further by generating multiple shots inside a single prompt and keeping the take coherent through the cuts. That is the workflow shift; everything else is table stakes now.
Receipts: Black Forest Labs launch post (Aug 4) · RuntimeWire launch analysis · FLUX 3 docs.
The following images were generated using Nano Banana 2:
Write your prompt as a shot list, not a sentence. Multi-shot only works when you tell the model where the cuts are. Describe every angle explicitly — Shot 1: overhead close-up on hands slicing bread. Shot 2: pull back to wide of the kitchen. Shot 3: over-the-shoulder as she carries the plate outside. FLUX 3 uses that structure to plan the take. A single-paragraph description gets you one continuous shot; a shot list gets you a real reel.
Draft first, render second. A 20-second draft costs $1.20 versus roughly $3.40 for the same length at full HD. Iterate on the draft until the subjects, composition, and motion feel right — then submit that draft for the full render. The final inherits the exact take you approved, so you are not gambling on a new roll of the dice.
Put the spoken line inside the prompt. You do not need a separate text-to-speech leg or a Wav2Lip pass. Write the dialogue in quotes inside the shot description — She looks straight at camera and says, "You are going to want to see this." — and FLUX 3 generates the voice AND the mouth movement in the same pass. Specify the language explicitly if it is not English; lip-sync accuracy depends on it.
Pin your face with a keyframe. For creator-avatar work, use image-to-video mode and set a locked reference frame as your first keyframe (an end frame is optional). Identity carries across the multi-shot cuts, which is the piece that used to break every longer AI reel. This is the workflow that keeps your avatar recognizable through 20 seconds of angle changes.

Draft vs Full HD prompt:
"Product still, two identical monitors side by side on a light-grey concrete desk. Left monitor shows an early-draft AI video frame of a woman in a grey jacket in a modern corridor, softer detail. Right monitor shows the same frame at full HD, crisper. Small printed paper price-tag stickers on each bezel. High-key daylight, minimalist depth, editorial."

Multi-shot prompt (shot list):
"Shot 1: warm indie talk-show set, medium shot of the host in a cream boucle chair, cinema camera visible in the foreground. Shot 2: pull back to wide as a second locked camera reveals the coffee table and books. Shot 3: over-the-shoulder cut from behind the host. Continuous take, warm amber+cream palette, hard warm key with soft fill. 20 seconds, HD."

Dialogue-in-prompt (Hindi lip-sync):
"Bright modern classroom, medium close-up on a 60-something South Asian male teacher in a dark grey blazer, gesturing mid-word. He looks straight at camera and says in Hindi: ‘यही वो बदलाव है जो हमें चाहिए।’ Natural daylight from a large window, cream and terracotta palette, precise lip-sync in Hindi. 8 seconds, HD."
The comparison to Seedance 2.0 matters less than the shape of the workflow change. Every other model in this class — Kling 2.6, Wan 2.6, Seedance 2.5, Runway Gen-4.5 — still asks you to render each shot separately and edit them together. FLUX 3 Video collapses shot planning, character continuity, and dialogue into one prompt. If Black Forest Labs ships the promised FLUX 3 Dev open-weight variant later this year, the same workflow lands on any desktop with a GPU. That is the shift; the benchmark preference rates are the noise.
Multi-shot in one prompt is the feature to try first, and the only way to feel it is on a reel where the cuts carry information — a recipe change, a room walk-through, a punchline reveal. Everything else FLUX 3 does is fair-game against Seedance 2.5, Kling, and Grok Imagine 1.5, but the shot-list workflow is genuinely new. If you are still shooting to a talking-head format and want to lock a face across those cuts, start with your avatar — identity consistency is the piece FLUX 3 assumes you already solved.
How is FLUX 3 Video priced?
Draft mode is $0.06 per second — $1.20 for a 20-second preview. Full HD render runs $0.17 per second, or $0.85 for a 5-second HD clip. Draft-to-render preserves the composition, so you only pay full price once per direction.
Can it actually lip-sync in Hindi or Japanese?
Yes. Supported languages at launch: English (multiple dialects), Chinese, Spanish, French, German, Japanese, Portuguese, Russian, Italian, Indonesian, Turkish, Hindi, and Punjabi — with precise lip-sync. Black Forest Labs is adding more.
How does it compare to Seedance 2.5?
Black Forest Labs claims FLUX 3 ties Seedance 2.0 in image-to-video and beats it in text-to-video, based on internal human-rater evals. Seedance 2.5's advantage is 30-second single-take length and region-level editing. FLUX 3's advantage is multi-shot inside one generation and native dialogue in the same model.
Is FLUX 3 Image also out?
Not yet. FLUX 3 Image is on the roadmap but has not shipped — video is the first generally-available modality from the multimodal FLUX 3 backbone. FLUX 3 Dev, an open-weight variant, is also planned.
What does Draft mode actually save me?
If you iterate through three drafts before committing to the final render, you pay $1.20 × 3 + $3.40 = $7.00 for a finished 20-second clip. Doing the same iteration at full HD would cost about $13.60. Roughly half the total spend for the same creative process.