Black Forest Labs launches FLUX 3, a multimodal model that generates 20-second video with native audio and predicts robotic actions
FLUX 3, announced by Black Forest Labs on July 23, 2026, trains on images, video, and audio simultaneously in a single architecture rather than routing each format through a separate pipeline. It is BFL's first video-capable model and the first to ship audio generated natively alongside the visual content, per the BFL announcement. The same underlying architecture also drives a robotics action module, FLUX 3 Action, already in production testing at Audi.
One model, four output types
FLUX 3 is built on Self-Flow, BFL's research framework for aligning multimodal generation and understanding within a shared transformer backbone. Dedicated encoders and decoders convert images, video, and audio into a unified internal representation and back into outputs, per the BFL announcement. The video component supports text-to-video, image-to-video, video-to-video transformation, keyframe-based transitions, multilingual dialogue, and agentic chaining of individual clips into longer multi-shot sequences across a wide range of visual styles.
The robotics component, FLUX 3 Action, finetunes the video backbone for physical action prediction. BFL built it with Mimic Robotics under the name FLUX-mimic. Per the BFL blog, FLUX-mimic is already in testing on production manipulation tasks at Audi.
What BFL's evaluations show
BFL ran its own early preference tests against eight video generation models using 10-second clips at 720p. No independent benchmarks have been published. In BFL's tests, FLUX 3 was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93%, per the BFL blog post. The margin narrowed against tougher competition: FLUX 3 edged Kling v3 Pro at 60%, Happy Horse v1 at 59%, and both Seedance 2.0 and Gemini Omni Flash at 52%. BFL described all results as preliminary and said further improvements are expected before the full release.
The Decoder reported that a 52% preference rate against Seedance would place FLUX 3 among the top video generation models, while noting BFL's caveats about the early state of the results.
Rollout schedule
FLUX 3 Video is in early access now. FLUX 3 Image, covering synthesis and editing across styles, aspect ratios, and languages, is due in early access within the coming weeks. FLUX 3 Action is available to selected research and commercial partners, starting with Mimic Robotics. FLUX 3 Dev, an open-weight release for both content creation and action prediction, is planned for later in 2026, per BFL.
Why it matters
Training audio into the same model that generates video is architecturally distinct from layering sound tracks onto finished clips after the fact. The model learns causal links between visual events and the sounds they produce during training. BFL says FLUX 3 Video is "particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities," per its announcement. Whether those learned associations produce a detectable output advantage over tools that composite audio separately is what community evaluations in the next few weeks will determine.
The robotics path is an unusual bet for an image-generation company. BFL is using the same trained backbone for creative media and for teaching physical robots new manipulation tasks. Audi is the first named production test site. If the FLUX 3 video backbone proves transferable to new tasks with limited finetuning data, BFL would hold a position in physical AI training pipelines, not only in media workflows.
Hugging Face community evaluations are expected within one to two weeks and will be the first independent check on BFL's self-reported preference margins against Runway, Kling, and Sora.
Sources
- FLUX 3 - Real World Models: Black Forest Labs blog, July 23, 2026 (primary)
- Flux 3 generates videos with native audio up to 20 seconds long: The Decoder, July 23, 2026 (secondary)
