Black Forest Labs, the team behind the FLUX image models, released FLUX 3, its first attempt at a genuinely multimodal foundation model rather than a still-image specialist. Instead of bolting a separate video or audio model onto an image generator, FLUX 3 is trained jointly across images, video, and audio from the ground up, and the headline capability is generating up to 20 seconds of video with audio that's synchronized to what's happening on screen, dialogue, sound effects, and ambient noise all generated together rather than dubbed on afterward. The model also handles multi-turn edits, keeps characters visually consistent across a sequence using reference images, and is described as particularly strong at rendering facial expressions and tying sounds to physical events. The most interesting piece for builders is the rollout strategy and the robotics angle. FLUX 3 Video and a companion FLUX 3 Action mode launched in gated early access, meaning you apply for API access rather than getting immediate self-serve availability, while FLUX 3 Image and an eventual open-weight FLUX 3 Dev release are promised for later in the year. Alongside that, Black Forest Labs partnered with mimic robotics to introduce FLUX-mimic, which reuses FLUX 3's video-prediction backbone with a lightweight decoder that turns predicted video frames into robot motion; it's already being tested by Audi for flexible door-seal installation on a production line. That's a meaningful signal about where world-model style video generation is heading, not just content generation but action prediction for physical systems, and it means teams evaluating FLUX 3 today are really evaluating an early-access, invite-gated product with a promised but not yet shipped open-weight tier, so plan integration timelines accordingly rather than assuming immediate self-hosted access.