Summary

  • Black Forest Labs has introduced FLUX 3 in early access, marking its first model capable of generating video clips of up to 20 seconds, complete with synchronized audio.
  • The underlying technology also supports FLUX-mimic, a robotics model currently being tested by Audi on its production lines.
  • Only the open-weight “Dev” version is expected to be available in late 2026; Video and Action features remain accessible through APIs and select partners for now, with Image capabilities following soon.

On Thursday, Black Forest Labs unveiled FLUX 3, which is notable for being the company’s first model to produce video rather than solely still images. The German AI laboratory, recognized for its FLUX series of image generators, has trained this new model on a combination of images, video, and audio within a unified framework.

This approach is referred to as multimodality, where a single model processes various types of information concurrently rather than relying on separate tools.

The highlight of FLUX 3 is its ability to generate video clips that last up to 20 seconds, with audio synchronized to the visuals—incorporating dialogue, sound effects, and background noise. Initial assessments indicate that human reviewers favored FLUX 3’s output over Runway Gen-4.5 in 77% of comparisons and outperformed Luma Ray 3.2 in 93% of cases. It also slightly surpassed Gemini Omni and Seedance, winning 52% of evaluations against those models.

It’s important to note that these results stem from preference tests, where evaluators watch two clips and select the one they find more convincing, with BFL tracking the number of wins for FLUX 3.

BFL emphasizes that this model serves a purpose beyond merely creating content. “A model that only learns images can only generate images,” remarked co-founder and CEO Robin Rombach. The company believes that by learning to predict video, it also acquires an understanding of the underlying physical principles—such as weight, contact, and timing—that are essential for machines to navigate the physical world effectively.

This concept is embodied in FLUX-mimic, developed in collaboration with Zurich-based mimic robotics. This system leverages FLUX 3’s video-prediction capabilities and integrates a lightweight “decoder” that translates the model’s internal understanding of movement into actual robotic actions. Audi is already testing this technology for tasks like installing flexible door seals, which traditional automation has found challenging.

“Audi represents the kind of manufacturing partner we designed FLUX-mimic for,” stated mimic co-founder Stephan-Daniel Gravert. Audi’s Christoph Schneider noted that the robots are now capable of “solving complex soft-body manipulation tasks” that older machines struggled with. BFL claims the entire system responds in approximately 101 milliseconds, comparable to human visual reflexes.

The emergence of FLUX has not occurred in isolation. Founded in August 2024 by experienced researchers from the original Stable Diffusion models at Stability AI, Black Forest Labs has released Flux models that have outperformed MidJourney and surpassed Stability’s less impressive Stable Diffusion 3.

The open-source Flux Dev and Schnell models were awarded the title of “best open source image generator,” a designation that AI artists had anticipated Stable Diffusion 3.5, Stability’s revised version, would eventually reclaim.

This did not happen. FLUX 1.1 Pro went on to dominate the Artificial Analysis image arena that October, though it was not open source.

BFL introduced FLUX.2 in November 2025, but it did not achieve the same level of popularity. The original Flux’s open-source status continued until Alibaba’s Z-Image Turbo overtook it in late 2025, matching its quality on lower-end consumer graphics cards. “This is what SD3 was supposed to be,” commented a CivitAI user at the time.

FLUX 3 represents BFL’s resurgence, although it is not yet fully accessible. Video and Action functionalities are currently in early access through APIs and select partners, including mimic robotics, while image generation capabilities are expected to follow “in the coming weeks,” according to BFL. The open-weight Dev version, which is the only tier planned for local use, is not expected to launch until late 2026.

Daily Debrief Newsletter

Start your day with the latest news highlights, along with original features, podcasts, videos, and more.