FLUX 3 hits 20-second video, and now wants to give your robot actual hands 🤖
Back to feed

FLUX 3 hits 20-second video, and now wants to give your robot actual hands 🤖

By our Markets Desk2 min read

Black Forest Labs released FLUX 3 on Thursday, marking the first time the German AI lab's flagship model generates video rather than still images. The new system was trained on images, video, and audio together inside one shared architecture, a setup known as multimodality. FLUX 3 produces clips up to 20 seconds long, with audio generated alongside the picture and synced to on-screen action, including dialogue, sound effects, and ambient noise.

In early evaluations, human reviewers preferred FLUX 3's output over Runway Gen-4.5 in 77% of head-to-head comparisons and over Luma Ray 3.2 in 93% of pairings. The model edged out Gemini Omni and Seedance, winning 52% of those matchups. Black Forest Labs noted the results are based on preference tests rather than a fixed scoring rubric, with evaluators simply choosing the clip that looked and sounded more convincing. The model also retains strong still-image capabilities, with BFL sharing samples it described as versatile across styles beyond photorealism.

Black Forest Labs, founded in August 2024 by researchers who helped build the original Stable Diffusion models at Stability AI, has framed the release as more than a content tool. "A model that only learns images can only generate images," said co-founder and CEO Robin Rombach. The company's bet is that learning to predict video also teaches the underlying physics—weight, contact, and timing—that a machine needs to operate in the physical world. That bet has a product name: FLUX-mimic, built with Zurich-based mimic robotics. It pairs FLUX 3's video-prediction engine with a lightweight decoder that translates the model's internal sense of motion into robot commands. Car maker Audi is already testing the system on tasks such as fitting flexible door seals. "Audi represents the kind of manufacturing partner we built FLUX-mimic for," said mimic co-founder Stephan-Daniel Gravert. Audi's Christoph Schneider said the robots now "solve complex soft-body manipulation work" that older machines couldn't handle. Black Forest Labs said the full system reacts in about 101 milliseconds, comparable to human visual reflexes.

The lab's earlier Flux models had already beaten MidJourney and topped the Artificial Analysis image arena by October with FLUX 1.1 Pro, while its open-source Dev and Schnell releases held the "best open-source image generator" title that Stability AI's Stable Diffusion 3.5 never reclaimed.

Share:
Publishercryptonewsroom.xyz
Published
CategoryMarkets

Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.

See our Terms of Service, Privacy Policy, and Editorial Policy.