One model, multiple senses
FLUX 3 learns images, videos and sounds together. According to the idea, movement, sight and sound are different impressions of the same reality, so they can be better harmonized in a common model.
What does the video part know?
The early version creates videos of up to 20 seconds with native audio from text, images or video references. It supports video continuation, keyframes, multi-language dialogue, and sequential cutscenes.
And the robot?
The FLUX‑mimic version uses the internal world knowledge of the same basic model to predict robot movements. The joint system of Black Forest Labs and mimic robotics is also being tested in Audi's production environment.
Can everyone use it now?
Not yet. FLUX 3 is in early access, image, video, sound and action features are planned to be introduced gradually. The published comparisons are preliminary and the manufacturer's measurements, so we leave the confetti cannon secured for the time being.
Why is this interesting for video creators?
When image, motion, and sound are produced in the same system, there can be fewer separate steps to create a scene. In theory, it can be easier to coordinate the movement of the character with the sound of the environment, to move a character through several settings, or to build a scene with sound from a keyframe. This can be especially useful for creators who today move the same material between several separate applications.
But early access also means demos don't equal day-to-day reliability. Character continuity, accurate speech, longer cutscenes, and small visual details still need a lot of testing. After the spectacular demo, it is therefore always worth asking: will the tenth generation succeed as well?
What will we watch in the public version?
In terms of real-world use, price, resolution, generation time, commercial rights, and fidelity to reference images will be at least as important as the prettiest demo video. It is also a crucial question to what extent it is possible to fix a faulty part of a scene without having to redo the entire clip.
Job interview with FLUX 3
"What position did you apply for?" – Image generator, film director, cinematographer, sound engineer and robotics intern. "That's five separate positions!" "I know, that's why I brought five resumes, all generated by me."
On the test day, he makes a twenty-second film, writes the sound, adjusts the lights, and then teaches a robot to make coffee. At first, the robot puts the mug next to the machine, but at least it makes mistakes in a cinematic way, with native sound effects and perfect camera movement.
