Black Forest Labs' FLUX 3 generates video with synchronized sound — and the same model can drive robots
Black Forest Labs announced FLUX 3 on July 23: one model trained jointly on images, video, and audio that can produce up to 20 seconds of video with synchronized dialogue, sound effects, and music in a single pass. The company reports early evaluators preferred it over rival video models in most comparisons, and a variant called FLUX-mimic, built with Mimic Robotics, extends the same base model to predicting robot actions, with manufacturing partners including Audi testing it. Video and robotics features start in limited access; image generation enters early access in the coming weeks.
Why it matters: Creative generation and robot control converging on one architecture is a strong signal for anyone planning tooling in either space.