← Back
SiTech Team⏱️ 3 წთ. საკითხავი

Flux 3 Generates AI Videos with Native Audio — Up to 20 Seconds Long

Flux 3 Generates AI Videos with Native Audio — Up to 20 Seconds Long

Black Forest Labs released Flux 3 — the first model that generates video and audio simultaneously, up to 20 seconds. This is a major milestone in AI media generation.

Flux 3 Generates AI Videos with Native Audio

On July 23, 2026, German AI company Black Forest Labs (BFL) officially unveiled Flux 3 — their first multimodal model capable of generating videos with native, synchronized audio up to 20 seconds long. This is a watershed moment in AI media generation history, marking the transition from single-modality outputs to true multimedia AI creation.

What Is Flux 3 and Why Does It Matter?

Flux 3 is the third generation of Black Forest Labs' Flux model family. While previous versions (Flux 1 and Flux 2) generated static images and silent videos, Flux 3 adds the crucial audio dimension — producing complete media content from a single text prompt. Flux 1 gained widespread adoption for image generation, and Flux 2 introduced video but without audio. Flux 3 changes the game entirely by generating both video and audio in one unified pass.

The Technical Innovation: Self-Flow Architecture

Flux 3 uses a new architecture called Self-Flow, which ensures audio-visual synchronization. The model simultaneously predicts both visual frames and audio waveforms, producing 1080p resolution videos with natural sound. This differs from other models that generate audio as a separate post-processing step after video creation, which often results in desynchronized outputs.

Key Capabilities

Flux 3 can generate videos up to 20 seconds long with full audio at 1080p resolution. The model handles diverse scenarios from city street scenes with ambient sounds to complex action sequences with synchronized effects. Black Forest Labs claims Flux 3 outperforms all 8 rival models on standard benchmarks for video generation quality and audio-visual synchronization accuracy.

Benchmarks: How Does It Compare?

Black Forest Labs trained Flux 3 on an NVIDIA DGX cluster and reports significant improvements over both its predecessors and competitors. The model achieves state-of-the-art results on video quality metrics, audio synchronization accuracy, and prompt adherence. It surpasses models from OpenAI, Google, and Meta in combined video+audio generation benchmarks — a notable achievement given those companies' larger resources.

Availability and Open Access

Black Forest Labs plans to make Flux 3 available via cloud API first, followed by local deployment. Notably, they will release an open-weight version called Flux 3 Dev, enabling third-party developers to build their own applications and fine-tune the model for specific use cases. This decision positions them as a serious contender in the open media model space, competing with Stability AI and others.

Conclusion

Flux 3 represents a major leap in AI content creation. The ability to generate complete video and audio from a single prompt opens possibilities for marketing, content creation, film pre-visualization, and even robotics training simulations. As Black Forest Labs releases the open-weight version, we can expect an explosion of innovation in AI media generation, similar to what Stable Diffusion did for AI image generation.

📖 Source