Black Forest Labs’ first video model is now available through its API and selected partners. Its most important feature may not be 20-second clips, but the attempt to generate video, sound, motion, and physical continuity from one multimodal system.
Picture an editor building a product demonstration from AI-generated clips. The video looks good, but the problems arrive in the gaps: the hammer strikes before the impact sound, a character’s voice slips out of sync with their mouth, and the camera changes position between shots. Each defect is small. Together, they make the footage unusable.
That is the production problem Black Forest Labs is targeting with FLUX 3 Video.
The company’s first video-generation model became generally available on August 4, 2026, through the Black Forest Labs API and selected partners. It can generate video clips up to 20 seconds long, with audio created alongside the image sequence. The initial release supports text-to-video, image-to-video, keyframes, first-and-last-frame generation, video continuation, multiple shots, dialogue, and multilingual speech. Black Forest Labs’ launch announcement describes it as the first generally available video capability in the wider FLUX 3 family.
That availability is a meaningful change from the company’s July 23 announcement, which introduced FLUX 3 as an early-access multimodal model. The broader FLUX 3 family still includes planned image-generation, action-prediction, and open-weight components. The general release today is specifically for FLUX 3 Video, not the complete model family. The original FLUX 3 announcement laid out that staged rollout.
One model for the whole shot
Most AI video workflows still treat a clip as a visual object first. Sound may be generated by another model, added later, or treated as a separate production layer. Character references, keyframes, and continuation tools may also depend on separate systems.
Black Forest Labs is presenting FLUX 3 differently. The company says the model was trained jointly across images, video, and audio within one architecture. Its premise is that these are not unrelated types of data. They are different observations of the same world: images describe spatial relationships, video adds time and movement, and audio reveals events such as impacts, friction, speech, and machinery. Black Forest Labs’ FLUX 3 overview explains the approach in those terms.
That does not mean FLUX 3 has solved physical simulation. It means the company is betting that a model exposed to these modalities together can learn stronger relationships between them.
A falling object should accelerate. A glass should make a sound when it breaks. A person’s voice should correspond to their mouth. A camera move should preserve the relative position of foreground and background objects. These relationships are what make generated video feel coherent rather than merely attractive frame by frame.
The underlying research direction is called Self-Flow. Black Forest Labs describes it as a self-supervised flow-matching framework that combines representation learning with generation. Its technical explanation argues that a generative model can learn stronger semantic representations when it is trained to infer missing information rather than only denoise a complete input. The research reports improvements across image, video, and audio generation in its own experiments. The Self-Flow research report provides the technical details.
For creators, the architectural story matters only if it produces better shots. The practical test is whether the model can preserve an event from beginning to end while keeping the image, motion, dialogue, and sound aligned.
FLUX 3 Video is here.
— Black Forest Labs (@bfl_ai) August 4, 2026
Serious, fun, creative, real, cinematic, whatever you need it to be.
Native audio, Text to Video, Image to Video with multiple frames, video continuation, dialogue in multiple languages.
Comes with Draft mode so you can explore ideas fast at a fraction of⦠pic.twitter.com/VPT8diaF4j
What FLUX 3 Video can do
The most basic workflow is text-to-video. Users can describe a scene in natural language, from a short idea to a detailed sequence. Black Forest Labs says the model can interpret complex prompts, move between scenes and camera angles, and generate video that is not limited to a polished cinematic style.
That stylistic range is part of the company’s positioning. The FLUX 3 model page shows examples spanning documentary footage, animation, comedy, underwater scenes, stop-motion, product imagery, and graphic design. The goal is not to produce one standardized AI-movie look. It is to handle different visual languages within the same system.
Image-to-video offers a more controlled starting point. A creator can supply a still image and describe the movement that should follow. That is useful for animating concept art, product photography, character designs, or a carefully composed first frame.
Keyframes extend that idea. Users can provide a starting frame, an ending frame, or multiple frames that define the intended progression of a shot. Rather than asking the model to invent every moment from a paragraph, the creator gives it visual anchors and asks it to generate the transitions.
The launch also supports continuation from an existing video and audio clip of up to four seconds. FLUX 3 is intended to carry forward movement, camera behavior, dialogue, and sound across the seam. That could make it useful for extending a performance, lengthening a product shot, or building a longer sequence from shorter generated pieces.
The model can also create multiple scenes and camera angles within one generation. This is different from asking for one continuous camera move. It gives the model room to construct a small sequence: an establishing shot, a close-up, an action beat, and a reaction. The feature is promising for explainers, advertisements, trailers, and short-form storytelling, although it also creates a larger consistency problem. Every cut introduces another opportunity for a character, prop, or lighting setup to drift.
Audio is generated with the video rather than added afterward. Black Forest Labs says FLUX 3 can produce dialogue, sound effects, and ambience alongside individual frames. The system supports multiple languages and dialects, with lip-syncing included in the company’s description of the release.
That makes the model more useful for a finished draft than a silent visual generator. A filmmaker can assess not only whether the shot looks right, but whether the action carries the right sonic meaning. A chef’s knife, a train passing a window, a street vendor’s wok, or a character speaking in a particular language can all be part of the initial generation.
The release also includes Draft Mode. A draft gives creators a faster, lower-cost preview. When a draft is selected, FLUX 3 can render it at full quality while preserving the same subjects, composition, and motion. fal’s FLUX 3 product page describes the same workflow as a draft preview followed by a full-quality enhancement.
That is a practical feature, not a cosmetic one. Video generation is expensive because many attempts are discarded. A preview system changes the creative loop from waiting for a final render and then discovering the shot is wrong to testing several directions, choosing one, and spending more on the version that works.
The 20-second question
Twenty seconds is long by the standards of many current video-generation systems, but it is not long-form video.
It is long enough for a complete social clip, a product beat, a short dialogue exchange, or a mini narrative with a beginning and an end. It is also long enough to expose continuity failures that remain hidden in five-second generations.
Black Forest Labs says longer pieces can be created by chaining clips together. That makes FLUX 3 more like a shot-generation system than an automatic replacement for an edit suite. The model may produce a self-contained shot, but the creator still needs to decide when to cut, how to arrange scenes, how to manage pacing, and whether the same character or environment remains convincing across generations.
The distinction between the model and the host platform also matters. Black Forest Labs describes the initial release as HD 720p, with Full HD output available through upscaling. On fal, the public FLUX 3 text-to-video endpoint lists 720p and 1080p settings, 5-to-20-second durations, and pricing of $0.17 per second at 720p and $0.29 per second at 1080p. Those are fal’s published rates and implementation details, not necessarily universal pricing for every FLUX 3 partner.
How it compares with other video models
FLUX 3 is entering a market that already includes models with overlapping capabilities.
Google’s Veo 3.1 also emphasizes native audio, dialogue, physical realism, and creative controls. Its public materials highlight scene extension, first-and-last-frame workflows, and internal human-rater comparisons. Google DeepMind’s Veo page presents Veo as a filmmaker-oriented system with integrated sound.
ByteDance’s Seedance 2.0 takes an especially broad reference-based approach. Its official announcement says the model accepts text, images, audio, and video as inputs, supports up to nine images, three video clips, and three audio clips, and can generate 15-second multi-shot audio-video output. ByteDance’s Seedance 2.0 announcement positions it as a multimodal editing and reference system.
Runway’s Gen-4.5 is more narrowly specified in its current public documentation. It supports text-to-video and image-to-video, with 2-to-10-second generations at 720p. Its documentation emphasizes prompt adherence, camera choreography, motion quality, and visual control. Runway’s Gen-4.5 documentation describes the model’s current public workflow.
The comparison is less about one universal quality ranking and more about production philosophy. FLUX 3 combines relatively long single generations, synchronized audio, keyframe control, continuation, and draft-to-final rendering in one API-facing workflow. Seedance 2.0 emphasizes multimodal references. Veo emphasizes audio and filmmaking controls. Runway emphasizes visual control and prompt-directed motion.
Black Forest Labs has published internal evaluations claiming that FLUX 3 performs strongly against named competitors. The company’s launch materials say internal human raters found FLUX 3 ahead in text-to-video and tied with Seedance 2.0 in one image-to-video comparison. Those findings are company-reported and should not be treated as an independent leaderboard. The results also do not answer practical questions about cost, reliability, generation speed, safety, or performance across difficult production footage.
What still needs to be tested
The public demonstrations establish a promising feature set. They do not establish that the model is ready for every production environment.
Independent testing needs to examine character identity across chained clips, the stability of hands and objects during complex actions, the reliability of spoken dialogue, and whether generated sound remains accurate when scenes become crowded or physically complicated.
Prompt adherence also needs a harder test. A model may understand the broad intention of “a chef working in a busy kitchen” while failing at the specific details that matter to a commercial: a logo, a product shape, a precise sequence of actions, or a legally approved line of dialogue.
Audio deserves its own evaluation. Native generation can produce better synchronization, but it also couples two failure modes. A visually strong clip with unusable speech or distracting ambience may require regeneration rather than a simple audio replacement.
The economics are equally important. Draft Mode could reduce wasted spend, but a serious production may require dozens of attempts per shot. Partner pricing, queue times, resolution workflows, API limits, content policies, and retry costs will determine whether FLUX 3 is useful at scale.
Safety and rights also remain part of the deployment question. Black Forest Labs says it worked with Cinder to evaluate risks including non-consensual intimate imagery and child sexual abuse material before release. That is evidence of a mitigation process, not a guarantee that misuse or ambiguous outputs have been eliminated. The launch announcement describes the company’s safety work.
Who should use it now?
FLUX 3 Video is worth trying now for developers building creative tools, agencies testing AI-assisted production, filmmakers creating storyboards and previs, marketers producing short-form variants, and teams that need dialogue or sound effects in the first generation.
It is also a strong candidate for workflows where a single coherent shot matters more than a long finished film: product demonstrations, social advertising, educational clips, visual explainers, stylized documentary sequences, and concept development.
Teams should wait before rebuilding an entire production pipeline around it if they need broadcast-grade guarantees, deterministic repeatability, stable identity across many minutes, exact brand compliance, or fully documented enterprise terms. Those requirements depend on evidence that has not yet been published at sufficient depth.
Verdict
FLUX 3 Video is more than a new entry in the text-to-video market. Its importance lies in the model’s design direction: Black Forest Labs is trying to make video, audio, motion, reference control, and physical continuity parts of one generative system.
That is a meaningful product bet. It could reduce the number of separate tools required to move from an idea to a usable shot, especially when sound and movement need to agree from the first frame.
But the release is still a beginning. The 20-second generation limit makes FLUX 3 useful for shots, not finished films. The internal evaluations are encouraging but not independent. The public demos show what the system can produce, not how reliably it will perform under production pressure.
FLUX 3 Video deserves serious attention because it treats audiovisual generation as one problem. Whether it becomes a dependable production foundation will be decided by the less glamorous tests: consistency, retries, costs, controls, and the quality of the tenth generation rather than the first.
Sources and methodology: This article was based on Black Forest Labs’ August 4 launch announcement, its July 23 FLUX 3 overview, the FLUX 3 model page, the Self-Flow technical report, fal’s FLUX 3 product and endpoint documentation, Google DeepMind’s Veo materials, ByteDance’s Seedance 2.0 announcement, and Runway’s Gen-4.5 documentation. Company evaluations and marketing claims have been identified as such; no independent hands-on testing was performed.
