
What Is Flux 3? Complete Review of BFL's Multimodal Model
Is Flux 3 worth it? This review covers what it is, how Self-Flow works, image/video/audio capabilities, benchmarks vs Runway and Sora, pricing, and access.
On July 24, 2026, Black Forest Labs released Flux 3, and for the first time the company's flagship model does far more than still images. Flux 3 is a multimodal foundation model that jointly learns from images, video, and audio inside one unified architecture, and it can generate a 20-second video with synced audio, edit a photo, render multilingual text, and even predict actions for robots on an assembly line.
For anyone who has followed the FLUX family (from the open-weight FLUX.1 Dev and Schnell that dominated the community in 2024, to FLUX 1.1 Pro that topped the Artificial Analysis image arena, to the quieter FLUX.2 in late 2025), Flux 3 is the lab's biggest pivot yet. It is also arriving into a fiercely competitive video-model market alongside Runway, Sora, Kling, and Veo. This guide breaks down what Flux 3 actually is, how its Self-Flow architecture works, what it can do, how it benchmarks against rivals, how to get access, and how it fits into 3D and creative workflows. Our short verdict: Flux 3 is the most ambitious multimodal launch of 2026, and its early numbers are genuinely strong, but it is still in Early Access, so whether it is "worth it" depends on whether you need unified image-video-audio now or can wait for the open-weight FLUX 3 Dev release later in 2026.
TL;DR
| You want… | Pick |
|---|---|
| One model for image, video, and audio from a single prompt | Flux 3 |
| Up to 20-second video clips with native synced sound | Flux 3 Video (Early Access) |
| Image generation and editing across many styles | Flux 3 Image (rolling out soon) |
| Run the multimodal backbone locally, free weights | FLUX 3 Dev (later in 2026) |
| Action prediction for robotics and physical AI | FLUX-mimic / FLUX 3 Action |
| A finished, fully open model you can use today | Wait: Flux 3 is still in Early Access |
What Is Flux 3?
Flux 3 is Black Forest Labs' multimodal AI model, released in July 2026. It generates images, video with synchronized audio (up to 20 seconds per clip), and robot action predictions from a single unified architecture called Self-Flow, succeeding the Flow Matching approach used in earlier FLUX image models. Video and Action are available now in Early Access, with the open-weight FLUX 3 Dev release planned for later in 2026.
Flux 3 is Black Forest Labs' new multimodal foundation model. The defining idea is that a single model should not learn images, video, and audio as three separate skills bolted together. Instead, it should learn one underlying representation of the world (how objects hold together, how things move, and how events sound), because each modality is just a different projection of the same reality.
Images capture spatial structure at a single moment. Video restores time and reveals physical dynamics. Audio exposes causal relationships between mechanical events that vision alone cannot detect. Language links all of it to goals and instructions. Black Forest Labs' argument is that when a model learns from all of these at once, the mutual constraints teach it more than any single modality ever could: the sound has to match the impact, the motion has to obey mass, the future has to follow from the past.
Flux 3 is the first model the lab has built entirely on that principle.
Flux 3 at a glance
| Spec | Detail |
|---|---|
| Maker | Black Forest Labs (BFL), Germany |
| Type | Multimodal foundation model |
| Modalities | Image, video, audio, action prediction |
| Max video length | 20 seconds per generation (with audio) |
| Eval resolution | 10-second clips at 720p |
| Architecture | Self-Flow (unified generation + understanding) |
| Image access | Early Access "in the coming weeks" |
| Video access | Early Access now (APIs + private weights) |
| Open weights | FLUX 3 Dev, planned for later in 2026 |
| Notable partner | mimic robotics (Audi production line) |
Who Made Flux 3? (Black Forest Labs)
Black Forest Labs is a German AI lab founded in August 2024 by veteran researchers who had helped build the original Stable Diffusion models at Stability AI, including co-founder and CEO Robin Rombach. The company quickly became known for the FLUX line of image generators.
The trajectory matters for understanding Flux 3:
- FLUX.1 Dev and Schnell (2024) grabbed the "best open source image generator" title that the community had expected Stability's own Stable Diffusion 3.5 to reclaim. It never did.
- FLUX 1.1 Pro topped the Artificial Analysis image arena in October 2024, though unlike the Dev/Schnell releases, it was not open source.
- FLUX.2 shipped in November 2025 but landed with less fanfare and was not as widely adopted.
- The open-source crown the original FLUX held was eventually taken by Alibaba's Z-Image Turbo in late 2025, which matched FLUX quality on lower-end consumer GPUs.
Against that backdrop, Flux 3 is BFL's comeback and a deliberate strategic shift away from still images toward video and physical AI. The lab's argument is that a model trained only on images can only ever produce images; to perceive and predict the physical world, it has to learn video too, because video encodes the physics (weight, contact, timing) that a machine needs to operate in the real world.
How Flux 3 Works: Self-Flow Architecture

The technical backbone of Flux 3 is Self-Flow, Black Forest Labs' approach for efficiently aligning multimodal generation and understanding within the same underlying architecture. This is the successor idea to the Flow Matching technique used in earlier FLUX image models.
Why move beyond Flow Matching? BFL's own comparison frames it clearly: when you train one modality at a time, you get a good model of that one projection. When you train image, video, and audio together, their mutual constraints act as cross-checks: the audio must fit the visual impact, the motion must obey physical laws. Self-Flow is designed to exploit those cross-modal constraints rather than ignore them.
According to BFL's internal evaluation charts, Self-Flow achieves lower generation error (measured as Fréchet distance, normalized so Flow Matching = 100) across each modality, and a higher success rate on manipulation tasks averaged over four task groups through fine-tuning. In plain terms: the unified approach is both a better generator and a better foundation for downstream tasks like robotics.
Based on Self-Flow, BFL "significantly scaled up compute and data resources to train Flux 3 across video, images, and audio at the same time." More technical details on the underlying approach are promised for later release.
Flux 3 Capabilities
Because Flux 3 is a single unified model, its capabilities bleed into one another. A video generation can carry an image reference for character consistency; an image edit can respect a multilingual text instruction; a robot action can be predicted from the same dynamics engine that animates a cinematic clip. Here is what the model can do today, according to the launch announcement, and note that all video outputs come with native audio generation.
Image generation
Flux 3 can synthesize and edit images across a wide variety of styles, aspect ratios, and resolutions. In preliminary mid-training evaluations BFL shared, Flux 3 already shows significant improvement over earlier FLUX versions in two areas in particular:
- Complex prompt handling: multi-subject, multi-constraint prompts are followed more faithfully.
- Text generation: the model renders high-accuracy text in multiple languages, a longstanding weak spot for image models.
The model also produces a broad range of output styles well beyond photorealism, extending into animation, design, and typographic compositions. An Early Access phase for Flux 3 Image is opening in the weeks following launch.
Video generation
This is the headline feature. Flux 3 can create highly diverse videos up to 20 seconds in length in a single generation, with synchronized audio attached natively. Core video capabilities include:
- Text-to-video generation from a pure prompt.
- Image-to-video, either continuing from a starting frame ("animation") or using images as visual references.
- Video-to-video, carrying central elements of a source clip (for instance the same character) into a new scene or context.
- Generative video-audio continuation from input video and audio.
- Keyframe-to-video generation for controlled transitions between defined moments.
- Agentic chaining of individual clips into longer, multi-shot sequences lasting several minutes, with visual references keeping characters consistent across scenes.
Style diversity is a stated strength: Flux 3 Video handles everything from candid camcorder footage to animation to cinematic output, and it includes strong typography generation and animated designs.
Audio synthesis
Audio is not an afterthought. Every video output ships with native audio, and the model explicitly associates sounds with physical events on screen: dialogue, sound effects, and ambient noise. BFL highlights that Flux 3 Video is "particularly strong in capturing human facial expressions, associating sounds with physical events, and multilingual capabilities." This audio-visual coupling is a direct payoff of the multimodal training: the model learned that an impact should produce a matching sound.
Action prediction (FLUX-mimic)
The most surprising capability is action prediction, using the same video-prediction engine to drive physical robots. Black Forest Labs took two routes:
- Native action prediction integrated directly into Flux 3, scaling up the initial Self-Flow work.
- A dynamics-aware foundation, using the pretrained video backbone as a base that specialized action models can be fine-tuned from with limited task-specific data.
The second route produced FLUX-mimic, built with Zurich-based mimic robotics. It takes Flux 3's video-prediction engine and adds a lightweight "decoder" that translates the model's internal sense of how things move into actual robot motions. The first major deployment is at Audi, where FLUX-mimic robots handle tasks like fitting flexible door seals, soft-body manipulation work that conventional automation has struggled with. BFL reports the full FLUX-mimic system reacts in about 101 milliseconds (BFL mimic blog, July 2026), in the neighborhood of human visual reflexes.
Flux 3 vs Runway, Sora, Kling, and Veo
Flux 3 does not enter an empty market. Runway, OpenAI's Sora, Kuaishou's Kling, and Google's Veo are all competing for the same text-to-video and image-to-video workflows. The clearest (though vendor-reported) comparison data BFL has published comes from human-preference evaluations, where reviewers watch two clips and pick the more convincing one.
The comparison matrix
In Black Forest Labs' early human-preference evaluations from July 2026 (10-second 720p text-to-video clips with audio), Flux 3 Video was preferred over:
| Rival model | Flux 3 win rate |
|---|---|
| Luma Ray 3.2 | 93% |
| Runway Gen-4.5 | 77% |
| Grok Imagine Video | up to 69% |
| Kling v3 Pro | 60% |
| Happy Horse v1 | 59% |
| Happy Horse 1.1 | 57% |
| Seedance 2.0 | 52% |
| Gemini Omni Flash | 52% |
A few honest caveats: these are preference tests, not a fixed scoring rubric, and BFL reports each as "up to" a given percentage (the peak win rate across comparison batches). They are reported by the model's own maker during an Early Access phase where BFL itself notes "we expect further improvements." Treat the numbers as directional, not definitive. Independent benchmarks from arenas like Artificial Analysis will be the real test once Flux 3 Video is broadly available.
When to pick Flux 3
- You want image, video, and audio from one model rather than stitching separate tools.
- You need synced native audio tied to on-screen physics.
- You want a path toward open weights (FLUX 3 Dev) for local and custom workflows.
- You are exploring physical AI / robotics and want a dynamics-aware backbone.
When to pick a rival
- You need a fully shipped, generally available product today; Runway and Veo are more mature on that front.
- You want the simplest possible interface with zero setup; cloud platforms still win on convenience.
- You need independent, third-party-verified quality numbers before committing.
How to Access Flux 3
Flux 3 is rolling out in stages, each behind an Early Access phase for safety testing and feedback. Here is what is available and what is coming.
Early Access (now)
- Flux 3 Video and audio generation and editing through APIs and private weight access.
- Action prediction through selected research and commercial partners, beginning with mimic robotics (FLUX-mimic and FLUX 3 Action).
You can request Early Access directly through Black Forest Labs' website. Third-party platforms such as fluxpro.ai and muapi.ai have also begun surfacing Flux 3 endpoints, though always verify you are hitting an authorized API.
FLUX 3 Image (coming weeks)
Image synthesis and editing through APIs and private weight access opens in the weeks following launch. This is the tier most relevant to users coming from FLUX 1.1 Pro or FLUX.2 image workflows.
FLUX 3 Dev open weights (later in 2026)
The open-weight multimodal backbone (for content creation: video, audio, image, and action prediction) is planned for release later in 2026. This is the only tier BFL intends to release for local use, and it is the one the open-source and ComfyUI communities are watching most closely.
Pricing
Specific pricing has not been finalized publicly at launch; Flux 3 is in Early Access and BFL is collecting feedback before locking in commercial terms. Expect API-based pricing to follow the per-generation or per-second-of-video model common across Runway, Kling, and Veo, with the open-weight Dev tier remaining free to download (you supply the hardware).
How to Use Flux 3 for 3D Workflows
This is where Flux 3 connects most directly to the kind of work Trify3D is built for. AI image and video models are incredibly useful as reference and concept generators that feed downstream 3D pipelines, and a multimodal model that understands physics, motion, and consistent characters is a particularly strong reference source.
A practical workflow looks like this:
- Generate a reference image with Flux 3 Image: a character, a product, a prop, or an environment concept, rendered in any style and aspect ratio.
- Iterate on the prompt until the silhouette, proportions, and material look right. Flux 3's improved complex-prompt handling makes this faster than older FLUX versions.
- Turn that reference into a 3D model. Feed the image into an image-to-3D pipeline to produce clean geometry with PBR textures. This is exactly the gap Trify3D fills, taking a 2D concept and generating a usable 3D asset without manual modeling.
- Optionally generate a short video with Flux 3 Video to study how the subject moves, which informs rigging, animation, and physics setup on the 3D side.
Because Flux 3 keeps characters consistent across chained clips, it is especially useful for maintaining visual continuity across a whole set of reference images destined for a 3D scene or game asset library.
Strengths, Limitations, and Who It's For
Strengths
- True multimodality: image, video, audio, and action from one model, not a bundle of separate tools.
- Native synced audio tied to on-screen physics, a capability many rivals still lack or bolt on separately.
- Strong human-preference numbers against established video models in early tests.
- A credible open-weight roadmap via FLUX 3 Dev for the local-run community.
- Physical AI credibility through the mimic robotics / Audi deployment.
Limitations
- Early Access only: not a finished, generally available product yet.
- Vendor-reported benchmarks: independent validation is still pending.
- 20-second clip ceiling per generation (though chaining extends this).
- Open weights delayed: FLUX 3 Dev is not due until later in 2026, so local users must wait.
- Action prediction is partner-gated: not something an individual developer can simply download.
Who it's for
- Creators who want image, video, and audio from a single prompt without stitching tools.
- Developers building on FLUX APIs who want a multimodal upgrade path.
- Open-source and ComfyUI users willing to wait for FLUX 3 Dev.
- Robotics and physical AI teams exploring dynamics-aware foundation models.
- 3D artists looking for a high-quality reference and concept generator for image-to-3D pipelines.
The Bigger Picture: Multimodal AI in 2026
Flux 3 is part of a clear industry-wide shift from single-modality specialists toward unified multimodal models. OpenAI's GPT-style omni models, Google's Gemini Omni, and now Black Forest Labs' Flux 3 all reflect the same thesis: that perception, generation, and eventually action should live in one model because they are all evidence about the same underlying reality.
BFL frames this as a step toward "real-world visual intelligence: models that perceive, predict, and act across physical and digital environments." Whether or not that grand vision lands fully, the practical effect for creators is concrete: fewer tools to stitch together, better audio-visual consistency, and a foundation that improves downstream tasks from 3D generation to robotics.
The open question is execution. Flux 3's early numbers are strong, the architecture is credible, and the lab has a track record, but the model still has to survive independent benchmarking, a broad rollout, and the eventual open-weight release before it can be fairly judged against Runway, Sora, Kling, and Veo. For now, it is the most interesting multimodal launch of 2026, and worth watching closely.
Frequently Asked Questions
What is Flux 3?
Flux 3 is Black Forest Labs' multimodal foundation model, released in Early Access in July 2026. Unlike earlier FLUX models that only generated images, Flux 3 jointly learns from images, video, and audio inside a single architecture and can generate image, video with synced audio, and action predictions for robotics.
Is Flux 3 free to use?
Flux 3 Video and Action are available through Early Access APIs and private weight access, which are not free. An open-weight FLUX 3 Dev release is planned for later in 2026, which will let developers run the multimodal backbone locally at no license cost, though you still need your own GPU hardware.
How long can Flux 3 videos be?
Flux 3 can generate video clips up to 20 seconds long in a single generation, with native synchronized audio including dialogue, sound effects, and ambient noise. Multiple clips can also be chained agentically into sequences lasting several minutes while keeping characters consistent.
How does Flux 3 compare to Runway and Sora?
In Black Forest Labs' early human-preference evaluations, Flux 3 Video was preferred over Runway Gen-4.5 in 77% of comparisons and over Luma Ray 3.2 in 93%. It also beat Kling v3 Pro (60%), Grok Imagine Video (up to 69%), and roughly tied Seedance 2.0 and Gemini Omni Flash at 52%. Note these are vendor-reported preference tests, not independent benchmarks.
What is Self-Flow in Flux 3?
Self-Flow is Black Forest Labs' method for aligning multimodal generation and understanding within the same architecture. Flux 3 is the first model built entirely on this principle, scaling compute and data to train on video, images, and audio simultaneously, unlike the older Flow Matching approach used in earlier FLUX versions.
Who made Flux 3?
Flux 3 was made by Black Forest Labs (BFL), a German AI lab founded in August 2024 by researchers who helped build the original Stable Diffusion at Stability AI, including co-founder and CEO Robin Rombach. BFL is the company behind the FLUX line of image generators including FLUX 1.1 Pro and FLUX.2.
When will FLUX 3 Dev open weights be released?
Black Forest Labs plans to release FLUX 3 Dev, the open-weight multimodal backbone for content creation and action prediction, later in 2026. Video and audio APIs plus FLUX 3 Image are rolling out through Early Access first, with the Dev tier being the only version planned for local use.
Can Flux 3 generate audio?
Yes. Every Flux 3 video output comes with native audio generation, including synced dialogue, sound effects tied to physical events on screen, and ambient noise. Flux 3 can also do generative video-audio continuation from input video and audio.
Next Steps
Flux 3 is still unfolding: Early Access for video now, image in the coming weeks, open weights later in 2026. If you want to put a multimodal reference generator to work in a 3D pipeline today, the fastest path is generating a concept image and turning it directly into a 3D model. Try Trify3D's image-to-3D workflow to convert your Flux 3 (or any AI) reference images into clean, textured 3D assets without manual modeling. For a broader look at where Flux 3 sits among other AI models, see our roundup of the best AI 3D model generators and our review of the Sulphur 2 open-source video model.
Run it yourself in Trify3D
Keep reading from this topic
What Is the Best 3D Model AI Generator in 2026? (Verdict)
The best 3D model AI generator in 2026 is Meshy for the full pipeline and Tripo for all-around quality. See our use-case verdicts, pricing, and how to pick.
ReviewsSulphur 2: Open-Source Uncensored AI Video Model
Sulphur 2 is an open-source uncensored AI video model on LTX 2.3. Compare text-to-video, image-to-video, hardware, and workflows vs Runway and Kling AI.
ReviewsBest AI 3D Model Generators in 2026 (Tested & Ranked)
Tested 10 AI 3D model generators in 2026. Compare Meshy, Tripo, Rodin, Hunyuan3D, Trellis, Sloyd and more on quality, price, export formats, and best use case.