VC
Value Add VC
⚡HomePulse⚡Helpful Apps📝Blog
Illustration for: Black Forest Labs Launches FLUX 3, a Multimodal Media Model
← Value Add PulseAI

Black Forest Labs Launches FLUX 3, a Multimodal Media Model

Black Forest Labs released FLUX 3, a multimodal model that generates images and 20-second videos with native audio and can predict robot actions from the same set of weights, in limited early-access release.

Up to 20 sec
Video length
77% preferred
vs Runway Gen-4.5
93% preferred
vs Luma Ray 3.2
Limited/API
Release
TC
Trace Cohen
Early-stage VC & angel · Founder, New York Venture Partners
July 23, 2026
3 min read
ShareXLinkedInEmail
THE RUNDOWN
1

FLUX 3 jointly learns from images, video and audio within a single unified architecture, generating up to 20 seconds of video with dialogue, sound effects and background music suited to the scene, plus image editing and readable text rendering from the same weights

2

The model extends beyond media generation into robot-action prediction, positioning Black Forest Labs to compete not just with video-generation rivals like Runway and Google's Veo but with physical-AI and robotics-foundation-model companies as well

3

In human evaluations, FLUX 3 recorded a 77% preference rate versus Runway Gen-4.5 and 93% versus Luma Ray 3.2, with a slight edge over Google Gemini Omni and Seedance in 52% of comparisons -- genuinely strong relative performance if the evaluation methodology holds up to scrutiny

4

The release is deliberately staggered: video generation and action features ship first via API and to select partners, image generation follows in coming weeks, and a locally usable open-weights version is planned only for the FLUX Dev model in the second half of 2026

TC
The VC Read · Trace's TakeTrace Cohen

A model that generates convincing video and predicts robot actions from the same weights is the clearest evidence yet that video-generation research and physical-AI research are converging into one field, not two. If you're building a robotics foundation model from scratch right now, you need a very good answer for why you're not just fine-tuning something like this instead -- the capital efficiency argument is getting harder to make every quarter.

Humanoid Robot Race → AI Landscape →

Black Forest Labs released FLUX 3 on July 23, a multimodal foundation model that jointly learns from images, video and audio within a single unified architecture -- generating up to 20 seconds of video complete with dialogue, sound effects and background music suited to the scene, alongside image editing, readable text rendering and, notably, robot-action prediction from the same set of model weights. The release is staged: video generation and action-prediction features are available first via API and to select partners, image-generation capability follows in the coming weeks, and a publicly downloadable, locally usable open-weights version is planned only for the smaller FLUX Dev model in the second half of 2026.

The headline capability -- extending the same base model from media generation into robot-action prediction -- is the most strategically significant detail. It positions Black Forest Labs to compete not just against video-generation rivals like Runway, OpenAI's Sora and Google's Veo, but against the growing field of physical-AI and robotics-foundation-model companies that have attracted enormous venture funding this year, including this same week's $1.7 billion Atoms round and the broader $55.8 billion in 2026 robotics funding. A single architecture that generates convincing video and predicts physical robot actions suggests the underlying representations learned from massive video datasets may transfer more directly into physical-world action prediction than previously assumed.

The competitive benchmarks Black Forest Labs disclosed are genuinely strong, if the underlying evaluation methodology holds up: a 77% preference rate in human evaluations against Runway's Gen-4.5, a 93% preference rate against Luma's Ray 3.2, and an edge over Google's Gemini Omni and Seedance in just over half of head-to-head comparisons. Self-reported benchmarks from any model developer deserve some skepticism regarding evaluation methodology and prompt selection, but the margins claimed here -- particularly against Luma -- are large enough to suggest a real capability gap rather than marginal, cherry-picked wins.

“The headline capability -- extending the same base model from media generation into robot-action prediction -- is the most strategically significant detail.”

The staggered release strategy is also worth noting for what it signals about the company's priorities. Shipping video and action-prediction capability first, ahead of image generation -- traditionally the more commercially proven and widely used capability -- suggests Black Forest Labs sees video and physical-world action prediction as the more strategically important frontier to establish leadership in first, even if image generation currently drives more near-term commercial revenue across the industry.

For founders and investors in creative tools, video production and increasingly physical AI, FLUX 3's cross-domain architecture is a signal that the boundary between "generative media company" and "robotics foundation model company" is blurring faster than most category definitions have kept pace with. Startups that assumed video generation and robot-action prediction were separate technical domains requiring entirely different specialized teams may need to revisit that assumption as unified architectures like FLUX 3 demonstrate meaningful transfer between the two.

The bear case: limited early-access release means FLUX 3's real-world performance and reliability outside curated demo and benchmark conditions remain unverified by independent developers, and Runway's newly launched Media Router -- designed explicitly to abstract away exactly this kind of model-versus-model comparison -- may blunt some of the competitive pressure FLUX 3's benchmark wins would otherwise create.

Watch how quickly FLUX 3 moves from limited access to broader API availability, whether independent developers replicate the disclosed preference-rate benchmarks once wider access rolls out, and whether the robot-action-prediction capability attracts direct partnership or investment interest from physical-AI and robotics companies looking to build on top of Black Forest Labs' architecture rather than training their own foundation models from scratch.

ShareXLinkedInEmail

Originally reported by VentureBeat. Analysis and editorial commentary by Value Add Pulse.

← Back to Pulse

THE WIRE in your inbox— Tech, startup & VC news with Trace's take. Free, no spam.

Read Next

AI· Jul 24, 2026

Inside the 25-Company Letter Backing Openweight AI Models

Illustration for: Inside the 25-Company Letter Backing Openweight AI Models
AI

Inside the 25-Company Letter Backing Openweight AI Models

Nvidia, Microsoft, Meta, a16z, IBM, Dell, Palantir, Mistral, Hugging Face and Y Combinator jointly argued open-weight AI models keep America competitive against China, while OpenAI and Anthropic declined to sign.

AI· Jul 24, 2026

OpenAI Brings New Voice Mode to ChatGPT Desktop App

Illustration for: OpenAI Brings New Voice Mode to ChatGPT Desktop App
AI

OpenAI Brings New Voice Mode to ChatGPT Desktop App

OpenAI rolled out its newer, more natural-sounding voice mode to the ChatGPT desktop app, extending a feature previously limited to mobile as the company races to keep pace with rivals on multimodal interaction.

AI· Jul 23, 2026

Microsoft's New In-House AI Models Cut Costs Up to 89%

Illustration for: Microsoft's New In-House AI Models Cut Costs Up to 89%
AI

Microsoft's New In-House AI Models Cut Costs Up to 89%

Microsoft released two new in-house AI models, MAI-Image-2.5-Pro and MAI-Voice-2-Flash, claiming GPU cost reductions of up to 89% for voice and 84% for image generation compared to OpenAI's equivalent models.

@Trace_Cohen·t@nyvp.com