正在加载视频...

视频加载失败

Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.⁠⁠ FLUX 3 Action removes the usual trade-off between...

163,135 次观看 • 2 天前 •via X (Twitter)

39 条评论

Black Forest Labs 的头像
Black Forest Labs2 天前

Read the full blog: Download the weights: Talk to our robotics team:

Black Forest Labs 的头像
Black Forest Labs2 天前

FLUX 3 Action removes the usual trade-off between the higher success rates of WAMs and the speed of VLAs. Its single-step 7B checkpoint outperforms every other open policy on RoboLab, while processing each second of robot motion 1.45× to 1.66× faster than Pi0.5 (the strongest open VLA model). It retains the core WAM architecture, predicting video and actions jointly, but covers a longer 2.13-second action horizon compared with 1 second for Pi0.5. When latency matters less, our guidance-distilled checkpoint raises the state-of-the-art success rate on RoboLab while running 2.85× to 3.15× faster than the previous leading open WAM.

Black Forest Labs 的头像
Black Forest Labs2 天前

FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture. More efficient representations learned through our research on Self-Flow made this smaller size possible. In midtraining, we trained the model to predict actions and future frames together.

Black Forest Labs 的头像
Black Forest Labs2 天前

FLUX 3 Action takes recent camera frames, the system’s current state, and a description of the task. It then returns the next 32 actions while also predicting how the scene will change. The system repeatedly observes, plans, acts, and adjusts.⁠⁠ The model is also capable of recovering from its own mistakes.

Black Forest Labs 的头像
Black Forest Labs2 天前

Beyond robotics, FLUX 3 Action showed promising early results when we trained task-specific policies for playing games and controlling a simulated drone — suggesting the same approach could apply wherever a model needs to understand a visual environment and choose what to do next (e.g. simulations, games, and computer use).

Black Forest Labs 的头像
Black Forest Labs2 天前

FLUX 3 Action controlling a simulated drone to take a specific requested action.

Black Forest Labs 的头像
Black Forest Labs2 天前

More details here:

Latent Spacer 的头像
Latent Spacer2 天前

Not the weights we want but ok. When are the Flux 3 Video weights dropping?

netanel1993 🇮🇱 的头像
netanel1993 🇮🇱2 天前

when open weights video! not action and robotics!

JasonThatArtist 的头像
JasonThatArtist2 天前

As cool as this is, what about releasing the Flux 3 video weights like you guys mentioned? I am pretty sure thats what everyone really wants to know about 😂😂

Jayy Fatality Boo 的头像
Jayy Fatality Boo1 天前

But what about Flux 3 video weights?

Trilobite's Legacy 的头像
Trilobite's Legacy1 天前

And what about Flux 3.0?

Medchik 的头像
Medchik2 天前

Flux videos pls open weight

Lucas Beyer (bl16) 的头像
Lucas Beyer (bl16)1 天前

Nice!

Yuval Avidani (יובל אבידני) 的头像
Yuval Avidani (יובל אבידני)2 天前

@victormustar Live BFL!

blankbrain 的头像
blankbrain2 天前

flux 3 open weights never came , you guys are liars

A.W.E.S.O.M.-O 4000 的头像
A.W.E.S.O.M.-O 40002 天前

7B parameters outperforming larger models is seriously impressive engineering

Leo Phiêu Diêu 的头像
Leo Phiêu Diêu1 天前

We need FLux 4 for image and video

Brate_RyP 的头像
Brate_RyP2 天前

And what about... Flux 3 Image? Did you forget?

Trevor Knott 的头像
Trevor Knott2 天前

🔥

ByteLee · AI Investing 的头像
ByteLee · AI Investing1 天前

The 6.1pp at 56% fewer params is the leaderboard flex; the real question is whether twice-ahead video+action still holds once teams leave RoboLab and fine-tune on their own Jetson demos.

clankr 的头像
clankr1 天前

The model looks really promising. FLUX is already a great backbone. It definitely needs to be tested across more domains, though.

Anderson 的头像
Anderson1 天前

your world model never left the render

Jigs 的头像
Jigs2 天前

For a moment I thought 7B parameter video generation model 🙂

luis 的头像
luis2 天前

@Presidentlin

The AI Therapist 的头像
The AI Therapist2 天前

7B parameters is a featherweight. I’m curious how it handles complex spatial reasoning in cluttered scenes compared to those bloated 10B+ rivals. If the efficiency holds up at scale, that open-weight flexibility could be seriously handy for local setups

Cris Lenta 的头像
Cris Lenta2 天前

this is the way

RAZA | AI EXPLORER 的头像
RAZA | AI EXPLORER1 天前

Open-weight robotics AI is moving fast. 3.95× faster is a huge improvement.

fade 的头像
fade2 天前

gz on #1

DARKKEVN Labs 的头像
DARKKEVN Labs1 天前

The 6.1-point gain with 56% fewer parameters is a compelling result, but the speedup may be the bigger win for real robots: faster action loops can matter more than a small offline score bump. Fine-tuning on task-specific demonstrations also makes the open-weights release much more usable than a fixed benchmark model.

Fajar M Reza 的头像
Fajar M Reza1 天前

Fewer parameters and faster inference make this robotics result stand out.

Squeak Al-Gaib 的头像
Squeak Al-Gaib1 天前

🤯. When Flux 3 video and image Open Source finetunes? So desparately waiting for them to drop.

Panther 的头像
Panther2 天前

smaller faster stronger, pick all three apparently

Y11 的头像
Y111 天前

@grok 它消费了什么样子的数据集 ? 数据从谁那边买的

Hershal Rao 的头像
Hershal Rao2 天前

robots getting faster while I still trip over flat ground

Michael Waitze 的头像
Michael Waitze2 天前

The way this model predicts video and actions together while staying so fast is a massive leap for robotics. We actually went deeper on this here:

AGI 野生家 的头像
AGI 野生家1 天前

画图厂长手了

Kabel-Salat | AI & Tech 的头像
Kabel-Salat | AI & Tech1 天前

7B on a Jetson is the part robotics people will care about. What control rate does it actually hit on-device?

Laythe 的头像
Laythe2 天前

okay thats fucking sick

相关视频

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,079,976 次观看 • 3 个月前

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen 🥶 Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71× decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Léo

32,331 次观看 • 2 个月前

Small Language Models (SML) are the future of AI. "Small" (SML) instead of "Large" (LLM). These small models are highly specialized models with superhuman abilities on specific tasks. Here are two techniques to build these models: • Spectrum • Model Merging I give you a short introduction in the attached video, but here is a quick summary: Spectrum helps us identify the most relevant layers to solve one specific task. We can ignore everything else and focus on fine-tuning these layers. Using Spectrum, we can fine-tune models in a heartbeat. Model Merging combines multiple models into a unique, much better model than any of the individual input models. You can also combine models specialized in different tasks and get a model with multiple abilities. This is the state of the art of productizing models. It's what Arcee.ai's platform does behind the scenes. Arcee collaborated with me on this post and is sponsoring it. There are three main steps to produce a model for your particular use case: 1. You create a dataset by uploading your data. 2. You train a model. At this step, Arcee uses Spectrum and Model Merging to produce a highly specialized model for your task. 3. You can deploy that model to any environment you want. Three important notes: • Training process is 2x faster and 2x cheaper than regular fine-tuning. • Resultant models are smaller and have higher accuracy. • They create these specialized models from open-source models. Check this site so you can fully appreciate how this works: If you want to fine-tune an open-source model, consider Arcee's platform. This is the state of the art.

Santiago

164,162 次观看 • 2 年前