正在加载视频...
视频加载失败
Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster. FLUX 3 Action removes the usual trade-off between... show more
163,135 次观看 • 2 天前 •via X (Twitter)
39 条评论

Read the full blog: Download the weights: Talk to our robotics team:

FLUX 3 Action removes the usual trade-off between the higher success rates of WAMs and the speed of VLAs. Its single-step 7B checkpoint outperforms every other open policy on RoboLab, while processing each second of robot motion 1.45× to 1.66× faster than Pi0.5 (the strongest open VLA model). It retains the core WAM architecture, predicting video and actions jointly, but covers a longer 2.13-second action horizon compared with 1 second for Pi0.5. When latency matters less, our guidance-distilled checkpoint raises the state-of-the-art success rate on RoboLab while running 2.85× to 3.15× faster than the previous leading open WAM.

FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture. More efficient representations learned through our research on Self-Flow made this smaller size possible. In midtraining, we trained the model to predict actions and future frames together.

FLUX 3 Action takes recent camera frames, the system’s current state, and a description of the task. It then returns the next 32 actions while also predicting how the scene will change. The system repeatedly observes, plans, acts, and adjusts. The model is also capable of recovering from its own mistakes.

Beyond robotics, FLUX 3 Action showed promising early results when we trained task-specific policies for playing games and controlling a simulated drone — suggesting the same approach could apply wherever a model needs to understand a visual environment and choose what to do next (e.g. simulations, games, and computer use).

FLUX 3 Action controlling a simulated drone to take a specific requested action.

More details here:

Not the weights we want but ok. When are the Flux 3 Video weights dropping?

when open weights video! not action and robotics!

As cool as this is, what about releasing the Flux 3 video weights like you guys mentioned? I am pretty sure thats what everyone really wants to know about 😂😂

But what about Flux 3 video weights?

And what about Flux 3.0?

Flux videos pls open weight

Nice!

@victormustar Live BFL!

flux 3 open weights never came , you guys are liars

7B parameters outperforming larger models is seriously impressive engineering

We need FLux 4 for image and video

And what about... Flux 3 Image? Did you forget?

🔥

The 6.1pp at 56% fewer params is the leaderboard flex; the real question is whether twice-ahead video+action still holds once teams leave RoboLab and fine-tune on their own Jetson demos.

The model looks really promising. FLUX is already a great backbone. It definitely needs to be tested across more domains, though.

your world model never left the render

For a moment I thought 7B parameter video generation model 🙂

@Presidentlin

7B parameters is a featherweight. I’m curious how it handles complex spatial reasoning in cluttered scenes compared to those bloated 10B+ rivals. If the efficiency holds up at scale, that open-weight flexibility could be seriously handy for local setups

this is the way

Open-weight robotics AI is moving fast. 3.95× faster is a huge improvement.

gz on #1

The 6.1-point gain with 56% fewer parameters is a compelling result, but the speedup may be the bigger win for real robots: faster action loops can matter more than a small offline score bump. Fine-tuning on task-specific demonstrations also makes the open-weights release much more usable than a fixed benchmark model.

Fewer parameters and faster inference make this robotics result stand out.

🤯. When Flux 3 video and image Open Source finetunes? So desparately waiting for them to drop.

smaller faster stronger, pick all three apparently

@grok 它消费了什么样子的数据集 ? 数据从谁那边买的

robots getting faster while I still trip over flat ground

The way this model predicts video and actions together while staying so fast is a massive leap for robotics. We actually went deeper on this here:

画图厂长手了

7B on a Jetson is the part robotics people will care about. What control rate does it actually hit on-device?

okay thats fucking sick
