Video wird geladen...
Video konnte nicht geladen werden
Reasoning VLAs can think. They just can't think fast. Until now. Introducing FlashDrive⚡ 🚀 716 ms → 159 ms on RTX PRO 6000 (up to 5.7×) ✅ Zero accuracy loss FlashDrive = streaming inference + DFlash speculative reasoning + ParoQuant W4A8 Real-time reasoning for autonomous driving is here!
191,948 Aufrufe • vor 5 Monaten •via X (Twitter)
32 Kommentare

Code + model checkpoints coming soon. The ingredients are already open-source: 🔥 DFlash → ⚡ ParoQuant →

Great work led by @Richard91316073 and @YihaoLiang01, with @hongfeizhang0xF and @jianchen1799.

Thanks for the paper, I've spent around 2 hours reading it and come up with my response. Screenshot because of twitter character limit.

Thanks for sharing your feedback! The baseline is under BF16.

awesome work 🫡

716ms to 159ms is the line real-time robotics has been waiting on. Every VLA roadmap priced in at least another year before reasoning models could sit inside the control loop. Quietly huge.

Trying this on an in-car 3090 with a comma soon

5.7x faster with zero accuracy loss sounds like the kind of magic trick that only works on paper until you actually try it in a car moving at 60mph

holy shit. need this deployed somewhere and served on an api right now

could be more compelling from safety perspective if the model used for speculative reasoning is RL-ed to be more cautious than the full reasoning. safety-first in the literal sense. :-)

@kstonekuan this looks interesting and in your field!

Sub-160ms VLAs that don't think in molasses. ParoQuant + DFlash cuts latency 5.7x with zero accuracy loss. Lab demos just became production reality. ⚡️

@yassineyousfi_

Nice work getting reasoning VLA down to 159ms on RTX PRO 6000 with no accuracy drop. Streaming + DFlash speculative decoding + ParoQuant W4A8 is a clean stack for closing the control loop. Looking forward to the checkpoints.

very cool !

Holy smokes this is really cool :O

this is very practical

716ms to 159ms. Real-time VLA reasoning is no longer a lab benchmark.

The latency gap is real for live video reasoning. The quality tradeoff at high token budgets is the other side - we found diminishing returns past ~700 thinking tokens for video understanding. Benchmark at

Really cool!

The latency wall was always the easier problem. Zero accuracy loss at 5.7x is impressive, but the moat in autonomous driving isn't inference speed — it's whether the system can handle the one corner case that no training distribution ever covered.

are they going after @comma_ai

@QuanquanPeng03 tql

wen weights tho

speculative decoding on rtx pro is one thing. what about consumer gpus that people actually use. a10s cost way less

We've also tested on many other GPUs. See our blog for more details!

Algorithm + System Collaboration

Some humans I know will need this to finally see the world with reasoning

😳

I need to update superlab to have support for playing around with VLAs and reasoning VLAs @FelionSpike

One thing I wonder as I watch this video, does the software do a delta between the prior frame to the next phrase to spot changes then analyze that to see what changed that's important?

flashdrive slicing latency. what's the real-world task it unlocks first
