Loading video...
Video Failed to Load
Most speculative decoding still drafts tokens one at a time. That's not parallel generation — it just hides the serial loop behind a smaller model. UC San Diego's z-lab just drew a clear line between the two. They released DFlash — a lightweight block diffusion model that drafts a... show more
23,328 views • 1 month ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
