正在加载视频...
视频加载失败
Most speculative decoding still drafts tokens one at a time. That's not parallel generation — it just hides the serial loop behind a smaller model. UC San Diego's z-lab just drew a clear line between the two. They released DFlash — a lightweight block diffusion model that drafts a... show more
23,328 次观看 • 1 个月前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
