
Mihir Prabhudesai
@mihirp98 • 2,761 subscribers
PhD student at Carnegie Mellon
Shorts
Videos

Images, video, audio, actions — generative modeling has converged on one recipe: compress into continuous latents, generate in latent space, decode. Language is the lone exception: still generated token by token, as a long discrete stream. Latent Thought Flows: We compress 256 text tokens into 8 continuous latents, generate them with a one-step flow model, and read them out as text using an autoregressive decoder. Result: a better inference-compute vs. generation-quality Pareto frontier than a tuned autoregressive baseline — thanks to compression and one-step generation. Co-led w/ Zhengyang Geng 🧵
Mihir Prabhudesai81,598 views • 1 month ago
No more content to load