
Xiaochuang Han
@XiaochuangHan • 1,061 subscribers
Research Scientist at Meta FAIR. Formerly @uwnlp.
Videos

Can we simplify video generation by decomposing it into interleaved text-video co-generation? Would explicit, repeated thinking in language improve generation in pixels? We introduce TV2TV: a unified model that jointly learns - language modeling (next-token prediction) - video flow matching (next-frame prediction) At inference, TV2TV dynamically alternates between textual thinking and video generation. Model generations below: interleaved text plans and video slices (~1–2s) are co-generated over time, conditioned on a single frame per sport. 📖
Xiaochuang Han28,925 görüntüleme • 6 ay önce
Daha fazla içerik yok.