正在加载视频...
视频加载失败
Variable-length compressive tokenization is a promising direction. Not just for efficiency, but for actually learning powerful representations. FlexTok scratched the surface of this direction with images. But real-world data has additional structure that can be tapped, such as the temporal structure of videos. VideoFlexTok develops this concept for video,... show more
17,775 次观看 • 4 个月前 •via X (Twitter)
1 条评论

Sivan Doveh4 个月前
Great paper Amir! Super interesting
