正在加载视频...
视频加载失败
Training world models needs egocentric video and dense action signals, synchronized. That data is genuinely hard to find. We built it from Counter-Strike 2 demos. CS2-10k: 600K+ player-round videos, 10K+ hours, per-frame annotations — keyboard state, mouse delta, 3D position, camera yaw/pitch. All paired to the visual stream. Why... show more
0 条评论
暂无评论
原始帖子的评论将显示在这里
