正在加载视频...
视频加载失败
Vision-language models (VLMs) can see well, but they struggle to reason. In this episode, Antonia Wüst (PhD researcher, TU Darmstadt) explains how combining VLMs with program synthesis yields more reliable visual reasoning, with fewer tokens than chain-of-thought.
22,130 次观看 • 6 个月前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里

