Loading video...
Video Failed to Load
Molmo 2 doesn't just answer questions about clips—it searches & points. The model returns coordinates & timestamps over videos + images, powering QA, counting, dense captioning, artifact detection, & subtitle-aware analysis. You can see exactly how it reasoned.
67,967 views • 5 months ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here

