Loading video...
Video Failed to Load
📢📢📢Introducing xGen-MM-Vid (BLIP-3-Video)! This highly efficient multimodal language model is laser-focused on video understanding. Compared to other models, xGen-MM-Vid represents a video with a fraction of the visual tokens (e.g., 32 vs. 4608 tokens). Paper: Website: Researcher’s 🧵:👇
12,550 views • 1 year ago •via X (Twitter)
0 Comments
No comments available
Comments from the original post will appear here
