Loading video...

Video Failed to Load

Go Home

🚀Introducing MiniCPM-V 2.6! 🔥 1、Surpassing GPT-4V in single image, multi-image and video understanding 📸🎥 2、Outperforms GPT-4o mini and Gemini 1.5 on OpenCompass 🏆 3、Real-time video analysis on iPad 📱💨 Try out the best on-device multimodal LLM here! 👑 GitHub: Huggingface: #MLLM #MiniCPM

196,382 views • 2 years ago •via X (Twitter)

10 Comments

Henry Sibanda's profile picture
Henry Sibanda2 years ago

Almost everything openai has been holding back because ‘safety’ is available opensource except the speech to speech. And we all still doing just fine…

kache's profile picture
kache2 years ago

great work guys

Pabli! 🔥💥💫's profile picture
Pabli! 🔥💥💫2 years ago

The previous one was great. This one is superb. 🙌✨

Hailey Collet's profile picture
Hailey Collet2 years ago

This + OWL-ViT (or Florence) + SAM2 is going to enable some crazy stuff

Tiezhen WANG's profile picture
Tiezhen WANG2 years ago

The demo is super cool!

J Sam🌐's profile picture
J Sam🌐2 years ago

🐐

Alex Volkov (Thursd/AI)'s profile picture
Alex Volkov (Thursd/AI)2 years ago

Going to absolutely mention this on the next @thursdai_pod tomorrow!

The Canaanite's profile picture
The Canaanite2 years ago

Wow, that's really something. Amazing work!

pallendromeda's profile picture
pallendromeda2 years ago

lol it is me

Soible_VR's profile picture
Soible_VR2 years ago

whats the status of llamacpp support for MiniCPM family?

Related Videos

🚀 🚀Excited to announce the technical report of MiniCPM-o 4.5! MiniCPM-o 4.5 transitions #AI interaction from traditional turn-based processing to a real-time, native full-duplex stream-based paradigm. 🌊 The Omni-Flow Framework Instead of traditional VAD-based workarounds, we introduce the #Omni-#Flow framework. This unified stream paradigm aligns video, audio, and text on a synchronized millisecond timeline. • Native Full-Duplex: Simultaneous perception and response. • Proactive Interaction: Natively manages turn-taking without external VAD, supports proactive reminding. 📉 9B Scale, SOTA Performance MiniCPM-o 4.5 demonstrates SOTA multimodal intelligence at its scale: • Multimodal Benchmarks: Comparable to #Gemini 2.5 Flash on MMBench EN (87.6) and MathVista (80.1). • Streaming Evaluation: 54.4% win rate on LiveSports-3K-CC, surpassing specialized models. 💻 The Ultimate Edge AI — Fully Functional without Network Connection We are providing one-click installers for Windows (12G VRAM,RTX 5070) and macOS (M1-M5 Max/ M5 Pro). • Local API Support: Deploy your own inference server to integrate native full-duplex into custom apps. • Free Access: We are offering free community API services for exploration. • 100% Private: Your data never leaves your machine. Deploy in under 10 minutes. 🛠️👇 👐 Join the Open Future The weights are open. The protocol is public. 📄 Technical Report: 💻 GitHub: 🤗 HuggingFace: 🌐 Web Demo: #MiniCPMo #OpenSourceAI #EdgeAI #MachineLearning #ComputerVision #LLM

OpenBMB

147,994 views • 3 months ago