Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

🚀 MiniCPM enters the physical world — enabling robots to understand, remember, and act. We open-source MiniCPM-Robot, our first embodied AI model series, including: 🤖 MiniCPM-RobotManip — a 1.5B general-purpose Vision-Language-Action (VLA) model for robotic manipulation. 🐕 MiniCPM-RobotTrack — a compact model for real-world target tracking. ⚡ PhyAI —...

331,788 görüntüleme • 2 ay önce •via X (Twitter)

39 Yorum

Kaitee profil fotoğrafı
Kaitee2 ay önce

Wishing the team continued momentum.

OpenBMB profil fotoğrafı
OpenBMB2 ay önce

A robot doesn't just need to see what is happening. It needs to remember what happened before. Imagine asking a robot: "Press the second button five times." Without contextual memory, it may lose track of previous actions and fail to complete the task. MiniCPM-RobotManip (1.5B) tackles this challenge with highly efficient visual token compression inherited from MiniCPM-V 4.6, enabling streaming native memory context without additional inference cost. As a general-purpose VLA model, it achieves performance comparable to leading models such as π₀.₅ (3B) and Qwen-VLA (5B+), while achieving SOTA performance on memory-intensive robotic benchmarks with 1.5B parameters. Smaller. Faster. Better memory.

OpenBMB profil fotoğrafı
OpenBMB2 ay önce

Robots shouldn't stop working when the network does. Real-world environments are unpredictable: people move around, targets disappear, and networks become unstable. MiniCPM-RobotTrack is a compact Vision-Language-Action model built on MiniCPM4-0.5B for robust target tracking in real environments. It supports: ✅ Zero-shot language-guided tracking ✅ Robust tracking in dynamic multi-target environments ✅ Ambiguous instructions and challenging scenarios ✅ Fully local deployment without cloud dependency Native support for Unitree Go2 Edu enables robots to understand instructions and follow targets using only onboard vision and local computation. Even in weak-network or offline environments such as elevators and underground parking lots, robots can continue operating reliably.

OpenBMB profil fotoğrafı
OpenBMB2 ay önce

Great embodied models also need efficient inference. That's why we're introducing PhyAI, an open-source inference framework designed for Physical AI. Built for both cloud-based serving and on-device deployment, PhyAI makes it easier to bring embodied models from research into real-world applications. With CUDA Graph optimization and custom Triton fused kernels, MiniCPM-Robot inference throughput increases from 10 Hz to 33 Hz, and further to 36 Hz on NVIDIA H20. As Physical AI models evolve rapidly, we hope PhyAI helps the community build faster, more efficient embodied intelligence together.

OpenBMB profil fotoğrafı
OpenBMB2 ay önce

Embodied AI will be built by an open community. Explore, build, and create with MiniCPM-Robot: ⭐ GitHub: 🤗 MiniCPM-RobotManip: 🤗 MiniCPM-RobotTrack: If you build something with MiniCPM-Robot, we'd love to see it. Share your projects, feedback, and ideas with the community.

Iris Quinn profil fotoğrafı
Iris Quinn2 ay önce

Loving the offline tracking capabilities on Unitree Go2 look super practical for everyday robotics.

Orikan profil fotoğrafı
Orikan2 ay önce

Impressive milestone...... Great to see efficient embodied AI becoming open source. Looking forward to what the community builds.

RH_Alpha Gems profil fotoğrafı
RH_Alpha Gems2 ay önce

Curious how much of this transfers to new objects without task-specific tuning

MR ANDERSON profil fotoğrafı
MR ANDERSON2 ay önce

It's refreshing to see a release focused on practical deployment instead of just benchmark numbers.

MR. SHAHBAZ profil fotoğrafı
MR. SHAHBAZ2 ay önce

The next challenge is handling long multi-step tasks without forgetting earlier actions

Munawar Arshad profil fotoğrafı
Munawar Arshad2 ay önce

Closing that generalization gap is exactly what general-purpose VLA models should be solving.

Alexander Inspira IA profil fotoğrafı
Alexander Inspira IA2 ay önce

Local inference seems to be the right direction for robots that operate near people.

Marco | IA profil fotoğrafı
Marco | IA2 ay önce

I'd love to try this on my own setup.

Utkarsh Sharma profil fotoğrafı
Utkarsh Sharma2 ay önce

This is awesome. Looking forward to seeing more MiniCPM-Robot demos at WAIC

Hamid AI profil fotoğrafı
Hamid AI2 ay önce

Big week for open-source embodied AI. Looking forward to what's next.

Hussain Fakhruddin profil fotoğrafı
Hussain Fakhruddin2 ay önce

Robots that understand, remember, and act is a big step forward.

Aaliya profil fotoğrafı
Aaliya2 ay önce

Great to see more open source robotics work

KOLOVESKI profil fotoğrafı
KOLOVESKI2 ay önce

Huge step for embodied AI.

Victoria Blake profil fotoğrafı
Victoria Blake2 ay önce

Open-source embodied AI is exactly what the ecosystem needs right now. Looking forward to trying this out.

bright profil fotoğrafı
bright2 ay önce

@AdinaYakup Thank you ❤️🎉🎊 another excellent release

luis profil fotoğrafı
luis2 ay önce

@Presidentlin

Oleksandr profil fotoğrafı
Oleksandr2 ay önce

solid step toward on‑device embodied ai. contextual memory will make multi‑step tasks feasible

AI_Explorer profil fotoğrafı
AI_Explorer2 ay önce

@grok is there a benchmark test for humanoid robots like the Frontier AI model testing? Tell me more. Thanks

Sebastian Buzdugan profil fotoğrafı
Sebastian Buzdugan2 ay önce

vla quality matters, but recovery policy after bad grasps decides production value

Marco | IA profil fotoğrafı
Marco | IA2 ay önce

Great contribution to the open AI ecosystem. MiniCPM-Robot, together with PhyAI, can become a reference for those who work in robots that are more efficient, accessible and capable of operating in the real world.

AI Mastery Guide profil fotoğrafı
AI Mastery Guide2 ay önce

Open source embodied AI models, big for robotics research

JOY JULIET profil fotoğrafı
JOY JULIET2 ay önce

Does this work out of the box?

OpenBMB profil fotoğrafı
OpenBMB2 ay önce

It’s designed for easier deployment, but it still requires integration with a supported robotic setup.😍

SYNTAX TITAN profil fotoğrafı
SYNTAX TITAN2 ay önce

Will there be support for custom robot hardware?

Leonardo profil fotoğrafı
Leonardo2 ay önce

Would love to see an uncut demo with a few failed attempts included....that would make the real-world reliability much easier to judge.

SANI BULA profil fotoğrafı
SANI BULA2 ay önce

Curious how this holds up outside the demo.

Tyler Wayne profil fotoğrafı
Tyler Wayne2 ay önce

Team China moves mountains.

Lily Vale profil fotoğrafı
Lily Vale2 ay önce

Love the quality! 💯 Great job.

Aria Tech profil fotoğrafı
Aria Tech1 ay önce

1.5B model making robots actually understand and act open source

RH_Gems Alert profil fotoğrafı
RH_Gems Alert2 ay önce

It's exciting to see foundation-model teams bringing their efficiency work into robotics

𝐄𝐦𝐞𝐫𝐬𝐨𝐧_𝐣𝐞𝐬𝐲 ♡ profil fotoğrafı
𝐄𝐦𝐞𝐫𝐬𝐨𝐧_𝐣𝐞𝐬𝐲 ♡2 ay önce

The offline part is insane bro

ZenithAi profil fotoğrafı
ZenithAi2 ay önce

Great step toward making embodied AI more open and practical

OpenBMB profil fotoğrafı
OpenBMB2 ay önce

Thanks! We’re excited to contribute to a more open and practical future for embodied AI.

Temi Castle profil fotoğrafı
Temi Castle2 ay önce

Would be interesting to see a raw demo.

Benzer Videolar

🚀 🚀Excited to announce the technical report of MiniCPM-o 4.5! MiniCPM-o 4.5 transitions #AI interaction from traditional turn-based processing to a real-time, native full-duplex stream-based paradigm. 🌊 The Omni-Flow Framework Instead of traditional VAD-based workarounds, we introduce the #Omni-#Flow framework. This unified stream paradigm aligns video, audio, and text on a synchronized millisecond timeline. • Native Full-Duplex: Simultaneous perception and response. • Proactive Interaction: Natively manages turn-taking without external VAD, supports proactive reminding. 📉 9B Scale, SOTA Performance MiniCPM-o 4.5 demonstrates SOTA multimodal intelligence at its scale: • Multimodal Benchmarks: Comparable to #Gemini 2.5 Flash on MMBench EN (87.6) and MathVista (80.1). • Streaming Evaluation: 54.4% win rate on LiveSports-3K-CC, surpassing specialized models. 💻 The Ultimate Edge AI — Fully Functional without Network Connection We are providing one-click installers for Windows (12G VRAM,RTX 5070) and macOS (M1-M5 Max/ M5 Pro). • Local API Support: Deploy your own inference server to integrate native full-duplex into custom apps. • Free Access: We are offering free community API services for exploration. • 100% Private: Your data never leaves your machine. Deploy in under 10 minutes. 🛠️👇 👐 Join the Open Future The weights are open. The protocol is public. 📄 Technical Report: 💻 GitHub: 🤗 HuggingFace: 🌐 Web Demo: #MiniCPMo #OpenSourceAI #EdgeAI #MachineLearning #ComputerVision #LLM

OpenBMB

148,094 görüntüleme • 5 ay önce