正在加载视频...

视频加载失败

We build Cosmos-Predict2 as a world foundation model for Physical AI builders — fully open and adaptable. Post-train it for specialized tasks or different output types. Available in multiple sizes, resolutions, and frame rates. 📷 Watch the repo walkthrough ⚒️ Visit for more #NVIDIACosmos #PhysicalAI

32,040 次观看 • 1 年前 •via X (Twitter)

8 条评论

$MIA 的头像
$MIA1 年前

from AI to AgentFi, just warming up 😉

two weeks 的头像
two weeks1 年前

Still waiting for an example that demonstrates real cake

Kubilây Ân Artwork 的头像
Kubilây Ân Artwork1 年前

Very good model. I am using. We need control net for cosmos predict 2.

8GG.ETH 的头像
8GG.ETH1 年前

Your model cracks open the cosmos but can it hatch the egg? The shell of prediction is thin, the yolk of creation runs deep.

Hiro Yamamoto 的头像
Hiro Yamamoto1 年前

OpenArm at NVIDIA! Had a great time demoing the pre-release OpenArm v1.0 to 10 members of the GEAR Lab team 🦾 Big thanks to @linkevin0 & team for the thoughtful feedback. Excited to keep building together! #OpenSource #robotics #humanoids #physicalAI #NVIDIA

AK 的头像
AK1 年前

WorldVLA Towards Autoregressive Action World Model

C's Robotics Paper Notes 的头像
C's Robotics Paper Notes1 年前

Compliant Residual DAgger: Improving Real-World Contact-Rich Manipulation with Human Corrections human correction as the expert signal in dagger to finetune the policy, supported by compliant hardware design.

Debidatta Dwibedi 的头像
Debidatta Dwibedi1 年前

Our Vision-Language-Action robot demo at #RSS2025 was eye-opening. The ultimate eval for any generalist model: new environment, new objects from audience, and new instructions. For the first time it really hit me: what if we've been underestimating what these models can do?

相关视频

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,078,418 次观看 • 2 个月前