Загрузка видео...
Не удалось загрузить видео
Introducing Cosmos 3: Our latest frontier model for Physical AI Cosmos 3 is the world’s first fully open omnimodel with native vision reasoning, world and action generation. Today we’re releasing Super (32B) and Nano (8B) variants.
429,077 просмотров • 4 месяцев назад •via X (Twitter)
Комментарии: 39

Cosmos 3 ties everything together. Previous releases separated world generation, physical understanding, and controlled scene generation. Cosmos 3’s MoT architecture unifies these capabilities by pairing an autoregressive reasoner tower with a diffusion-based generator tower.

Delivering leading results on physical AI benchmarks among open models, it ranks first across @ArtificialAnlys, Physics-IQ, PAI-Bench and R-Bench for world generation accuracy, RoboLab and RoboArena for action policy and the VANTAGE-Bench and TAR leaderboards for vision understanding.

In addition to understanding and reasoning across modalities, Cosmos 3 excels at simulating physical environments, predicting future world states, and helping train robots to perform specific tasks. It can do subsecond vision reasoning, large scale synthetic data generation and simplified robot learning policy development.

Image-to-video generation is just as impressive. Input image: "Generate a 16:9 image from a dashcam view of a formula 1 racing event" Video prompt: "A high-speed racing event where a car navigates multiple winding turns" 🔊 Sound on - generated by Cosmos 3.

Trained on billions of samples across modalities, the model provides developers with a powerful pretrained foundation for building physical AI systems with less data and lower training costs. Read more in our technical blog:

As always, Cosmos 3 is fully open. This includes model weights and post-training recipes. Available now on @huggingface

Physical AI produces the data that robots learn from, like simulations, sensors, etc. It needs storage that’s verifiable and persistent.

Tested NVIDIA's new Cosmos 3 myself. Easy tier first: zero-shot gate turnaround. Single inference of 1.5min video, 0.5fps ingestion, 6 segmented states. Thinking on, no fine-tuning, prompted it to focus on cargo in the cropped ROI. Harder tests below. 🧵

World and action generation models are moving faster than most infra teams have tooling to consume them. The gap between research release and production agent integration is still wide.

This changes who can build for physical AI!

Actually cool Tried out the Super model on @nebiustf It's Pretty cool

Physical AI is where AI starts interacting with the real world, not just understanding it. Teaching machines and robots to perceive, reason, and act efficiently in dynamic environments is one of the most important challenges in technology today. The closer we get to bridging intelligence and action, the closer we get to AI that can genuinely assist humans in the physical world.

NVIDIA Cosmos 3 is a massive game-changer for overcoming the sim-to-real gap. Personally, I would use its Cosmos-Transfer tool to build an automated pipeline that generates infinite, ultra-realistic edge-case data (like rare factory hazards or autonomous driving accidents).

There is only one @cosmos woof!

Must try!!

LFG

You should check out the open source Project telos. We’re working on training ml policies via a webcam for teleoperation rather than using hardware rigs . The technology is still developing but it’s a promising venture .

#1 across every major physical ai leaderboard + already running on dgx spark within hours. heard here:

NVIDIA is putting frontier model capabilities on the table for free.

Fully open omnimodel is the kind of phrase that means one thing to the engineers and another to the lawyers.

32B looks solid. For real robot control the question is latency: how far can it predict in the 50ms between sensing and acting? Model size is not the bottleneck. Latency is.

Open-source AI is moving so fast that my weekend project is outdated before Monday.

Amazing!!

Wow this is one of a kind model !! Will look into it ASAP..

Nice to see NVIDIA embracing open source for physical AI, that's the kind of competition that keeps markets honest

Wanna try out the Nano model on <16GB memory? get your quants here! 🤗

everyone building world models we're building AI companions people actually remember

这么小的参数,就可以做到这样,太厉害了吧

Cool.

This is awesome. Thank you!

The world's first fully open physical AI omnimodel — the 32B Super version is now open-sourced. Things are about to heat up in the physical AI community 🔥

The future of AI isn't another chatbot. It's AI that can see the world, understand the world, and act within the world. That's why nearly every major AI company is suddenly talking about robotics.

By plugging this directly into a CI/CD loop, we can safely stress-test and train drones or robotic arms in software before they ever touch the physical world. I truly believe this is going to save so much expensive hardware from crashing! 🤖🏭🚗

I'm really curious about two things here: 1) Why not use the term "world model"? 2) How does the training of this model enforce the learning of physics? Does it use masked videos like V-JEPA?

The combination of vision reasoning, world generation, and action generation in a fully open model is the part that stands out here.

NVIDIA changing the world 🙏🏾

The next frontier of AI isn't just understanding text, images, or code. It's understanding and interacting with the physical world. Models that can reason about vision, actions, environments, and outcomes may become foundational building blocks for robotics, autonomous systems, and embodied AI.

very excited to see how the traces work for this omni model

Models like Cosmos 3 are getting incredibly powerful. Now it is all about handling prompts correctly to get the best out of them. That's exactly why I built PromptCentral — a hub to craft, organize and reuse well-structured prompts across frontier models so you're always getting consistent, quality outputs. Here's one to get started:



