Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Cosmos 3: Our latest frontier model for Physical AI Cosmos 3 is the world’s first fully open omnimodel with native vision reasoning, world and action generation. Today we’re releasing Super (32B) and Nano (8B) variants.

429,077 просмотров • 4 месяцев назад •via X (Twitter)

Комментарии: 39

Фото профиля NVIDIA AI
NVIDIA AI4 месяцев назад

Cosmos 3 ties everything together. Previous releases separated world generation, physical understanding, and controlled scene generation. Cosmos 3’s MoT architecture unifies these capabilities by pairing an autoregressive reasoner tower with a diffusion-based generator tower.

Фото профиля NVIDIA AI
NVIDIA AI4 месяцев назад

Delivering leading results on physical AI benchmarks among open models, it ranks first across @ArtificialAnlys, Physics-IQ, PAI-Bench and R-Bench for world generation accuracy, RoboLab and RoboArena for action policy and the VANTAGE-Bench and TAR leaderboards for vision understanding.

Фото профиля NVIDIA AI
NVIDIA AI4 месяцев назад

In addition to understanding and reasoning across modalities, Cosmos 3 excels at simulating physical environments, predicting future world states, and helping train robots to perform specific tasks. It can do subsecond vision reasoning, large scale synthetic data generation and simplified robot learning policy development.

Фото профиля NVIDIA AI
NVIDIA AI4 месяцев назад

Image-to-video generation is just as impressive. Input image: "Generate a 16:9 image from a dashcam view of a formula 1 racing event" Video prompt: "A high-speed racing event where a car navigates multiple winding turns" 🔊 Sound on - generated by Cosmos 3.

Фото профиля NVIDIA AI
NVIDIA AI4 месяцев назад

Trained on billions of samples across modalities, the model provides developers with a powerful pretrained foundation for building physical AI systems with less data and lower training costs. Read more in our technical blog:

Фото профиля NVIDIA AI
NVIDIA AI4 месяцев назад

As always, Cosmos 3 is fully open. This includes model weights and post-training recipes. Available now on @huggingface

Фото профиля Filecoin
Filecoin4 месяцев назад

Physical AI produces the data that robots learn from, like simulations, sensors, etc. It needs storage that’s verifiable and persistent.

Фото профиля Erik Kokalj
Erik Kokalj4 месяцев назад

Tested NVIDIA's new Cosmos 3 myself. Easy tier first: zero-shot gate turnaround. Single inference of 1.5min video, 0.5fps ingestion, 6 segmented states. Thinking on, no fine-tuning, prompted it to focus on cargo in the cropped ROI. Harder tests below. 🧵

Фото профиля Ofek Shaked | AI Engineer
Ofek Shaked | AI Engineer4 месяцев назад

World and action generation models are moving faster than most infra teams have tooling to consume them. The gap between research release and production agent integration is still wide.

Фото профиля PrismaX
PrismaX4 месяцев назад

This changes who can build for physical AI!

Фото профиля Arindam Majumder 𝕏
Arindam Majumder 𝕏4 месяцев назад

Actually cool Tried out the Super model on @nebiustf It's Pretty cool

Фото профиля Rise-Raise
Rise-Raise4 месяцев назад

Physical AI is where AI starts interacting with the real world, not just understanding it. Teaching machines and robots to perceive, reason, and act efficiently in dynamic environments is one of the most important challenges in technology today. The closer we get to bridging intelligence and action, the closer we get to AI that can genuinely assist humans in the physical world.

Фото профиля cv usk
cv usk4 месяцев назад

NVIDIA Cosmos 3 is a massive game-changer for overcoming the sim-to-real gap. Personally, I would use its Cosmos-Transfer tool to build an automated pipeline that generates infinite, ultra-realistic edge-case data (like rare factory hazards or autonomous driving accidents).

Фото профиля Chihuahua ($HUAHUA)
Chihuahua ($HUAHUA)4 месяцев назад

There is only one @cosmos woof!

Фото профиля CrazyAI Tech
CrazyAI Tech4 месяцев назад

Must try!!

Фото профиля JSFILMZ
JSFILMZ4 месяцев назад

LFG

Фото профиля Xavier
Xavier4 месяцев назад

You should check out the open source Project telos. We’re working on training ml policies via a webcam for teleoperation rather than using hardware rigs . The technology is still developing but it’s a promising venture .

Фото профиля Alina Fomina
Alina Fomina4 месяцев назад

#1 across every major physical ai leaderboard + already running on dgx spark within hours. heard here:

Фото профиля Nitin Bisht
Nitin Bisht4 месяцев назад

NVIDIA is putting frontier model capabilities on the table for free.

Фото профиля Patrick
Patrick4 месяцев назад

Fully open omnimodel is the kind of phrase that means one thing to the engineers and another to the lawyers.

Фото профиля Ferbin
Ferbin4 месяцев назад

32B looks solid. For real robot control the question is latency: how far can it predict in the 50ms between sensing and acting? Model size is not the bottleneck. Latency is.

Фото профиля Subhan Tariq
Subhan Tariq4 месяцев назад

Open-source AI is moving so fast that my weekend project is outdated before Monday.

Фото профиля Amit Jain
Amit Jain4 месяцев назад

Amazing!!

Фото профиля Suyash Kelvin Savant
Suyash Kelvin Savant4 месяцев назад

Wow this is one of a kind model !! Will look into it ASAP..

Фото профиля Macro Bombastic
Macro Bombastic4 месяцев назад

Nice to see NVIDIA embracing open source for physical AI, that's the kind of competition that keeps markets honest

Фото профиля Reza Sayar
Reza Sayar4 месяцев назад

Wanna try out the Nano model on <16GB memory? get your quants here! 🤗

Фото профиля DopaMint
DopaMint4 месяцев назад

everyone building world models we're building AI companions people actually remember

Фото профиля breezePeak
breezePeak4 месяцев назад

这么小的参数,就可以做到这样,太厉害了吧

Фото профиля Saquib Mehmood
Saquib Mehmood4 месяцев назад

Cool.

Фото профиля Nate Brown
Nate Brown4 месяцев назад

This is awesome. Thank you!

Фото профиля Iris Blake
Iris Blake4 месяцев назад

The world's first fully open physical AI omnimodel — the 32B Super version is now open-sourced. Things are about to heat up in the physical AI community 🔥

Фото профиля Ivan Lim
Ivan Lim4 месяцев назад

The future of AI isn't another chatbot. It's AI that can see the world, understand the world, and act within the world. That's why nearly every major AI company is suddenly talking about robotics.

Фото профиля cv usk
cv usk4 месяцев назад

By plugging this directly into a CI/CD loop, we can safely stress-test and train drones or robotic arms in software before they ever touch the physical world. I truly believe this is going to save so much expensive hardware from crashing! 🤖🏭🚗

Фото профиля Léo
Léo4 месяцев назад

I'm really curious about two things here: 1) Why not use the term "world model"? 2) How does the training of this model enforce the learning of physics? Does it use masked videos like V-JEPA?

Фото профиля Subhan Tariq
Subhan Tariq4 месяцев назад

The combination of vision reasoning, world generation, and action generation in a fully open model is the part that stands out here.

Фото профиля Mit Perin
Mit Perin4 месяцев назад

NVIDIA changing the world 🙏🏾

Фото профиля MicroDegree
MicroDegree4 месяцев назад

The next frontier of AI isn't just understanding text, images, or code. It's understanding and interacting with the physical world. Models that can reason about vision, actions, environments, and outcomes may become foundational building blocks for robotics, autonomous systems, and embodied AI.

Фото профиля Kingston Kuan
Kingston Kuan4 месяцев назад

very excited to see how the traces work for this omni model

Фото профиля Kaustubh Joshi
Kaustubh Joshi4 месяцев назад

Models like Cosmos 3 are getting incredibly powerful. Now it is all about handling prompts correctly to get the best out of them. That's exactly why I built PromptCentral — a hub to craft, organize and reuse well-structured prompts across frontier models so you're always getting consistent, quality outputs. Here's one to get started:

Похожие видео

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,080,061 просмотров • 4 месяцев назад