Loading video...

Video Failed to Load

Go Home

Woow Nvidia has just released a 2.6B open-source world model 🔥 You can turn a single image, text prompt and trajectory into controllable worlds... And on a single GPU! - Code available on GitHub - Paper as well on arxiv You can use it for many things like embodied...

174,447 views • 4 months ago •via X (Twitter)

41 Comments

Paul Couvert's profile picture
Paul Couvert4 months ago

You can learn more about it here: - Project: - GitHub: - Paper: World models are evolving so fast! Soon we'll have a Genie 3 equivalent running at home.

Dennis McLeod 🇺🇸's profile picture
Dennis McLeod 🇺🇸4 months ago

lol. “Everyone” that has a 4000.00 video card…..

Paul Couvert's profile picture
Paul Couvert4 months ago

Still a consumer GPU. You can rent one online for like $0.40/hr.

Weever's profile picture
Weever4 months ago

@dennis_mcleod How?

Paul Couvert's profile picture
Paul Couvert4 months ago

@dennis_mcleod Using platforms like vast ai, runpod or novita

DO Chain's profile picture
DO Chain4 months ago

@amerukraine Does it have free api calls too 😬

Paul Couvert's profile picture
Paul Couvert4 months ago

Brand new, so maybe we'll have a cheap API soon if a neo cloud is hosting it.

DO Chain's profile picture
DO Chain4 months ago

Super clutch! Thank you !

Alpha Batcher's profile picture
Alpha Batcher4 months ago

Nvidia cooking as well in last days big W

Paul Couvert's profile picture
Paul Couvert4 months ago

They're doing a lot for the open source ecosystem!

Richard Deetlefs's profile picture
Richard Deetlefs4 months ago

Did the release the model? I cannot find it

Supersocks's profile picture
Supersocks4 months ago

@llmgram

OmnipotentCEO's profile picture
OmnipotentCEO4 months ago

wow

Eclipse 🌖's profile picture
Eclipse 🌖4 months ago

Interesting release. 2.6B params on single GPU is a meaningful efficiency benchmark for embodied AI research. Curious how the trajectory control generalizes across unseen environments.

Paul Couvert's profile picture
Paul Couvert4 months ago

Yes exactly. The efficiency is quite incredible. Usually you need powerful cloud capacities to run this kind of model.

Sirfer - ducking under the lip of life's profile picture
Sirfer - ducking under the lip of life4 months ago

@C4ndide Everyone who has 5 or 6 thousand dollars for a GPU

Parsi AI 🇮🇷's profile picture
Parsi AI 🇮🇷4 months ago

Not for most

Junior's profile picture
Junior4 months ago

Nano banana + this is probably goated

Overton Pusher's profile picture
Overton Pusher4 months ago

At some pointpoint NVIDIA will go back to targeting prosumers. If the can open source mini models, they could give the big dogs something to worry about

Lemin's profile picture
Lemin4 months ago

Crazy how these models getting more accessible each day

Oskar Minor's profile picture
Oskar Minor4 months ago

It looks awesome!

Auro Zera's profile picture
Auro Zera4 months ago

This is great. 2.6B can even run on a 250 dollar 8GB VRAM card

Dimi's profile picture
Dimi4 months ago

Another day, another world model I'll probably spend a weekend failing to integrate into something useful. My GPU is ready for the abuse. One tiny project closer.

Peter Meyer 💙🥝🕊️'s profile picture
Peter Meyer 💙🥝🕊️4 months ago

Nothing is released. Where are the Models? Inference Code?

geo ppls's profile picture
geo ppls4 months ago

That’s huge

ReAiLity Labs's profile picture
ReAiLity Labs4 months ago

Will it run on RTX 4090?

Tim's profile picture
Tim4 months ago

The models aren’t released yet 😢

Steven's profile picture
Steven4 months ago

This is truly groundbreaking news from Nvidia. Releasing a 2.6B open-source world model that can run efficiently on a single GPU is a massive leap forward for democratization in AI. The potential applications for robotics, immersive simulations, and interactive environments are absolutely immense. Thank you for sharing this exciting update!

Artificial Unboxed's profile picture
Artificial Unboxed4 months ago

This is wild. Single GPU world models just became real

Enzo's profile picture
Enzo4 months ago

everyone? bro, no one I know has a h100 just sitting around

Arslan Yousaf's profile picture
Arslan Yousaf4 months ago

Running an open-source world model on a single GPU is a huge accessibility moment for AI research

HouseOfUkiro's profile picture
HouseOfUkiro4 months ago

Road to VR immersion

Paul Couvert's profile picture
Paul Couvert4 months ago

Holodeck confirmed

AI Mastery Guide's profile picture
AI Mastery Guide4 months ago

Image plus text prompt plus trajectory into a controllable world is a sentence that would have sounded like sci-fi two years ago

Toucan's profile picture
Toucan4 months ago

@machinelearnflx So cool 😎

Stjepan's profile picture
Stjepan4 months ago

are you a shill ... "things like embodied AI and robotics research" .... did you ever do any robotic research dude .. no.. you just repeat like a parrot

TheOrder's profile picture
TheOrder4 months ago

Mesmerized

CrazyAI Tech's profile picture
CrazyAI Tech4 months ago

Magic time

NFAdancer777's profile picture
NFAdancer7774 months ago

@grok

Alan Silva's profile picture
Alan Silva4 months ago

This is amazing!

Saeed Anwar's profile picture
Saeed Anwar1 month ago

Controllable world generation from a single image on one GPU being open-source removes the biggest simulation bottleneck for robot training. What's the visual fidelity ceiling?

Related Videos

We made a thing! Very happy to announce sqlcoder-pro and the Defog Alignment Platform. Available to use immediately without a wait-list, weights will be open-sourced very soon. The video does a quick show and tell comparison against ChatGPT (with gpt-4o). Read on for more details! TLDR 💪 equal (or better) performance on text-to-SQL as the most capable Claude-3.5 or GPT-4 models 🤝 You can use it today on a free plan/free trial, without a waitlist 🪽 self-hostable on a single RTX4090, with 2 second median generation times for SQL queries 🔁 exactly the same output every time, give the same prompt 👨🏻‍🏫 teachable and steerable: show the model what you want it to do 🛞 debuggable – you can understand WTF is going on inside the model, instead of treating it like a black box Let's dig into each of these one-by-one! Performance SQLCoder-8b-pro significantly exceeds the performance of our previous sqlcoder-8b model on Postgres text-to-SQL (from 88.2% to 90.2% accuracy - gpt-4o is at 87.6%, for reference). It is also better at following instructions. This was done via self-merges, hand crafted fine-tuning data, and adapting the training data to fit our tokenizer. Cost You can host this on the model on a single $3,500 RTX4090, and support ~5 requests/second via VLLM. If you're looking to host on the cloud instead, you can run it on a single L4 GPU that costs $300/mo on GCP Repeatability We have a dense 8b model with no MoE shenanigans. For the same prompt with temperature=0, you'll always get the same answer – which is critical in BI. Teachable In our alignment and feedback modes, you can give the model feedback on how it answered certain questions, and it will automatically adapt to the feedback. Debuggable You can use logprobs and attention scores to determine where, exactly is the model paying attention to inside a prompt + what it's getting confused by when generating outputs. Available today You can use Defog on the cloud today by going to docs[dot]defog[dot]ai, and getting an API key. Excited to hear what you think!

Rishabh Srivastava

13,469 views • 2 years ago