正在加载视频...

视频加载失败

1/5 🚀 Thrilled to open-source OSCAR 🤖 — an action-conditioned world model for robotics, led by the visiting student in my group Zhuoyuan Wu! It generalizes across different robot embodiments with precise action controllability. All trained on a single GH200 GPU, and outperforms existing open-sourced baselines, which have larger...

104,967 次观看 • 3 个月前 •via X (Twitter)

16 条评论

Jun Gao 的头像
Jun Gao3 个月前

2/5 📊 Data. Data is the key to success🔑 We built a large-scale, standardized data pipeline that curates, filters, and dedups broad public robotics and egocentric human datasets. This gives us a clean training dataset with diverse tasks, scenarios, actions, and embodiments.

Jun Gao 的头像
Jun Gao3 个月前

3/5 🦾 Method. We leverage a 2D kinematic skeleton rendering as a unified control signal. From robot arms to human hands ✋, every embodiment is rendered as an image-aligned 2D skeleton. This enables high-precision action conditioning for video models. 🎯

Jun Gao 的头像
Jun Gao3 个月前

4/5 📈 Results. With our cleaned data and kinematic skeleton conditioning, we finetune the Cosmos2.5-2B model on a single GH200 GPU ⚡. Our method delivers superior quality on action following, appearance, and motion consistency, pushing SOTA performance with a fraction of the compute. 🏆

Jun Gao 的头像
Jun Gao3 个月前

5/5 🧪 Robot policy evaluation. Finally, we deploy OSCAR for real-world robot policy evaluation using the episodes from RoboArena. Our extensive experiments show a high correlation ✅ between virtual assessment in our world model and real-world results, paving the way for a future where robot policies are evaluated entirely within generated environments!!

Haotian Xue 的头像
Haotian Xue3 个月前

@wuzy2115 Impressive, it seems much more efficient. Is that because it introduces no new modules?

Jun Gao 的头像
Jun Gao3 个月前

@wuzy2115 No new modules is one part of the efficiency. The most useful things are: (i) curating a balanced and diverse dataset, (ii) skeleton conditioning as a unified action representation, and (iii) the base model is only a 2B model (Cosmos 2.5)

Vai Viswanathan 的头像
Vai Viswanathan3 个月前

@wuzy2115 Have you looked into multi-camera generation?

Jun Gao 的头像
Jun Gao2 个月前

@wuzy2115 We haven't looked into it yet, but it's a great future direction!

Oscar Brisset 的头像
Oscar Brisset3 个月前

@wuzy2115 Love the name

Albert Zhai 的头像
Albert Zhai3 个月前

@wuzy2115 the skeleton is an interesting representation, i wonder if there are any ambiguity problems when the three red arrows are colinear (their plane intersects camera center)?

Jun Gao 的头像
Jun Gao2 个月前

@wuzy2115 Good question! Yes, this will indeed cause the ambiguity; one potential way of fixing it can be to use three planes to indicate the gipper (e.g., the triplane idea, and each plane has three red arrows).

Virgil Maro 的头像
Virgil Maro3 个月前

@wuzy2115 a world model is the agent rehearsing what happens if i do x before committing. video prediction is just what that rehearsal looks like from outside

AiDevCraft 的头像
AiDevCraft3 个月前

The 2D skeleton-as-control choice is the clever part — it forces the conditioning channel to be embodiment-shape-invariant, so the dynamics head cannot smuggle embodiment cues through the action stream. Curious how that holds up under DoF mismatches, e.g. a parallel gripper and a five-finger hand sharing the same skeleton abstraction.

sora 的头像
sora3 个月前

@wuzy2115 I understand that you guys trained the model with inputs of the 2d skeleton and the first frame to generate video outputs. My question is that during inference when you have to do a robot action for a certain input frame, you need the 2D skeleton right! How do you accurately get that ??

Tatsuya Matsushima #ICRA2026 @Tokyo Bay Area 🍣 的头像
Tatsuya Matsushima #ICRA2026 @Tokyo Bay Area 🍣3 个月前

@wuzy2115 Great work! Thanks for using our dataset, AIRoA MoMa dataset @airoa_org !

EB1A Experts 的头像
EB1A Experts3 个月前

@wuzy2115 The combination of cross-embodiment generalization and full reproducibility stands out here. It's encouraging to see strong robotics results paired with open data, open weights, and a compute budget that remains accessible to the broader research community.

相关视频

JUST IN: Dyna Robotics just published one of the most important research papers in robotics this year. It could fundamentally change how robot foundation models are trained. A scaling law that transfers from human video to robot performance. Dyna-2 is out and it's 🔥 Here's what that means in plain terms. Dyna-2 was pre-trained on ONE MILLION hours of egocentric human video, 170 years of continuous human experience, cooking, folding, assembling, cleaning. And as that human data scaled, robot performance improved. Predictably. Monotonically. Across 39 tasks on two different robot embodiments the model had never seen. → 1,000 hours pre-training → 20% normalised task performance → 10,000 hours → 28% → 100,000 hours → 45% → 1,000,000 hours → 53% Human video exists at effectively unlimited scale. Every cook, every factory worker, every craftsperson wearing a camera is generating training data for future robots. But the finding that stunned even the researchers, world modeling is what makes the transfer work. A model trained to predict future video AND actions massively outperforms one trained on actions alone. Video is the new scaling axis for robotics. One more jaw-dropping data point. 13 minutes of teleoperation data was enough to fine-tune Dyna-2 to open a bottle cap using two five-fingered robot hands. The robots are coming, and they're learning from us directly :D Read more here: Congrats Jason Ma and team! ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

23,576 次观看 • 1 个月前

🔥 JUST IN: Open-source robotics dataset from 100% real-world scenarios! 🤯 Chinese robotics company AGIBOT just released AGIBOT WORLD 2026, an open-source dataset systematically covering key embodied AI research directions. Built entirely from real-world environments: commercial spaces, and homes. Collected using AGIBOT G2 robots in free-form collection mode, providing structured, accurately annotated, high-quality data. Digital twin technology creates 1:1 scale replicas in simulation matching the real environments. Both real-world and simulation data are open-sourced. The AGIBOT G2 platform collects multiple data types simultaneously: RGB(D) cameras, tactile sensors, force sensors, LiDAR, IMU, and full-body joint states. Whole-body control coordinates arms, waist, and hands for complex tasks. First-person teleoperation lets operators control the robot from its perspective. The tasks covered are fine-grained manipulation, ultra-long-horizon tasks, spatial navigation, dual-arm coordination, and multi-agent/human-robot collaboration. The dataset includes error-recovery trajectories with annotations. Most datasets only show successful demonstrations. AGIBOT includes failures and how the robot recovers, teaching models how to handle mistakes. After collection, data is tested through policy training and real-robot deployment to ensure quality. Then processed through industrial quality control with multiple screening and cleaning rounds. Making it open-source accelerates embodied AI research by giving researchers access to high-quality real-world robot data at scale. 🇨🇳 Learn more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

40,583 次观看 • 5 个月前

Synthetic data will provide the next trillion tokens to fuel our hungry models. I'm excited to announce MimicGen: massively scaling up data pipeline for robot learning! We multiply high-quality human data in simulation with digital twins. Using 50,000 training episodes across 18 tasks, multiple simulators, and even in the real-world! The idea is simple: 1. Humans tele-operate the robot to complete a task. It is extremely high-quality but also very slow and expensive. 2. We create a digital twin of the robot and the scene in high-fidelity, GPU-accelerated simulation. 3. We can now move objects around, replace with new assets, and even change the robot hand - basically augment the training data with procedural generation. 4. Export the successful episodes, and feed that to a neural network! You now have an near-infinite stream of data. One of the key reasons that robotics lags far behind other AI fields is the lack of data: you cannot scrape control signals from the internet. They simply don't exist in-the-wild. MimicGen shows the power of synthetic data and simulation to keep our scaling laws alive. I believe this principle apply beyond robotics. We are quickly exhausting the high-quality, real tokens from the web. Artificial intelligence from artificial data will be the way forward. We are big fans of the OSS community. As usual, we open-source everything, including the generated dataset! - Website: - Paper: - Dataset is hosted on HuggingFace (thanks AK!!): - Code: MimicGen is led by Ajay Mandlekar, deep dive in the thread:

Jim Fan

332,238 次观看 • 2 年前