正在加载视频...

视频加载失败

Figure’s Helix 2.5 shows how pretraining on human behavior data can improve a humanoid robot’s ability to work in unfamiliar environments. In Figure’s test, the Index-pretrained model achieved 56% zero-shot success across 30 unseen homes, compared with just 9% for a model trained from scratch. The key takeaway is...

58,425 次观看 • 11 天前 •via X (Twitter)

5 条评论

モモフ好き 日本には酸っぱい🍋がいっぱいw 的头像
モモフ好き 日本には酸っぱい🍋がいっぱいw10 天前

Helixは、米Figure AIが開発したヒューマノイドロボット向けの最先端のVision-Language-Action(VLA)AI基盤

Patriota Volta 的头像
Patriota Volta11 天前

56% vs 9% just from pretraining on human behavior. the gap is ridiculous

lulu 的头像
lulu10 天前

Treat the robot with respect. Also it's very slow.

Nick Champrenault 的头像
Nick Champrenault10 天前

56% against 9% is a real gap and still a coin flip in someone's kitchen. the number that decides deployment is which 44% failed and whether those failures cluster. random failure and dark-countertop failure are different problems.

Crogon 的头像
Crogon11 天前

i cannot be the only one who finds that ai voice the most annoying among them all.

相关视频

JUST IN: Dyna Robotics just published one of the most important research papers in robotics this year. It could fundamentally change how robot foundation models are trained. A scaling law that transfers from human video to robot performance. Dyna-2 is out and it's 🔥 Here's what that means in plain terms. Dyna-2 was pre-trained on ONE MILLION hours of egocentric human video, 170 years of continuous human experience, cooking, folding, assembling, cleaning. And as that human data scaled, robot performance improved. Predictably. Monotonically. Across 39 tasks on two different robot embodiments the model had never seen. → 1,000 hours pre-training → 20% normalised task performance → 10,000 hours → 28% → 100,000 hours → 45% → 1,000,000 hours → 53% Human video exists at effectively unlimited scale. Every cook, every factory worker, every craftsperson wearing a camera is generating training data for future robots. But the finding that stunned even the researchers, world modeling is what makes the transfer work. A model trained to predict future video AND actions massively outperforms one trained on actions alone. Video is the new scaling axis for robotics. One more jaw-dropping data point. 13 minutes of teleoperation data was enough to fine-tune Dyna-2 to open a bottle cap using two five-fingered robot hands. The robots are coming, and they're learning from us directly :D Read more here: Congrats Jason Ma and team! ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

23,681 次观看 • 1 个月前

NEWS: Humanoid robotics company Figure has released Helix 02, what they claim in their most capable humanoid model yet. "A single neural system that controls the full body directly from pixels, enabling dexterous, long horizon autonomy across an entire room: • Autonomous, long‑horizon loco-manipulation: Helix 02 unloads and reloads a dishwasher across a full-sized kitchen - a four-minute, end-to-end autonomous task that integrates walking, manipulation, and balance with no resets and no human intervention. We believe this is the longest horizon, most complex task completed autonomously by a humanoid robot to date. • All sensors in. All actuators out: Helix 02 connects every onboard sensor - vision, touch, and proprioception - directly to every actuator through a single unified visuomotor neural network. • Human-like whole body control from human data: All results are enabled by System 0, a learned whole‑body controller trained on over 1,000 hours of human motion data and sim‑to‑real reinforcement learning. System 0 replaces 109,504 lines of hand‑engineered C++ with a single neural prior for stable, natural motion. • New classes of dexterity: With Figure 03’s embedded tactile sensing and palm cameras, Helix 02 performs manipulation that was previously out of reach: extracting individual pills, dispensing precise syringe volumes, and singulating small, irregular objects from clutter despite self‑occlusion. Helix 02 is trained on over 1,000 hours of human motion data and integrates vision, touch, and proprioception."

Sawyer Merritt

624,910 次观看 • 8 个月前