Загрузка видео...

Не удалось загрузить видео

На главную

NVIDIA AI Open-Sources ViPE (Video Pose Engine): A Powerful and Versatile 3D Video Annotation Tool for Spatial AI ViPE integrates bundle adjustment with dense optical flow, sparse keypoint tracking, and metric depth priors to estimate camera intrinsics, poses, and dense depth maps at 3–5 FPS on a single GPU....

217,453 просмотров • 11 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Trained a humanoid entirely in a 3D scan of the office. Zero real-world fine-tuning. It just walked in and worked. RL needs hundreds of thousands of attempts, and real robots can't afford to crash. A misjudged gap or a glass door collision breaks hardware and costs hours resetting. So you train in a sim. But sim policies usually train on randomized, untextured geometry; depth is easy to fake. The robot learns structure, not the real world: no materials, no lighting, no idea what anything actually is. RGB cameras carry all of that but training RGB policies in generic fake worlds won’t generalize to the real world. Niantic Spatial 🌎 Scaniverse reconstructs your scan of the real deployment site. One 360° camera walkthrough → photorealistic 3D Gaussian splat at metric scale → collision mesh pulled from the same reconstruction, so vision and physics match exactly. Drops straight into NVIDIA Isaac Sim/Lab, no manual conversion. Flexion simulation-first approach then seamlessly enables the training of RGB-only nav policies inside that reconstruction. With added domain randomization + large image encoders for robustness, this deploys straight to hardware. No real-world fine-tuning. Deployment: months of on-site adaptation → days. Tune into the NVIDIA livestream on 12 August to hear how these companies are closing the sim2real gap: NVIDIA Robotics ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

119,356 просмотров • 28 дней назад