Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

NVIDIA AI Open-Sources ViPE (Video Pose Engine): A Powerful and Versatile 3D Video Annotation Tool for Spatial AI ViPE integrates bundle adjustment with dense optical flow, sparse keypoint tracking, and metric depth priors to estimate camera intrinsics, poses, and dense depth maps at 3–5 FPS on a single GPU....

217,453 görüntüleme • 11 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Trained a humanoid entirely in a 3D scan of the office. Zero real-world fine-tuning. It just walked in and worked. RL needs hundreds of thousands of attempts, and real robots can't afford to crash. A misjudged gap or a glass door collision breaks hardware and costs hours resetting. So you train in a sim. But sim policies usually train on randomized, untextured geometry; depth is easy to fake. The robot learns structure, not the real world: no materials, no lighting, no idea what anything actually is. RGB cameras carry all of that but training RGB policies in generic fake worlds won’t generalize to the real world. Niantic Spatial 🌎 Scaniverse reconstructs your scan of the real deployment site. One 360° camera walkthrough → photorealistic 3D Gaussian splat at metric scale → collision mesh pulled from the same reconstruction, so vision and physics match exactly. Drops straight into NVIDIA Isaac Sim/Lab, no manual conversion. Flexion simulation-first approach then seamlessly enables the training of RGB-only nav policies inside that reconstruction. With added domain randomization + large image encoders for robustness, this deploys straight to hardware. No real-world fine-tuning. Deployment: months of on-site adaptation → days. Tune into the NVIDIA livestream on 12 August to hear how these companies are closing the sim2real gap: NVIDIA Robotics ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

119,356 görüntüleme • 28 gün önce