
Hao Zhao
@HaoZhao_AIRSUN • 1,402 subscribers
https://t.co/l52lTpdO16 Computer vision is good, have fun.
Videos

🚨 A single image → a fully editable UE5 world. Introducing Lumera: a turnkey pipeline that reconstructs an engine-native 3D scene directly from ONE image — not just pixels, point clouds, or Gaussians. It gives you: → object-level 3D meshes → editable 3D layouts → engine-native lights → HDR environment lighting → a scene ready for Unreal Engine 5 In other words: Image → 3D world → UE5. We believe the future of image-to-3D is not “generate something that looks 3D.” It is generate the actual world project. 🌐 📄
Hao Zhao51,740 просмотров • 13 дней назад

🔥 #ICRA2026 Best Paper Finalist The era of "robot VLA = single-arm gripper" is ending. Introducing Dexora — the first open-source Vision-Language-Action system for dual-arm, dual-hand, 36-DoF dexterous manipulation. 🦾 Dual Arms 🖐️ Dual Hands 🎯 36 DoF Control 🌍 Open Source Trained on: • 100K simulated trajectories • 10K real-world demonstrations Dexora achieves: ✓ 90%+ success on basic manipulation ✓ Strong dexterous manipulation performance ✓ Cross-embodiment generalization Our key hypothesis: Train on the hardest embodiment. Transfer to simpler robots later. Instead of scaling up gripper policies, we train directly in the most expressive action space and project downward to simpler embodiments. This may be a practical path toward universal robot controllers. 🎥 Demos: 📄 Paper:
Hao Zhao17,474 просмотров • 3 месяцев назад

#CVPR2026 Can frontier LLMs write PhD-level 3D vision code? We introduce GeoCodeBench, a benchmark that asks models to read real 3D geometric vision papers and implement core functions. Best result so far: GPT-5 reaches only 36.6%. This suggests that scientific coding in 3D vision remains far from solved. Paper: Project:
Hao Zhao18,273 просмотров • 3 месяцев назад

If you’re excited by Tesla’s new world model, meet OmniNWM—our research take on panoramic, controllable driving world models • Ultra-long demos • precise camera control • RGB/semantics/depth/occupancy • intrinsic closed-loop rewards Arxiv: Watch: #WorldModel #AutonomousDriving #GenerativeAI #Diffusion #3DVision
Hao Zhao33,186 просмотров • 10 месяцев назад

🚀 Agility Meets Stability (AMS) — one unified policy for humanoids that can dance, run, and balance like Ip Man 🥋. By learning from heterogeneous data (human MoCap + synthetic balance motions), AMS achieves both dynamic agility and extreme stability in a single controller. 👉
Hao Zhao12,106 просмотров • 9 месяцев назад
Больше нет контента для загрузки