Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Very excited to share our first public release after I joined Robbyant! We present Lingbot-Depth 👀 — a state-of-the-art depth foundation model trained with RGB-D MAE on millions of real & simulated RGBD pairs. 🔹 Camera depths as natural masks for RGB-D MAE modeling 🔹 Large-scale real + sim...

22,787 Aufrufe • vor 8 Monaten •via X (Twitter)

10 Kommentare

Profilbild von Yinghao Xu
Yinghao Xuvor 8 Monaten

-Website: -Code: -Tech Report: -HuggingFace: Stay tuned!

Profilbild von David
Davidvor 8 Monaten

@robbyant_brain great work!

Profilbild von Ze Liu
Ze Liuvor 7 Monaten

@robbyant_brain great work!

Profilbild von Xiatao Sun
Xiatao Sunvor 7 Monaten

@robbyant_brain Really exciting release! RGB‑D MAE at scale with “sensor depth as masks” feels like a strong recipe for a depth foundation model, and the gains on transparent/reflective + thin structures are exactly the pain points for robotics.

Profilbild von Zwecharki of Immense AGI fundamentalism
Zwecharki of Immense AGI fundamentalismvor 7 Monaten

@robbyant_brain @grok does this apply to SfM for camera finding and such for gaussian splats?

Profilbild von Reza Sayar
Reza Sayarvor 7 Monaten

@robbyant_brain 👏👏👏👏👏👏😻

Profilbild von David Branca
David Brancavor 7 Monaten

@robbyant_brain Do you have any info on performance? Latency, etc

Profilbild von Philip McBride
Philip McBridevor 8 Monaten

@robbyant_brain Great work!

Profilbild von Nick Landolfi
Nick Landolfivor 8 Monaten

@robbyant_brain cool!!

Profilbild von Junlin Chang
Junlin Changvor 7 Monaten

@robbyant_brain Hi! Very interesting work! I wanna know how to address flying point.

Ähnliche Videos

In just one week, Binh and I trained a full-body Unitree G1. Here's a recap: 1. Secured a Unitree G1 humanoid through a LinkedIn post 2. Deployed TWIST2 full-body teleoperation pipelines 3. Adapted TWIST2 for Zed stereo camera & collected full-body teleoperation samples (carried by Binh ) 4. Adapted & fine-tuned NVIDIA Gr00T N1.5 VLA on the TWIST2 public datasets, which I fine-tuned on an 8xNVIDIA H100 Cluster. We picked Gr00T N1.5 as it was trained with Unitree G1 embodiment data. 5. Adapted the TWIST2 codebase to stream in the actions from Gr00T via ZMQ using a co-located NVIDIA H100 for ~200ms inference latency 6. Tested the model in sim, then deployed to the real-world Unitree G1. We streamed a training sample observation to the VLA (as we didn't want to break robot in case real observations were OOD) We were the first team in the world to deploy the full TWIST2 data collection pipeline to the unitree g1 :) Much more work ahead though, which I'll work on as a side-project over the next months: 1. Exploring the various types of 'world models': video backbones, dynamics models, v-jepa-2 models. I believe these will generalize better & train much more data-efficiently than VLM backbones 2. Speeding up inference - I believe low-latency robotics inference will be a big challenge. There are many works in video diffusion which I'd like to test (e.g. SageAttention, SparseAttention, Drifting Models). Perhaps also writing custom CUDA kernels. 3. Economics of inference scaling :) What will be the compute demands as we scale inference up to millions of humanoids? Will it run on edge or on distributed 'co-located' inference clusters? These are questions I'd like to answer. Adapted TWIST2 codebase: Adapted Gr00T-N1.5 codebase: The ETH Robotics Club are doing a cool GTC Golden ticket competition with NVIDIA , so this is my submission :) The DGX Spark compute will get me a long way with initial prototyping & especially working on inference optimization for next-gen Blackwell GPUs #NVIDIAGTC #GOLDENTICKET #ETHRC

Arnie Ramesh

23,236 Aufrufe • vor 7 Monaten