Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

🚀 Introducing LeVERB, the first 𝗹𝗮𝘁𝗲𝗻𝘁 𝘄𝗵𝗼𝗹𝗲-𝗯𝗼𝗱𝘆 𝗵𝘂𝗺𝗮𝗻𝗼𝗶𝗱 𝗩𝗟𝗔 (upper- & lower-body), trained on sim data and zero-shot deployed. Addressing interactive tasks: navigation, sitting, locomotion with verbal instruction. 🧵

96,584 görüntüleme • 1 yıl önce •via X (Twitter)

11 Yorum

Haoru Xue profil fotoğrafı
Haoru Xue1 yıl önce

(1/9) Old approaches made humanoid robots follow hand-crafted action commands (like setting a walking speed or an arm pose) from the language module. This limited them to a small, predefined skill set and made complex whole-body motions hard to achieve.

Haoru Xue profil fotoğrafı
Haoru Xue1 yıl önce

(2/9) LeVERB instead learns a latent action space, shout out to PULSE and MaskedMimic. The high-level VL model outputs a latent “verb”, the low-level controller decodes it into joint motion (seperatedly trained). Result: a far richer expressive skill set.

Haoru Xue profil fotoğrafı
Haoru Xue1 yıl önce

(3/9) Two-system brain: System 2 thinks at 10 Hz (vision + language). System 1 reacts at 50 Hz (balance + contacts). Slow reasoning + fast reflexes = stable, expressive whole-body control.

Haoru Xue profil fotoğrafı
Haoru Xue1 yıl önce

(4/9) Training data is the real bottleneck, so we built LeVERB-Bench: 154+ photorealistic sim scenes with heavy randomization—lighting, textures, clutter, camera angles. This diversity is what lets LeVERB generalize. Tasks: visual navigation, sitting, reaching, locomotion, etc.

Haoru Xue profil fotoğrafı
Haoru Xue1 yıl önce

(5/9) LeVERB sees 80% zero-shot success rate on simple visual navigation tasks, and 58.5% across the board. This is 7.8 times better than a naive hierarchical VLA implementation with no latent regularization, highlighting the unique challenge brought by async decoupled WBC loop.

Haoru Xue profil fotoğrafı
Haoru Xue1 yıl önce

(6/9) Generalization: LeVERB sees “take a seat”, “sit down”, or “sit on blue chair” and knows they mean the same thing. It also reasons about space: if the chair is in front, it turns first, then sits.

Haoru Xue profil fotoğrafı
Haoru Xue1 yıl önce

(7/9) Related inspiration: Helix, NaVILA, LangWBC, VBC, … each pushes latent or hierarchical VLA in its own niche (upper-body, legged nav, etc.). LeVERB adds whole body latent control plus an open benchmark for everyone to build on.

Haoru Xue profil fotoğrafı
Haoru Xue1 yıl önce

(8/9) Current status: dynamics-level sim2real ✔️ vision sim2real - to be released • LeVERB-Bench dataset is already open-sourced in LeRobot format • Full code release is coming

Haoru Xue profil fotoğrafı
Haoru Xue1 yıl önce

(9/9) Dive deeper 👉 Collaborators: @x_h_ucb @Dantong_Niu @qiayuanliao @tjomiii Jan Tommy Gravdahl @xbpeng4 @GuanyaShi @trevordarrell @KoushilSreenath Shankar Sastry

clem 🤗 profil fotoğrafı
clem 🤗1 yıl önce

so cool! integrated in @LeRobotHF ?

Haoru Xue profil fotoğrafı
Haoru Xue1 yıl önce

@LeRobotHF only the dataset so far haha would like to see if model can be integrated too 😉

Benzer Videolar

Tencent presents GameGen-O Open-world Video Game Generation We introduce GameGen-O, the first diffusion transformer model tailored for the generation of open-world video games. This model facilitates high-quality, open-domain generation by simulating a wide array of game engine features, such as innovative characters, dynamic environments, complex actions, and diverse events. Additionally, it provides interactive controllability, thus allowing for the gameplay simulation. The development of GameGen-O involves a comprehensive data collection and processing effort from scratch. We collect and build the first Open-World Video Game Dataset (OGameData), amassed extensive data from over a hundred of next-generation open-world games, employing a proprietary data pipeline for efficient sorting, scoring, filtering, and decoupled captioning. This robust and extensive OGameData forms the foundation of our model's training process. GameGen-O undergoes a two-stage training process, consisting of foundation model pretraining and instruction tuning. In the first phase, the model is pre-trained on the OGameData via the text-to-video and video continuation, endowing GameGen-O with the capability for open-domain video game generation. In the second phase, the pre-trained model is frozen, and we fine-tuned using a trainable InstructNet, which enables the production of subsequent frames based on multimodal structural instructions. This whole training process imparts the model with the ability to generate and interactively control content. In summary, GameGen-O represents a notable initial step forward in the realm of open-world video game generation via generative models. It underscores the potential of generative models to serve as an alternative to rendering techniques, which can efficiently combine creative generation with interactive capabilities.

AK

367,000 görüntüleme • 1 yıl önce

🔬 Exciting News! Our manuscript, "scGPT: toward building a foundation model for single-cell multi-omics using generative AI" is now finally published in Nature Methods (Nature Methods) 🎉 !!! (Re-)Introducing scGPT: A transformative foundation model engineered for single-cell omics analysis. Developed through the analysis of over 33 million human cells, scGPT sets a new benchmark for application versatility, offering both fine-tuning and zero-shot capabilities. Since its preprint in May 2023, scGPT has significantly impacted the field, evidenced by 13K+ installations, 600+ GitHub stars 🌟, and 40+ citations before its official publication! scGPT has been validated by numerous benchmark studies as a leading foundation model in single-cell analysis. Its pre-trained embeddings extend its utility beyond single-cell studies, enhancing a variety of downstream tasks including protein enrichment and genetic perturbation predictions. Some key updates lately: ---Expanded zero-shot applications for efficient reference mapping and integration, now with CellXGene census integration. ---Advanced perturbation analysis capabilities, including genome-scale perturb-seq data analysis and bulk sequencing data generalization. ---Upgraded scGPT package, offering versatile model loading compatible with PyTorch and flash-attn, for both GPU and CPU. ---Cloud-based scGPT applications for reference mapping, cell annotation, and gene regulatory network inference are available on ---Integration with Hugging Face for easier model training. Limitations: scGPT is an early foray into foundation models for single-cell omics, facing challenges like limited zero-shot learning in some tasks, pretraining constraints, data quality issues, and evaluation limitations. See our Supplementary Notes for details. 🚀 Future Work? Short-Term Goals: 1. Releasing a Mouse Model for broader analysis. 2. Developing a comprehensive evaluation suite for foundation models in single-cell analysis. 3. Creating a foundation model for single-cell spatial omics. 4. Enhancing zero-shot capacity by integrating scGPT with RAG (e.g., knowledge graphs). Long-Term Goals: 1. Expanding scGPT for comprehensive single-cell multi-omics analysis. 2. Developing an in-silico perturbation model for predicting genetic perturbation effects. 3. Merging scGPT with multi-modal genomic sequence models for a deeper understanding of cell biology. 📚 Access the paper on Nature Methods: 🔬Preprint in Bioarixv: 💻 All our codes/data/weights are open source: Wholehearted congratulations to all the authors, especially the two co-first authors, Haotian (Haotian Cui ) and Chloe (ChloeXWang), who are really the emerging superstars in AI and biology! Vector Institute Peter Munk Cardiac Centre AI U of T Department of Computer Science Department of Laboratory Medicine & Pathobiology University Health Network University of Toronto #scGPT #GenerativeAI #AI4Science #Combio #opensource

Bo Wang

199,708 görüntüleme • 2 yıl önce

Today, we're joined by Nikita Rudin, co-founder and CEO of Flexion to discuss the gap between current robotic capabilities and what’s required to deploy fully autonomous robots in the real world. Nikita explains how reinforcement learning and simulation have driven rapid progress in robot locomotion—and why locomotion is still far from “solved.” We dig into the sim2real gap, and how adding visual inputs introduces noise and significantly complicates sim-to-real transfer. We also explore the debate between end-to-end models and modular approaches, and why separating locomotion, planning, and semantics remains a pragmatic approach today. Nikita also introduces the concept of "real-to-sim", which uses real-world data to refine simulation parameters for higher fidelity training, discusses how reinforcement learning, imitation learning, and teleoperation data are combined to train robust policies for both quadruped and humanoid robots, and introduces Flexion's hierarchical approach that utilizes pre-trained Vision-Language Models (VLMs) for high-level task orchestration with Vision-Language-Action (VLA) models and low-level whole-body trackers. Finally, Nikita shares the behind-the-scenes in humanoid robot demos, his take on reinforcement learning in simulation versus the real world, the nuances of reward tuning, and offers practical advice for researchers and practitioners looking to get started in robotics today. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 04:07 - Is robot locomotion solved? 06:04 - Sim-to-real gap 08:58 - Adding semantics to policies 09:42 - Modular vs end-to-end architectures 10:29 - Planner model 12:21 - Adapting RL techniques from quadrupeds to humanoids 15:39 - Behind robot demos 18:09 - Humanoid robots in home environments 22:03 - Training approach 23:56 - VLA models 27:59 - Closing the sim-to-real gap 32:55 - Task orchestration using VLMs 36:38 - Tool use 38:10 - Model hierarchy 43:37 - Simulator versus simulation environment 44:57 - Combining imitation learning and reinforcement learning 46:42 - RL in real world versus RL in simulation 52:58 - Reward tuning and value functions in robotics 56:38 - Predictions 1:00:10 - Humanoids, quadropeds, and wheeled platforms 1:02:45 - Advice, recommended robot kits, and community pla

The TWIML AI Podcast

22,533 görüntüleme • 6 ay önce