Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

๐Ÿš€ Introducing LeVERB, the first ๐—น๐—ฎ๐˜๐—ฒ๐—ป๐˜ ๐˜„๐—ต๐—ผ๐—น๐—ฒ-๐—ฏ๐—ผ๐—ฑ๐˜† ๐—ต๐˜‚๐—บ๐—ฎ๐—ป๐—ผ๐—ถ๐—ฑ ๐—ฉ๐—Ÿ๐—” (upper- & lower-body), trained on sim data and zero-shot deployed. Addressing interactive tasks: navigation, sitting, locomotion with verbal instruction. ๐Ÿงต

96,617 Aufrufe โ€ข vor 1 Jahr โ€ขvia X (Twitter)

11 Kommentare

Profilbild von Haoru Xue
Haoru Xuevor 1 Jahr

(1/9) Old approaches made humanoid robots follow hand-crafted action commands (like setting a walking speed or an arm pose) from the language module. This limited them to a small, predefined skill set and made complex whole-body motions hard to achieve.

Profilbild von Haoru Xue
Haoru Xuevor 1 Jahr

(2/9) LeVERB instead learns a latent action space, shout out to PULSE and MaskedMimic. The high-level VL model outputs a latent โ€œverbโ€, the low-level controller decodes it into joint motion (seperatedly trained). Result: a far richer expressive skill set.

Profilbild von Haoru Xue
Haoru Xuevor 1 Jahr

(3/9) Two-system brain: System 2 thinks at 10 Hz (vision + language). System 1 reacts at 50 Hz (balance + contacts). Slow reasoning + fast reflexes = stable, expressive whole-body control.

Profilbild von Haoru Xue
Haoru Xuevor 1 Jahr

(4/9) Training data is the real bottleneck, so we built LeVERB-Bench: 154+ photorealistic sim scenes with heavy randomizationโ€”lighting, textures, clutter, camera angles. This diversity is what lets LeVERB generalize. Tasks: visual navigation, sitting, reaching, locomotion, etc.

Profilbild von Haoru Xue
Haoru Xuevor 1 Jahr

(5/9) LeVERB sees 80% zero-shot success rate on simple visual navigation tasks, and 58.5% across the board. This is 7.8 times better than a naive hierarchical VLA implementation with no latent regularization, highlighting the unique challenge brought by async decoupled WBC loop.

Profilbild von Haoru Xue
Haoru Xuevor 1 Jahr

(6/9) Generalization: LeVERB sees โ€œtake a seatโ€, โ€œsit downโ€, or โ€œsit on blue chairโ€ and knows they mean the same thing. It also reasons about space: if the chair is in front, it turns first, then sits.

Profilbild von Haoru Xue
Haoru Xuevor 1 Jahr

(7/9) Related inspiration: Helix, NaVILA, LangWBC, VBC, โ€ฆ each pushes latent or hierarchical VLA in its own niche (upper-body, legged nav, etc.). LeVERB adds whole body latent control plus an open benchmark for everyone to build on.

Profilbild von Haoru Xue
Haoru Xuevor 1 Jahr

(8/9) Current status: dynamics-level sim2real โœ”๏ธ vision sim2real - to be released โ€ข LeVERB-Bench dataset is already open-sourced in LeRobot format โ€ข Full code release is coming

Profilbild von Haoru Xue
Haoru Xuevor 1 Jahr

(9/9) Dive deeper ๐Ÿ‘‰ Collaborators: @x_h_ucb @Dantong_Niu @qiayuanliao @tjomiii Jan Tommy Gravdahl @xbpeng4 @GuanyaShi @trevordarrell @KoushilSreenath Shankar Sastry

Profilbild von clem ๐Ÿค—
clem ๐Ÿค—vor 1 Jahr

so cool! integrated in @LeRobotHF ?

Profilbild von Haoru Xue
Haoru Xuevor 1 Jahr

@LeRobotHF only the dataset so far haha would like to see if model can be integrated too ๐Ÿ˜‰

ร„hnliche Videos

Tencent presents GameGen-O Open-world Video Game Generation We introduce GameGen-O, the first diffusion transformer model tailored for the generation of open-world video games. This model facilitates high-quality, open-domain generation by simulating a wide array of game engine features, such as innovative characters, dynamic environments, complex actions, and diverse events. Additionally, it provides interactive controllability, thus allowing for the gameplay simulation. The development of GameGen-O involves a comprehensive data collection and processing effort from scratch. We collect and build the first Open-World Video Game Dataset (OGameData), amassed extensive data from over a hundred of next-generation open-world games, employing a proprietary data pipeline for efficient sorting, scoring, filtering, and decoupled captioning. This robust and extensive OGameData forms the foundation of our model's training process. GameGen-O undergoes a two-stage training process, consisting of foundation model pretraining and instruction tuning. In the first phase, the model is pre-trained on the OGameData via the text-to-video and video continuation, endowing GameGen-O with the capability for open-domain video game generation. In the second phase, the pre-trained model is frozen, and we fine-tuned using a trainable InstructNet, which enables the production of subsequent frames based on multimodal structural instructions. This whole training process imparts the model with the ability to generate and interactively control content. In summary, GameGen-O represents a notable initial step forward in the realm of open-world video game generation via generative models. It underscores the potential of generative models to serve as an alternative to rendering techniques, which can efficiently combine creative generation with interactive capabilities.

AK

367,110 Aufrufe โ€ข vor 1 Jahr

๐Ÿ”ฌ Exciting News! Our manuscript, "scGPT: toward building a foundation model for single-cell multi-omics using generative AI" is now finally published in Nature Methods (Nature Methods) ๐ŸŽ‰ !!! (Re-)Introducing scGPT: A transformative foundation model engineered for single-cell omics analysis. Developed through the analysis of over 33 million human cells, scGPT sets a new benchmark for application versatility, offering both fine-tuning and zero-shot capabilities. Since its preprint in May 2023, scGPT has significantly impacted the field, evidenced by 13K+ installations, 600+ GitHub stars ๐ŸŒŸ, and 40+ citations before its official publication! scGPT has been validated by numerous benchmark studies as a leading foundation model in single-cell analysis. Its pre-trained embeddings extend its utility beyond single-cell studies, enhancing a variety of downstream tasks including protein enrichment and genetic perturbation predictions. Some key updates lately: ---Expanded zero-shot applications for efficient reference mapping and integration, now with CellXGene census integration. ---Advanced perturbation analysis capabilities, including genome-scale perturb-seq data analysis and bulk sequencing data generalization. ---Upgraded scGPT package, offering versatile model loading compatible with PyTorch and flash-attn, for both GPU and CPU. ---Cloud-based scGPT applications for reference mapping, cell annotation, and gene regulatory network inference are available on ---Integration with Hugging Face for easier model training. Limitations: scGPT is an early foray into foundation models for single-cell omics, facing challenges like limited zero-shot learning in some tasks, pretraining constraints, data quality issues, and evaluation limitations. See our Supplementary Notes for details. ๐Ÿš€ Future Work? Short-Term Goals: 1. Releasing a Mouse Model for broader analysis. 2. Developing a comprehensive evaluation suite for foundation models in single-cell analysis. 3. Creating a foundation model for single-cell spatial omics. 4. Enhancing zero-shot capacity by integrating scGPT with RAG (e.g., knowledge graphs). Long-Term Goals: 1. Expanding scGPT for comprehensive single-cell multi-omics analysis. 2. Developing an in-silico perturbation model for predicting genetic perturbation effects. 3. Merging scGPT with multi-modal genomic sequence models for a deeper understanding of cell biology. ๐Ÿ“š Access the paper on Nature Methods: ๐Ÿ”ฌPreprint in Bioarixv: ๐Ÿ’ป All our codes/data/weights are open source: Wholehearted congratulations to all the authors, especially the two co-first authors, Haotian (Haotian Cui ) and Chloe (ChloeXWang), who are really the emerging superstars in AI and biology! Vector Institute Peter Munk Cardiac Centre AI U of T Department of Computer Science Department of Laboratory Medicine & Pathobiology University Health Network University of Toronto #scGPT #GenerativeAI #AI4Science #Combio #opensource

Bo Wang

199,718 Aufrufe โ€ข vor 2 Jahren

Today, we're joined by Nikita Rudin, co-founder and CEO of Flexion to discuss the gap between current robotic capabilities and whatโ€™s required to deploy fully autonomous robots in the real world. Nikita explains how reinforcement learning and simulation have driven rapid progress in robot locomotionโ€”and why locomotion is still far from โ€œsolved.โ€ We dig into the sim2real gap, and how adding visual inputs introduces noise and significantly complicates sim-to-real transfer. We also explore the debate between end-to-end models and modular approaches, and why separating locomotion, planning, and semantics remains a pragmatic approach today. Nikita also introduces the concept of "real-to-sim", which uses real-world data to refine simulation parameters for higher fidelity training, discusses how reinforcement learning, imitation learning, and teleoperation data are combined to train robust policies for both quadruped and humanoid robots, and introduces Flexion's hierarchical approach that utilizes pre-trained Vision-Language Models (VLMs) for high-level task orchestration with Vision-Language-Action (VLA) models and low-level whole-body trackers. Finally, Nikita shares the behind-the-scenes in humanoid robot demos, his take on reinforcement learning in simulation versus the real world, the nuances of reward tuning, and offers practical advice for researchers and practitioners looking to get started in robotics today. ๐Ÿ—’๏ธ For the full list of resources for this episode, visit the show notes page: ๐Ÿ“– CHAPTERS =============================== 00:00 - Introduction 04:07 - Is robot locomotion solved? 06:04 - Sim-to-real gap 08:58 - Adding semantics to policies 09:42 - Modular vs end-to-end architectures 10:29 - Planner model 12:21 - Adapting RL techniques from quadrupeds to humanoids 15:39 - Behind robot demos 18:09 - Humanoid robots in home environments 22:03 - Training approach 23:56 - VLA models 27:59 - Closing the sim-to-real gap 32:55 - Task orchestration using VLMs 36:38 - Tool use 38:10 - Model hierarchy 43:37 - Simulator versus simulation environment 44:57 - Combining imitation learning and reinforcement learning 46:42 - RL in real world versus RL in simulation 52:58 - Reward tuning and value functions in robotics 56:38 - Predictions 1:00:10 - Humanoids, quadropeds, and wheeled platforms 1:02:45 - Advice, recommended robot kits, and community pla

The TWIML AI Podcast

22,592 Aufrufe โ€ข vor 6 Monaten