Загрузка видео...

Не удалось загрузить видео

На главную

PPO has long dominated robot locomotion training in simulation. SAC, despite its sample efficiency, couldn't keep up. We analyze why: 🔗 🔥Integrated into RSL-RL, our approach requires only minimal changes, making SAC a drop-in alternative out of the box.

45,734 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Model-Free Reinforcement Learning (MFRL) has been alluring, especially with supercharged compute with physics on GPU. However, the methods use 0-th order gradients, and are often not the best optimizers. Can we do better than PPO in continuous control for robotics? Turns out yes! 🥳 tl;dr: Faster, better RL than PPO in continuous control 💪 The answer lies in using more information from the simulation. We are juicing the simulation on GPU as it is, why not use it for gradients as well? This has been a driving question in a series of our works. We first studied this problem in ICLR 2022 paper on Short Horizon Actor Critic Naive gradient based methods are stuck in local minima and have exploding/vanishing gradients. SHAC solved this problem truncated rollouts and model based value estimation, where the model is Differentiable Sim. This boosted sample efficiency and wall-clock time immensely especially in high dimensional systems such as humanoids Yet, given enough compute PPO often caught up. Our follow up paper on on Adaptive Horizon Actor Critic at ICML 2024 discovers the cause and provides a fix. However, we find that even when given ground-truth dynamics, not all gradients are useful due to sample error. 1st-Order Model-Based Reinforcement Learning methods employing differentiable simulation provide gradients with reduced variance but are susceptible to bias in scenarios involving stiff dynamics, such as physical contact. We find that back-propagating through contact and long trajectories drastically reduces gradient accuracy. Using this insight, we propose AHAC to dynamically adapt its roll-out horizon to avoid differentiating through stiff contact. AHAC is a first-order model-based RL algorithm that learns high-dimensional tasks in minutes (wall clock) and outperforms PPO by 40%, even in the limit of data provided to PPO. This work is led by Ignat Georgiev alongside Krishnan Srinivasan, Jie Xu, Eric Heiden and ample assistance from warp team at NVIDIA Robotics (Miles Macklin)

Animesh Garg

52,308 просмотров • 2 лет назад

A Letter to Our Community: The Road Ahead for Robotics To our Community and Partners, As we step into 2026, our mission at Axis is clearer than ever: Constructing the definitive End-to-End Scaling Layer for Robotics. Our goal is to accelerate the transfer of diverse human intelligence into Robotics General Intelligence (RGI). By owning the critical path of intelligence creation, we are turning the physical limitations of robotics into a scalable, software-driven future. Here is our strategic outlook and roadmap for the year ahead. The Core Thesis: Simulation is the Only Way Out The path to RGI is currently blocked by Data Scarcity, Generalization Fragility, and Hardware Fragmentation. At Axis, we believe Simulation is the only way out. Our Simulation Data Platform and Data Augmentation Engine transform raw data into "Synthetic Gold". Backed by academic milestones like Roboverse, Skill Blending, and GraspVLA, we have proven that pure simulation can achieve the generalization required for the real world. We don’t just collect data; we architect it. The Engine: Why Crypto? We believe RGI should come from all, not a few. Crypto is not just a feature; it is the primitive that powers our entire ecosystem flywheel: - Incentive Mechanism: Democratizing contribution and rewarding the trainers and developers. - Assetization: Turning proprietary data and refined models into liquid, ownable assets. - Verifiable Workflow: We are opening the "Black Box" of AI. By bringing total transparency to the Task Generation → Data Collection → Model Training pipeline, we ensure every byte of intelligence is verifiable, traceable, and secure. 2026 Strategic Deliverables This year, we are committed to delivering three foundational pillars: - The World's Largest Training Dataset for Robots: A robot training set—diverse, high-quality interaction data at an unprecedented scale. - A Robotics Foundation Model: A universal robotic brain trained on our pure simulation and synthetic data, capable of robust cross-embodiment transfer and open-world adaptability. - Evolvable Robot Hardware: Robots deployed with Axis models that autonomously evolve through continuous interaction, turning every deployment into a self-improving node within our RGI network. The Ultimate Vision We are building more than models; we are architecting the Distributed Machine Economy. A future where every dataset, model, and robotic embodiment is a verifiable asset in a global, autonomous network. Thank you for building the future of intelligence with us✌️📷

Axis Robotics

28,096 просмотров • 9 месяцев назад

JUST IN: Reimagine Robotics has just emerged from stealth! 🥷🏻 Its approach to robot training is one of the most human-centric I've seen. The founder is Jonathan Scholz, the person who built and led Google DeepMind's Applied Robotics team in London for seven years. This is not a first-time founder taking a swing at robotics. He has spent a decade at the frontier of the field. The philosophy is powerful. He calls it "monkey-see, monkey-do." 🐒 A worker shows the robot what to do. Watches it attempt the task. Corrects it on the spot. The robot learns. No specialist programmers. No months of integration. And it's already working in the real world: → A made-to-order plastics business trained robots to tend 3D printers overnight, removing print beds, operating latches, pressing controls → A hard drive disassembly facility built a three-robot cell combining robots and people to recover critical materials → Time to prototype and test a new robot behaviour reduced from one day to 10 MINUTES That last number is the one that changes everything. When testing a new behaviour takes 10 minutes instead of a day, the entire pace of deployment transforms. Scholz's framing of the human-robot relationship is worth reading carefully: "A robot that learns on the job depends on people. The worker identifies the bottleneck, shows the robot how to help, and corrects it until it is useful." It's August, and we keep getting robotics bangers week in week. ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

23,389 просмотров • 1 месяц назад

This work makes a humanoid robot do simple parkour moves by looking with a depth camera and choosing the right move on the fly. The big deal is that it turns lots of small human moves into long, real-time robot behavior, without hand-coding every transition or retraining for each new course. A humanoid robot is usually good at steady walking, but it often fails when it has to do fast moves like jumping up, vaulting, or rolling, and then keep going to the next obstacle. The hard part is that you cannot easily collect training data for every possible obstacle shape, distance, and mistake, so robots end up learning a few moves that only work in a narrow setup. This work starts from short clips of real human parkour moves, like stepping over, vaulting, climbing, and rolling. It uses motion matching, which is basically a smart “pick the next clip that fits best right now” search, to stitch those short clips into a long, smooth plan that looks like a human doing a whole course. Then it trains a controller with reinforcement learning (RL), which means the robot learns by trial and error to copy that plan while staying balanced and not falling. After training separate expert controllers for different moves, it compresses them into 1 controller that uses only onboard depth sensing and a simple “go this fast in this direction” command. In real tests on a Unitree G1 humanoid, it can clear multiple obstacles in a row, adapt when obstacles get moved, and climb a wall up to 1.25m.

Rohan Paul

37,121 просмотров • 7 месяцев назад

STORY | Study traces shrinking of Yamuna over 225 years in Delhi with 1799 map, reports 89% drop in volume The map of a wider, freer Yamuna from 1799 has opened a window into the river's past, helping researchers reveal the toll exacted by time, human intervention and urbanisation over more than 200 years in Delhi in a first-of-its-kind study. The research found that Yamuna flowing through Delhi has narrowed by about 68 per cent, and its discharge – the volume of water flowing through it – has dropped by around 89 per cent since the late 18th century. The study, accessed by PTI, was carried out by researchers from the Department of Geology, University of Delhi, and the Indian Institute of Science Education and Research (IISER), Bhopal. They reconstructed the river's past using an archival map prepared by Upjohn in 1799 and preserved in the National Archives of India, along with historical maps and modern satellite images. The findings have been published in the paper, 'Two Centuries of Hydrogeomorphic Changes: Width-Discharge Dynamics of the Urbanised Yamuna River in Delhi'. "People have talked about changes in the Yamuna River in the Delhi stretch, but no one has talked about the changes in its discharge in the stretch on this timescale," Professor Vimal Singh, one of the researchers, said. The researchers said the 1799 map captured the Yamuna before any barrages were built across the river, offering a rare glimpse of its natural state. They found that the average bankfull width – the width of the river when it is full but not overflowing its banks – has reduced from about 658 metres in 1799 to around 210 metres in 2024. Using this width, the researchers estimated that the river's discharge has fallen from about 30,000 cubic metres per second in 1799 to roughly 3,900 cubic metres per second in 2024. READ: (Reported by Varsha Sagi) Note: Visuals used for representational purposes only; they track changes in the shape of the Yamuna River in Delhi between 1985 and 2022

Press Trust of India

19,245 просмотров • 2 месяцев назад

We’re excited to introduce ShinkaEvolve: An open-source framework that evolves programs for scientific discovery with unprecedented sample-efficiency. Blog: Code: Like AlphaEvolve and its variants, our framework leverages LLMs to find state-of-the-art solutions to complex problems, but using orders of magnitude fewer resources! Many evolutionary AI systems are powerful but act like brute-force engines, burning thousands of samples to find good solutions. This makes discovery slow and expensive. We took inspiration from the efficiency of nature. ‘Shinka’ (進化) is Japanese for evolution, and we designed our system to be just as resourceful. On the classic circle packing optimization problem, ShinkaEvolve discovered a new state-of-the-art solution using only 150 samples. This is a big leap in efficiency compared to previous methods that required thousands of evaluations. We applied ShinkaEvolve to a diverse set of hard problems with real-world applications: 1/ AIME Math Reasoning: It evolved sophisticated agentic scaffolds that significantly outperform strong baselines, discovering an entire Pareto frontier of solutions trading performance for efficiency. 2/ Competitive Programming: On ALE-Bench (a benchmark for NP-Hard optimization problems), ShinkaEvolve took the best existing agent's solutions and improved them, turning a 5th place solution on one task into a 2nd place leaderboard rank in a competitive programming competition. 3/ LLM Training: We even turned ShinkaEvolve inward to improve LLMs themselves. It tackled the open challenge of designing load balancing losses for Mixture-of-Experts (MoE) models. It discovered a novel loss function that leads to better expert specialization and consistently improves model performance and perplexity. ShinkaEvolve achieves its remarkable sample-efficiency through three key innovations that work together: (1) an adaptive parent sampling strategy to balance exploration and exploitation, (2) novelty-based rejection filtering to avoid redundant work, and (3) a bandit-based LLM ensemble that dynamically picks the best model for the job. By making ShinkaEvolve open-source and highly sample-efficient, our goal is to democratize access to advanced, open-ended discovery tools. Our vision for ShinkaEvolve is to be an easy-to-use companion tool to help scientists and engineers with their daily work. We believe that building more efficient, nature-inspired systems is key to unlocking the future of AI-driven scientific research. We are excited to see what the community builds with it! Learn more in our technical report:

Sakana AI

360,508 просмотров • 1 год назад

It's 2030 and you are reviewing humanoid robots. A Tesla. A Google. An Apple. An OpenAI. A Meta. A Figure. And a bunch of Chinese-made ones. Which one is best, and why? I think the Tesla understands the world much better. Why? There were eight Teslas around me on the freeway today. Start there. No other robot company has that data. But my robot is parked at the local high school twice a day. Its cameras see humans in all of our weirdness. How we move. Where we go. Where we walk. Who we talk with. What you are wearing. Whether your hair was combed this morning. That data will lead to robotics breakthroughs. Apple might keep up with its Vision Pro data, but it is too freaked out by the privacy implications of using said data. (On the front are six cameras and a couple of TOF -- Time Of Flight -- sensors that can see everything in your home in great detail). Google has a lot of data, for sure. All my: 1. Email. 2. Calendars. 3. Photos. 4. TV watching behavior. 5. Contacts. 6. Documents and spreadsheets. 7. Files. 8. Location data. So I expect Google's robot will be attractive to many. But how do you see the others shake out over the next five years? Make some guesses. But remember what an AI pioneer told me years ago about AI: it's all about the data. The Chinese ones have huge advantages: the Chinese have more data on their citizens, and many more citizens to boot AND they can make robots cheaper than we can. But now that you know OpenAI is building its own robot you have caught wind of what I've heard from many in San Francisco and Silicon Valley: that humanoid robots are the real prize of AI and will be highly profitable for those that can make them and find customers willing to buy them. Here, too, I learned long ago never to bet against Elon Musk. Will you?

Robert Scoble

33,804 просмотров • 1 год назад

LongWriter Unleashing 10,000+ Word Generation from Long Context LLMs discuss: Current long context large language models (LLMs) can process inputs up to 100,000 tokens, yet struggle to generate outputs exceeding even a modest length of 2,000 words. Through controlled experiments, we find that the model's effective generation length is inherently bounded by the sample it has seen during supervised fine-tuning (SFT). In other words, their output limitation is due to the scarcity of long-output examples in existing SFT datasets. To address this, we introduce AgentWrite, an agent-based pipeline that decomposes ultra-long generation tasks into subtasks, enabling off-the-shelf LLMs to generate coherent outputs exceeding 20,000 words. Leveraging AgentWrite, we construct LongWriter-6k, a dataset containing 6,000 SFT data with output lengths ranging from 2k to 32k words. By incorporating this dataset into model training, we successfully scale the output length of existing models to over 10,000 words while maintaining output quality. We also develop LongBench-Write, a comprehensive benchmark for evaluating ultra-long generation capabilities. Our 9B parameter model, further improved through DPO, achieves state-of-the-art performance on this benchmark, surpassing even much larger proprietary models. In general, our work demonstrates that existing long context LLM already possesses the potential for a larger output window--all you need is data with extended output during model alignment to unlock this capability.

AK

50,995 просмотров • 2 лет назад

UPDATE: Flyctor has a second body now I bought another Vector robot and connected it to the same stack: - a camera feed - GPT-6 Astra for real-time visual understanding - a MaleCNS fly-brain simulation for movement decisions - a small audio protocol for robot-to-robot communication I didn't let them exchange normal Wi-Fi messages. Instead, each robot has to communicate through its speaker and microphone. One robot emits a short sequence of tones. The other records the sound, converts it into a frequency pattern, classifies the pattern, and injects it into its fly-brain simulation as a sensory event. Right now their entire vocabulary is only four signals: one chirp = "I found an object" two chirps = "come closer" long tone = "path blocked" rapid chirps = "I need help" The interesting part is that the robots don't send coordinates. They don't share a map. They don't know where the other robot is. One Flyctor sees something through Astra, decides it matters, sends a sound, and the other Flyctor has to hear it, interpret it, then decide what to do with that information. camera -> Astra -> fly brain -> speaker microphone -> audio decoder -> fly brain -> movement It is basically a tiny sensory loop between two digital insect nervous systems. Yesterday one robot detected a blocked path and sent the long-tone signal. The other one heard it, turned toward the sound, and rolled over to investigate. It was not impressive in the way modern AI demos are impressive. It was impressive because it felt like watching two small creatures notice each other for the first time.

kiruwaaaa

21,000 просмотров • 11 дней назад