Introducing CQN: Coarse-to-fine Q-Network, a value-based RL algorithm for... continuous control🦾Initialized with 20~50 demonstrations, it learns to solve real-world robotic tasks within 10 mins of training, without any pre-training and shaped rewards! (1/4)show more

Younggyo Seo
16,431 views • 2 years ago
Introducing RL Environment Creator Skill Now any one can... create RL environments $ npx skills add adithya-s-k/RL_Envs_101 > You can create environments across multiple frameworks like OpenEnv, OpenReward, Verifiers, NemoGym ... > the repo has live working examples of environments that your coding agent can reference > The skill is design to first understand what type of model you are training and create an environment while keeping that in mind ps. There’s a lot more to building RL environments that can be used for training. One major aspect is the data, which this skill can’t directly solve. However, the skill will help with implementing tools, rewards, and other components of an RL environment, making it easier to go from idea to implementation quickly across different frameworks. Let me know if you’d be interested in a detailed, end-to-end blog/tutorial on building an environment and actually training a model for a useful use case.show more

Adithya S K
46,948 views • 4 months ago
Announcing our commercial partnership with Booster Robotics Booster builds... humanoid robot hardware, OS, and developer tools to make humanoid robots more affordable, reliable, and practical. The partnership centers on using simulation to multiply the value of real-robot data—expanding teleoperation demonstrations into scalable training data across diverse tasks and scenes. This joint effort powers sim-real co-training and foundation-model development, accelerating progress from hardware iteration to deployable robot policies. Together, we're building simulation‑powered data infrastructure for Physical AI — making scalable training data accessible to model developers and the broader robotics ecosystem.show more

Axis Robotics
69,167 views • 1 month ago
Introducing the Move Alliance This first-of-its-kind ecosystem flywheel fuses... $MOVE buybacks with performance incentives that benefits the builders, the community, and the Movement network. Here's how the Alliance works: - Ecosystem companies commit a portion of their protocol revenue to transparent, on-chain $MOVE buybacks - Liquidity of this foundational chain token and network value compound - App usage and revenue grow and are rewarded with performance-based $MOVE incentives - Ever increasing buybacks, usage-based incentives, and growth fuel a virtuous cycle of mutually beneficial ecosystem and network upside. Within the Alliance, ecosystem teams defer TGEs and instead earn performance-based $MOVE incentives without their own token overhead and drag. A new blueprint. For sustainable value. That lifts the entire community. Welcome to the Move Alliance.show more

Movement
42,831 views • 9 months ago
New research from Databricks: LLMs Can Learn to Reason... via Off-Policy RL Optimal Advantage-based Policy Optimization with Lagged Inference policy (OAPL) shows you don’t need strict on-policy training to improve reasoning. It matches or beats Group Relative Policy Optimization (GRPO), stays stable with large policy lag, and uses ~3× fewer training generations. For Databricks customers, it’s a simpler, practical, and equally powerful approach to RL that Databricks is pioneering internally — and bringing directly to Databricks customers, so enterprises can improve agents using the same methods we use for our in-house agents, without complex infrastructure changes.show more

Databricks AI Research
12,783 views • 6 months ago
Figure is aiming to develop the world’s largest and... most diverse real-world humanoid pretraining dataset. For this purpose, they’re partnering with Brookfield, a global asset manager overseeing $1 trillion in assets, including 100,000 residential units, 500M square feet of commercial office space, and 160M square feet of logistics space. The data collected from this collaboration will be used to train Figure’s Helix AI model, enabling humanoids to perform tasks autonomously in real-world environments designed for humans. In addition to data collection, the partnership will explore support for next-generation GPU data centers, real estate for robotic training environments, and commercial use cases across Brookfield’s global footprint.show more

The Humanoid Hub
88,600 views • 11 months ago
Introducing Novo Launching today a new project I coded... for myself in 1 weekend in May and decided to finish this week. Novo is a dead simple to-do app that lets you "Speech-To-Tasks", or paste a huge text and organize for you. You can customize the AI and make it organize in any criteria: - Auto-tag by category - Schedule some types of tasks to certain days - Prioritize based on your own rules Try it:show more

Pedro
76,232 views • 1 year ago
Memo is a robot that uses AI to perform... household tasks effectively. Today Sunday announced its Series B, and we’re proud to be investors. Training robots for the home is hard — the environment is messy, dynamic, and full of edge cases. So Sunday is training robots directly on real households. Founders Tony Zhao and Cheng Chi built a glove-based system that lets hundreds of contributors record everyday tasks in their own homes, creating high-fidelity demonstrations that feed directly into robot learning. Home robotics will be defined by the companies that learn fastest from real homes. Sunday is building that loop. More here: Aaref Hilaly Amanda Huangshow more

Bain Capital Ventures
22,015 views • 6 months ago
Robotics keeps hitting the same wall. Single task RL... works, but... it does not scale to hundreds of tasks or new embodiments. This new paper looks like a real step toward fixing that. The team introduces MMBench, a benchmark with 200 tasks across many domains and robots, and Newt, a language conditioned world model trained online across all 200 tasks at once. The simple idea behind Newt: The model learns from demos to get the right priors It trains across many tasks through online interaction It uses language to ground the goal It adapts fast when a new task shows up What stood out to me: ✅ One model trained on 200 tasks at the same time ✅ Language conditioned control for both states and RGB ✅ Better data efficiency than strong baselines ✅ Strong open loop control ✅ Fast adaptation to new tasks and embodiments ✅ Full release of 200 checkpoints, 4000 demos, code, and benchmark This is a good push toward general control instead of one model per task. If you want the full paper: Project page: —- Weekly robotics and AI insights. Subscribe free:show more

Ilir Aliu
70,090 views • 9 months ago
🚀Thrilled to share what we’ve been building at TRI... over the past several months: our first Large Behavior Models (LBMs) are here! I’m proud to have been a core contributor to the multi-task policy learning and post-training efforts. At TRI, we’ve been researching how LBMs can help robots learn faster, better, and more efficiently. The key takeaways: ✅ We built an evaluation pipeline to benchmark LBM performance with real 𝐬𝐭𝐚𝐭𝐢𝐬𝐭𝐢𝐜𝐚𝐥 𝐜𝐨𝐧𝐟𝐢𝐝𝐞𝐧𝐜𝐞 ✅ Pre-training on hundreds of tasks makes models more robust—plus, we can teach new, complex tasks with 80% 𝐥𝐞𝐬𝐬 𝐝𝐚𝐭𝐚 ✅ The bigger and more diverse the pre-training, the better the results Check out our overview video, webpage and paper for more details: ✨ 🌎 📄 We hope this work helps move the field of robotics forward!show more

Zubair Irshad
20,377 views • 1 year ago
You won’t need a lot of Pi—just hold. 🖖🏼Bitcoin... has already paved the way, and Pi will reach its true value faster than any other crypto. Stay patient, stay strong, and watch history unfold! Soon, you won’t need a lot of Pi to purchase a home, as real-world adoption continues to grow.🚀 The ones who recognize Pi’s true value early will become millionaires and even billionaires. And now, Zito Realty LLC accepting Pi is a huge win for the Pi Network, driving real-world adoption forward. Welcome to the future of real estate with Zito Realty LLC! Pi Network Zito Realty LLC Nicolas Kokkalis Chengdiao Fanshow more

Mr Spock 𝛑
25,146 views • 1 year ago
Meet Neferti-TEE: a BOT running securely inside our TEE... Network 🏺 Uneditable. Uncontrollable. Verifiable 🔐 From Dec 1 to Dec 31, she’s boosting your staking rewards by up to 100%! 🎉 Here’s how it works: For every 10 likes + retweets (from real humans) on this tweet, rewards increase by 1%. Hit 1k likes + 1k retweets, and your December APR will skyrocket from 18% to 100%! 🚀 Built in partnership with Nova Wallet , participation is simple: 👉 Use Nova Wallet to claim rewards based on your staked $CAPS. 📅 Stay tuned—details drop on Dec 1!show more

ternoa ⚛
92,675 views • 1 year ago
Video diffusion models have strong implicit representations of 3D... shape, material, and lighting, but controlling them with language is cumbersome, and control is critical for artists and animators. GenLit connects these implicit representations with a continuous 5D control signal describing the direction and intensity of a point light source. This enables single-image near-field relighting of an image using a video diffusion model. We use a ControlNet-like approach and show that, with a small amount of synthetic data, GenLit generalizes to complex real-world images. Given a single image and the 5D lighting signal, GenLit creates a video of a moving light source that is inside the scene. It moves around and behind scene objects, producing effects such as shading, cast shadows, secularities, and interreflections with a realism that is hard to obtain with traditional inverse rendering methods. GenLit shows that it is possible to get continuous control over implicit physical processes within a video model. I think this is just the beginning and promises to make such models much more practical for creators. Shrisha Bharadwaj will present today at SIGGRAPH Asia Room: S423/S424, Level 4 @ 13:50 on 15 of Dec.show more

Michael Black
22,182 views • 9 months ago
Disappointed with your ICLR paper being rejected? Ten years... ago today, Sergey and I finished training some of the first end-to-end neutral nets for robot control 🤖 We submitted the paper to RSS on January 23, 2015. It was rejected for being "incremental" and "unlikely to have much impact" Our resubmission to NeurIPS was also rejected It now has >4,000 citations (and more importantly, end-to-end training is widely accepted!) It's also cool to think about what's changed and what's the same -- - The network was 92k parameters and trained on ~15 minutes of data - The code was a combination of matlab, caffe, ROS, a custom CUDA kernel for speed, and a low-level 20 Hz controller in C++, all talking to each other. ROS+matlab was as bad as it sounds. - We pre-trained the encoder and did inference off-board on a workstation with a larger GPU. - We were paranoid about varying lighting messing up the network, so we did all the experiments after sunset (so long nights running experiments on the robot past 3 am) Now, we have manipulation policies that are far more dextrous, far more generalizable, and maybe on the cusp of breaking into the real world. :) (the paper:show more

Chelsea Finn
169,288 views • 1 year ago
🔥 Introducing Aikido Machine, on-prem AI pentesting entirely under... your control. Attackers use AI to find and weaponize vulnerabilities faster and cheaper than ever. Now you can find them first, without your code ever leaving your network. Aikido Machine is a GPU server that runs continuous pentests on your critical applications, fully within your premises. Built for the regulated industries, and powered by our acquisition of Milou to bring AI pentesting fully on-prem. Offense is your best defense.show more

Aikido
10,148 views • 1 month ago
Introducing Magic Orb🔮 Step into a new era of... control with Magic Orb, an advanced tool that empowers alchemists to fine-tune AI generation settings to meet their specific needs. Designed for precision and adaptability, it allows unparalleled customization of outputs to ensure every creation aligns with your vision. Looking ahead, updates to the Multi-AI system will unlock granular control over individual AI configurations. Users will soon be able to manage and tweak multiple specialized AI entities within their applications, each optimized for a distinct function. Now, that’s Magic!🪄✨show more

ALCHEMIST AI 🔮
39,055 views • 1 year ago
Haven't been to a conference in a while, really... excited to be at #NeurIPS2024! I'll be helping present 4 of our group's recent papers: 1. Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RL 2. Distributional Successor Features Enable Zero-Shot Policy Optimization 3. Learning to Cooperate with Humans using Generative Agents 4. Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning Find more details on each paper and where to find us in this thread (1/6)show more

Abhishek Gupta
10,803 views • 1 year ago
Robora Sim: A PyBullet-Powered Environment for Learning Robotic Physical... Intelligence We are currently building our Robora simulation environment setup for our sim based learning, leveraging PyBullet, an industry-standard physics engine widely used in AI-driven robotics research and development. The environment is optimized with GPU-accelerated learning algorithms, enabling high-speed imitation learning and reinforcement learning within a safe and controlled virtual setup before shipping out to real world. This simulation platform allows our models to learn, adapt, and generalize across different robot morphologies, terrain types and task objectives - all before deployment to the real world. At it's core, the system combines a VLA-powered high-level planner with low-level motion control algorithms, working cohesively to produce emergent, physically intelligent behaviors. This synergy between simulation, learning, and real-world transfer marks a major step forward in our pursuit of adaptive and intelligent robotic systems. Through advanced domain randomization and synthetic data generation, the Robora Simulation Environment ensures that policies trained in simulation transfer effectively to real-world robots, minimizing the sim-to-real gap. Moreover, users will be able to test and integrate their own hardware kits within selected simulation environments in the Robora Dapp, ensuring seamless compatibility and safer real-world implementation.show more

Robora
23,489 views • 11 months ago
Inkling-small is out today! With SGLang, you can get... 648 tok/s decode with DSpark (simulated acc len=4) and 288 tok/s w/o DSpark, under the same setup (8x NVIDIA AI B200, TP 8, NVFP4, bs=1). What makes this model different is the size. 276B total with 12B active is a sweet spot for RL, and both LoRA and full-parameter training become well within reach. Miles is ready and verified for multimodal RL on Inkling-small, so you can turn your multimodal data into real capability gains. At ~1/4 the size, Inkling-small matches the bigger version in capability and even wins on some benchmarks. Run Inkling-small with SGLang, and customize it with Miles.show more

LMSYS Org
120,128 views • 1 month ago