Here’s a pretty weird and surprising result - retrieval-augmented... generation works unreasonably well for robot learning – but only when parameterized using difference vectors! We introduce Difference-Aware Retrieval Policies for Imitation Learning (DARP), a simple, semi-parametric RAG architecture for imitation learning that achieves gains of up to 200% over standard behavior cloning. No additional assumptions beyond BC, just a little architecture switch! The theory backing it up is pretty cool too and it works on real robots! :) Play with our website to understand better: 🧵(1/7)show more

Abhishek Gupta
21,078 Aufrufe • vor 1 Monat
Robora Sim: A PyBullet-Powered Environment for Learning Robotic Physical... Intelligence We are currently building our Robora simulation environment setup for our sim based learning, leveraging PyBullet, an industry-standard physics engine widely used in AI-driven robotics research and development. The environment is optimized with GPU-accelerated learning algorithms, enabling high-speed imitation learning and reinforcement learning within a safe and controlled virtual setup before shipping out to real world. This simulation platform allows our models to learn, adapt, and generalize across different robot morphologies, terrain types and task objectives - all before deployment to the real world. At it's core, the system combines a VLA-powered high-level planner with low-level motion control algorithms, working cohesively to produce emergent, physically intelligent behaviors. This synergy between simulation, learning, and real-world transfer marks a major step forward in our pursuit of adaptive and intelligent robotic systems. Through advanced domain randomization and synthetic data generation, the Robora Simulation Environment ensures that policies trained in simulation transfer effectively to real-world robots, minimizing the sim-to-real gap. Moreover, users will be able to test and integrate their own hardware kits within selected simulation environments in the Robora Dapp, ensuring seamless compatibility and safer real-world implementation.show more

Robora
23,489 Aufrufe • vor 9 Monaten
𝗥𝗼𝗯𝗼𝘁𝘀 𝗱𝗼𝗻’𝘁 𝗻𝗲𝗲𝗱 𝗺𝗼𝗿𝗲 𝗱𝗲𝗺𝗼𝗻𝘀𝘁𝗿𝗮𝘁𝗶𝗼𝗻𝘀. 𝗧𝗵𝗲𝘆 𝗻𝗲𝗲𝗱 𝘁𝗼 𝗹𝗲𝗮𝗿𝗻... 𝗳𝗿𝗼𝗺 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 — 𝗮𝗳𝘁𝗲𝗿 𝘄𝗮𝘁𝗰𝗵𝗶𝗻𝗴 𝗵𝘂𝗺𝗮𝗻𝘀. Most robot learning systems assume failure is the end of learning. In our new work, we study whether robots can improve after deployment by learning from their own failures, without any human intervention, teleoperation, or corrective labels. The key idea is simple: human videos contain structure about how the world works. We use them to learn cross-embodiment representations of action, dynamics, and value, enabling a shared predictive space between human behavior and robot experience. This allows a new learning loop: 👉 pretrain on human videos 👉 deploy robot policy 👉 observe failures 👉 reinterpret failures using human priors 👉 improve autonomously We evaluate this across 7 real-world manipulation tasks, showing: 📈 40% → 81% success rate 🏆 Strong improvements over π0.6 RECAP and RISE ✔️ Zero human intervention during post-deployment improvement 🧬 Generalizes across robot embodiments and policy backbones A key finding is that explicit failure repair significantly outperforms failure reweighting, yielding substantially larger gains under identical data conditions (+25 pts vs +5 pts on the same π0.5 base policy). Overall, the results suggest a shift in how we think about robot learning: Human videos are not only for pretraining policies. They can provide the structure needed for continual self-improvement after deployment. 📄 Paper: 🌐 Project: I am grateful for working with the fantastic leads Hanzhi Chen and Anran Zhang, and our collaborators Simon Schaefer, Kejia Chen, Shi Chen, Daniel Cremers. Special thanks to Stefan Leutenegger for co-advising this project with me. ETH Zürich TU München Microsoft Check out Hanzhi's 🧵 for more detailsshow more

Oier Mees
11,985 Aufrufe • vor 25 Tagen
Today we’re announcing our new Mentava Basics curriculum, aimed... at kids who are still a bit too young for the standard Mentava curriculum. Here’s what people don’t understand about teaching a 2 or 3 year old to read: The necessary skills for reading don’t all develop at the same time. The ability to associate letters with sounds happens first, and at a pretty early age. If you think about it, learning “this funny shaped animal says moo” is pretty similar to learning “this funny shaped line says aaa.” However, the second necessary skill for reading is blending those letter sounds together, and kids often aren’t developmentally capable of doing that until at least 6-12 months later. Until today, our recommendation has been to wait until the child is developmentally ready to blend sounds. and then we just go full speed and teach them everything all at once, as fast as possible. But sometimes we have students who start a little younger. And then their parents are confused, because they see that their kids are having a ton of fun learning letter sounds super fast, but then are being gatekept from additional learning because they aren't yet developmentally ready to blend those sounds. That’s why we created the new Mentava Basics curriculum. Mentava Basics lets our youngest students focus on letter/sound pairings until they're developmentally ready to begin blending them. Mentava Basics takes the fun and delight of the core Mentava experience, but applies it to a curriculum that’s developmentally appropriate for even younger children. Mentava’s standard curriculum is still the fastest way to go from zero to reading, but with Mentava Basics now we can give kids a head start by helping them learn their letter sounds in advance. If your child is struggling with blending and you think it may just be a developmental readiness issue, you can use the grownup menu to switch into the Mentava Basics curriculum. We save your progress on both pathways, so you can switch back to our standard curriculum whenever you want.show more

Niels Hoven 🐮
10,497 Aufrufe • vor 1 Jahr
For Ghost Of Yōtei we rebuilt our frame to... use a more general dynamic resolution algorithm, with upsampling, partly so we could take advantage of PSSR. PSSR works very well for us with only a few tweaks, including running conservative rasterization for small particles. By comparison, our standard upsampling algorithm requires significantly more hinting and authoring help to achieve good results. Here’s a side-by-side look. If you zoom in very close (16x before gif compression), you can see that PSSR does a better job reconstructing fine details on architecture and foliage. It is also more stable under motion.show more

@Zuby_Tech
10,357 Aufrufe • vor 7 Monaten
This work makes a humanoid robot do simple parkour... moves by looking with a depth camera and choosing the right move on the fly. The big deal is that it turns lots of small human moves into long, real-time robot behavior, without hand-coding every transition or retraining for each new course. A humanoid robot is usually good at steady walking, but it often fails when it has to do fast moves like jumping up, vaulting, or rolling, and then keep going to the next obstacle. The hard part is that you cannot easily collect training data for every possible obstacle shape, distance, and mistake, so robots end up learning a few moves that only work in a narrow setup. This work starts from short clips of real human parkour moves, like stepping over, vaulting, climbing, and rolling. It uses motion matching, which is basically a smart “pick the next clip that fits best right now” search, to stitch those short clips into a long, smooth plan that looks like a human doing a whole course. Then it trains a controller with reinforcement learning (RL), which means the robot learns by trial and error to copy that plan while staying balanced and not falling. After training separate expert controllers for different moves, it compresses them into 1 controller that uses only onboard depth sensing and a simple “go this fast in this direction” command. In real tests on a Unitree G1 humanoid, it can clear multiple obstacles in a row, adapt when obstacles get moved, and climb a wall up to 1.25m.show more

Rohan Paul
37,121 Aufrufe • vor 5 Monaten
As a newly appointed 𝗔𝘀𝘀𝗶𝘀𝘁𝗮𝗻𝘁 𝗣𝗿𝗼𝗳𝗲𝘀𝘀𝗼𝗿 at Imperial College... London, I'm thrilled to announce the 𝗦𝗮𝗳𝗲 𝗪𝗵𝗼𝗹𝗲-𝗯𝗼𝗱𝘆 𝗜𝗻𝘁𝗲𝗹𝗹𝗶𝗴𝗲𝗻𝘁 𝗥𝗼𝗯𝗼𝘁𝗶𝗰𝘀 𝗟𝗮𝗯 (𝗦𝗪𝗜𝗥𝗟) at 𝗜𝗺𝗽𝗲𝗿𝗶𝗮𝗹 𝗖𝗼𝗹𝗹𝗲𝗴𝗲 𝗟𝗼𝗻𝗱𝗼𝗻. 𝗦𝗮𝗳𝗲 𝗪𝗵𝗼𝗹𝗲-𝗯𝗼𝗱𝘆 𝗜𝗻𝘁𝗲𝗹𝗹𝗶𝗴𝗲𝗻𝘁 𝗥𝗼𝗯𝗼𝘁𝗶𝗰𝘀 𝗟𝗮𝗯 (𝗦𝗪𝗜𝗥𝗟) ( is a new research lab focused on the intersection of safety and intelligence in next-generation robotics. We're hiring exceptional PhD students who are passionate about pushing the boundaries of robot learning. 𝗪𝗵𝗮𝘁 𝗺𝗮𝗸𝗲𝘀 𝗦𝗪𝗜𝗥𝗟 𝘂𝗻𝗶𝗾𝘂𝗲? We operate at the exciting convergence of: • Online & offline reinforcement learning • Imitation learning & human demonstrations • Sample-efficient learning methods • Whole-body and soft robotics systems We're 𝗹𝗼𝗼𝗸𝗶𝗻𝗴 𝗳𝗼𝗿 𝗽𝗿𝗼𝘀𝗽𝗲𝗰𝘁𝗶𝘃𝗲 𝗣𝗵𝗗 𝘀𝘁𝘂𝗱𝗲𝗻𝘁𝘀 interested in: • Developing safe exploration algorithms for robotic systems • Creating sample-efficient learning methods that minimize real-world trials • Building foundation models for robotics with safety guarantees • Advancing soft robotics and compliant human-robot interaction • Bridging theory and practice in embodied AI Why now? As robots become more capable and work closer with humans, we need systems that are both intelligent enough to handle complex tasks 𝗔𝗡𝗗 safe enough for real-world deployment. Traditional approaches treat safety and intelligence as competing priorities, we believe they're synergistic. If you're a motivated researcher who wants to develop the theoretical foundations and practical algorithms for tomorrow's safe, intelligent robots, I'd love to hear from you. Want to join? Apply viashow more

Stephen James
16,552 Aufrufe • vor 9 Monaten
China is smart. China learned from the US's successes,... using only the best parts. It's our turn. We should be learning from China. China's economy is growing three times faster than that of the USA. China leads in almost all areas of science, and has the best infrastructure the world has ever seen. The number one priority of China is tell help their people become more prosperous, especially the least well off. Is there still work to be done. Hell yeah. But there difference is that they are trying and we are not. I've been to hundreds of Chinese villages, for more than a decade, and people are becoming better off each year. They call it 'Socialism with Chinese Characteristics.' We don't need to replicate their system. But we should be learning from it. For those ignoring it, you are hurting America. Pretending China's model isn't successful is preventing the US from improving and learning. The China-Haters should be ashamed of the damage that they are doing to the West. Instead of slander and Sinophobia, we need to be learning from the successes of other nations and their peoples.show more

Jason Smith - 上官杰文
82,767 Aufrufe • vor 1 Jahr
Introducing VL-JEPA: Vision-Language Joint Embedding Predictive Architecture for streaming,... live action recognition, retrieval, VQA, and classification tasks with better performance and higher efficiency than large VLMs. • VL-JEPA is the first non-generative model that can perform general-domain vision-language tasks in real-time, built on a joint embedding predictive architecture. • We demonstrate in controlled experiments that VL-JEPA, trained with latent space embedding prediction, outperforms VLMs that rely on data space token prediction. • We show that VL-JEPA delivers significant efficiency gains over VLMs for online video streaming applications, thanks to its non-autoregressive design and native support for selective decoding. • We highlight that our VL-JEPA model, with an unified model architecture, can effectively handle a wide range of classification, retrieval, and VQA tasks at the same time. by Delong Chen (陈德龙) Mustafa Shukor Théo Moutakanni Willy Jade Lei Yu Tejaswi Kasarla Allen Bolourchi Yann LeCun Pascale Fungshow more

Pascale Fung
90,144 Aufrufe • vor 7 Monaten
Most imitation learning policies break when the camera moves... or the robot changes. NOT THIS ONE 👇 [📍 Bookmark for later ] A new 3D scene representation encoder, tackles this by enabling zero-shot generalization to unseen embodiments and viewpoints… And it works with any IL algorithm. The trick? •Use a 2D foundation model to extract semantic features •Lift them into 3D space for localization (not semantics) •Condition the IL policy on this spatially grounded vector Across 93 simulated and 6 real tasks, Adapt3R: ✅ Maintains IL performance on LIBERO & MimicGen benchmarks ✅ Outperforms DP3 and 3D Diffuser Actor in most settings ✅ Holds >80% success on LIBERO even with large camera rotations Thanks for sharing this, Animesh Garg & Albert Wilcox! 📍Paper: Website: Code:show more

Ilir Aliu
12,178 Aufrufe • vor 11 Monaten
Streaming iPhone data in real-time directly to Rerun 🚀... The collection process is one of the most frustrating parts of building imitation-learning datasets. I’ve got a little army of sensors—📱 iPhone, iPad, Quest 3—but getting them temporally aligned, spatially aligned, AND seeing real-time feedback while recording is tough. I stumbled on a great library from Cake Lab (WPI) called ARFlow. It’s a thin client built on Unity’s ARFoundation that connects over gRPC to a server running Rerun for live data logging. I forked it to: - Log the SLAM translation poses, and - Upgrade rerun to v0.23 for my use case. So far, it works well, but there are still a few hitches: 1. Right now, it’s solid on iPhone and iPad; my Quest 3 client is still slow and not super reliable. 2. I’m using an older ARFlow branch focused on real-time streaming only—no spatial or temporal sync yet. Unity builds for iOS keep failing. 🛠️ 3. Nothing is saved locally to the client, so packet loss is a risk on shaky networks. There’s huge potential in tapping the ubiquitous sensors we carry around every day, and ARFlow is a big step toward making that easyshow more

Pablo Vela
27,383 Aufrufe • vor 1 Jahr
🤖 Robotic hands are evolving faster than we realize... — and they’re learning to master our world. I find this both fascinating and a little unsettling. At first glance, these robotic hands still look awkward - stiff fingers, clumsy motion. But behind the scenes, every iteration brings: ➡️ Tighter sensors ➡️ Smarter control loops ➡️ Better training on human-designed objects Week by week, they move closer to something astonishing — not imitation, but mastery. They’re learning to grip, twist, and adapt to tools made for our anatomy. And once they reach human-level dexterity, they won’t stop there. They’ll surpass it — steadier than surgeons, faster than assembly lines, tireless and precise. This is more than robotics. It’s the moment machines stop mimicking us… and start outperforming us. How do we redefine “human skill” in a world where robots master it too? #AI #Robotics #Innovation #Technology #Automation #FutureOfWork #Engineeringshow more

Pascal Bornet
97,867 Aufrufe • vor 6 Monaten
Ok, that's amazing. 🦄 A side by side three.js... comparison - with GitMCP and without. GitMCP's story originally started from a tweet by Three.js about their large, non-LLM-digestable, documentation file. Ido Salomon and I ended up creating a cool generic documentation MCP server, but we didn't forget threejs and mrdoob. So we put it to the test. Using GitMCP, I gave this prompt to Cursor + Claude 3.7: "Build a Three.js scene featuring a controllable realistic person navigating a textured dynamic urban environment with realistic lighting and subtle bloom effects. Ensure keyboard controls (WASD) for movement." This is the result. To the right is the one-shot result with our MCP server. To the left is the result without it, based just on Cursor's limited knowledge. The video is pretty conclusive. 🤯 It's really amazing since Cursor already knows a lot about three.js, but with our dedicated documentation MCP server it just outputs better results. And for all other libraries out there - that are not included in Cursor - this is a total game changer. A huge shout out to Cloudflare Developers and Anni Wang, who worked with us yesterday to quickly insert new capabilities in their AutoRAG feature (that enables indexing large amount of documentation) according to our specific needs, and to troubleshoot issues. It was very helpful. Knowledge is power, and that applies for every coding assitant. And not just for three.js! Check out GitMCP for any library you're using - link in the next comment.show more

Liad Yosef
41,576 Aufrufe • vor 1 Jahr
World modeling and imitation learning have largely been considered... two disparate worlds. In our recent work, Unified World Models, just accepted to #RSS2025, Chuning Zhu provides a dead-simple unifying solution: just train a joint diffusion model over actions and future states, but with *decoupled* diffusion time steps across these modalities. Manipulating these decoupled time steps then allows for marginalization or conditioning on actions or states; a single model can serve as a policy, forward dynamics model, video prediction model, or inverse dynamics model by simply setting diffusion timesteps carefully. The resulting model can leverage video datasets along with robot training data much more effectively, and shows improved robustness, generalization, and flexibility. This is exciting because it is frustratingly simple, scalable, and shows strong improvement on real-world robotics problems. Please refer to Chuning Zhu 's excellent thread for more details! More details/code can be found on our website and in the paper -show more

Abhishek Gupta
11,430 Aufrufe • vor 1 Jahr
We’ve seen humanoid robots walk around for a while,... but when will they actually help with useful tasks in daily life? The challenge here is the diversity and complexity of real-world scenes. Our new work tackles this problem via 3D visuomotor policy learning. Using data from only 1 scene, our Improved 3D Diffusion Policy (iDP3) enables a full-sized humanoid robot to autonomously pick&place objects, pour water, and wipe tables, in the wild open world. (and all these skills are useful, right?) Web: Fully open-sourced code:show more

Yanjie Ze
75,248 Aufrufe • vor 1 Jahr
The real story of our success in Port Coquitlam... isn’t just that we now have the lowest average property taxes and utility fees of any city in the region, it’s also that we achieved it while improving the services people rely on every day. Take curbside glass pickup. Not long ago, residents had to load up their cars and drive to a depot. It was inconvenient, and the result was predictable: too much glass ended up in blue bins or the garbage. So we fixed it. Today, curbside glass pickup is delivered directly to homes that receive City waste services. It’s done efficiently, reliably, and by our own crews. It’s a better service, it’s easier for residents, and it’s the kind of practical improvement that makes a real difference in people’s lives. It’s just one example of what’s possible happens when you stay relentlessly focused on core municipal services, and you’re willing to constantly assess, improve, and adapt. We believe that good government isn’t about doing less or doing more — it’s about doing what matters, and doing it well.show more

Brad West
27,095 Aufrufe • vor 2 Monaten
I watched Robot on the Road and honestly I... did not really know what to expect going in At the beginning it just feels like a pretty normal little road trip setup with this robot hitchhiking around Japan and it almost gives off a chill slice of life vibe But it does not stay like that for long The more it goes on the more you realize the robot is doing some really questionable stuff and the way it focuses on women without their consent makes the whole thing feel really uncomfortable By the end I was not really sure if I was supposed to find it funny or just feel weird about it Overall it feels like one of those shorts that tries to mix dark humor with awkward situations but it mostly just ends up being unsettlingshow more

Anime Posts
6,555,239 Aufrufe • vor 1 Monat
learn the difference between when you need rest and... when there is still energy to show up for yourself. it’s easy to let our emotions take over our day and how we feel, but there is still work to be done. you may not feel better going to the gym, but you’ll feel great knowing you did a workout when it was mentally/emotionally hard. 💙 *i am also okay. i have hard days mixed in with a lot of good ones. i decided to be vulnerable and post this to encourage showing up when it’s hard.*show more

Miranda Cohen
14,195 Aufrufe • vor 1 Jahr
Free ClawdBot is printing money on Polymarket The script... made about $460,000 - pretty crazy result It auto-trades BTC and ETH, focusing on 15-minute up and down moves and catching small delays and inefficiencies on centralized exchanges. No insider tools Just a smart automated script Profile: Copytrade: How it works Short-term moves The bot focuses on very short time frames - mostly 15-minute price direction on major coins. This is not long-term investing, it’s fast micro-trading. Exchange inefficiencies It watches centralized exchanges for tiny price gaps or delays and reacts instantly when an opportunity appears. Scale and frequency 500+ trades per week. Each trade aims for a very small profit, but over time those small gains add up and compound. The surprising part Most of the logic was created by prompting AI to generate the code - no developer degree needed. Bottom line High frequency Small edges Strong compounding Automation over emotionshow more

winkle.
38,648 Aufrufe • vor 5 Monaten
Exciting Milestone: Our First CEX Listing with MEXC! We’re... thrilled to announce that VentureMind AI ($VNTR) is officially listed on mexc_listings ! Choosing our first centralized exchange partner was a big decision, and we’re excited to team up with MEXC to bring $VNTR to a wider audience. This is a historic moment for us as the first AI incubator project launched on the Seedify platform to achieve a CEX listing. We couldn’t have reached this point without Seedify’s guidance, helping us navigate the journey with confidence and clarity. This partnership is just the beginning! It gives easier access to $VNTR and introduces the first onramp to purchase our token using fiat stablecoins, making it more accessible to a global audience. And we’re just getting started! Additional exchange listings are already in the works, and we can’t wait to share more milestones as we expand the reach and utility of $VNTR. To our incredible community, thank you for your unwavering support. This is only the start of an exciting journey, and we’re so grateful to have you with us every step of the way!show more

VentureMind AI
22,406 Aufrufe • vor 1 Jahr