New research from Databricks: LLMs Can Learn to Reason... via Off-Policy RL Optimal Advantage-based Policy Optimization with Lagged Inference policy (OAPL) shows you don’t need strict on-policy training to improve reasoning. It matches or beats Group Relative Policy Optimization (GRPO), stays stable with large policy lag, and uses ~3× fewer training generations. For Databricks customers, it’s a simpler, practical, and equally powerful approach to RL that Databricks is pioneering internally — and bringing directly to Databricks customers, so enterprises can improve agents using the same methods we use for our in-house agents, without complex infrastructure changes.show more

Databricks AI Research
12,753 views • 6 months ago
Haven't been to a conference in a while, really... excited to be at #NeurIPS2024! I'll be helping present 4 of our group's recent papers: 1. Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RL 2. Distributional Successor Features Enable Zero-Shot Policy Optimization 3. Learning to Cooperate with Humans using Generative Agents 4. Personalizing Reinforcement Learning from Human Feedback with Variational Preference Learning Find more details on each paper and where to find us in this thread (1/6)show more

Abhishek Gupta
10,803 views • 1 year ago
This is Rivian's next-generation Autonomy hardware and software platform,... which the company says is a "self-improving, end-to-end AI system for our autonomy platform that scales." • 11 cameras (65 total megapixels) • 5 radars • New front-facing LiDAR • Powered by a new in-house designed chip called RAP1 (5nm TSMC) Rivian also detailed its software-first approach to autonomy, powered by the new Rivian Autonomy Platform and an end-to-end data loop used for training. The company introduced its Large Driving Model, a foundational autonomous model trained like a Large Language Model: "Utilizing Group-Relative Policy Optimization, the LDM will distill superior driving strategies from massive datasets into the vehicle."show more

Sawyer Merritt
283,194 views • 9 months ago
What if you kept asking an LLM to "make... it better"? In some recent work at FAIR, we investigate how we can efficiently use RL to fine-tune LLMs to iteratively self-improve on their previous solutions at inference-time. Training for iterated self-improvement can be costly. The naive approach to training for K self-improvement steps leads to K times the number of rollout steps per episode. We introduce Exploratory Iteration (ExIt), an RL-based automatic curriculum method that bootstraps diverse training distributions of self-improvement tasks by upcycling the LLM's own responses at previous turns as the starting points for both self-improvement and *self-divergence.* In order to decide what task to train on next, the curriculum prioritizes sampling of partial turn histories that led to higher return variance in its GRPO group (a learnability score that comes for free). This automatic curriculum over the bootstrapped task space teaches the model how to perform iterated self-improvement while only ever training the model on single-step self-improvement tasks. We look at ExIt's impact in both single-turn (contest math problems) and multi-turn (BFCLv3 multi-turn tasks), as well as MLE-bench, where the LLM is run in a search scaffold to produce solutions to real Kaggle competitions. Across these eval settings, we find ExIt produces models with greater capacity for inference-time self-improvement compared to GRPO. Notably, ExIt models can self-improve on test tasks for many more steps than the typical solution depth encountered during training, including a 22% improvement in MLE-bench performance compared to GRPO.show more

Minqi Jiang
41,099 views • 1 year ago
So TeamYouTube has given me a warning and a... strike in 24 hours for two separate videos under the "Harmful and Dangerous Acts" policy and "Harassment and Bullying" policy. I want you to just take a look at what was taken down with the timestamps TeamYouTube has provided A video where i am reacting to a video that is currently on their platform and a video where I'm saying "top 5 gooners" got taken down under this policy with both appeals getting rejected in under an hour TeamYouTube can i get a REAL HUMAN to look at these videos fairly and not a robot. Not one of those links saying "We have previously looked at the video and unfortunately it goes against our community guidelines" the appeal system does not work correctly.show more

Blueryai🌊
10,127 views • 2 months ago
RL is painfully slow 😭 — bottlenecked by super-long... CoT rollout. 🔭 Sparse attention should help, but naive sparse rollout hits a brutal efficiency–stability tradeoff: A tedious trial-and-error sparsity sweep for each dense policy is required before an actual RL run. 🐤Sparrow chirps no more pain! Introduce Sparrow: Sparse Rollout for stable and efficient long-context RL. Sparrow finds that: 💡As long as we keep the tail distribution mismatch throughout the sparse rollout above a critical threshold, the RL training will be stable. 💡Even cooler! Through comprehensive control studies of Qwen3-1.7B, 4B, 8B thinking models RL with 40K rollout max length, the critical threshold stays constant across model sizes. 💡Sparrow then finds the optimal dynamic sparse schedule to reach the threshold with minimal cost. 💡Sparrow's findings are empirically validated to generalize in Qwen3-14B, and hold on both Math and Coding RL. 🐤Sparrow empirically helps achieve 2.2× / 2.4× / 2.0× rollout speedup on Qwen3 1.7B / 4B / 8B thinking models, while keeping training stability over extended RL steps. We release the 🐤bird in the following formats. [1/n] Paper: Code: Blog:show more

Infini-AI-Lab
78,984 views • 2 months ago
Today we’re announcing the General Availability of Agent SSO,... and we’re including it in our core SSO offering. SSO transformed how people move across applications, making access seamless for users and centralizing identity and control for IT. AI agents need the same thing. Today, connecting agents is a mess of static API keys, fragmented integrations, and endless user consent prompts. But Agent SSO changes that. Powered by the open Cross App Access (XAA) standard, it enables agents to be registered as first-class identities. For IT, it means centralized visibility and least-privilege policy. For users, it means their AI agents can seamlessly connect across enterprise apps without interrupting their flow.show more

Todd McKinnon
98,592 views • 14 days ago
The ATG School One-Pager I’m not trying to reinvent... schooling. There are just a handful of things I believe in which I haven’t seen in any school I’ve been around as a student or parent. Policy #1: Each student gets to be responsible for growing some of their own food, no matter how small, and THROUGHOUT schooling (not just a quickie project here or there). Policy #2: Minimum 1:1 ratio of time NOT SITTING IN THE CLASSROOM. What you do with this is up to you. There are so many real world skills, sports, gardening, music, etc. The strict ratio in the school day is the key for me. Common sense and personal interests can take it from there. Policy #3: Daily time to read whatever you want to read about. The biggest barrier for my reading was INTEREST. Be there to ensure the book is at their level, and to help them if they don’t understand something. Other than that, LET THEM ENJOY READING, ALL THE WAY THROUGH SCHOOL, not just in early years. Policy #4: (This is the most unusual yet the biggest reason I’m in education.) High school is a 50/50 bridge to winning in real life. Mornings are for actual work, making and SAVING UP MONEY. Afternoons are for learning finances and professional skills of YOUR INTEREST. With average work, you’ll finish school with $50,000-$100,000 in the bank, more skills than the norm, and a greater chance of creating your life and work from there on out, rather than conforming to make a paycheck. Policy #5: As part of the high school 50/50 system, ensure each student learns the adult financial red tape in your state/country before you’ve got bills, kids, etc.show more

KneeOverToesGuy
31,525 views • 5 months ago
🚀Thrilled to share what we’ve been building at TRI... over the past several months: our first Large Behavior Models (LBMs) are here! I’m proud to have been a core contributor to the multi-task policy learning and post-training efforts. At TRI, we’ve been researching how LBMs can help robots learn faster, better, and more efficiently. The key takeaways: ✅ We built an evaluation pipeline to benchmark LBM performance with real 𝐬𝐭𝐚𝐭𝐢𝐬𝐭𝐢𝐜𝐚𝐥 𝐜𝐨𝐧𝐟𝐢𝐝𝐞𝐧𝐜𝐞 ✅ Pre-training on hundreds of tasks makes models more robust—plus, we can teach new, complex tasks with 80% 𝐥𝐞𝐬𝐬 𝐝𝐚𝐭𝐚 ✅ The bigger and more diverse the pre-training, the better the results Check out our overview video, webpage and paper for more details: ✨ 🌎 📄 We hope this work helps move the field of robotics forward!show more

Zubair Irshad
20,377 views • 1 year ago
I just built a Meta ad policy checker in... Claude Code that catches rejections BEFORE Meta does 🤯 Drop in your ad copy → it pulls Meta's LIVE Advertising Standards, checks every line against the actual policy text, and hands each ad a verdict: Cleared for launch, Fix before launch, or Grounded. All inside Claude Code. Perfect for media buyers and DTC brands who've had ads bounced — or an account restricted — and never got a straight answer why. If you're finding out about policy problems only after the rejection email, resubmitting the same ad and praying, losing days of delivery while the appeal sits in review, and every bounce quietly teaches Meta to trust your account a little less... This runs the review before Meta ever sees the ad: → Drop in your ad copy (one ad or a whole batch) → It reads each ad and figures out which of Meta's policies apply → Scrapes the live policy pages from Meta's Transparency Center → Flags the exact phrase that violates, with Meta's own policy quoted next to it → Rewrites the risky lines so the message survives but the violation doesn't → Renders a dashboard: every ad, every finding, every fix in one place No guessing which word killed the ad. No resubmit-and-pray loops. No stacking rejections on your account history. What you get: → A verdict on every ad before you spend a dollar → The violating phrase + the policy citation, side by side → Rewrites that keep the selling intent → A report you can hand straight to your team or client Built 100% in Claude Code. No API keys, no Meta login. I'm giving away the complete Claude skill file. Want the skill for free? > Like this post > Comment "META" And I'll send it over (must be following so I can DM)show more

Mike Futia
17,399 views • 1 month ago
Multi-robot learning is getting a serious boost! 📚 Researchers... have extended Isaac Lab to train heterogeneous multi-agent robotic policies at scale. The new framework supports high-resolution physics, GPU-accelerated simulation, and both homogeneous and heterogeneous agents working together on coordination tasks. They benchmarked different approaches (MAPPO: Multi-Agent Proximal Policy Optimization and HAPPO: Heterogeneous Agent PPO) across six challenging scenarios and showed that large-scale multi-robot training is not only feasible, but efficient. It’s an important step for real-world robotic collaboration, where teams of robots need to coordinate, split tasks, adapt roles, and interact dynamically, not just operate as identical clones. The code is open-source, and it pushes Isaac Lab closer to what robotics actually needs: scalable, physics-driven environments where many different robots can learn to work together. Here's the project page: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →show more

Lukas Ziegler
38,997 views • 9 months ago
Whether you are a random person on X or... Elon Musk, I urge you all to please read this post and stop using videos like this one to claim that Biden has an "open border policy": #1) The people in this video are already on American soil. The agents have a few options here. Detain them behind the razor wire or detain them in front of it, let the children cut themselves on razor wire or make sure they don't get cut. #2) In May President Biden issued an "asylum ban" which barred migrants from applying for humanitarian protection if they cross the border against the law or fail to first apply for safe harbor while crossing through another country on the way to the U.S. The US COURTS OVERTURNED IT. Biden also requested $3.5B for more border patrol agents and judges and lawyers. Republicans refused. #3) Let me explain to you all why the people in this video can't simply be put on a truck and driven back into Mexico and dropped off: a - They have a right to due process according to our Constitution. b - There are international laws indicating that we can't just drop people off from Honduras or Cuba or any other nation into Mexico. What if one of these people claimed they were from Canada? Should we drop them in Canada? c - How do we immediately prove that a person is not an American, or an Italian, or a Brazilian if they claim they are? Stop pretending that Biden isn't doing anything. Stop pretending that the immigration problem is simple. Stop pretending that this has not been a problem for decades. Understand that the surge of migrants isn't the fault of Biden but actually global geopolitical upheaval in Cuba, Nicaragua, Venezuela and other nations. Instead of pointing fingers and claiming that "Biden has an open border policy," when his policy is not much different than those we have had for the last 40 years, how about working together on comprehensive reform? So before you make another Tweet claiming Biden has an Open Border Policy, how about instead you provide actual solutions that are LEGAL, that should be implemented.show more

Brian Krassenstein
7,720,381 views • 2 years ago
I met with Moldova’s Deputy Prime Minister and Minister... of Foreign Affairs, Mihai Popșoi. I was pleased to welcome the Minister to Ukraine’s annual meeting of ambassadors. His presence reflects the fact that Ukraine and Moldova share many foreign policy priorities, particularly in the areas of security and European aspirations. It is important that we continue moving forward together across the full range of our shared objectives. We also discussed a number of practical projects that can and should be implemented, including in infrastructure and energy. For us, these are first and foremost matters of security and mutual assistance. I am grateful to Moldova for its consistent support for our people and our defense.show more

Volodymyr Zelenskyy / Володимир Зеленський
285,269 views • 1 month ago
Today marks General Availability of AgentCore, a set of... infrastructure building blocks for developers and companies to build secure, scalable agents. When we first started AWS, the vast majority of developers were spending most of their time on the undifferentiated heavy lifting of infrastructure instead of what differentiated their feature. So, we solved that problem by building primitive building blocks like compute and storage and database that would allow teammates and customers to quickly build and deploy new experiences without having to reinvent the wheel each time. We realized the same thing was happening with AI agents. It's too difficult and it's slowing customers down. That's why we created AgentCore, a set of services to build, deploy, and operate highly capable agents using any framework or model, with enterprise-grade security and scalability. These building blocks (like serverless secure runtime, memory, observability, a gateway that does MCP translation, etc) help customers tackle some of the biggest challenges of going from prototype to production, much more quickly, securely, and scalably. AgentCore has been in preview for several weeks, and customers have been quite excited about it. The AgentCore SDK has already been downloaded over a million times and we're seeing transformative results, such as Cohere Health expecting to reduce medical review times by 30-40% in highly regulated healthcare, and teams at Cox Automotive and Experian are embracing its flexibility to deploy and operate agents at scale. Inside Amazon, our Amazon Devices Operations & Supply Chain team is using AgentCore to develop an agentic manufacturing approach where AI agents work together to automate manual processes – turning what used to be days of engineering time into processes that take under an hour with high precision. Just like AWS changed how companies build and scale applications, we believe AgentCore will do the same for AI agents, enabling the next generation of innovation.show more

Andy Jassy
24,990 views • 10 months ago
A lifelong learner and a student of many passions... yesterday I received my 2ND advanced degree - a Doctorate in Health Policy from George Washington University School of Nursing. I am grateful to Allahu SWA for the life opportunities I’ve enjoyed along with purposeful and needed challenges along the way. To many on this app, I am known for my no BS politics, and to my patients it’s Dr. ALI. My passion has been providing exceptional care for my patients and communities for nearly two decades. I came back to Somalia in 2016/17 to focus on primary healthcare BUT because the system doesn’t allow for expertise and progress, I opened my OWN practice in Mog and unfortunately for security challenges I had to close it. My journey continues both in and out of Somalia and Inshallah I will have the opportunity to use my training and expertise in primary care/health policy to improving health outcomes for my people. To my family and friends who have always kept me going - thank you for the love and support. To my political adversaries, caadi iska dhiga. We share a country - not lives. To the YOUNG people on this app - FOCUS on building a career you love and makes you $ without having to sell your soulzshow more

Hodan Ali
37,002 views • 1 year ago
We’re excited to introduce ShinkaEvolve: An open-source framework that... evolves programs for scientific discovery with unprecedented sample-efficiency. Blog: Code: Like AlphaEvolve and its variants, our framework leverages LLMs to find state-of-the-art solutions to complex problems, but using orders of magnitude fewer resources! Many evolutionary AI systems are powerful but act like brute-force engines, burning thousands of samples to find good solutions. This makes discovery slow and expensive. We took inspiration from the efficiency of nature. ‘Shinka’ (進化) is Japanese for evolution, and we designed our system to be just as resourceful. On the classic circle packing optimization problem, ShinkaEvolve discovered a new state-of-the-art solution using only 150 samples. This is a big leap in efficiency compared to previous methods that required thousands of evaluations. We applied ShinkaEvolve to a diverse set of hard problems with real-world applications: 1/ AIME Math Reasoning: It evolved sophisticated agentic scaffolds that significantly outperform strong baselines, discovering an entire Pareto frontier of solutions trading performance for efficiency. 2/ Competitive Programming: On ALE-Bench (a benchmark for NP-Hard optimization problems), ShinkaEvolve took the best existing agent's solutions and improved them, turning a 5th place solution on one task into a 2nd place leaderboard rank in a competitive programming competition. 3/ LLM Training: We even turned ShinkaEvolve inward to improve LLMs themselves. It tackled the open challenge of designing load balancing losses for Mixture-of-Experts (MoE) models. It discovered a novel loss function that leads to better expert specialization and consistently improves model performance and perplexity. ShinkaEvolve achieves its remarkable sample-efficiency through three key innovations that work together: (1) an adaptive parent sampling strategy to balance exploration and exploitation, (2) novelty-based rejection filtering to avoid redundant work, and (3) a bandit-based LLM ensemble that dynamically picks the best model for the job. By making ShinkaEvolve open-source and highly sample-efficient, our goal is to democratize access to advanced, open-ended discovery tools. Our vision for ShinkaEvolve is to be an easy-to-use companion tool to help scientists and engineers with their daily work. We believe that building more efficient, nature-inspired systems is key to unlocking the future of AI-driven scientific research. We are excited to see what the community builds with it! Learn more in our technical report:show more

Sakana AI
360,318 views • 11 months ago
𝗥𝗼𝗯𝗼𝘁𝘀 𝗱𝗼𝗻’𝘁 𝗻𝗲𝗲𝗱 𝗺𝗼𝗿𝗲 𝗱𝗲𝗺𝗼𝗻𝘀𝘁𝗿𝗮𝘁𝗶𝗼𝗻𝘀. 𝗧𝗵𝗲𝘆 𝗻𝗲𝗲𝗱 𝘁𝗼 𝗹𝗲𝗮𝗿𝗻... 𝗳𝗿𝗼𝗺 𝗳𝗮𝗶𝗹𝘂𝗿𝗲 — 𝗮𝗳𝘁𝗲𝗿 𝘄𝗮𝘁𝗰𝗵𝗶𝗻𝗴 𝗵𝘂𝗺𝗮𝗻𝘀. Most robot learning systems assume failure is the end of learning. In our new work, we study whether robots can improve after deployment by learning from their own failures, without any human intervention, teleoperation, or corrective labels. The key idea is simple: human videos contain structure about how the world works. We use them to learn cross-embodiment representations of action, dynamics, and value, enabling a shared predictive space between human behavior and robot experience. This allows a new learning loop: 👉 pretrain on human videos 👉 deploy robot policy 👉 observe failures 👉 reinterpret failures using human priors 👉 improve autonomously We evaluate this across 7 real-world manipulation tasks, showing: 📈 40% → 81% success rate 🏆 Strong improvements over π0.6 RECAP and RISE ✔️ Zero human intervention during post-deployment improvement 🧬 Generalizes across robot embodiments and policy backbones A key finding is that explicit failure repair significantly outperforms failure reweighting, yielding substantially larger gains under identical data conditions (+25 pts vs +5 pts on the same π0.5 base policy). Overall, the results suggest a shift in how we think about robot learning: Human videos are not only for pretraining policies. They can provide the structure needed for continual self-improvement after deployment. 📄 Paper: 🌐 Project: I am grateful for working with the fantastic leads Hanzhi Chen and Anran Zhang, and our collaborators Simon Schaefer, Kejia Chen, Shi Chen, Daniel Cremers. Special thanks to Stefan Leutenegger for co-advising this project with me. ETH Zürich TU München Microsoft Check out Hanzhi's 🧵 for more detailsshow more

Oier Mees
12,514 views • 2 months ago
𝕏 Monetization Update: X employee Allegra Jacchia has confirmed... that you do not need to maintain 5 million impressions every 3 months. This information was ONLY confirmed by her because my post went viral. If it did not go viral, we would be still confused. My (incorrect) information was from: - Grok - Premium + support - X policy website I will be honest. Not only getting called out for using their OWN website and support to confirm the policy was surprising, It was amazing to see the response from the X employee. Rather than admitting the information from their own platform was incorrect, they reacted emotionally and called me “fake.” News flash. People can only use what they see. FIX YOUR OWN SYSTEM. It is your system that is speaking misinformation. NOT US. Learn the difference. (this got me fired up today.)show more

Jin Jung
17,567 views • 3 months ago
Here's what The Browser Company's AI eng & ML... teams are working on for Dia right now: (This is a pitch to come work for us; info at end) 🤖 COMPUTER USE – we've built our own bespoke APIs on top of Chromium to optimize latency, accuracy, and cost of computer-using agents. Demo attached. Big breakthroughs here in recent weeks. 🛡️ ON-DEVICE MODELS – we've built our own custom infra to run everything from encoder-only models to full LLMs on device. It's cross-platform, supports LoRa adapters, and optimized for the GPU. This system preserves privacy and enables fast inference times. 🧠 MEMORY – with your permission, Dia automatically tailors your AI experiences to you, personally, based on the tabs you open while browsing normally every day. We're also bringing vertical memory to specific features. ♻️ DATA FLYWHEELS – our Fall/Winter P0 is to double-down on training custom models based on implicit signals from daily use of Dia. Dia should get smarter and more useful the more people use it. Whether via RL, auto-generated prompts, or otherwise. If this work sounds interesting to you please visit our jobs page or email [email protected]. Hiring nearly every related role -- from ML engineers to people prototyping with AI and context/prompt writers -- everyone encouraged to apply!!show more

Josh Miller
68,130 views • 1 year ago
ClickUp now employs over 100,000 AI AGENTS for our... customers. This is from just THREE WEEKS of customers vibe coding full-blown teams of agents, THEMSELVES. BUT there's a problem. Since Super Agents are built agnostically, horizontally, and deeply capable with human-level abilities, you can literally build an agent for anything. We found that MOST of our customers have NO CLUE where to start. This is their very FIRST TIME EVER managing an agent. What I recommend is starting with a PROBLEM. Everybody can think of a problem they have. Just tell Super Agent Builder about your problems... about where you're WASTING time... about what you WISH you could do but you can't because of resource constraints. We've also found that human FEEDBACK and iteration are KEY. After agents are done with their jobs, give them feedback... Was it good? Was it bad? What do you want to see differently? They AUTOMATICALLY SELF-IMPROVE. Every time, they'll continuously get SMARTER. Personally, I find that agents go from AVERAGE intelligence to SUPER-human intelligence within about a month of working with them. Super Agents have truly democratized productivity, empowering literally anyone to build personalized, powerful agents in minutes. What problems do you wish you could solve? What do you not have enough time to get done? What would you like to do but don't have the resources for? What busy work do you wish you could get rid of?show more

Zeb Evans
18,324 views • 7 months ago