Does off-policy value-based RL scale? In LLMs, larger scale... predictably improves performance. Value-based RL learns from arbitrary data and is sample-efficient, but folk wisdom says it doesn't scale 🧵⬇️We show predictability for scaling value-based RL!show more

Oleg Rybkin
23,994 views • 1 year ago
🚨Current scalable RL algos train a policy w/o value... func, which is limiting with learning in open-ended, non-stationary, dynamic environments. But, how to scale value-based RL with more data/compute is unclear... Not anymore: presenting scaling laws for value-based RL 🧵⬇️show more

Aviral Kumar
37,377 views • 1 year ago
Introducing CQN: Coarse-to-fine Q-Network, a value-based RL algorithm for... continuous control🦾Initialized with 20~50 demonstrations, it learns to solve real-world robotic tasks within 10 mins of training, without any pre-training and shaped rewards! (1/4)show more

Younggyo Seo
16,431 views • 2 years ago
New work: The Value Axis 🎯 How do LLMs... choose which path to take mid-task? We find they internally track the chance of reaching their goal along a linear axis, akin to a value function in RL. We show it modulates confidence in math & coding and can be reshaped with DPO and SFT.show more

Nick Jiang
29,450 views • 2 months ago
New research from Databricks: LLMs Can Learn to Reason... via Off-Policy RL Optimal Advantage-based Policy Optimization with Lagged Inference policy (OAPL) shows you don’t need strict on-policy training to improve reasoning. It matches or beats Group Relative Policy Optimization (GRPO), stays stable with large policy lag, and uses ~3× fewer training generations. For Databricks customers, it’s a simpler, practical, and equally powerful approach to RL that Databricks is pioneering internally — and bringing directly to Databricks customers, so enterprises can improve agents using the same methods we use for our in-house agents, without complex infrastructure changes.show more

Databricks AI Research
12,783 views • 6 months ago
This figure from HIL-SERL is one of the clearest... visualisations of how RL learns differently from imitation learning. The difference comes down to this: imitation learning treats each (state, action) pair as independent. A correction at timestep 20 teaches nothing about timestep 19 or 21. RL propagates reward backward through time. One successful insertion updates the value estimate of every state along the trajectory. So RL builds a full map of "which states lead to success"; imitation learning just memorizes individual snapshots. Setup: a robot inserting a RAM stick into a motherboard slot. Each dot is an end-effector position (Y = lateral, Z = height). Starting position is randomized. Left to right = training progressing. Top row (RL): the policy builds a funnel. Broad at the top, narrowing into the target. It systematically fills in the state space, learning which paths lead to success from many different starting positions. Bottom row (imitation learning / HG-DAgger, same human data): sparse, diffuse, no funnel. The policy only learns near states the human demonstrated. Both have access to the same data, including human corrections, but a completely different structure emerges.show more

Dominique Paul
24,433 views • 7 months ago
GPT-5.5 by Reasoning Effort: I've asked it in Codex... to create a physics-based visualisation of RL cycles for different sized models (70b, 1t, 10t), to demonstrate how the amount of RL you can do differs by model size. My assessment of each: - Low: weird slop - Medium: kinda cooked - High: sort of tried but ultimately incoherent - Extra High: elite - really nice idea and well executed Obviously this is just one shot, but worth trying different reasoning levels for the new models, medium seems to be pretty good for GPT-5.5 and it was really bad for many previous GPT models.show more

Peter Gostev (SF: 22-26 June)
209,258 views • 4 months ago
Model-Free Reinforcement Learning (MFRL) has been alluring, especially with... supercharged compute with physics on GPU. However, the methods use 0-th order gradients, and are often not the best optimizers. Can we do better than PPO in continuous control for robotics? Turns out yes! 🥳 tl;dr: Faster, better RL than PPO in continuous control 💪 The answer lies in using more information from the simulation. We are juicing the simulation on GPU as it is, why not use it for gradients as well? This has been a driving question in a series of our works. We first studied this problem in ICLR 2022 paper on Short Horizon Actor Critic Naive gradient based methods are stuck in local minima and have exploding/vanishing gradients. SHAC solved this problem truncated rollouts and model based value estimation, where the model is Differentiable Sim. This boosted sample efficiency and wall-clock time immensely especially in high dimensional systems such as humanoids Yet, given enough compute PPO often caught up. Our follow up paper on on Adaptive Horizon Actor Critic at ICML 2024 discovers the cause and provides a fix. However, we find that even when given ground-truth dynamics, not all gradients are useful due to sample error. 1st-Order Model-Based Reinforcement Learning methods employing differentiable simulation provide gradients with reduced variance but are susceptible to bias in scenarios involving stiff dynamics, such as physical contact. We find that back-propagating through contact and long trajectories drastically reduces gradient accuracy. Using this insight, we propose AHAC to dynamically adapt its roll-out horizon to avoid differentiating through stiff contact. AHAC is a first-order model-based RL algorithm that learns high-dimensional tasks in minutes (wall clock) and outperforms PPO by 40%, even in the limit of data provided to PPO. This work is led by Ignat Georgiev alongside Krishnan Srinivasan, Jie Xu, Eric Heiden and ample assistance from warp team at NVIDIA Robotics (Miles Macklin)show more

Animesh Garg
52,308 views • 2 years ago
We’re excited to announce our integration with SKALE, a... high-performance, zero-gas blockchain purpose-built for speed, scale, and security. This partnership strengthens our infrastructure as we continue building transparent, trust-based systems for decentralized science. We’re excited about what this unlocks for researchers, contributors, and the future of data integrity in DeSci. 👀 Look out for more on how we’re using SKALE in the AxonDAO ecosystem.show more

AxonDAO
33,926 views • 1 year ago
High-resolution image and video generation is hitting a wall... because attention in DiTs scales quadratically with token count. But does every pixel need to be in full resolution? Introducing Foveated Diffusion: a new approach for efficient diffusion-based generation that allocates compute where it matters most. 1/7🧵show more

Gordon Wetzstein
171,136 views • 5 months ago
This one sentence from Mark Zuckerberg proves he's serious... about ending the censorship on his platforms. "We're going to move our trust and safety and content moderation teams out of California. And our US-based content review is going to be based in Texas." This means Silicon Valley liberals will no longer have their thumb on the scale and pick and choose what we, the peasants, get to post & view. I still don't forgive Zuck for rigging the 2020 election, but I think he means what he says.show more

George
431,555 views • 1 year ago
March 18, 2025 marked the public launch of OptimAI.... In one year, it has evolved from a lightweight node layer into a decentralized intelligence infrastructure powering real-time data, compute, and reinforcement for agentic systems. Not just nodes. Not just data. A continuously learning, network-driven intelligence layer. This is infrastructure for a new class of software: autonomous agents that persist, adapt, and operate across environments. Year one established the network. Year two is where it compounds into coordination and value flow. Personal agents. Reinforcement at network scale. Emerging primitives for AgentFi. New layers coming online. 2026 won’t just be about scale, it’s where the network starts to operate. Keep building!show more

OptimAI Network
34,396 views • 5 months ago
AI has had exactly two scaling axes that worked... so far, and the second one is starting to look finite too the first one was pretraining: with scaling parameters and data, we got world knowledge (i.e. ChatGPT had read enough to know things), but it started saturating a while ago the second one was RL, and people had been doing RL the whole time before that: RLHF is RL but it never scaled far because it was trying to control the exact output, which tokens come out, how the text reads, but you can only push that so far before you’re just polishing RLVR dropped that constraint: giving the model a task, then checking whether the final answer is right, and ignoring everything in between -- so the model does whatever it wants in the middle and only the endpoint gets graded, and that’s much closer to actual RL and it’s what bought us planning and reasoning (arguably, tool use sits around 2.5 on this list -- while useful, it's not a different kind of thing) so one axis gave knowledge, the other gave reasoning, and both of them are one model working alone the next axis is how many models you can get working on the same problem, which is a different kind of axis than the previous two we know that multi-agent RL has always been the harder problem: I spent years in that literature and the gap between single-agent and multi-agent is definitely not incremental -- it’s a whole different class of difficulty! which is also why the derivatives are steep at the start, nobody has picked the easy wins yet... and the thing that gates this multi-agent coordination is communication: models can only coordinate as well as they can exchange information, and right now they do that by writing sentences to each other imagine what could we possibly achieve if we properly open that third axis development by letting models to exchange information in their native "language" without loosing any computational data that they produce during inferenceshow more

Sasha Malysheva
12,177 views • 1 month ago
Pretty human-like hand Beijing-based SynapX will unveil its tendon-driven... OctoH-Hand at WRC. Human-scale in size, the hand uses a hybrid architecture combining fully tendon-driven actuation with direct-drive motors in the forearm. It integrates 28 independently controllable actuators and 23 active DoF, along with tactile sensors embedded in the palm. Interestingly…SynapX is building more than just a dexterous hand. It has also developed a World model(SYNWorld), and a data collection system(OctoSense), creating a loop from data collection to world understanding and policy generation, and finally to real-world execution and feedback. Another physical AI bridge for humanoid robots.show more

CyberRobo
49,646 views • 26 days ago
LongWriter Unleashing 10,000+ Word Generation from Long Context LLMs... discuss: Current long context large language models (LLMs) can process inputs up to 100,000 tokens, yet struggle to generate outputs exceeding even a modest length of 2,000 words. Through controlled experiments, we find that the model's effective generation length is inherently bounded by the sample it has seen during supervised fine-tuning (SFT). In other words, their output limitation is due to the scarcity of long-output examples in existing SFT datasets. To address this, we introduce AgentWrite, an agent-based pipeline that decomposes ultra-long generation tasks into subtasks, enabling off-the-shelf LLMs to generate coherent outputs exceeding 20,000 words. Leveraging AgentWrite, we construct LongWriter-6k, a dataset containing 6,000 SFT data with output lengths ranging from 2k to 32k words. By incorporating this dataset into model training, we successfully scale the output length of existing models to over 10,000 words while maintaining output quality. We also develop LongBench-Write, a comprehensive benchmark for evaluating ultra-long generation capabilities. Our 9B parameter model, further improved through DPO, achieves state-of-the-art performance on this benchmark, surpassing even much larger proprietary models. In general, our work demonstrates that existing long context LLM already possesses the potential for a larger output window--all you need is data with extended output during model alignment to unlock this capability.show more

AK
50,995 views • 2 years ago
We are at NeurIPS Conference for our 4th #MyoChallenge... and 7th #MyoSymposium! What started as a discussion with Vittorio Caggiano is a global community now MyoSuite 💪 In 2022, we started with two key hypotheses - 1⃣𝑺𝒄𝒂𝒍𝒊𝒏𝒈 𝑯𝒚𝒑𝒐𝒕𝒉𝒆𝒔𝒊𝒔: can we scale data driven learnings to achieve human level motor control? 2⃣𝑬𝒎𝒃𝒐𝒅𝒊𝒎𝒆𝒏𝒕 𝑯𝒚𝒑𝒐𝒕𝒉𝒆𝒔𝒊𝒔: Akin to Neurons's inspiration behind NN, are there embodied priors that will form the critical substrate to get to human performance? After 3 years, both these hypotheses are running strong. But in different ways than we anticipated -scaling hypothesis predicted vanilla RL algorithms (developed over OpenAI Gym and robotics tasks) will scale & realize human level motor control. RL did scale with better simulation & compute infrastructure but the curse of dimensionality became the limiter for high dimensional MSK systems. This is where our 2nd Embodiment Hypothesis kicked in. Spatial as well as morphological embodied priors facilitated development of next generation of algorithms at the intersection of representation and reinforcement learning - (DepRL from Pierre Schumacher et al, MyoDex & SAR from Cameron Vittorio Caggiano et al, Kinesis from Alberto Chiappa, muscleVAE from Yusen et al, etc) Current #MyoChallenge result evaluations phase was humbling realization - what started with a team of two evolved as a global community, the entry barriers has been lowered enough for even high school students, and underrepresented groups with limited resources to participate. It's incredible to realize the progress we have seen in 3 years. All the same time idiosyncrasies of the behaviors leave us quite unsatisfied and wanting more. Ahead of us there are exciting challenges on all frontiers -- embodiment, proprioceptive+exteroceptive sensing and control, validation -- presenting large real world potentials in health, wellness, sports, robotics. While there is a lot for us to be proud of, open challenges in understanding, as well as emulating human level motor intelligence remains. Join MyoSuite team for the awaited #MyoSymposium in discussing these frontiers on Saturday, Dec. 6th from 8-11 am: Ballroom 6D.show more

Vikash Kumar
11,145 views • 9 months ago
An experiment by frenpet dev (Adam) You send ETH,... the contract sends it straight back in the same tx; what persists is the record: an onchain, permissionless allowlist of wallets that provably held real capital at real times. ~$90M in transaction value has been deposited and received back. > Earn onchain points > Points are based on progressive ETH volume Using Austin Griffith interface makes it easier. Make sure smart wallet is turned off for EOA delegation.show more

〽️ᄃムt 🐾
50,714 views • 25 days ago
MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers paper... page: Recent advances in generative AI have significantly enhanced image and video editing, particularly in the context of text prompt control. State-of-the-art approaches predominantly rely on diffusion models to accomplish these tasks. However, the computational demands of diffusion-based methods are substantial, often necessitating large-scale paired datasets for training, and therefore challenging the deployment in practical applications. This study addresses this challenge by breaking down the text-based video editing process into two separate stages. In the first stage, we leverage an existing text-to-image diffusion model to simultaneously edit a few keyframes without additional fine-tuning. In the second stage, we introduce an efficient model called MaskINT, which is built on non-autoregressive masked generative transformers and specializes in frame interpolation between the keyframes, benefiting from structural guidance provided by intermediate frames. Our comprehensive set of experiments illustrates the efficacy and efficiency of MaskINT when compared to other diffusion-based methodologies. This research offers a practical solution for text-based video editing and showcases the potential of non-autoregressive masked generative transformers in this domain.show more

AK
25,449 views • 2 years ago