Does LLM RL post-training need to be on-policy?

Kianté Brantley
114,174 views • 6 months ago
LLM training on RTX 5090

ℏεsam
144,115 views • 1 year ago
Does off-policy value-based RL scale? In LLMs, larger scale... show more

Oleg Rybkin
23,994 views • 1 year ago
As a supporter of the open weights ecosystem, we're... show more

Applied Compute
25,234 views • 1 month ago
Rare serious post on here but it need to... show more

Alexandra Pembroke
10,899 views • 3 months ago
Tutorial Time: Run any open-source LLM locally. Now we... show more

Linus ✦ Ekenstam
915,956 views • 2 years ago
Reinforcement learning should be able to improve upon behaviors... show more

Vivek Myers
79,523 views • 1 year ago
New project! Flow Policy Gradients for Robot Control tldr;... show more

Brent Yi
92,732 views • 7 months ago
Need my house to look like a RL store

Prep Propaganda 👔
17,803 views • 3 days ago
How thick does ice need to be before you... show more

AlphaFox
43,053 views • 7 months ago
Introducing autoresearch for arXiv papers: replicate and experiment on... show more

alphaXiv
60,572 views • 12 days ago
Nature does not always need venom to be terrifying.

NECA
36,412 views • 1 month ago
Does your bed board need to be replaced 😏

sytoys-us1
49,674 views • 11 months ago
How does high-fidelity tactile simulation help robots nail the... show more

Binghao Huang
46,982 views • 10 months ago
How it feels to be an LLM

Beff (e/acc)
36,669 views • 1 year ago
i need to post more on here 😩

MS. F!NEE $HITT
124,925 views • 1 year ago
🚨Current scalable RL algos train a policy w/o value... show more

Aviral Kumar
37,377 views • 1 year ago
RL is painfully slow 😭 — bottlenecked by super-long... show more

Infini-AI-Lab
78,984 views • 2 months ago
🤔 How to fine-tune an Imitation Learning policy (e.g.,... show more

Tongzhou Mu 🤖🦾🦿
17,018 views • 1 year ago