Does LLM RL post-training need to be on-policy?

Kianté Brantley
114,174 görüntüleme • 6 ay önce
LLM training on RTX 5090

ℏεsam
144,115 görüntüleme • 1 yıl önce
Does off-policy value-based RL scale? In LLMs, larger scale... show more

Oleg Rybkin
23,994 görüntüleme • 1 yıl önce
As a supporter of the open weights ecosystem, we're... show more

Applied Compute
25,234 görüntüleme • 1 ay önce
Rare serious post on here but it need to... show more

Alexandra Pembroke
10,899 görüntüleme • 3 ay önce
Tutorial Time: Run any open-source LLM locally. Now we... show more

Linus ✦ Ekenstam
915,956 görüntüleme • 2 yıl önce
Reinforcement learning should be able to improve upon behaviors... show more

Vivek Myers
79,523 görüntüleme • 1 yıl önce
New project! Flow Policy Gradients for Robot Control tldr;... show more

Brent Yi
92,732 görüntüleme • 7 ay önce
Need my house to look like a RL store

Prep Propaganda 👔
17,980 görüntüleme • 3 gün önce
How thick does ice need to be before you... show more

AlphaFox
43,053 görüntüleme • 7 ay önce
Introducing autoresearch for arXiv papers: replicate and experiment on... show more

alphaXiv
60,572 görüntüleme • 12 gün önce
Nature does not always need venom to be terrifying.

NECA
36,412 görüntüleme • 1 ay önce
Does your bed board need to be replaced 😏

sytoys-us1
49,674 görüntüleme • 11 ay önce
How does high-fidelity tactile simulation help robots nail the... show more

Binghao Huang
46,982 görüntüleme • 10 ay önce
How it feels to be an LLM

Beff (e/acc)
36,669 görüntüleme • 1 yıl önce
i need to post more on here 😩

MS. F!NEE $HITT
124,925 görüntüleme • 1 yıl önce
🚨Current scalable RL algos train a policy w/o value... show more

Aviral Kumar
37,377 görüntüleme • 1 yıl önce
RL is painfully slow 😭 — bottlenecked by super-long... show more

Infini-AI-Lab
78,984 görüntüleme • 2 ay önce
🤔 How to fine-tune an Imitation Learning policy (e.g.,... show more

Tongzhou Mu 🤖🦾🦿
17,018 görüntüleme • 1 yıl önce