Does LLM RL post-training need to be on-policy?

Kianté Brantley
114,174 次观看 • 6 个月前
LLM training on RTX 5090

ℏεsam
144,115 次观看 • 1 年前
Does off-policy value-based RL scale? In LLMs, larger scale... show more

Oleg Rybkin
23,994 次观看 • 1 年前
As a supporter of the open weights ecosystem, we're... show more

Applied Compute
25,234 次观看 • 1 个月前
Rare serious post on here but it need to... show more

Alexandra Pembroke
10,899 次观看 • 3 个月前
Tutorial Time: Run any open-source LLM locally. Now we... show more

Linus ✦ Ekenstam
915,956 次观看 • 2 年前
Reinforcement learning should be able to improve upon behaviors... show more

Vivek Myers
79,523 次观看 • 1 年前
New project! Flow Policy Gradients for Robot Control tldr;... show more

Brent Yi
92,732 次观看 • 7 个月前
Need my house to look like a RL store

Prep Propaganda 👔
17,980 次观看 • 3 天前
How thick does ice need to be before you... show more

AlphaFox
43,053 次观看 • 7 个月前
Introducing autoresearch for arXiv papers: replicate and experiment on... show more

alphaXiv
60,572 次观看 • 12 天前
Nature does not always need venom to be terrifying.

NECA
36,412 次观看 • 1 个月前
Does your bed board need to be replaced 😏

sytoys-us1
49,674 次观看 • 11 个月前
How does high-fidelity tactile simulation help robots nail the... show more

Binghao Huang
46,982 次观看 • 10 个月前
How it feels to be an LLM

Beff (e/acc)
36,669 次观看 • 1 年前
i need to post more on here 😩

MS. F!NEE $HITT
124,925 次观看 • 1 年前
🚨Current scalable RL algos train a policy w/o value... show more

Aviral Kumar
37,377 次观看 • 1 年前
RL is painfully slow 😭 — bottlenecked by super-long... show more

Infini-AI-Lab
78,984 次观看 • 3 个月前
🤔 How to fine-tune an Imitation Learning policy (e.g.,... show more

Tongzhou Mu 🤖🦾🦿
17,018 次观看 • 1 年前