Does LLM RL post-training need to be on-policy?

Kianté Brantley
114,174 Aufrufe • vor 6 Monaten
LLM training on RTX 5090

ℏεsam
144,115 Aufrufe • vor 1 Jahr
Does off-policy value-based RL scale? In LLMs, larger scale... show more

Oleg Rybkin
23,994 Aufrufe • vor 1 Jahr
As a supporter of the open weights ecosystem, we're... show more

Applied Compute
25,234 Aufrufe • vor 1 Monat
Rare serious post on here but it need to... show more

Alexandra Pembroke
10,899 Aufrufe • vor 3 Monaten
Tutorial Time: Run any open-source LLM locally. Now we... show more

Linus ✦ Ekenstam
915,956 Aufrufe • vor 2 Jahren
Reinforcement learning should be able to improve upon behaviors... show more

Vivek Myers
79,523 Aufrufe • vor 1 Jahr
New project! Flow Policy Gradients for Robot Control tldr;... show more

Brent Yi
92,732 Aufrufe • vor 7 Monaten
Need my house to look like a RL store

Prep Propaganda 👔
17,803 Aufrufe • vor 3 Tagen
How thick does ice need to be before you... show more

AlphaFox
43,053 Aufrufe • vor 7 Monaten
Introducing autoresearch for arXiv papers: replicate and experiment on... show more

alphaXiv
60,572 Aufrufe • vor 12 Tagen
Nature does not always need venom to be terrifying.

NECA
36,412 Aufrufe • vor 1 Monat
Does your bed board need to be replaced 😏

sytoys-us1
49,674 Aufrufe • vor 11 Monaten
How does high-fidelity tactile simulation help robots nail the... show more

Binghao Huang
46,982 Aufrufe • vor 10 Monaten
How it feels to be an LLM

Beff (e/acc)
36,669 Aufrufe • vor 1 Jahr
i need to post more on here 😩

MS. F!NEE $HITT
124,925 Aufrufe • vor 1 Jahr
🚨Current scalable RL algos train a policy w/o value... show more

Aviral Kumar
37,377 Aufrufe • vor 1 Jahr
RL is painfully slow 😭 — bottlenecked by super-long... show more

Infini-AI-Lab
78,984 Aufrufe • vor 2 Monaten
🤔 How to fine-tune an Imitation Learning policy (e.g.,... show more

Tongzhou Mu 🤖🦾🦿
17,018 Aufrufe • vor 1 Jahr