Does LLM RL post-training need to be on-policy?

Kianté Brantley
114,174 просмотров • 6 месяцев назад
LLM training on RTX 5090

ℏεsam
144,115 просмотров • 1 год назад
Does off-policy value-based RL scale? In LLMs, larger scale... show more

Oleg Rybkin
23,994 просмотров • 1 год назад
As a supporter of the open weights ecosystem, we're... show more

Applied Compute
25,234 просмотров • 1 месяц назад
Rare serious post on here but it need to... show more

Alexandra Pembroke
10,899 просмотров • 3 месяцев назад
Tutorial Time: Run any open-source LLM locally. Now we... show more

Linus ✦ Ekenstam
915,956 просмотров • 2 лет назад
Reinforcement learning should be able to improve upon behaviors... show more

Vivek Myers
79,523 просмотров • 1 год назад
New project! Flow Policy Gradients for Robot Control tldr;... show more

Brent Yi
92,732 просмотров • 7 месяцев назад
Need my house to look like a RL store

Prep Propaganda 👔
17,980 просмотров • 3 дней назад
How thick does ice need to be before you... show more

AlphaFox
43,053 просмотров • 7 месяцев назад
Introducing autoresearch for arXiv papers: replicate and experiment on... show more

alphaXiv
60,572 просмотров • 12 дней назад
Nature does not always need venom to be terrifying.

NECA
36,412 просмотров • 1 месяц назад
Does your bed board need to be replaced 😏

sytoys-us1
49,674 просмотров • 11 месяцев назад
How does high-fidelity tactile simulation help robots nail the... show more

Binghao Huang
46,982 просмотров • 10 месяцев назад
How it feels to be an LLM

Beff (e/acc)
36,669 просмотров • 1 год назад
i need to post more on here 😩

MS. F!NEE $HITT
124,925 просмотров • 1 год назад
🚨Current scalable RL algos train a policy w/o value... show more

Aviral Kumar
37,377 просмотров • 1 год назад
RL is painfully slow 😭 — bottlenecked by super-long... show more

Infini-AI-Lab
78,984 просмотров • 3 месяцев назад
🤔 How to fine-tune an Imitation Learning policy (e.g.,... show more

Tongzhou Mu 🤖🦾🦿
17,018 просмотров • 1 год назад