Does LLM RL post-training need to be on-policy?

Kianté Brantley's profile picture

Kianté Brantley

114,174 görüntüleme • 6 ay önce

LLM training on RTX 5090

ℏεsam's profile picture

ℏεsam

144,115 görüntüleme • 1 yıl önce

Does off-policy value-based RL scale? In LLMs, larger scale...

Oleg Rybkin's profile picture

Oleg Rybkin

23,994 görüntüleme • 1 yıl önce

As a supporter of the open weights ecosystem, we're...

Applied Compute's profile picture

Applied Compute

25,234 görüntüleme • 1 ay önce

Rare serious post on here but it need to...

Alexandra Pembroke's profile picture

Alexandra Pembroke

10,899 görüntüleme • 3 ay önce

Tutorial Time: Run any open-source LLM locally. Now we...

Linus ✦ Ekenstam's profile picture

Linus ✦ Ekenstam

915,956 görüntüleme • 2 yıl önce

Reinforcement learning should be able to improve upon behaviors...

Vivek Myers's profile picture

Vivek Myers

79,523 görüntüleme • 1 yıl önce

New project! Flow Policy Gradients for Robot Control tldr;...

Brent Yi's profile picture

Brent Yi

92,732 görüntüleme • 7 ay önce

Need my house to look like a RL store

Prep Propaganda 👔's profile picture

Prep Propaganda 👔

17,980 görüntüleme • 3 gün önce

How thick does ice need to be before you...

AlphaFox's profile picture

AlphaFox

43,053 görüntüleme • 7 ay önce

Introducing autoresearch for arXiv papers: replicate and experiment on...

alphaXiv's profile picture

alphaXiv

60,572 görüntüleme • 12 gün önce

Nature does not always need venom to be terrifying.

NECA's profile picture

NECA

36,412 görüntüleme • 1 ay önce

Does your bed board need to be replaced 😏

sytoys-us1's profile picture

sytoys-us1

49,674 görüntüleme • 11 ay önce

How does high-fidelity tactile simulation help robots nail the...

Binghao Huang's profile picture

Binghao Huang

46,982 görüntüleme • 10 ay önce

How it feels to be an LLM

Beff (e/acc)'s profile picture

Beff (e/acc)

36,669 görüntüleme • 1 yıl önce

i need to post more on here 😩

MS. F!NEE $HITT's profile picture

MS. F!NEE $HITT

124,925 görüntüleme • 1 yıl önce

🚨Current scalable RL algos train a policy w/o value...

Aviral Kumar's profile picture

Aviral Kumar

37,377 görüntüleme • 1 yıl önce

RL is painfully slow 😭 — bottlenecked by super-long...

Infini-AI-Lab's profile picture

Infini-AI-Lab

78,984 görüntüleme • 2 ay önce

🤔 How to fine-tune an Imitation Learning policy (e.g.,...

Tongzhou Mu 🤖🦾🦿's profile picture

Tongzhou Mu 🤖🦾🦿

17,018 görüntüleme • 1 yıl önce