Why am I working on RL for LLMs, when... show more

Shane Gu
12,440 Aufrufe • vor 8 Monaten
Does off-policy value-based RL scale? In LLMs, larger scale... show more

Oleg Rybkin
23,979 Aufrufe • vor 1 Jahr
Excited to share CrystalReasoner, a reasoning model for crystal... show more

Sherry Yang
11,013 Aufrufe • vor 3 Monaten
Some part of our ongoing work on the news... show more

Joonho Lee
18,007 Aufrufe • vor 1 Jahr
I've gotten a mujoco sim RL training loop for... show more

kache
30,112 Aufrufe • vor 2 Monaten
1/ While most RL methods use shallow MLPs (~2–5... show more

Kevin Wang
155,514 Aufrufe • vor 1 Jahr
Thanks AK! Finally, robot can do continuous, agile, autonomous,... show more

Guanya Shi
32,155 Aufrufe • vor 1 Jahr
Mullin: There’s a lot of work we can do... show more

Republicans against Trump
37,539 Aufrufe • vor 5 Monaten
I am looking for this job, I can work... show more

Wafula
1,880,138 Aufrufe • vor 1 Monat
These are not CGI. Reinforcement learning is so back.... show more

Jim Fan
356,840 Aufrufe • vor 1 Jahr
A few selects of Animation / Motion Design work... show more

Erv
15,261 Aufrufe • vor 2 Jahren
i am on the right path and everything will... show more

pulga
3,001,711 Aufrufe • vor 2 Jahren
I agree. That’s why we’ve been working on something... show more

Joonas Virtanen
156,850 Aufrufe • vor 5 Monaten
🐶 there's only a month left for me to... show more

🐶
74,098 Aufrufe • vor 16 Tagen
Hmm what am I working on?!

Dan Greenheck
11,539 Aufrufe • vor 4 Monaten
"Why am i just seeing me and Bae on... show more

WeenUpdates
195,134 Aufrufe • vor 1 Jahr
Am I the only one who plays Test Track... show more

Derek Bell
10,392 Aufrufe • vor 1 Jahr