正在加载视频...

视频加载失败

It's getting better, but progress is pretty slow. The reward here is for staying upright while matching a rest post. This one has about 65k parameters trained with PPO for half an hour.

33,353 次观看 • 1 年前 •via X (Twitter)

11 条评论

Alex (Cama) 的头像
Alex (Cama)1 年前

Poor little guy

Brian Jordan 的头像
Brian Jordan1 年前

@brandon_xyzw

Clifton Sellers 的头像
Clifton Sellers1 年前

One day I just woke up with the ability to post and grow my business 3+ years of posting Constant testing/refining Committing to the process It doesn’t happen quick, but once you build out your personal brand the leverage is insane

Makan Gilani 的头像
Makan Gilani1 年前

30 mins training on just one agent or multiple?

Dennis Gustafsson 的头像
Dennis Gustafsson1 年前

It trains in 16 environments concurrently to better utilize the CPU cores, but still very inefficient because I'm training in python and using a proxy environment communicating with the game over a network link. I should really get all this running in c++ instead..

T 的头像
T1 年前

Have you ever tried letting it run for long periods of time? One of the biggest mistakes I made when experimenting with RL was trying to iterate too quickly and not allowing models time to converge

Dennis Gustafsson 的头像
Dennis Gustafsson1 年前

Interesting, I have almost the opposite experience. If I leave it training over night, it usually develops very weird, specific behavior. Maybe I should lower the learning rate..

🇮🇹🌴 F e n n y e n n 🌴🇮🇹 的头像
🇮🇹🌴 F e n n y e n n 🌴🇮🇹1 年前

poor lil guy getting abused

typeofalex 的头像
typeofalex1 年前

30min sounds amazing, but I hope not 30min and dozens of GPUs.

CrazyDescent 🍉 的头像
CrazyDescent 🍉1 年前

Great job accelerating climate change for your virtual punishment simulator

cryptobiot 的头像
cryptobiot1 年前

okay now needs sound effects and user control of the cube cannon

相关视频