正在加载视频...
视频加载失败
It's getting better, but progress is pretty slow. The reward here is for staying upright while matching a rest post. This one has about 65k parameters trained with PPO for half an hour.
11 条评论

Poor little guy

@brandon_xyzw

One day I just woke up with the ability to post and grow my business 3+ years of posting Constant testing/refining Committing to the process It doesn’t happen quick, but once you build out your personal brand the leverage is insane

30 mins training on just one agent or multiple?

It trains in 16 environments concurrently to better utilize the CPU cores, but still very inefficient because I'm training in python and using a proxy environment communicating with the game over a network link. I should really get all this running in c++ instead..

Have you ever tried letting it run for long periods of time? One of the biggest mistakes I made when experimenting with RL was trying to iterate too quickly and not allowing models time to converge

Interesting, I have almost the opposite experience. If I leave it training over night, it usually develops very weird, specific behavior. Maybe I should lower the learning rate..

poor lil guy getting abused

30min sounds amazing, but I hope not 30min and dozens of GPUs.

Great job accelerating climate change for your virtual punishment simulator

okay now needs sound effects and user control of the cube cannon
相关视频
Sensitive content
This is truly evil. Holding a toddler hostage for over half an hour with a knife to her stomach.
Ian Miles Cheong
60,812 次观看 • 11 个月前

