Loading video...

Video Failed to Load

Go Home

Some teams use sweeps, heuristics, or scaling laws to determine their training LR. At Character, we just have Noam Shazeer dial it to the right value.

169,651 views • 2 years ago •via X (Twitter)

10 Comments

Ying Xiao's profile picture
Ying Xiao2 years ago

It turns out that you need to spend a lot of compute and care to beat hyperparameters that Noam just makes up. It's absurd how often that happens.

Naman Goyal's profile picture
Naman Goyal2 years ago

lol bet that’s better than most if not all schedules

Yi Tay's profile picture
Yi Tay2 years ago

Haha nice!

Lucas Beyer (bl16)'s profile picture
Lucas Beyer (bl16)2 years ago

Why on earth would you plot /avg/ of lr though

Stephen Roller's profile picture
Stephen Roller2 years ago

it’s not. for all metrics we track avg and denominator. for LR, denominator is always 1. just a side effect of our templating

Zeming Lin's profile picture
Zeming Lin2 years ago

Fake news. That's definitely your hand doing it.

cooking's profile picture
cooking2 years ago

ser I needed this so bad

RBPrest's profile picture
RBPrest2 years ago

what I see is a training of advanced models I think that is but it is good

manoj's profile picture
manoj2 years ago

How do I get this for @dippy_ai

ghoulie's profile picture
ghoulie2 years ago

can you remove the filter next or implement a switch? I’ll actually send you money to atleast ATTEMPT to try to. or give the community some actual reason as to why we’re being ignored. please do something this is overall just disgusting.

Related Videos