Loading video...
Video Failed to Load
Some teams use sweeps, heuristics, or scaling laws to determine their training LR. At Character, we just have Noam Shazeer dial it to the right value.
169,651 views • 2 years ago •via X (Twitter)
10 Comments

It turns out that you need to spend a lot of compute and care to beat hyperparameters that Noam just makes up. It's absurd how often that happens.

lol bet that’s better than most if not all schedules

Haha nice!

Why on earth would you plot /avg/ of lr though

it’s not. for all metrics we track avg and denominator. for LR, denominator is always 1. just a side effect of our templating

Fake news. That's definitely your hand doing it.

ser I needed this so bad

what I see is a training of advanced models I think that is but it is good

How do I get this for @dippy_ai

can you remove the filter next or implement a switch? I’ll actually send you money to atleast ATTEMPT to try to. or give the community some actual reason as to why we’re being ignored. please do something this is overall just disgusting.
Related Videos
Sensitive content
WHO ARE YOU TO DICTATE WHERE THE PALESTIANS GO‼️ What right does Israel or its Western allies have to dictate where the Palestinians should go and to determine their right to return. WHY DON'T ISRAELIS JUST FUCK OFF TO POLAND GERMANY AND AMERICA‼️ Rahma ✊🏼🇵🇸
Earth Hippy 🌎🕊️💚
73,057 views • 1 year ago
Sensitive content
Women like this are the reason fertility clinics are pursuing 18 year old girls for their eggs. At some point we are going to have to ask if this is right or fair.
SurrogacyConcern
29,605 views • 5 months ago

