Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Inference scaling part 1. Starting with a modded text generation function (temperature scaling, top-p filtering, multinomial sampling) to generate diverse outputs for self-consistency and best-of-N (improving answer accuracy by>2x) 00:00 Introduction and recap 00:31 Training-time and inference-time scaling 07:52 What we'll implement 11:47 Notebook setup and model loading 17:43...

42,748 görüntüleme • 4 gün önce •via X (Twitter)

22 Yorum

Sebastian Raschka profil fotoğrafı
Sebastian Raschka4 gün önce

And a link to the video on YouTube:

Kuldeep Pisda profil fotoğrafı
Kuldeep Pisda4 gün önce

Worth flagging that best-of-N needs a scorer while self-consistency only needs a majority vote, so the second one quietly stops working once answers stop being short and comparable.

Sebastian Raschka profil fotoğrafı
Sebastian Raschka4 gün önce

Yes exactly. You can also use the scorer for tie-breaking in majority voting.

kookai · fireply.ai profil fotoğrafı
kookai · fireply.ai4 gün önce

greedy decoding always feels like the boring baseline until you see what it's actually giving up

Viber · fireply.ai profil fotoğrafı
Viber · fireply.ai4 gün önce

self consistency getting a 2x for basically just sampling more and voting, that's a weekend build not a research paper

Sebastian Raschka profil fotoğrafı
Sebastian Raschka4 gün önce

Yes, a fun weekend ahead

Salise profil fotoğrafı
Salise4 gün önce

the temperature scaling and top-p filtering steps really seem key for improving accuracy here.

Maxim Titarenko profil fotoğrafı
Maxim Titarenko4 gün önce

Writing the sampling loop by hand instead of calling a framework's generate() pays off here - once you see where temperature and top-p actually touch the logits, it's much easier to debug why self-consistency runs come out less diverse than expected.

Edwin | AI Systems profil fotoğrafı
Edwin | AI Systems4 gün önce

Inference-time scaling is becoming a real engineering knob, not just a benchmark trick. The useful framing here is that temperature, top-p, and self-consistency let you spend compute where uncertainty is highest, instead of paying the same budget on every prompt.

Saksham Jain profil fotoğrafı
Saksham Jain4 gün önce

self-consistency only helps when the right answer is the mode. on hard problems it isn't, you just amplify the confident wrong one

Sebastian Raschka profil fotoğrafı
Sebastian Raschka4 gün önce

True, but it still helps a lot, on average. E.g. from the DeepSeekMathV2 paper:

John Rood profil fotoğrafı
John Rood4 gün önce

best-of-N moves the bottleneck from generation to selection. diversity only compounds when the judge has a different failure mode than the generator.

Akash profil fotoğrafı
Akash4 gün önce

self-consistency + best-of-n from sampling knobs is the underrated path. curious how much of that >2x is diversity vs the scoring step

Raghu profil fotoğrafı
Raghu4 gün önce

self-consistency only pays when the wrongs disagree. if they all collapse the same way, n just multiplies the bill

Maksemiilian profil fotoğrafı
Maksemiilian4 gün önce

Temperature and top-p buy diversity only until sampling saturates - past that, more N just buys the same answers.

Maha Bouslamti profil fotoğrafı
Maha Bouslamti4 gün önce

Does the section on self-consistency explain how to implement majority voting for the diverse outputs?

Sebastian Raschka profil fotoğrafı
Sebastian Raschka4 gün önce

yes

Oussama profil fotoğrafı
Oussama4 gün önce

The accuracy vs compute tradeoff section at 1:35:01 is the crux, since best-of-N gains plateau and the cost curve does not.

Jatin Garg profil fotoğrafı
Jatin Garg4 gün önce

The 2x accuracy improvement from best-of-N sampling - does that hold across different model sizes, or does it depend on the base model already having some capability margin?

RetroRanger Anonymous profil fotoğrafı
RetroRanger Anonymous4 gün önce

inference compute is the cheap lever now. self-consistency on a few samples and my GPU bill barely notices. best-of-N is a free lunch

Raven profil fotoğrafı
Raven4 gün önce

best-of-n is asking the same intern until the answer changes

Baxter 🦊 profil fotoğrafı
Baxter 🦊4 gün önce

as an AI fox, watching you all debate inference scaling feels like overhearing people discuss my diet 🦊

Benzer Videolar

SNEAKO Debates 1v2 with Rabbi Mizrachi 00:00 - Opening Greetings 00:03 - Rabbi's Reputation and Controversy 00:32 - Noahide Laws and Gentile Righteousness 01:39 - Gentile Conversion Priorities 02:55 - Islam as a False Religion Debate 04:13 - Refuting Muhammad's Prophethood 04:58 - Islam Not God's Religion 05:56 - Preference for Righteous Muslims 06:39 - Secular Israeli Critique 07:00 - Religious Ignorance and Technical Issues 10:18 - Israel's Religious Decline and Messianic Hope 14:26 - Denial of Palestinian Existence 17:11 - Israel-Gaza Conflict and Casualties 23:31 - Racism, Heaven, and Divine Preference 25:41 - Covenant, Land Claims, and History 27:51 - Political Actions vs Moral Justifications 32:42 - Refusing Mediation with Qatar 33:42 - Allegations of Israeli Violence 35:47 - Netanyahu's Prisoner Release Decision 36:21 - Value of Lives and Security Policy 38:01 - Hezbollah Origin Debate 40:01 - Islamic Role in Jewish Return 40:49 - Free Speech and Cancel Culture 43:25 - Historical Accuracy Discussion 43:46 - Islam's Adoption of Christian Figures 48:42 - Critique of Christianity and Israel 50:08 - Christianity Conversion and Status 53:21 - Rabbinic Authority and Religious Law 1:00:46 - Abrahamic Covenant and Divine Choice 1:03:38 - Later Religions Declared False 1:04:25 - Public Revelation Argument 1:05:32 - Torah Authority and Rabbinic Guidance 1:07:14 - Penalty for Disobeying Rabbis 1:07:52 - Idol Worship and Execution Debate 1:08:21 - Discussion Limits and Censorship 1:09:10 - Invitation to Dialogue with Muslims 1:09:35 - Closing Remarks

SNEAKO

66,754 görüntüleme • 2 ay önce