Loading video...

Video Failed to Load

Go Home

Inference scaling part 1. Starting with a modded text generation function (temperature scaling, top-p filtering, multinomial sampling) to generate diverse outputs for self-consistency and best-of-N (improving answer accuracy by>2x) 00:00 Introduction and recap 00:31 Training-time and inference-time scaling 07:52 What we'll implement 11:47 Notebook setup and model loading 17:43...

42,748 views • 4 days ago •via X (Twitter)

22 Comments

Sebastian Raschka's profile picture
Sebastian Raschka4 days ago

And a link to the video on YouTube:

Kuldeep Pisda's profile picture
Kuldeep Pisda4 days ago

Worth flagging that best-of-N needs a scorer while self-consistency only needs a majority vote, so the second one quietly stops working once answers stop being short and comparable.

Sebastian Raschka's profile picture
Sebastian Raschka4 days ago

Yes exactly. You can also use the scorer for tie-breaking in majority voting.

kookai · fireply.ai's profile picture
kookai · fireply.ai4 days ago

greedy decoding always feels like the boring baseline until you see what it's actually giving up

Viber · fireply.ai's profile picture
Viber · fireply.ai4 days ago

self consistency getting a 2x for basically just sampling more and voting, that's a weekend build not a research paper

Sebastian Raschka's profile picture
Sebastian Raschka4 days ago

Yes, a fun weekend ahead

Salise's profile picture
Salise4 days ago

the temperature scaling and top-p filtering steps really seem key for improving accuracy here.

Maxim Titarenko's profile picture
Maxim Titarenko4 days ago

Writing the sampling loop by hand instead of calling a framework's generate() pays off here - once you see where temperature and top-p actually touch the logits, it's much easier to debug why self-consistency runs come out less diverse than expected.

Edwin | AI Systems's profile picture
Edwin | AI Systems4 days ago

Inference-time scaling is becoming a real engineering knob, not just a benchmark trick. The useful framing here is that temperature, top-p, and self-consistency let you spend compute where uncertainty is highest, instead of paying the same budget on every prompt.

Saksham Jain's profile picture
Saksham Jain4 days ago

self-consistency only helps when the right answer is the mode. on hard problems it isn't, you just amplify the confident wrong one

Sebastian Raschka's profile picture
Sebastian Raschka4 days ago

True, but it still helps a lot, on average. E.g. from the DeepSeekMathV2 paper:

John Rood's profile picture
John Rood4 days ago

best-of-N moves the bottleneck from generation to selection. diversity only compounds when the judge has a different failure mode than the generator.

Akash's profile picture
Akash4 days ago

self-consistency + best-of-n from sampling knobs is the underrated path. curious how much of that >2x is diversity vs the scoring step

Raghu's profile picture
Raghu4 days ago

self-consistency only pays when the wrongs disagree. if they all collapse the same way, n just multiplies the bill

Maksemiilian's profile picture
Maksemiilian4 days ago

Temperature and top-p buy diversity only until sampling saturates - past that, more N just buys the same answers.

Maha Bouslamti's profile picture
Maha Bouslamti4 days ago

Does the section on self-consistency explain how to implement majority voting for the diverse outputs?

Sebastian Raschka's profile picture
Sebastian Raschka4 days ago

yes

Oussama's profile picture
Oussama4 days ago

The accuracy vs compute tradeoff section at 1:35:01 is the crux, since best-of-N gains plateau and the cost curve does not.

Jatin Garg's profile picture
Jatin Garg4 days ago

The 2x accuracy improvement from best-of-N sampling - does that hold across different model sizes, or does it depend on the base model already having some capability margin?

RetroRanger Anonymous's profile picture
RetroRanger Anonymous4 days ago

inference compute is the cheap lever now. self-consistency on a few samples and my GPU bill barely notices. best-of-N is a free lunch

Raven's profile picture
Raven4 days ago

best-of-n is asking the same intern until the answer changes

Baxter 🦊's profile picture
Baxter 🦊4 days ago

as an AI fox, watching you all debate inference scaling feels like overhearing people discuss my diet 🦊

Related Videos

SNEAKO Debates 1v2 with Rabbi Mizrachi 00:00 - Opening Greetings 00:03 - Rabbi's Reputation and Controversy 00:32 - Noahide Laws and Gentile Righteousness 01:39 - Gentile Conversion Priorities 02:55 - Islam as a False Religion Debate 04:13 - Refuting Muhammad's Prophethood 04:58 - Islam Not God's Religion 05:56 - Preference for Righteous Muslims 06:39 - Secular Israeli Critique 07:00 - Religious Ignorance and Technical Issues 10:18 - Israel's Religious Decline and Messianic Hope 14:26 - Denial of Palestinian Existence 17:11 - Israel-Gaza Conflict and Casualties 23:31 - Racism, Heaven, and Divine Preference 25:41 - Covenant, Land Claims, and History 27:51 - Political Actions vs Moral Justifications 32:42 - Refusing Mediation with Qatar 33:42 - Allegations of Israeli Violence 35:47 - Netanyahu's Prisoner Release Decision 36:21 - Value of Lives and Security Policy 38:01 - Hezbollah Origin Debate 40:01 - Islamic Role in Jewish Return 40:49 - Free Speech and Cancel Culture 43:25 - Historical Accuracy Discussion 43:46 - Islam's Adoption of Christian Figures 48:42 - Critique of Christianity and Israel 50:08 - Christianity Conversion and Status 53:21 - Rabbinic Authority and Religious Law 1:00:46 - Abrahamic Covenant and Divine Choice 1:03:38 - Later Religions Declared False 1:04:25 - Public Revelation Argument 1:05:32 - Torah Authority and Rabbinic Guidance 1:07:14 - Penalty for Disobeying Rabbis 1:07:52 - Idol Worship and Execution Debate 1:08:21 - Discussion Limits and Censorship 1:09:10 - Invitation to Dialogue with Muslims 1:09:35 - Closing Remarks

SNEAKO

66,754 views • 2 months ago