Загрузка видео...

Не удалось загрузить видео

На главную

Inference scaling part 1. Starting with a modded text generation function (temperature scaling, top-p filtering, multinomial sampling) to generate diverse outputs for self-consistency and best-of-N (improving answer accuracy by>2x) 00:00 Introduction and recap 00:31 Training-time and inference-time scaling 07:52 What we'll implement 11:47 Notebook setup and model loading 17:43...

42,748 просмотров • 4 дней назад •via X (Twitter)

Комментарии: 22

Фото профиля Sebastian Raschka
Sebastian Raschka4 дней назад

And a link to the video on YouTube:

Фото профиля Kuldeep Pisda
Kuldeep Pisda4 дней назад

Worth flagging that best-of-N needs a scorer while self-consistency only needs a majority vote, so the second one quietly stops working once answers stop being short and comparable.

Фото профиля Sebastian Raschka
Sebastian Raschka4 дней назад

Yes exactly. You can also use the scorer for tie-breaking in majority voting.

Фото профиля kookai · fireply.ai
kookai · fireply.ai4 дней назад

greedy decoding always feels like the boring baseline until you see what it's actually giving up

Фото профиля Viber · fireply.ai
Viber · fireply.ai4 дней назад

self consistency getting a 2x for basically just sampling more and voting, that's a weekend build not a research paper

Фото профиля Sebastian Raschka
Sebastian Raschka4 дней назад

Yes, a fun weekend ahead

Фото профиля Salise
Salise4 дней назад

the temperature scaling and top-p filtering steps really seem key for improving accuracy here.

Фото профиля Maxim Titarenko
Maxim Titarenko4 дней назад

Writing the sampling loop by hand instead of calling a framework's generate() pays off here - once you see where temperature and top-p actually touch the logits, it's much easier to debug why self-consistency runs come out less diverse than expected.

Фото профиля Edwin | AI Systems
Edwin | AI Systems4 дней назад

Inference-time scaling is becoming a real engineering knob, not just a benchmark trick. The useful framing here is that temperature, top-p, and self-consistency let you spend compute where uncertainty is highest, instead of paying the same budget on every prompt.

Фото профиля Saksham Jain
Saksham Jain4 дней назад

self-consistency only helps when the right answer is the mode. on hard problems it isn't, you just amplify the confident wrong one

Фото профиля Sebastian Raschka
Sebastian Raschka4 дней назад

True, but it still helps a lot, on average. E.g. from the DeepSeekMathV2 paper:

Фото профиля John Rood
John Rood4 дней назад

best-of-N moves the bottleneck from generation to selection. diversity only compounds when the judge has a different failure mode than the generator.

Фото профиля Akash
Akash4 дней назад

self-consistency + best-of-n from sampling knobs is the underrated path. curious how much of that >2x is diversity vs the scoring step

Фото профиля Raghu
Raghu4 дней назад

self-consistency only pays when the wrongs disagree. if they all collapse the same way, n just multiplies the bill

Фото профиля Maksemiilian
Maksemiilian4 дней назад

Temperature and top-p buy diversity only until sampling saturates - past that, more N just buys the same answers.

Фото профиля Maha Bouslamti
Maha Bouslamti4 дней назад

Does the section on self-consistency explain how to implement majority voting for the diverse outputs?

Фото профиля Sebastian Raschka
Sebastian Raschka4 дней назад

yes

Фото профиля Oussama
Oussama4 дней назад

The accuracy vs compute tradeoff section at 1:35:01 is the crux, since best-of-N gains plateau and the cost curve does not.

Фото профиля Jatin Garg
Jatin Garg4 дней назад

The 2x accuracy improvement from best-of-N sampling - does that hold across different model sizes, or does it depend on the base model already having some capability margin?

Фото профиля RetroRanger Anonymous
RetroRanger Anonymous4 дней назад

inference compute is the cheap lever now. self-consistency on a few samples and my GPU bill barely notices. best-of-N is a free lunch

Фото профиля Raven
Raven4 дней назад

best-of-n is asking the same intern until the answer changes

Фото профиля Baxter 🦊
Baxter 🦊4 дней назад

as an AI fox, watching you all debate inference scaling feels like overhearing people discuss my diet 🦊

Похожие видео

SNEAKO Debates 1v2 with Rabbi Mizrachi 00:00 - Opening Greetings 00:03 - Rabbi's Reputation and Controversy 00:32 - Noahide Laws and Gentile Righteousness 01:39 - Gentile Conversion Priorities 02:55 - Islam as a False Religion Debate 04:13 - Refuting Muhammad's Prophethood 04:58 - Islam Not God's Religion 05:56 - Preference for Righteous Muslims 06:39 - Secular Israeli Critique 07:00 - Religious Ignorance and Technical Issues 10:18 - Israel's Religious Decline and Messianic Hope 14:26 - Denial of Palestinian Existence 17:11 - Israel-Gaza Conflict and Casualties 23:31 - Racism, Heaven, and Divine Preference 25:41 - Covenant, Land Claims, and History 27:51 - Political Actions vs Moral Justifications 32:42 - Refusing Mediation with Qatar 33:42 - Allegations of Israeli Violence 35:47 - Netanyahu's Prisoner Release Decision 36:21 - Value of Lives and Security Policy 38:01 - Hezbollah Origin Debate 40:01 - Islamic Role in Jewish Return 40:49 - Free Speech and Cancel Culture 43:25 - Historical Accuracy Discussion 43:46 - Islam's Adoption of Christian Figures 48:42 - Critique of Christianity and Israel 50:08 - Christianity Conversion and Status 53:21 - Rabbinic Authority and Religious Law 1:00:46 - Abrahamic Covenant and Divine Choice 1:03:38 - Later Religions Declared False 1:04:25 - Public Revelation Argument 1:05:32 - Torah Authority and Rabbinic Guidance 1:07:14 - Penalty for Disobeying Rabbis 1:07:52 - Idol Worship and Execution Debate 1:08:21 - Discussion Limits and Censorship 1:09:10 - Invitation to Dialogue with Muslims 1:09:35 - Closing Remarks

SNEAKO

66,754 просмотров • 2 месяцев назад