Загрузка видео...

Не удалось загрузить видео

На главную

why do language models think 9.11 > 9.9? at @transluceAI we stumbled upon a surprisingly simple explanation - and a bugfix that doesn't use any re-training or prompting. turns out, it's about months, dates, September 11th, and... the Bible?

375,507 просмотров • 1 год назад •via X (Twitter)

Комментарии: 12

Фото профиля Kevin Meng
Kevin Meng1 год назад

even frontier models like Claude-3.5-Sonnet and GPT-4o are puzzlingly bad at comparing numbers. 9.11 is less than 9.9(0), but they seem to think the opposite.

Фото профиля Kevin Meng
Kevin Meng1 год назад

(the new Claude-3.5-Sonnet is good at comparing numbers, but it's still not very good at these sorting problems)

Фото профиля Kevin Meng
Kevin Meng1 год назад

do models *really* not know this simple stuff? or is there something else going on? the community offered guesses like "it's because of software versioning" or "dates," but re-prompting didn't fix the problem. what a weird mystery. is there really nothing we can do about it?

Фото профиля Kevin Meng
Kevin Meng1 год назад

at @transluceAI, we've been working on Monitor, an interpretability interface that reveals the internal computations in language models and lets you control them. i was looking for use-cases to test our interface on when i thought to run the (infamous) 9.9 < 9.11 example.

Фото профиля Kevin Meng
Kevin Meng1 год назад

i typed the query in, noticed the incorrect "bigger" token, and ran attribution to find neurons influencing that mistake. we found concepts related to: - the september 11th attacks - biblical verses - dates and months okay, dates, sure, but bible verses? 9/11?

Фото профиля Kevin Meng
Kevin Meng1 год назад

i was curious about the biblical verses, so i clicked on that cluster and looked at where those neurons fire highly. turns out they really like verse numbers. notice the numbering system - Matthew 8:5 comes before 8:13. i'd have never guessed this!

Фото профиля Kevin Meng
Kevin Meng1 год назад

i tried prompting the model not to think of these things, but that didn't help. i guess models aren't very good at controlling what they think

Фото профиля Kevin Meng
Kevin Meng1 год назад

so why don't we try direct neuron interventions? we can directly stop the model from interpreting these numbers as dates or biblical verses, by setting those neuron activations to 0. and that works! zeroing out september 11th attack neurons also works.

Фото профиля Kevin Meng
Kevin Meng1 год назад

the dates + bible verses intervention also works reliably for the sort problem above that Claude struggles with ( but i couldn't get the trickier "sort these numbers backwards" case below. if anyone can figure it out, i'd love to know!

Фото профиля Kevin Meng
Kevin Meng1 год назад

we also ran some simple evals: llama-3.1 8b instruct is only 54% accurate on our test set; about random guessing. just by steering out bible verse neurons, we can get up to 76%. no further training, and no prompting. just getting *rid* of things!

Фото профиля Kevin Meng
Kevin Meng1 год назад

i wonder what else our models are "capable" of, that we think they can't do. perhaps y'all can find some examples - try out the interface if you'd like! writeup: interface demo:

Фото профиля Kevin Meng
Kevin Meng1 год назад

i am indeed excited to see further progress in this direction. :)

Похожие видео

David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” david friedberg: “I have had experiences where we've asked some fairly novel scientific questions, and (the AI model) identifies it as a novel insight. It's like, ‘Oh, never thought about that, interesting, blah, blah, blah.’ And then using a different account, asking the next version (of the model) later, I've now experienced this. It's like, ‘Oh, well, you could do this,’ and it actually just describes this exact thing that we had in our chat in the previous version. Now, these are a handful of anecdotal experiences, but I know the domain that we work in, and the niche of it, and the ideation of this stuff, and the novelty of this stuff, and the lack of papers being published, and so on. So I know that there isn't some new corpus of information out there that's training the new model. So all I can say at that point is that my conversation or our analyses have been used for training.” David Sacks: “Okay, this does raise a really good question. What does it mean that the model is allowed to train on unidentifiable data?” Friedberg: “Well, that's my point. So it doesn't use any of my personal information, but it can use an insight derived from our chat, which it can then say is some training data that is unrelated. But the truth is, it's actually a piece of IP that's our organization’s IP, and our engagement back and forth. We don't have any NDA or confidentiality provisions or protections with them being a service provider back to us. This is why I care a lot about open source because I don't want them having my chat logs because they can use it for training to create an IP advantage that is now diffused to the rest of the market.”

The All-In Podcast

56,501 просмотров • 20 дней назад