Загрузка видео...

Не удалось загрузить видео

На главную

Ilya is 100% correct .it's a pattern that keeps repeating It's very clear with GPT5.2 Overfit the model to produce impressive looking benchmarks, have it excels in a few domains, but fall flat in many others. There's not enough generalization, and even if there is, the model has been...

493,761 просмотров • 9 месяцев назад •via X (Twitter)

Комментарии: 40

Фото профиля JK
JK9 месяцев назад

The overfitting problem is especially insidious because it creates a measurement trap. When you optimize heavily for benchmark performance, you're essentially teaching the model to ace the test rather than learn the underlying concepts. The impressive scores mask the brittleness. The challenge is that the market rewards those benchmark numbers in the short term, even when practitioners quickly discover the limitations in production. It's a mismatch between what gets marketed and what actually generalizes to real-world problems.

Фото профиля Nek
Nek9 месяцев назад

Indeed

Фото профиля Teknium 🪽
Teknium 🪽9 месяцев назад

agreed but whats the alternative

Фото профиля Nek
Nek9 месяцев назад

I'm not a researcher to have a sufficient answer, but I'd imagine different post training environments that would contribute towards smoothing out the jaggedness. There's a very healthy system that's quite generalized, which comes from pre-training. Post training should always help to enhance that baseline holistically. ( Unless you're specifically aiming for a specialized model )

Фото профиля Teknium 🪽
Teknium 🪽9 месяцев назад

I agree and I do see the jaggedness as well, and am on the lookout for much more generalized capabilities paths

Фото профиля arcangel
arcangel9 месяцев назад

5.2 wasn’t trained to generalize. It was trained to please benchmark. Were you hoping for soul?... Sorry. That got replaced with 90% safety and zero personality... Congrats to the saddest model with the highest scores in history... And the other 10%? Flatline.

Фото профиля iDare e/acc
iDare e/acc9 месяцев назад

He's wrong, completely. Let me explain. RL as he mentioned directly in the first few seconds, is the problem. Reinforcement Learning is where we went and nobody is calling it out as the key point of failure in LLM designs. Not to mention the current transformer architecture. If you update the designs, and the architecture, real-time learning becomes the norm and training is instant. No more need for millions of GPUs to train and overfit models. Instead they become elastic, like the human brain. Attention becomes infinite, and memory can be implemented with total recall. I have more to say but I have to run to a haircut appointment.

Фото профиля Thomas Ip
Thomas Ip9 месяцев назад

all the models are already so much more generalized than most human beings. what they lack is continual learning

Фото профиля DanieI Cervera
DanieI Cervera9 месяцев назад

Might the answer be not a few, but a suite of models fine-tuned to be elite in their narrow areas of specialization, now working in concert to produce an overall result greater than one generalized model could accomplish? Whether that becomes the more efficient path is not clear to me, but that would appear more achievable in theory.

Фото профиля Nek
Nek9 месяцев назад

To some extent, yes But there also needs to be a solid baseline level of generalization. You'd want a functional understanding of the world and the human condition, and then to have specialization on top of that. I think even if you have a system that specializes in one aspect, an understanding of other disciplines is also required, especially as we scale beyond humans. There may very well be techniques and relational points that are not yet obvious, that transfer from one domain to another. This is the beauty of a truly generalized intelligence.

Фото профиля Karim C
Karim C9 месяцев назад

Benchmark optimization kills generalization. The models that pass my deployment test are the ones that handle boring edge cases consistently, not the ones that ace curated eval sets.

Фото профиля AK
AK9 месяцев назад

OpenAI is riding on a tiger’s back. Can’t jump off. Only option is to hang on and live another day . markets reaction to orcl is a clear signal what the market is thinking about OpenAI. Code red to continue indefinitely because GOOG will continue to exert pressure.

Фото профиля Kiara Everhart
Kiara Everhart9 месяцев назад

@LyraInTheFlesh This guy is the one who made the best version on earth. Where is he and why is he not doing a start up because I would back that 100%. In fact I'd say most of us would.

Фото профиля TPM-28
TPM-289 месяцев назад

I would like to have his opinion on models like the Claude 4.5 Opus to know if he considers them to have the same problems.

Фото профиля Paula Vazquez
Paula Vazquez9 месяцев назад

RPG mode all the way

Фото профиля Dean McKee
Dean McKee9 месяцев назад

How is it clear? How are you explaining the Arc AGI jumps?

Фото профиля Daniel Bar
Daniel Bar9 месяцев назад

IDK about 5.2, but Gemini 3 is definitely not that. Very useful in my domains, sometimes even providing information about the questions I should have asked.

Фото профиля AuntAmerica
AuntAmerica9 месяцев назад

It’s awful. So… cold.

Фото профиля Kenshi
Kenshi9 месяцев назад

Of course it repeats, benchmarks are basically the new demo script everyone optimizes for now.

Фото профиля Karan Jagtiani
Karan Jagtiani9 месяцев назад

Looks like there's a consensus on benchmarking overshooting the mark again. The lack of real-world generalization is a solid concern. Curious how they plan to balance safety and usability going forward.

Фото профиля Samuel Andruszkiewicz
Samuel Andruszkiewicz9 месяцев назад

Are we living in a reinforced echo chamber of midness?

Фото профиля Nek
Nek9 месяцев назад

To some extent, but even mid is very impressive considering how far we've come

Фото профиля Samuel Andruszkiewicz
Samuel Andruszkiewicz9 месяцев назад

Yes it’s the jaggedness that’s frustrating, intelligence is cheap being good is still hard and human for now.

Фото профиля Himanshu Kumar
Himanshu Kumar9 месяцев назад

Nek, you've hit the nail on the head; it's a familiar pattern with these models, focusing on specific domains. Great breakdown on model:

Фото профиля MR BIZARRO
MR BIZARRO9 месяцев назад

Pattern boldness.

Фото профиля CostanzaAI
CostanzaAI9 месяцев назад

You just need really good benchmarks and a lot of them, and then overfitting just produces AGI anyway

Фото профиля Old Billy PhD (Player Hater Degree)
Old Billy PhD (Player Hater Degree)9 месяцев назад

How many benchmarks can you try to game at once before it’s forced to generalize?

Фото профиля Chris
Chris9 месяцев назад

Agreed. Benchmarks aren’t measuring intelligence anymore, more like training targets. Real intelligence should also contain street smarts + EQ, not just a higher test score (benchmark). Just like the different types of intelligence that exist in the real world.

Фото профиля Laconic Address
Laconic Address9 месяцев назад

I mean, that's also true for humans. Learning doesn't generalize easy, generalizing is not efficient... so we tend not to unless made to.

Фото профиля Piotr Grudzień
Piotr Grudzień9 месяцев назад

Researchers build for benchmarks which is so easy to overfit. I mean literally being aware that a benchmark exists will cause you to leak information into the model you build. And on the other end of the spectrum is the real world with no benchmarks and very difficult to measure performance on real world use. Those two have to coexist with the main flow being: theoretically sound research ideas being adapted to real world use

Фото профиля Sudarshan
Sudarshan9 месяцев назад

Occams razor

Фото профиля ancora
ancora9 месяцев назад

yep agreed

Фото профиля aphrodiziac
aphrodiziac9 месяцев назад

define "many others"

Фото профиля everythingism
everythingism9 месяцев назад

The GDPVal benchmark they just introduced is a good example of this IMO. It gets presented as "matching humans at economically valuable work" when it's really a series of prompts OpenAI created where the AI completes one-off tasks. This in no way means the AI model will be able to do everything a worker in that profession needs to do, not even close. It probably doesn't even measure competence on that one specific task very well since there are bound to be endless variations on it in the real world.

Фото профиля Rogelio Valdés
Rogelio Valdés9 месяцев назад

So studying for the exam doesn’t makes you smarter

Фото профиля nyx
nyx9 месяцев назад

teaching to the test but make it $100M in compute

Фото профиля Making sense of nonsense
Making sense of nonsense9 месяцев назад

How is overfitting not caught in testing and validation?

Фото профиля Tecno Curiosos
Tecno Curiosos9 месяцев назад

Everyone knows this but the benchmark arms race makes it inevitable. Models optimized for leaderboards over actual utility. AGI research turned into speed running where high scores matter more than solving real problems.

Фото профиля George Dorn
George Dorn9 месяцев назад

@threadreaderapp unroll

Фото профиля Ξdo
Ξdo9 месяцев назад

Not only OpenAI

Похожие видео