Загрузка видео...

Не удалось загрузить видео

На главную

🚨 Shocking: Frontier LLMs score 85-95% on standard coding benchmarks. We gave them equivalent problems in languages they couldn't have memorized. They collapsed to 0-11%. Presenting EsoLang-Bench. Accepted to the Logical Reasoning and ICBINB workshops at ICLR 2026 🧵

1,265,900 просмотров • 6 месяцев назад •via X (Twitter)

Комментарии: 44

Фото профиля Lossfunk
Lossfunk6 месяцев назад

1/ Here's the intuition. When you learn Fibonacci in Python, you can write it in Java tomorrow without years of Java training. You transfer the logic. The loop, the state, the termination condition. Syntax is just a costume. LLMs claim to do this. We wanted to see if they actually can.

Фото профиля Lossfunk
Lossfunk6 месяцев назад

2/ Our method: test them on esoteric programming languages. Brainfuck. Befunge-98. Whitespace. Unlambda. Shakespeare. All Turing-complete. All requiring identical reasoning to Python. All with 1,000-100,000x fewer GitHub repos than mainstream languages. Same problems. Radically less training data.

Фото профиля Lossfunk
Lossfunk6 месяцев назад

3/ 80 problems across 4 difficulty tiers. Easy: sum two integers, reverse a string. Medium: Fibonacci, factorial. Hard: count primes, balanced parentheses. Extra-Hard: longest increasing subsequence, Josephus problem. Trivial in Python. A different story in esoteric languages.

Фото профиля Lossfunk
Lossfunk6 месяцев назад

4/ We tested GPT-5.2, O4-mini, Gemini 3 Pro, Qwen3-235B, and Kimi K2 across 5 prompting strategies. Models scoring 85-95% on HumanEval scored 0-11% on equivalent esoteric tasks. And every model, every language, every strategy scored 0% beyond the Easy tier. Not 2%. Not 5%. Zero.

Фото профиля Lossfunk
Lossfunk6 месяцев назад

5/ We threw everything at it to try to close the gap. Few-shot examples. Self-reflection. ReAct pipelines. Coder-critic pairs. Average improvement from few-shot: +0.8 percentage points. Statistically insignificant. ICL works by activating knowledge that already exists from pretraining. When that knowledge isn't there to begin with, a few examples in the context window can't substitute for it.

Фото профиля Lossfunk
Lossfunk6 месяцев назад

6/ The error profiles make the data coverage story concrete. Brainfuck and Befunge-98 have more online presence → models get syntax right but fail on logic. They understand the grammar, not the meaning. Unlambda and Shakespeare have almost none → 88-95% compile failures across every model. Models can't even produce valid syntax from scratch. Performance tracks data coverage remarkably cleanly.

Фото профиля Lossfunk
Lossfunk6 месяцев назад

7/ After the paper was finalized, we ran agentic systems that mimic how humans would learn to solve problems in esoteric languages. We supplied our agents with a custom harness + tools on the same benchmark. They absolutely crushed the benchmark. Stay tuned 👀

Фото профиля Lossfunk
Lossfunk6 месяцев назад

8/ This work was done by @inceptmyth under the supervision of @paraschopra at @lossfunk. If you work on evals or OOD generalization, we'd love to hear what you think. Pinging: @karpathy @fchollet @GaryMarcus @ylecun @AndrewYNg @demishassabis @drfeifei @goodfellow_ian @MelMitchell1 @percyliang @pmddomingos

Фото профиля Lossfunk
Lossfunk6 месяцев назад

@inceptmyth @paraschopra @karpathy @fchollet @GaryMarcus @ylecun @AndrewYNg @demishassabis @drfeifei @goodfellow_ian 9/ We're releasing everything: 🌐 Website: 📄 Paper: 🤗 Dataset: 💻 Code:

Фото профиля Bronson Schoen
Bronson Schoen6 месяцев назад

Do you have a human baseline? Solving problems in brainfuck is just empirically harder. I’m skeptical you’re actually testing memorization vs “esolangs made to be hard are hard”.

Фото профиля Mags
Mags6 месяцев назад

How do you expect anyone or anything to know something that wasn’t taught? This is nonsense.

Фото профиля Alex Nichol
Alex Nichol6 месяцев назад

Is o4-mini seriously the only reasoning model you tried? The word reasoning literally appears 44 times in your paper; seems like this should matter a *lot*.

Фото профиля Karan Handa
Karan Handa6 месяцев назад

Curious how it would perform if instructed to write a transpiler to these esoteric languages in a language that it's familiar with

Фото профиля Lossfunk
Lossfunk6 месяцев назад

After the paper was done, we tried modern agentic tools like claude code, gave them tools and instructed them to explore/learn We found it actually wrote something like this by itself (without instructing) Stay tuned for this update.

Фото профиля sdmat
sdmat6 месяцев назад

For this to be meaningful you need a human baseline with representatice median programmers unfamiliar with the languages. And I guarantee you they would not do well either as these languages suck by design. Theoretical Turing equivalence is irrelevant, LLMs aren't alien beings of pure logic. They work much more like humans, heavily leaning on pattern recognition and scaffolding. What would be much more interesting here is creating a de novo practical language that the median human programmers do well on and seeing how far models get with ICL (giving same reference docs for each).

Фото профиля The American Sun
The American Sun6 месяцев назад

@phl43 I bet they suck at writing haikus in Klingon too

Фото профиля Jenia Jitsev 🏳️‍🌈 🇺🇦 🇮🇱 🇮🇷
Jenia Jitsev 🏳️‍🌈 🇺🇦 🇮🇱 🇮🇷6 месяцев назад

I am afraid this amounts to sensationalist clickbait without proper background. Check for exotic code syntax comprehension in first place. Eg, after training on physics problems in english, not solving them in russian does not mean generalization deficits for physics.

Фото профиля Matt Arderne 🌊
Matt Arderne 🌊6 месяцев назад

should have included Excel

Фото профиля Paul Calcraft
Paul Calcraft6 месяцев назад

What reasoning level(s) did you try for GPT 5.2?

Фото профиля 7oponaut
7oponaut6 месяцев назад

this is unreasonable. humans don't generalize to brainfuck either

Фото профиля Sam
Sam6 месяцев назад

Do you test human baseline

Фото профиля Alex Rozinov
Alex Rozinov6 месяцев назад

Disappointed ArnoldC didn’t make the cut in EsoLang-Bench: GET TO THE CHOPPER

Фото профиля Szymon Teżewski
Szymon Teżewski6 месяцев назад

“Unseen language = collapse” is a neat headline, not a law of nature. I’ve tested models on Aver and NanoLang too. When the language is small, regular, and backed by a strong corrective loop, they do far better than this framing suggests. The real variable is not just memorization. It’s feedback.

Фото профиля shiv
shiv6 месяцев назад

Recently trained a small transformer (~472M tokens) with BPE, RoPE, GQA, etc., and now I’m exploring applying similar ideas to Indian classical music. Specifically looking at representing ragas as sequence data (starting with MIDI, possibly moving to audio later). Curious about whether transformers can actually capture deeper raga structure , not just note sequences, but progression, mood, and inherent constraints. do you think this is something transformers can learn with scale, or would it require a different modeling approach / inductive bias? Would love your perspective.

Фото профиля hnp
hnp6 месяцев назад

Bruh. Someone could publish a paper where the test dataset consists of tribal languages which LLMs would not have been trained on, claim credit, and put a 🚨 emoji in a post.

Фото профиля Mudit Srivastava
Mudit Srivastava6 месяцев назад

Really liked this from @paraschopra. Same logic, new surface form, and LLMs collapse. Tells you a ton about where fluency ends and reasoning begins. It's similar to why we're differently approaching reasoning at @pathway_com. Recent result:

Фото профиля Matt Shumer
Matt Shumer6 месяцев назад

😂

Фото профиля Arcani Venator
Arcani Venator6 месяцев назад

oh no, the llm has to take a few seconds and write up a transpiler anyway

Фото профиля Joseph Garvin
Joseph Garvin6 месяцев назад

It's too bad this will only work once, they'll just generate a ton of training data for esolangs if this result becomes too prominent. You'd need to generate esolangs.

Фото профиля James Miller
James Miller6 месяцев назад

Now test college students (1) based on problems they have seen before vs (2) problems that are basically the same but worded in a slightly different manner from what they have ever encounter before.

Фото профиля Yagao Dirac
Yagao Dirac6 месяцев назад

let me propose a test. Rename all the keywords of a mainstream, py or anything. Let's say, if into what if, else into otherwise, for into for them all, or something similar. What's the accuracy drop?

Фото профиля Somers
Somers6 месяцев назад

Whats the human baseline? Lots of tests could be created that are nonsensical

Фото профиля GOY SUPERSTAR
GOY SUPERSTAR6 месяцев назад

Yeah, because LLMs can't generalize Let the scam go on tho, maybe some of us peasants will get an opportunity to buy into it at some point, too

Фото профиля JS
JS6 месяцев назад

We called it 'intelligence' when it memorized training data. The real test was always: can it learn what it hasn't seen? Most models just failed that test.

Фото профиля LBM_LXXVIII
LBM_LXXVIII6 месяцев назад

Were you expecting that models wd nail languages they have basically never seen? What human being wd be able to master a programming language they barely saw before?

Фото профиля Sam Elliott
Sam Elliott6 месяцев назад

Beyond retarded. Comparing Python to Java and then use Brainfuck. Which looks like this

Фото профиля Vladyslav Hunt
Vladyslav Hunt6 месяцев назад

benchmarks were never reasoning tests, they were memorization tests with nicer charts

Фото профиля aizk ✡️
aizk ✡️6 месяцев назад

@theo probably up your alley

Фото профиля BeastTitanHunter
BeastTitanHunter6 месяцев назад

New Hill Identified Time to climb @gdb @roon @apples_jimmy

Фото профиля Christopher
Christopher6 месяцев назад

no shit, same with humans. nobody cares about languages that practically don't exist 😂

Фото профиля Ben Eng
Ben Eng6 месяцев назад

0% accuracy speaking a language that it does not know sounds like you have achieved AGI. 0% is what a human would achieve for that test.

Фото профиля sugnerhan
sugnerhan6 месяцев назад

ok well that's 100% better than me. eso is not real lmao

Фото профиля Marcin
Marcin6 месяцев назад

Isn't this like asking an English speaking person questions in Chinese? 👀 Of course they wouldn't answer correctly even if they are super intelligent. Shocking sure, but LLM are pattern matching engines, not intelligence engines

Фото профиля Midas 👑
Midas 👑6 месяцев назад

this is the non bait tweet that shows that harness are the next frontier

Похожие видео

Introducing ALE-Bench, ALE-Agent! Towards Automating Long-Horizon Algorithm Engineering for Hard Optimization Problems Blog: Paper: ALE-Bench is a coding benchmark primarily focused on hard optimization (NP-hard) problems. We developed this benchmark with AtCoder Inc., a leading coding contest platform company. What makes ALE-Bench unique is its focus on hard optimization problems that demand long-horizon and creative reasoning. It’s open-ended, in the sense that true optima are out of reach (NP-hard) and scores can continuously improve. We believe this benchmark has the potential to become one of the key benchmarks for reasoning and coding in the next generation. ALE-Agent is our end-to-end agent that we specifically designed for this challenging domain. In fact, our ALE-Agent has already built an impressive track record in the wild! In May 2025, our agent participated in a live AtCoder Heuristic Competition (AHC), alongside 1,000 other participants in real-time. AHC is considered to be one of the most challenging coding competitions in this domain. Our ALE-Agent achieved an impressive ranking of 21st out of 1,000 human participants in the competition (top 2%), marking a turning point for AI discovery of solutions to hard optimization problems with a wide spectrum of important real world applications such as logistics, routing, packing, factory production planning, power-grid balancing. We look forward to applying this technology to real industrial optimization opportunities. Building on the insights from this study, Sakana AI will continue to tackle the challenge of developing AI with even greater algorithm engineering capabilities. ALE-Bench Dataset: ALE-Bench Code: This research was conducted in collaboration with AtCoder Inc. (AtCoder). We are deeply grateful for their outstanding expertise and contributions in optimization and algorithms, which were invaluable in providing data, analyzing results, and enabling our AI agent’s participation in their contests.

Sakana AI

237,195 просмотров • 1 год назад