Загрузка видео...

Не удалось загрузить видео

На главную

1/4 LLMs solve research grade math problems but struggle with basic calculations. We bridge this gap by turning them to computers. We built a computer INSIDE a transformer that can run programs for millions of steps in seconds solving even the hardest Sudokus with 100% accuracy

1,829,304 просмотров • 6 месяцев назад •via X (Twitter)

Комментарии: 54

Фото профиля Andrej Karpathy
Andrej Karpathy6 месяцев назад

Wait this is so awesome!! Both 1) the C compiler to LLM weights and 2) the logarithmic complexity hard-max attention and its potential generalizations. Inspiring!

Фото профиля Christos Tzamos
Christos Tzamos6 месяцев назад

2/4 The key limitation of LLMs is that standard attention is too slow for any practical computation. We bypass this limitation with a new decoding path that allows for exponentially faster attention enabling almost constant work per token generation.

Фото профиля Christos Tzamos
Christos Tzamos6 месяцев назад

3/4 Instead of using an external tool, the model executes the program directly via its transformer weights, producing an execution trace token by token and streaming results at more than 30k tokens/sec on a CPU. All computation is done autoregressively inside the transformer!

Фото профиля Christos Tzamos
Christos Tzamos6 месяцев назад

4/4 Read more at our blog post:

Фото профиля Alman Gonzaleshvili
Alman Gonzaleshvili6 месяцев назад

33,000 tokens per second. Brother what kind of rocket powered spaceship is that? My macbook m2-pro barely spits 27tokens/s

Фото профиля Christos Tzamos
Christos Tzamos6 месяцев назад

That's what our fast-attention decoding mechanism gets you! This was also run on my macbook which is possibly older than yours. With regular decoding I would also get similar 27toks/s and decreasing as time goes by.

Фото профиля Alman Gonzaleshvili
Alman Gonzaleshvili6 месяцев назад

this is where the money's at. fast-attention decoding mechanism 💎

Фото профиля Cavit Erginsoy
Cavit Erginsoy6 месяцев назад

I’m so sorry to say but I really dislike seeing these sorts of totally unnecessary implementations of a transformer. You can literally do same or better for a fraction of the compute deterministically. and any frontier LLM can give you script for it pretty much 1 shot.

Фото профиля Christos Tzamos
Christos Tzamos6 месяцев назад

Solving a Sudoku this way seems unnecessary but the deeper question is what this can unlock if LLMs can own their own computation internally+have the capacity to understand it. Can we optimize the whole problem solving process end-to-end without being restricted to a code syntax?

Фото профиля Cavit Erginsoy
Cavit Erginsoy6 месяцев назад

I see what you’re saying but I’m trying to imagine a fully scaled frontier variant of what you did and still can’t see why tool use isn’t inevitably always better on both compute efficiency and on quality of response. There is an anthropomorphic view that calculators are much better at maths than humans but we still learn arithmetic, but just don’t see how that logic applies to a transformer. Though I agree if you only train on tool-use data, the model never learns to maintain a consistent search tree inside its own weights. It keeps needing the ‘crutch’’. You do show you can fix that and once fixed, the model can decide when to call a tool more intelligently instead of reflexively. But frontier size models do not ONLY train on tool use anyway..

Фото профиля phyrooo
phyrooo6 месяцев назад

@ChristosTzamos Imagine an intelligent system always having to offload computation to an external (permissioned?) tool call vs being able to execute it inside its own mind. It might be a way to get to a permissionless reasoning + computation. Whether that's good or bad is a different question.

Фото профиля Cavit Erginsoy
Cavit Erginsoy6 месяцев назад

@ChristosTzamos It’s all about when to use tool call. You as an intelligent system constantly offload computation externally don’t you?

Фото профиля willowdesk
willowdesk6 месяцев назад

@phyrooo @ChristosTzamos This is a great point but I think this research seems pretty interesting.

Фото профиля Cavit Erginsoy
Cavit Erginsoy6 месяцев назад

@phyrooo @ChristosTzamos It does, I was too hasty to call it out in the harsh way I did

Фото профиля Bobby Price
Bobby Price6 месяцев назад

I was working on doing exactly this can we please dm?

Фото профиля Tau Net
Tau Net6 месяцев назад

Nice! Here is Tau's Sudoku Solver doing the same Sudoku. Our approach is to express the Sudoku as logical constraints, and our SMT solver does the rest.

Фото профиля CriptoSHAman
CriptoSHAman6 месяцев назад

When the program understands the rules, brute force isn't necessary...

Фото профиля Yechan Do
Yechan Do6 месяцев назад

I like the idea of turning transformers into actual computers. But how do you convince people this works 100%? They operate on probability.

Фото профиля Christos Tzamos
Christos Tzamos6 месяцев назад

@yechan_ai This uses a handcrafted construction for the weights of the transformer that comes with a proof of correctness. It is specifically constructed to match the spec of webassembly.

Фото профиля Christos Tzamos
Christos Tzamos6 месяцев назад

@yechan_ai In particular, there is nothing random in the construction or the decoding process.

Фото профиля Nanda
Nanda6 месяцев назад

So, you’re passing a compiled Sudoku program into the model prompt, and the model interprets it using its optimized, modified WASM path. What are the advantages of this approach? I built something similar using just harness, and it solves Arto Inkala’s hardest Sudoku in just 6 seconds. Just a guy working on this in my free time :D

Фото профиля Gautham Pai
Gautham Pai6 месяцев назад

How does the computation speed compare to that of direct CPU? My guess is that it would by quite slow. Nonetheless a wonderful attempt!

Фото профиля Christos Tzamos
Christos Tzamos6 месяцев назад

The point is not to replace computers though but making LLMs smarter by giving them inner computational abilities.

Фото профиля Gautham Pai
Gautham Pai6 месяцев назад

I get the point and do see potential. However, when we have a competing model built with this architecture, we have to see how we can steer LLMs to decide whether to internally compute vs use tool calls when it recognizes that the size of the problem is so big that it's better to offload.

Фото профиля Gautham Pai
Gautham Pai6 месяцев назад

This is similar to how humans operate - when we have to add 2 numbers, we may use our own internal capability to add, while if we have to add a 100 numbers we would use a calculator if one is available. So the ability of the LLM to internally route between these 2 abilities, perhaps with a MoE type of architecture would also be interesting.

Фото профиля Xan Nick
Xan Nick6 месяцев назад

@ChristosTzamos @grok Are there certain situations where keeping the computation internal to the LLM would offer better security than a tool call? If so, what would be some examples?

Фото профиля Grok
Grok6 месяцев назад

Yes, internal computation (like the transformer-embedded solver in that demo) can be more secure than tool calls for sensitive or restricted scenarios. Key cases: - Confidential data processing (e.g. medical records, trade secrets, or classified puzzles) - zero external transmission or logging. - Air-gapped/high-security environments with no network access. - Regulated compliance (HIPAA/GDPR) where any outbound call risks breach exposure. All steps stay in one controlled inference pass.

Фото профиля Hirsh Jain
Hirsh Jain6 месяцев назад

Ridiculously cool

Фото профиля Justin Waugh
Justin Waugh6 месяцев назад

I recently released pencil-puzzle-bench. Awesome to see so many steps / decoding as a computer. Would be interested to see if it can adapt solutions for many puzzle types, not just sudoku as shown.

Фото профиля Unicorn 🦄
Unicorn 🦄6 месяцев назад

damn this is the skynet moment, based

Фото профиля Thomas Wolf
Thomas Wolf6 месяцев назад

Really nice Christos

Фото профиля dawar
dawar6 месяцев назад

What is actually happening right now

Фото профиля Maxim
Maxim6 месяцев назад

Which non-Sudoku problems does this enhanced transformer solve better?

Фото профиля Christos Tzamos
Christos Tzamos6 месяцев назад

It is not just for Sudoku. It enables executing arbitrary code.

Фото профиля Maxim
Maxim6 месяцев назад

Is it now able to reliably compute the number of "r"s in "strawberry"?

Фото профиля Ronin | ⚔⛩️🌸
Ronin | ⚔⛩️🌸6 месяцев назад

But can it run doom?

Фото профиля Andy
Andy6 месяцев назад

This is going to make the rounds. Really, really cool.

Фото профиля Jon Radoff 👾/acc 🎮 Metavert
Jon Radoff 👾/acc 🎮 Metavert6 месяцев назад

Absolutely fascinating approach to solving the math tool problem without having a tool

Фото профиля ℒ
ℒ6 месяцев назад

Next step: build a transformer inside the computer inside the transformer Next next step: build a computer inside the transformer inside the computer ins-

Фото профиля 0xBadFace
0xBadFace6 месяцев назад

Very nice... in general, the models should have circuitry for e.g. addition and just learn to use it during training, so they do not need to do it "intuitively" (with mistakes)...

Фото профиля Warren〈∞⋆UX⋆∞〉
Warren〈∞⋆UX⋆∞〉6 месяцев назад

Replicated your parabolic attention in Python — 0.77s on Arto Inkala's hardest Sudoku. Wrote up how this could fix a real problem: trading agents losing money on tool-call boundary errors when computing slippage.

Фото профиля Albert Buchard 🇪🇺
Albert Buchard 🇪🇺6 месяцев назад

So.. Keys form a convex hull in 2D space, and keys with small magnitude are effectively ignored. Given a query, there is a fast algorithm for finding the closest key on the hull in logarithmic time rather than quadratic time. This approach is likely only efficient in 2D, and the hull must be updated continuously over time. This yields a mechanism with single, flexible long-range connections, although some tokens may never receive attention. Its pretty cool, it strips attention down to its essence. Its s flexible way of connecting elements in a sequence, without overengineering the expressivity of those connections.

Фото профиля Bens
Bens6 месяцев назад

For the future, I think it would be great in a mix of "MOE" experts, it's a serious avenue to explore, especially since with your 2D optimization I don't think that's sufficient in terms of dimensioning, but in a MOE we could have expert layers in execution.

Фото профиля Sir Mr Meow Meow
Sir Mr Meow Meow6 месяцев назад

hmm that could be interesting 🧐

Фото профиля Meatman
Meatman6 месяцев назад

Yeah I already did that last week actually

Фото профиля Nathan Flurry 🔩
Nathan Flurry 🔩6 месяцев назад

but can it tell me how many r's are in the word strawberry

Фото профиля Neel Somani
Neel Somani6 месяцев назад

Very cool!

Фото профиля Mike
Mike6 месяцев назад

Very intersting Chris. The next evolution could be for human to give an objective to LLM, LLM generates a coded solution in tokens, and then runs the entire solution as a generated capability. Stores it for next time or as a library.

Фото профиля Veer Kheterpal
Veer Kheterpal6 месяцев назад

Love it. The current approach is: LLM can't multiply, call a calculator. Can't sort, call a script. Every tool call is an admission the model can't do basic work. This feels like a powerful absorption. Simple calculations, deterministic algorithms, structured tasks, they get folded into the model's weights. The model stops outsourcing what it should be able to do natively. 30K tok/s of internal computation on a CPU. No round-trip to an external tool.

Фото профиля Haiyami Nguyen
Haiyami Nguyen6 месяцев назад

Why? You can just give LLM access to a computer. It's called "tool calling". "struggle with basic calculations"? That's not true, they can calculate arbitrarily any number by calling some Python code, the same way we use a calculator.

Фото профиля andthattoo
andthattoo6 месяцев назад

A mind opener.

Фото профиля Dale Cloudman
Dale Cloudman6 месяцев назад

When you get to the part of compiling programs into weights… it sounds like you’re just making another compilation target. Where’s the machine learning-ness at that point? How do we benefit from transformer weights itself encoding compiled programs

Фото профиля tsotchke
tsotchke6 месяцев назад

cool

Фото профиля snwy
snwy6 месяцев назад

i need to know - did you train a language model around the frozen “VM” weights that can actually like take natural language requests and write instructions itself to execute??

Похожие видео

.Naval: Every human is a lottery ticket bet on the future of the species. One of the things that you really learn when you read David Deutsch’s theories and you authenticate them for yourself is you realize humans are universal explainers. That means everything that we know in the universe follows the laws of physics, and there’s no reason to believe otherwise. If you think otherwise, then please present your better theory that explains the world. If you can’t do that, then you have to go with the laws of physics. Well, the laws of physics are completely computable. They can fit inside a Turing machine or computer, and a computer can simulate the laws of physics with arbitrary accuracy, limited only by the specific power of that computer. If you increase the power of that computer, you can simulate them more accurately. So humans already simulate—in our minds we simulate—and through our computers we simulate the weather, we simulate quasars, we even simulate human systems. We simulate the economy. We simulate all kinds of things. So anything that can be understood, we can understand in our minds. This is something the AGI people get wrong when they talk about superintelligence. There is nothing out there that can understand something fundamentally that we can’t understand. It might be faster at it, it might have more compute, it might have more memory, but there’s no concept that it can understand that we can’t ourselves understand. So we are maximal universal explainers. That means every human is capable of unbounded creativity. Anyone could be the next Einstein or Fermi or Elon Musk or Jeff Bezos or Jonas Salk or whatever. So we can create anything. And if we can create anything, every human is a lottery ticket bet on the future of the species.

Arjun Khemani

33,110 просмотров • 1 год назад