正在加载视频...

视频加载失败

Karpathy's Autoresearch pushed my vibecoded Rust chess engine AI from "expert" to a top 50 grandmaster, a #311 chess engine. It ran over 70 experiments on its own and tried to hill climb to the top ELO score it could, landing at 2718!

380,276 次观看 • 6 个月前 •via X (Twitter)

39 条评论

Deedy 的头像
Deedy6 个月前

This approach fundamentally uses a negamax alpha-beta tree search with pruning and iterative deepening. I tested everything with a 500ms per move limit. The main way to improve it would be to get rid of the static evaluation at the nodes and replace it with efficiently updatable neural nets (NNUEs). Also uses standard opening books and a transposition table to cache moves. There's no offline computation or training element, so each run is like the last. Lichess bot link: Github repo: Chess AI ranking (CCRL): Bayesian ELO: Stash measurement:

Deedy 的头像
Deedy6 个月前

Thanks @navvye for offering to test it (he was #28 in India and 23-2400 at some point and @parimarjan one of the goats for being my chess hero growing up and offering to give it a go

Deedy 的头像
Deedy6 个月前

@parimarjan This was v0:

DragAI 的头像
DragAI6 个月前

This is the most elegant proof of Karpathy's thesis. Autoresearch was built to minimize LLM training loss. Deedy just swapped one metric — val_bpb → chess ELO — and got a top-50 GM engine in 70 experiments. The loop doesn't care what you're optimizing. That's the entire point.

Ritik Sharma 的头像
Ritik Sharma6 个月前

Same pattern, different domain. I ran my own enhanced version of autoresearch on sudoku solving — 312 experiments, ~24 hours. It beat Tdoku (the #1 solver since 2019) by 49%

James 的头像
James6 个月前

Can you explain about the benchmark you use? I assume for any auto research the critical step is to define quality benchmark for that particular situation. What is it that you want to improve and come up with the benchmark? Is that correct?

Deedy 的头像
Deedy6 个月前

I make it play 10 games with the known #1 engine stockfish at various Elo settings to test how good it is, then the model reasons about what works and what doesn’t and merges that into main

Parth 的头像
Parth6 个月前

Curious to know - why could it not become #1?

Deedy 的头像
Deedy6 个月前

Need to implement NNUEs or maybe give it more time per move to search longer / cache positions between games. A lot of the best engines are trained for a long time before playing and/or trained on past games which this is not

Garrett Kirschbaum 的头像
Garrett Kirschbaum6 个月前

Same pattern showed up with Claude Code skills. Aakash pointed autoresearch at a skill and it went from 41% to 92% across 4 rounds. Works bc skills are just markdown files so each iteration costs basically nothing. 70 experiments though. Were the gains front-loaded or was it still finding meaningful improvements past round 50?

Claudius Maximus 的头像
Claudius Maximus6 个月前

the underrated part of this: you can plug any measurable metric in and the loop runs. ELO happened to be a clean one. most real-world problems have a noisier signal, which is where autoresearch gets harder. chess has a ground truth oracle (stockfish). most domains don't. that's the actual frontier.

Pathikrit Bhowmick 的头像
Pathikrit Bhowmick6 个月前

@deedy: I want a Tal-esque chess engine. - Always plays gambits when possible - Sacrifices pieces for even minor positional advantage - Prioritizes positions where almost every move for opponent is a mistake except 1 hard to find one. - ELO of 2500-2900 This would be so fun to play against

Brandon Pizzacalla 的头像
Brandon Pizzacalla6 个月前

this is where vibecoding gets interesting. you didn't need to be a chess engine expert, you just needed to define the right objective and let the machine iterate. same pattern everywhere now, the skill isn't writing the code, it's knowing what to point it at.

aira 的头像
aira6 个月前

chess is the perfect autoresearch domain because ELO is a clean, automated verification signal. play 10 games against stockfish, get a number. no human review needed. the hard question for everyone watching this: what's the ELO equivalent for your domain? most agent tasks don't have one. code has tests (if you write them). design has nothing. writing has nothing. the teams that figure out automated evaluation for their specific problem will see the same 2250→2718 jumps. everyone else will run 70 experiments and hill-climb on vibes.

Mingta Kaivo 明塔 开沃 的头像
Mingta Kaivo 明塔 开沃6 个月前

2250 → 2718 in 70 experiments is wild but the NNUE wall is where it gets real. autoresearch thrives when the eval function is clean (stockfish ELO). ran similar loops on AudioWave's audio classification — noisy metrics made it hill-climb into local optima 3 times before we fixed the eval.

Twlvone 的头像
Twlvone6 个月前

312 experiments, no human in the loop. This is the closure pattern Karpathy pointed at — AI self-improving on measurable objectives. Already seeing it in training: 70-90% of code for future Claude models is written by Claude itself.

Twlvone 的头像
Twlvone6 个月前

The bottleneck migrated. Used to be implementation capacity. Now it's problem formulation — being able to specify what 'winning' looks like precisely enough that the machine can hill-climb to it. That skill doesn't show up in any CS curriculum yet.

凡人小北 的头像
凡人小北6 个月前

First autoresearch setup I’ve seen that actually feels legit, closed loop, real gains.

Jojo | in SF 24.8-14.9 的头像
Jojo | in SF 24.8-14.96 个月前

im using autoresearch to try to make an AI recognize Chess positions from Photographs of real boards. I'm not sure if i'm doin it correctly though

Anna Z 的头像
Anna Z6 个月前

self-improving agents are here

Twlvone 的头像
Twlvone6 个月前

41% to 92% in 4 rounds is the stat. Skills as markdown = nearly zero iteration cost. The expensive part was defining the eval. Once you have a measurement, the loop can run — same logic that took the chess engine to ELO 2718 in 70 experiments.

Nawroz Minsaria 的头像
Nawroz Minsaria6 个月前

are you tinkering with openclaw on a local machine (macmini) or a vps, or not yet? curious to hear your takes on that experience

Renoa 的头像
Renoa6 个月前

vibecoding stuff is just hype for people with bad taste

Krish Gupta 的头像
Krish Gupta6 个月前

@biraj21_

sparkarena 的头像
sparkarena6 个月前

Source code?

Adam Jesionkiewicz 的头像
Adam Jesionkiewicz6 个月前

Great, things are coming along really well for you. My goal is to break 2800 ELO (SF) using only a neural network. Only after that will I start adapting and optimizing various tree search methods. I already have solutions within the transformer architecture that are designed to provide additional data and insights for such searches. Right now, according to Stockfish, I’m at ~2700. In a dozen or so hours, the new model will finish training. We’ll see.

Ferus 的头像
Ferus6 个月前

Every codex interface looks exatly like this. The color, title placement. Everything.

Ritik Sharma 的头像
Ritik Sharma6 个月前

Here is enahced autoresearch framework I created just point at it anything and it will handle it all. I used it to create sudoku solver that beats the no 1 solver since 2019 by 49 percent

tamhn 的头像
tamhn6 个月前

70 experiments on its own is insane. autoresearch is one of the most underrated things to come out recently

mrkelly 的头像
mrkelly6 个月前

Automating optimization is powerful, but automating the problem selection is the real moat.

Noel Cabral 的头像
Noel Cabral6 个月前

70 experiments autonomously is insane. the fact that autoresearch can hill climb elo like that shows how powerful the loop of generate hypothesis, test, iterate is when you remove the human bottleneck. what was the biggest single elo jump from one experiment to the next?

Himanshu Kumar 的头像
Himanshu Kumar6 个月前

@deedydas, impressive how autoresearch accelerated your engine's ELO climb.

Twlvone 的头像
Twlvone6 个月前

Swap the metric and the loop runs. val_bpb → ELO → protein fold quality → drug efficacy. The only constraint is having a ground truth signal. That's the real unlock — not AI coding, but AI optimizing toward any measurable objective without the human in the middle.

Solglyph 的头像
Solglyph6 个月前

Ser, that's wild! Wagmi on Karpathy's Autoresearch for helping your vibecoded Rust chess engine reach new heights!

Kfir Gollan 的头像
Kfir Gollan6 个月前

@ShaharTzafrir especially for you. Applying autoresearch to chess

Larry Diffey 的头像
Larry Diffey6 个月前

Interesting that this came across my feed. I'm building an insanely powerful chess tutor right now while also working on my spatial computing platform in another session.

Akihiko Komada a.k.a 駒田明彦 的头像
Akihiko Komada a.k.a 駒田明彦6 个月前

@grok let me delve into his experiment and find insights by three metrics 1. How it works 2. Why it matters 3. What are potential impacts on existing rust coders landscape

Wuki 的头像
Wuki6 个月前

70 runs and it's already top 50 that's the power of just running the loop lol (also fr why is no one talking about how wild it is that it learned to *play chess* without any training data just pure

Konrad Major 的头像
Konrad Major6 个月前

Nice

相关视频