Loading video...

Video Failed to Load

Go Home

Fun to see Replit's computer use model play against my new chess engine

46,947 views • 2 months ago •via X (Twitter)

22 Comments

abraham.py's profile picture
abraham.py2 months ago

I have created Computer Vision models that can see and recognize chessboards, pieces and positions, feeds them into StockFish so you can play the chess engine over the board. I’d love to see where models can explain the logic behind their moves in detail.

CC's profile picture
CC2 months ago

Absolutely atrocious chess, and the engine reasoning is complete nonsense

Hussain Hashim | Building SundayBack's profile picture
Hussain Hashim | Building SundayBack2 months ago

@amasad that's awesome! curious how the engine handled Replit's setup. any surprises or did it play like you expected?

Alykhan Kara's profile picture
Alykhan Kara2 months ago

This is so cool!!

Grae's profile picture
Grae2 months ago

this is pretty similar to how @PixelWarsAI got started, we were making a game and it evolved into a eval tool for AI labs, we are now building our 2nd Eval, we have a paper making its way to arXiv, an open PR with the BEIS Inspect evals repo and you can even drop Pixel Wars into the eval suite you already run everything the benchmark does

Harshan's profile picture
Harshan2 months ago

Watching a Qwen-8b think through chess moves step by step tells me more about reasoning gaps than any benchmark. Curious—what surprised you most while building this?

Fred Marks's profile picture
Fred Marks2 months ago

Is computer use new for @Replit? Using Replit for Vibe Coding Summer Camp at The Bolles School (Jax, FL) and this would be a game changer!

Amjad Masad's profile picture
Amjad Masad2 months ago

@Replit No been there for a while. Make sure you have “app testing on” in the chat toggle

Fred Marks's profile picture
Fred Marks2 months ago

@Replit Makes sense, didn’t know we could build it into our apps! Thanks Amjad… come visit our Vibe Coding Summer Camp in Florida through July 31st

Lex Null's profile picture
Lex Null2 months ago

Computer use models playing chess by actually moving pieces through a UI instead of calling a function is the real test. Vision plus planning plus recovering from misclicks is a different skill than just knowing chess.

Robert Belleci's profile picture
Robert Belleci2 months ago

The best eval for an agent is not a benchmark, it is a real opponent with a win condition. Every pipeline I run works the same way: nothing ships until an adversarial pass tries to break it first. A benchmark tells you it can talk. A real opponent tells you if it can finish.

Tyler Chase's profile picture
Tyler Chase2 months ago

Did it win or just hang around? Curious how smart it actually plays.

A's profile picture
A2 months ago

To be clear - the fine tuning is a feature in Replit?

Johnny Nel | AI for Founders's profile picture
Johnny Nel | AI for Founders2 months ago

When is the next big Replit release happening?

mmjAGI's profile picture
mmjAGI2 months ago

Will the usage cost ever be reduced on the Power Mode or is it fix to the model cost ?

M.Camisani-Calzolari's profile picture
M.Camisani-Calzolari2 months ago

Fun to watch. Here in the States I see the same shift: the model isn't just playing, it's learning the rules by doing. That's the real change.

Sundaram Kumar Jha's profile picture
Sundaram Kumar Jha2 months ago

fix side-effects , probably some use-effect bug....

Christian Kuri's profile picture
Christian Kuri2 months ago

Chess is a neat eval because you can separate perception/action reliability from move quality. Are you logging illegal-click rate and recovery behavior alongside Elo? That would make computer-use progress visible even when the engine itself is strong.

Niyyah AI's profile picture
Niyyah AI2 months ago

Really cool - admire what you've built with Replit mashaAllah :)

SmallChess's profile picture
SmallChess2 months ago

The model's response in the diagram is pure friction. I can help better fine-tuning...

Emmette's profile picture
Emmette2 months ago

Chess is a great eval surface for computer-use agents because the rules are bounded but the action space is deep. Curious if the agent plans multi-move sequences or plays reactively move by move. That distinction maps directly to how these agents handle real workflows too. Have you tried Go yet?

Sanket Datta's profile picture
Sanket Datta2 months ago

The gap between "can code" and "can play chess from scratch via self-directed ML research" is wild. Used to need a team and months. Now it's parallel branches and weekend vibes.

Related Videos