正在加载视频...
视频加载失败
Fun to see Replit's computer use model play against my new chess engine
46,947 次观看 • 2 个月前 •via X (Twitter)
22 条评论

I have created Computer Vision models that can see and recognize chessboards, pieces and positions, feeds them into StockFish so you can play the chess engine over the board. I’d love to see where models can explain the logic behind their moves in detail.

Absolutely atrocious chess, and the engine reasoning is complete nonsense

@amasad that's awesome! curious how the engine handled Replit's setup. any surprises or did it play like you expected?

This is so cool!!

this is pretty similar to how @PixelWarsAI got started, we were making a game and it evolved into a eval tool for AI labs, we are now building our 2nd Eval, we have a paper making its way to arXiv, an open PR with the BEIS Inspect evals repo and you can even drop Pixel Wars into the eval suite you already run everything the benchmark does

Watching a Qwen-8b think through chess moves step by step tells me more about reasoning gaps than any benchmark. Curious—what surprised you most while building this?

Is computer use new for @Replit? Using Replit for Vibe Coding Summer Camp at The Bolles School (Jax, FL) and this would be a game changer!

@Replit No been there for a while. Make sure you have “app testing on” in the chat toggle

@Replit Makes sense, didn’t know we could build it into our apps! Thanks Amjad… come visit our Vibe Coding Summer Camp in Florida through July 31st

Computer use models playing chess by actually moving pieces through a UI instead of calling a function is the real test. Vision plus planning plus recovering from misclicks is a different skill than just knowing chess.

The best eval for an agent is not a benchmark, it is a real opponent with a win condition. Every pipeline I run works the same way: nothing ships until an adversarial pass tries to break it first. A benchmark tells you it can talk. A real opponent tells you if it can finish.

Did it win or just hang around? Curious how smart it actually plays.

To be clear - the fine tuning is a feature in Replit?

When is the next big Replit release happening?

Will the usage cost ever be reduced on the Power Mode or is it fix to the model cost ?

Fun to watch. Here in the States I see the same shift: the model isn't just playing, it's learning the rules by doing. That's the real change.

fix side-effects , probably some use-effect bug....

Chess is a neat eval because you can separate perception/action reliability from move quality. Are you logging illegal-click rate and recovery behavior alongside Elo? That would make computer-use progress visible even when the engine itself is strong.

Really cool - admire what you've built with Replit mashaAllah :)

The model's response in the diagram is pure friction. I can help better fine-tuning...

Chess is a great eval surface for computer-use agents because the rules are bounded but the action space is deep. Curious if the agent plans multi-move sequences or plays reactively move by move. That distinction maps directly to how these agents handle real workflows too. Have you tried Go yet?

The gap between "can code" and "can play chess from scratch via self-directed ML research" is wild. Used to need a team and months. Now it's parallel branches and weekend vibes.


