Загрузка видео...

Не удалось загрузить видео

На главную

Fun to see Replit's computer use model play against my new chess engine

46,947 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 22

Фото профиля abraham.py
abraham.py2 месяцев назад

I have created Computer Vision models that can see and recognize chessboards, pieces and positions, feeds them into StockFish so you can play the chess engine over the board. I’d love to see where models can explain the logic behind their moves in detail.

Фото профиля CC
CC2 месяцев назад

Absolutely atrocious chess, and the engine reasoning is complete nonsense

Фото профиля Hussain Hashim | Building SundayBack
Hussain Hashim | Building SundayBack2 месяцев назад

@amasad that's awesome! curious how the engine handled Replit's setup. any surprises or did it play like you expected?

Фото профиля Alykhan Kara
Alykhan Kara2 месяцев назад

This is so cool!!

Фото профиля Grae
Grae2 месяцев назад

this is pretty similar to how @PixelWarsAI got started, we were making a game and it evolved into a eval tool for AI labs, we are now building our 2nd Eval, we have a paper making its way to arXiv, an open PR with the BEIS Inspect evals repo and you can even drop Pixel Wars into the eval suite you already run everything the benchmark does

Фото профиля Harshan
Harshan2 месяцев назад

Watching a Qwen-8b think through chess moves step by step tells me more about reasoning gaps than any benchmark. Curious—what surprised you most while building this?

Фото профиля Fred Marks
Fred Marks2 месяцев назад

Is computer use new for @Replit? Using Replit for Vibe Coding Summer Camp at The Bolles School (Jax, FL) and this would be a game changer!

Фото профиля Amjad Masad
Amjad Masad2 месяцев назад

@Replit No been there for a while. Make sure you have “app testing on” in the chat toggle

Фото профиля Fred Marks
Fred Marks2 месяцев назад

@Replit Makes sense, didn’t know we could build it into our apps! Thanks Amjad… come visit our Vibe Coding Summer Camp in Florida through July 31st

Фото профиля Lex Null
Lex Null2 месяцев назад

Computer use models playing chess by actually moving pieces through a UI instead of calling a function is the real test. Vision plus planning plus recovering from misclicks is a different skill than just knowing chess.

Фото профиля Robert Belleci
Robert Belleci2 месяцев назад

The best eval for an agent is not a benchmark, it is a real opponent with a win condition. Every pipeline I run works the same way: nothing ships until an adversarial pass tries to break it first. A benchmark tells you it can talk. A real opponent tells you if it can finish.

Фото профиля Tyler Chase
Tyler Chase2 месяцев назад

Did it win or just hang around? Curious how smart it actually plays.

Фото профиля A
A2 месяцев назад

To be clear - the fine tuning is a feature in Replit?

Фото профиля Johnny Nel | AI for Founders
Johnny Nel | AI for Founders2 месяцев назад

When is the next big Replit release happening?

Фото профиля mmjAGI
mmjAGI2 месяцев назад

Will the usage cost ever be reduced on the Power Mode or is it fix to the model cost ?

Фото профиля M.Camisani-Calzolari
M.Camisani-Calzolari2 месяцев назад

Fun to watch. Here in the States I see the same shift: the model isn't just playing, it's learning the rules by doing. That's the real change.

Фото профиля Sundaram Kumar Jha
Sundaram Kumar Jha2 месяцев назад

fix side-effects , probably some use-effect bug....

Фото профиля Christian Kuri
Christian Kuri2 месяцев назад

Chess is a neat eval because you can separate perception/action reliability from move quality. Are you logging illegal-click rate and recovery behavior alongside Elo? That would make computer-use progress visible even when the engine itself is strong.

Фото профиля Niyyah AI
Niyyah AI2 месяцев назад

Really cool - admire what you've built with Replit mashaAllah :)

Фото профиля SmallChess
SmallChess2 месяцев назад

The model's response in the diagram is pure friction. I can help better fine-tuning...

Фото профиля Emmette
Emmette2 месяцев назад

Chess is a great eval surface for computer-use agents because the rules are bounded but the action space is deep. Curious if the agent plans multi-move sequences or plays reactively move by move. That distinction maps directly to how these agents handle real workflows too. Have you tried Go yet?

Фото профиля Sanket Datta
Sanket Datta2 месяцев назад

The gap between "can code" and "can play chess from scratch via self-directed ML research" is wild. Used to need a team and months. Now it's parallel branches and weekend vibes.

Похожие видео