Загрузка видео...
Не удалось загрузить видео
Deepseek-v4.1 local vs Opus 5 chess: 2nd match Thinking on Reasoning low DeepSeek wins with checkmate in 22. So far thinking off Opus won Reasoning low Deepseek won Tiebreaker with thinking high underway. I’ll make a tournament style match off with more models. Best out of 3
26,540 просмотров • 2 дней назад •via X (Twitter)
Комментарии: 16

this is so cool. LLM chess tournament. nice one. you could form 2 teams - open vs closed 😎

I’m thinking one tree open, one tree closed and the two winners then against each other

ok best of both clash at the end, nice!

Checkmate in 22 under reasoning low is a clean signal: local models don't always need max compute. Best of 3 with thinking high will test if Opus 5's edge survives.

Last match I think deepseek started winning but lost the queen and went downhill from there

DeepSeek sacrificing a bishop early in the game, kinda ballsy.

Cool idea!

Thank you! Thinking on matches take ours tho, but working on more now

You should play 15 - 20 matches as sample and collect results

pls publish token consumed and total cost.

How about you let the model choose thinking effort but have a time limit, so they cannot spend too much time thinking! Just like Humans chess.

我认为你应该测试围棋,围棋变化更多

This is so cool

Nice work! I had made something very similar a couple of months ago where anyone can pit any models against each other

Testing models with chess? Yes!!!!!

Fascinating. Was one of the matches with both having thinking off (or low)? Amazing this close. Also says something about the thinking/reasoning off or low... perhaps more fiddling with related flags.
