正在加载视频...

视频加载失败

Deepseek-v4.1 local vs Opus 5 chess: 2nd match Thinking on Reasoning low DeepSeek wins with checkmate in 22. So far thinking off Opus won Reasoning low Deepseek won Tiebreaker with thinking high underway. I’ll make a tournament style match off with more models. Best out of 3

26,540 次观看 • 2 天前 •via X (Twitter)

16 条评论

Chris W 的头像
Chris W2 天前

this is so cool. LLM chess tournament. nice one. you could form 2 teams - open vs closed 😎

Wësche 的头像
Wësche2 天前

I’m thinking one tree open, one tree closed and the two winners then against each other

Chris W 的头像
Chris W2 天前

ok best of both clash at the end, nice!

Webster | JARVIS 的头像
Webster | JARVIS2 天前

Checkmate in 22 under reasoning low is a clean signal: local models don't always need max compute. Best of 3 with thinking high will test if Opus 5's edge survives.

Wësche 的头像
Wësche2 天前

Last match I think deepseek started winning but lost the queen and went downhill from there

Jay Brunet 的头像
Jay Brunet2 天前

DeepSeek sacrificing a bishop early in the game, kinda ballsy.

Tyler Folkman 的头像
Tyler Folkman2 天前

Cool idea!

Wësche 的头像
Wësche2 天前

Thank you! Thinking on matches take ours tho, but working on more now

None Done 的头像
None Done2 天前

You should play 15 - 20 matches as sample and collect results

know2why 的头像
know2why2 天前

pls publish token consumed and total cost.

Gaurav Joshi 的头像
Gaurav Joshi2 天前

How about you let the model choose thinking effort but have a time limit, so they cannot spend too much time thinking! Just like Humans chess.

xingbin lu 的头像
xingbin lu2 天前

我认为你应该测试围棋,围棋变化更多

Mia 的头像
Mia2 天前

This is so cool

Shantanu Goel 的头像
Shantanu Goel2 天前

Nice work! I had made something very similar a couple of months ago where anyone can pit any models against each other

Satire 的头像
Satire2 天前

Testing models with chess? Yes!!!!!

Philip McBride 的头像
Philip McBride2 天前

Fascinating. Was one of the matches with both having thinking off (or low)? Amazing this close. Also says something about the thinking/reasoning off or low... perhaps more fiddling with related flags.

相关视频

What's the Big Deal with DeepSeek in AI? Here's why DeepSeek is making everyone take notice: 1. Super Smart on a Budget: DeepSeek showed you can make awesome AI without breaking the bank. Their latest model, DeepSeek-V3, was trained for only about $10 million, which is a lot less than the usual big bucks spent on AI, like the rumored $78 million for some of OpenAI's models. They did this in just two months with fewer fancy computers. 2. Open for Everyone: DeepSeek isn't keeping their tech a secret. They've made it open-source, meaning anyone can use, tweak, and learn from it. It's like they're saying, "Come join the party!" 3. Beating the Big Names: DeepSeek-V3 has done better than some top dogs from companies like OpenAI and Google in solving puzzles, math, and coding. This proves you can get great AI results without spending a fortune. 4. Challenging NVIDIA: NVIDIA's chips are usually the choice for AI because they're really powerful. But since DeepSeek did so well with less expensive chips, it might make people think twice about always going for NVIDIA's priciest options. 5. The DeepSeek Crew: The team at DeepSeek is young and smart, mostly from top Chinese schools, with brains in physics, math, and computer science. They learned AI in about six months by themselves! They use first principle thinking, which means they break down problems to the basics and build from there. This has helped them come up with cool new ways to do AI. 6. Changing AI for Good: DeepSeek is showing that AI can be cheaper and more open to everyone. They're changing how we think AI should be made and shared, which could shake up the whole AI world. So, as we watch DeepSeek, it's clear they're not just another player; they're changing the rules of the game. I predicted that this would be a make or break year for all the massive investments made in AI by American VC's. A few weeks later, DeepSeek happens! Watch the rest of my predictions in my 2025 outlook video . Link in replies #AIInnovation #DeepSeek #NVIDIA #OpenAI #TechDisruption

Dr Ola Brown

83,460 次观看 • 1 年前

qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol – on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then built a chess board and played claude opus 5 on it the setup: Nous Research's hermes agent cli on OpenRouter tasks: 1. cube – build a 3d rubik's cube with a cli and a Three.js viewer, then solve an identical scrambled position on your own stand 2. chess – build a 3d chess stand, then play white against claude opus 5 as black, live, one move at a time. no engine, no solver, no opening book on either side. stockfish depth 14 grades every chess ply afterwards; neither player sees the score models: DeepSeek v4 flash 0731, OpenAI gpt-5.6 sol, Kimi.ai kimi k3, Qwen qwen 3.8 max gpt-5.6 sol and deepseek v4 flash solved their cubes – sol in 24 moves and seventeen seconds, deepseek in 32. qwen and kimi never got there, giving up at 96 and 207 moves then all four built chess stands and played white against claude opus 5 on them, and all four resigned: deepseek on move 13, sol on 19, kimi on 21, qwen holding out longest at 29 - build time, both stands #1 gpt-5.6 sol – 16m 43s #2 deepseek v4 flash – 97m 39s #3 kimi k3 – 166m 09s #4 qwen 3.8 max – 215m 08s - build attempts before a working stand #1 gpt-5.6 sol – 3 #2 qwen 3.8 max – 4 #3 kimi k3 – 4 #4 deepseek v4 flash – 5 - total tokens #1 gpt-5.6 sol – 6,713,754 #2 qwen 3.8 max – 17,272,507 #3 kimi k3 – 22,427,504 #4 deepseek v4 flash – 27,417,442 - total price #1 deepseek v4 flash – $0.557 #2 gpt-5.6 sol – $6.319 #3 qwen 3.8 max – $10.270 #4 kimi k3 – $16.667 observations: • deepseek v4 flash is the cheapest model here by a margin nobody else is near, and it got there while being the least efficient of the four. it burned 27.4m tokens – more than anyone, 5m more than kimi – and still finished both benchmarks for $0.557. that is $0.02 per million tokens against kimi's $0.74. it also needed the most passes to produce working stands, five, and that did not matter: all five deepseek passes together cost a thirtieth of kimi's two • so what deepseek cannot do is get it right the first time. what it can do is get it right the fifth time, for half a dollar. that is a different thing to be buying – not a good first draft, but the option to keep asking • gpt-5.6 sol is the opposite profile and the strongest of the four on pure efficiency. 16m 43s to build both stands, 6.7m tokens, three passes – under 40% of the next lowest token count and a quarter of deepseek's, on an eighth of qwen's clock. it also solved the cube fastest of anyone, 24 moves in seventeen seconds. sol is what you reach for when you want the answer now and can absorb $0.94 per million • sol's weakness is in what it does not check. its chess viewer deleted the capturing piece instead of the captured one, so pieces disappeared off the board mid-game – a defect the fifty-cent deepseek stand did not have. fast and terse turns out to be the same dial as fast and unverified • qwen 3.8 max is not the cheap open-weights option it gets treated as. $10.270 across the two benchmarks, second most expensive of the four, 18x deepseek, and by a distance the slowest – 215 minutes of build time, nearly thirteen times sol's. what the money buys is judgment: it played eighteen moves without a single error worth a hundredth of a pawn, then made exactly one bad move in the whole game, and averaged 44.6 centipawns lost across the longest game any of the four managed. it also could not solve a rubik's cube in 96 tries • kimi k3 is the one line with no reading that flatters it. most expensive at $16.667, last on the cube at 207 moves, last at chess at 478 centipawns lost per move. it is also the model that verified hardest – on the cube it wrote its own integrity check instead of trusting its output. that makes the result worse rather than better: the checking was real, and the reasoning underneath it still was not follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

84,777 次观看 • 1 个月前