Video wird geladen...
Video konnte nicht geladen werden
Ran FlappyBench on DeepSeek V4.1-Flash, GLM 5.3, and Kimi K3 with the same /design prompt. 🔹 DSV4.1-Flash: 9/10 · $0.0089 one-shot with the lowest cost 🔹 GLM-5.3: 9.5/10 · $0.0184 · nailed the UI but cost 2x more 🔹 Kimi K3: 8/10 · $0.0740 · hard-coded gameplay and 8x more
18,397 Aufrufe • vor 2 Tagen •via X (Twitter)
17 Kommentare

Our engineering and design team has been testing 26+ side-by-side comparisons across frontier and open models. All runs are public and open source. Benchmark for this demo here:

Chinese model are doing great

I remember when the cost was in full $. Now it’s so little we have to use different currency soon 😂

GLM looks sharp

Fastest and most efficient model so far by @deepseek_ai

n=10, one prompt: 9 vs 9.5 is inside noise, so the real result is 'all three pass'. And $0.0089 is one-shot spend, not cost per accepted run - log the failed attempt too. Retries are where the 8x gap shrinks. My 85-task run: $9 base -> $13 with 30% retries.

Our very own FlappyBench 😉

The cost gap is the interesting result, but I’d want to see the same prompt rerun across seeds before calling it a win. Did you keep the tool and model settings identical besides the model?

I have the GOAT plan, I ran out of weekly credits, I bought additional credits, but now I realize I can't use them until the end of the month... Why? It would be more practical to be able to use them when I need them, like today...

I honestly think this is one of the most useless bench, but one thing is clear, kimi somehow is uniquely good in every benchmark and its impressive

看来gemini连flash模型都要out了

DSV4.1-Flash is the rational pick: 9/10 at $0.0089. GLM-5.3's +0.5 at 2x cost is demo polish, not production value. Kimi's hard-coded gameplay at 8x is the tell.

please improve your permissions settings and auto approve settings. ........................................... very problematic

why is ds v4.1 flash not available on /model list?

Update CLI

请修复issue#829

Same prompt, published cost. 9/10 at $0.009 is the routing argument — leaderboards without unit economics are vibes.
