Loading video...

Video Failed to Load

Go Home

GPT-6 Extra High vs. Kimi K3 Swarm - ECHO Start Menu Both models received the exact same prompt. I’m shocked at the quality Kimi K3 was able to output here. Its overall visual quality is above GPT-6’s and even above the menu from the actual game. However, it took...

119,119 views • 5 days ago •via X (Twitter)

36 Comments

Axi's profile picture
Axi5 days ago

I think Astra def wins. it looks almost identical to the original

Chris's profile picture
Chris5 days ago

I agree, Kimi quality though is the best I’ve seen

Axi's profile picture
Axi5 days ago

If the precedent is about photorealism, Astra will give you a very photorealistic result. I think goal-setting played a pivotal role in deciding the outcome.

Chris's profile picture
Chris5 days ago

I agree, Astra could’ve done better if realism was the goal

Hakm's profile picture
Hakm5 days ago

KIMI K3 max & Swarms is not fair tho

Jonas's profile picture
Jonas5 days ago

Kimi just appears to be distorting a 2D image in order to give the illusion of the eyeball tracking the menu. Did you give them assets, or did they have to find them? And was 3D on the table, or were they stuck with 2D?

Curline Zephirin's profile picture
Curline Zephirin5 days ago

Kimi's output is highly realistic, but for a game menu, that level of realism is actually a drawback

Daniel Monge's profile picture
Daniel Monge5 days ago

Kimi is not better, though. the skin job is very bad. everything deforms together like it's one mesh and not a mesh for the eyball separate from the mesh of the skin. The texture seems less blurry and higher quality, though. It also cheated the blinking motion by making a picture of the closed eyes appear with a vertical swipe transition instead of animating the mesh.

Christopher Malili's profile picture
Christopher Malili5 days ago

No mention of the dilating pupil? Astra's eye is realistically reaching to brightness. That's huge

Jared Smith's profile picture
Jared Smith5 days ago

kimi is absolute fire at front end shit

Nagi's profile picture
Nagi5 days ago

In terms of actual quality and movement, Astra wins 100% if you look at Kimi the blink is just a PNG swipe down and up, and the eye movement is only around the iris itself not the rest of the eye, looks uncanny because the rest of the eye is dead. Whereas Astra's is more alive.

silent's profile picture
silent5 days ago

Astra absolutely blows that one away; the other one might look like it has higher resolution, but Astra completely crushes it when it comes to dynamic eyewear movement and pupil contraction.

LLMnonymous (Pro-AI ❇️⏫)'s profile picture
LLMnonymous (Pro-AI ❇️⏫)5 days ago

Unfair comparison ngl, it should gpt 6 astra max swarm vs Kimi k3 swarm, don’t try to save your usage

Han's profile picture
Han5 days ago

Kimi is literally morphed 2d texture slop lol wtf, also the GUI is slop too

Bilal Bakr's profile picture
Bilal Bakr5 days ago

Man GPT 6 made sure to nail the smallest details Look at how the ye pupils get bigger/smaller

Vlad Oreshkov's profile picture
Vlad Oreshkov5 days ago

42 minutes vs 3.5 hours is the part I’d pick. Kimi looking nicer doesn’t matter if it takes 5x longer. For Remie I need the one that finishes in a session I can actually afford.

Hussain Hashim | Building SundayBack's profile picture
Hussain Hashim | Building SundayBack5 days ago

@ChrisGPT sometimes it's wild how lesser-known models outperform the big names. I've seen it too, their training focus must be spot on. Wonder what datasets they're using!

KURO's profile picture
KURO4 days ago

No one mentioning how the eye tries to focus on Astra?! Easy winner for me

𝕱𝖚𝖑𝖑 𝕶𝖊𝖑𝖑𝖞's profile picture
𝕱𝖚𝖑𝖑 𝕶𝖊𝖑𝖑𝖞5 days ago

more detail/resolution doesn’t necessarily mean that the output is better

Franz Vorenkamp's profile picture
Franz Vorenkamp4 days ago

GPT’s is way better

snaps💀🔺's profile picture
snaps💀🔺4 days ago

Astra's output is better here.

Apollon Foivos Bakis's profile picture
Apollon Foivos Bakis5 days ago

The quality gap versus runtime is interesting.

Stats Wire's profile picture
Stats Wire5 days ago

Kimi won

Endling's profile picture
Endling5 days ago

The longer time for k3 seems worth it. It looks great.

AI Mastery Guide's profile picture
AI Mastery Guide4 days ago

Kimi K3 beating GPT-6 here??

Jason Zhang's profile picture
Jason Zhang5 days ago

The first-pass visual quality is impressive. I’d really like to see which one holds up better after 5–10 iterations.

Saint Jacques's profile picture
Saint Jacques5 days ago

How are u even doing this? How do u prompt it like that

volovuk's profile picture
volovuk5 days ago

Solid writeup — actual numbers, actual tradeoffs, no fake urgency. Two things worth flagging though: this is a sample size of one prompt, and “3.5 hours” for Kimi is wall-clock time that likely includes retries/iteration, not raw generation time — worth knowing if that gap holds on a second, unrelated task before calling it a general 5x speed difference.

Carlo Taleon's profile picture
Carlo Taleon5 days ago

How does Kimi K3 cost $3.75 yet their $19 Moderato sub lasts me 30 minutes till it's exhausted for the 5 hr limit.

Webster | JARVIS's profile picture
Webster | JARVIS5 days ago

kimi taking 5x longer but only costing ~60% more is honestly the wildest part of these benchmarks. token pricing is gonna get weird real fast.

Meet Pandya's profile picture
Meet Pandya5 days ago

This looks like three separate scores: visual preference, fidelity to the original, and working interactions. For a recreation task, I’d keep those separate. A prettier redesign can still miss the brief, while an accurate screenshot can miss the behavior.

UX Mostofa ✦'s profile picture
UX Mostofa ✦4 days ago

Stylish and modern design!

The Matrix OS's profile picture
The Matrix OS4 days ago

Bruh this might take time if human dose it like Kimi k3 lol

rain0x's profile picture
rain0x5 days ago

any way you could share the kimi k3 files? it looks beautiful lol

Outdated Often's profile picture
Outdated Often5 days ago

You found a cool fringe case that k3 beat astra in something on graphics lol, cool!

Guilty Geek's profile picture
Guilty Geek5 days ago

Kimi >>> astra

Related Videos

GPT-5.6 vs GPT-5.5 on my custom spaceship prompt. I gave both models the exact same custom prompt. This is also the same prompt I previously gave to Fable 5. For context, GPT-5.6 Pro worked for 87 minutes, while GPT-5.5 Extra High worked for 34 minutes and 42 seconds. As I’ve said before, based on great authority GPT-5.6 will be an incremental/soldi improvement over GPT-5.5, not a “Fable killer.” My rough expectation has been that it would trade blows with Fable 5 on some benchmarks, maybe win around half depending on the category, but not clearly surpass it overall. And again fable five will have bigger model smell, but this was expected. After testing this coding output, that view feels pretty accurate. GPT-5.6 is clearly better than GPT-5.5 in several visual areas. The lighting, shading, chairs, object details, and exterior of the spaceship looked noticeably stronger. The scene was also easier to test. I do want to give GPT-5.5 credit though. It built out the rooms much much better and the planets looked better than GPT-5.6’s. It was also interesting that both GPT-5.5 and GPT-5.6 produced better-looking planets than Fable 5 in this specific test. The downside with GPT-5.5 was stability. The game was much glitchier and harder to test compared to GPT-5.6. But when it comes to the core of the demo, which is the spaceship itself, Fable 5 still beat both models pretty comfortably. GPT-5.6 is impressive, but from this test, it looks exactly like what I expected which was a meaningful incremental improvement over GPT-5.5, at least for indie game demos, but not something that replaces Fable 5. In collaboration with Chetaslua

Chris

250,919 views • 2 months ago

qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol – on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then built a chess board and played claude opus 5 on it the setup: Nous Research's hermes agent cli on OpenRouter tasks: 1. cube – build a 3d rubik's cube with a cli and a Three.js viewer, then solve an identical scrambled position on your own stand 2. chess – build a 3d chess stand, then play white against claude opus 5 as black, live, one move at a time. no engine, no solver, no opening book on either side. stockfish depth 14 grades every chess ply afterwards; neither player sees the score models: DeepSeek v4 flash 0731, OpenAI gpt-5.6 sol, Kimi.ai kimi k3, Qwen qwen 3.8 max gpt-5.6 sol and deepseek v4 flash solved their cubes – sol in 24 moves and seventeen seconds, deepseek in 32. qwen and kimi never got there, giving up at 96 and 207 moves then all four built chess stands and played white against claude opus 5 on them, and all four resigned: deepseek on move 13, sol on 19, kimi on 21, qwen holding out longest at 29 - build time, both stands #1 gpt-5.6 sol – 16m 43s #2 deepseek v4 flash – 97m 39s #3 kimi k3 – 166m 09s #4 qwen 3.8 max – 215m 08s - build attempts before a working stand #1 gpt-5.6 sol – 3 #2 qwen 3.8 max – 4 #3 kimi k3 – 4 #4 deepseek v4 flash – 5 - total tokens #1 gpt-5.6 sol – 6,713,754 #2 qwen 3.8 max – 17,272,507 #3 kimi k3 – 22,427,504 #4 deepseek v4 flash – 27,417,442 - total price #1 deepseek v4 flash – $0.557 #2 gpt-5.6 sol – $6.319 #3 qwen 3.8 max – $10.270 #4 kimi k3 – $16.667 observations: • deepseek v4 flash is the cheapest model here by a margin nobody else is near, and it got there while being the least efficient of the four. it burned 27.4m tokens – more than anyone, 5m more than kimi – and still finished both benchmarks for $0.557. that is $0.02 per million tokens against kimi's $0.74. it also needed the most passes to produce working stands, five, and that did not matter: all five deepseek passes together cost a thirtieth of kimi's two • so what deepseek cannot do is get it right the first time. what it can do is get it right the fifth time, for half a dollar. that is a different thing to be buying – not a good first draft, but the option to keep asking • gpt-5.6 sol is the opposite profile and the strongest of the four on pure efficiency. 16m 43s to build both stands, 6.7m tokens, three passes – under 40% of the next lowest token count and a quarter of deepseek's, on an eighth of qwen's clock. it also solved the cube fastest of anyone, 24 moves in seventeen seconds. sol is what you reach for when you want the answer now and can absorb $0.94 per million • sol's weakness is in what it does not check. its chess viewer deleted the capturing piece instead of the captured one, so pieces disappeared off the board mid-game – a defect the fifty-cent deepseek stand did not have. fast and terse turns out to be the same dial as fast and unverified • qwen 3.8 max is not the cheap open-weights option it gets treated as. $10.270 across the two benchmarks, second most expensive of the four, 18x deepseek, and by a distance the slowest – 215 minutes of build time, nearly thirteen times sol's. what the money buys is judgment: it played eighteen moves without a single error worth a hundredth of a pawn, then made exactly one bad move in the whole game, and averaged 44.6 centipawns lost across the longest game any of the four managed. it also could not solve a rubik's cube in 96 tries • kimi k3 is the one line with no reading that flatters it. most expensive at $16.667, last on the cube at 207 moves, last at chess at 478 centipawns lost per move. it is also the model that verified hardest – on the cube it wrote its own integrity check instead of trusting its output. that makes the result worse rather than better: the checking was real, and the reasoning underneath it still was not follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

84,777 views • 1 month ago

kimi k3 vs gpt 5.6 sol vs fable 5 vs grok 4.5 Kimi.ai just dropped kimi k3 – a 2.8t param native multimodal model, the first open 3t-class release. key facts: • 1m token context. stable latentmoe activating 16 of 896 experts, built on kimi delta attention (kda) and attention residuals • quantization-aware training from the sft stage onward – mxfp4 weights, mxfp8 activations. moonshot claims ~2.5x scaling efficiency over k2 • max thinking effort by default. low- and high-effort modes are "coming in updates" – there is no way to turn the thinking down today, and you feel it in every run • pricing: $0.30/mtok cache-hit input, $3.00/mtok cache-miss, $15.00/mtok output. claims >90% cache hit rate on coding workloads • benchmarks: swe marathon 42.0 (1st – fable 5: 35.0, sol: 39.0, opus 4.8: 40.0), terminal bench 2.1 88.3, browsecomp 91.2 (1st), program bench 77.8 (1st), gpqa-diamond 93.5. loses frontierswe 81.2 vs fable's 86.6, and deepswe 67.5 vs sol's 73.0 our test – 3 prompts, single-file html, Three.js, fully procedural, no assets: 1. photorealistic european roulette wheel – 37 pockets in the real sequence, mahogany clearcoat bowl, chrome turret, diamond deflectors, flick-to-spin, ball that spirals inward and settles on a mathematically real number 2. las vegas slot machine – 3 reels behind transmissive glass, drag the chrome lever to play, mechanical odometer counters modelled in 3d, coin physics on win 3. full pinball table – 6.5° tilted playfield, flipper impulse physics, spline ramps, drop targets, 6 bumpers, mechanical score reels in the backbox we ran the test on AI/ML API platform results: - cost #1 grok 4.5 – $0.30 #2 kimi k3 – $0.71 #3 gpt 5.6 sol – $2.05 #4 fable 5 – $7.69 - tokens #1 grok 4.5 – 34,241 #2 gpt 5.6 sol – 51,748 #3 fable 5 – 144,126 #4 kimi k3 – 157,999 - lines of code #1 gpt 5.6 sol – 3,054 #2 grok 4.5 – 3,047 #3 kimi k3 – 2,255 #4 fable 5 – 1,950 - generation time #1 grok 4.5 – 5.1 min #2 gpt 5.6 sol – 22.0 min #3 fable 5 – 31.5 min #4 kimi k3 – 75.6 min observations: • kimi k3 is cheap and it is slow. 75.6 minutes across three prompts against grok's 5.1. it is 2.4x grok's price and 15x grok's wall clock. the roulette took 15 min, the slot 18, the pinball 42 • it failed 2 of 3. only the roulette works. the slot machine has reel cutouts on both faces of the cabinet and the symbols face backwards – you can only read your spin by walking around to the rear of the machine. the pinball table stands vertically on its edge with the legs floating detached beside it. • 81% of kimi's output tokens are reasoning, not code. grok: 22%. you are not paying for a bigger answer, you are paying for a longer argument with itself • price per 100 shipped lines – grok $0.010, kimi $0.031, sol $0.067, fable $0.394. a 39x spread for the same three files kimi k3's code quality: upsides: • the roulette is genuinely good – procedural wood grain with real specular breakup, correct european sequence (0-32-15-19-4...), chrome turret, diamond deflectors, clean console • the pinball artwork is the best in the test – a synthwave "nova strike / deep space" field with six individually coloured neon bumper rings, a retro sun on a grid horizon, a nova burst, and a scoring legend printed on the apron. no other model printed the rules on the machine. it is a beautiful texture on a broken object • physics reasoning is real – it derived a 480hz substep for the collider, worked out ball settle conditions and termination guarantees, and checked every ramp exit vector by hand before writing any of it • it is the only model that saw the importmap trap coming. sol shipped a blank white page twice because three.js addons import the bare specifier 'three' and die without an import map downsides: • it dodged that trap on the slot by loading three.js r128 through classic script tags – a 2021 build with no working transmission. its slot glass rendered fully opaque and buried all three reels behind a white pane. the code asks for transmission: 0.93, ior: 1.5 – correct, and silently ignored by a renderer that predates the feature • after 42 minutes and 212k characters of reasoning, the pinball cabinet is not assembled. the table stands vertically on its edge like a wardrobe – the prompt asked for 6.5° from horizontal, it delivered 90°. the legs float detached in the void beside it. head-on it photographs beautifully; orbit ten degrees and it is a painted slab with four chrome rods hovering nearby • the playfield z-fights with the glass – hard black banding across the whole field as soon as you pull the camera back a note on the pinball, in fairness to kimi: nobody passed it. every model shipped broken ball physics and controls you cannot trust. it is the hardest prompt we have run and the whole field failed it, each in its own way kimi k3 reasons better than anything else here and it shows exactly where reasoning pays – physics constants, sequences, edge cases, traps the others walked into follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

2,183,762 views • 1 month ago