Loading video...

Video Failed to Load

Go Home

🤯KIMI K3 ABSOLUTELY MOGS! BEATING Opus 4.8, GPT 5.5, and even Fable 5 in multiple benchmarks. They scored 1688 on GDPval-AA v2 🔥 This is a completely different breed of open-source models Kimi creates better games and front-end designs than Fable 5, but it's 8x cheaper! The results coming...

128,202 views • 2 months ago •via X (Twitter)

29 Comments

Mihnea Peteu's profile picture
Mihnea Peteu2 months ago

true

Mark Santos's profile picture
Mark Santos2 months ago

post some more tests

AshutoshShrivastava's profile picture
AshutoshShrivastava2 months ago

Incredible man..

Mark Santos's profile picture
Mark Santos2 months ago

Thanks! I posted a few more results here:

Ronaldo Ballecer's profile picture
Ronaldo Ballecer2 months ago

It's really good. But I saw a post the 20 dollar sub can one shot the 5 hour limit and still not finish. So it should be plan for the first 5 hour limit then implement in the next and add only features or sub-features for more complex tasks.

linyishan's profile picture
linyishan2 months ago

早上立马订阅了一年,K3正式进入生产线

Shubham Sharma | Video Editor's profile picture
Shubham Sharma | Video Editor2 months ago

8x cheaper while beating that many benchmarks is a strong claim to watch play out

Nulledge's profile picture
Nulledge2 months ago

我不信任国产的,肯定又是蒸馏

Mark Santos's profile picture
Mark Santos2 months ago

try it out yourself

AI Mastery Guide's profile picture
AI Mastery Guide2 months ago

8x cheaper while beating that many benchmarks is a strong claim to watch play out

Joy Haines's profile picture
Joy Haines2 months ago

cool~The smoothness of the interactions is what stands out to me.

The Full Read's profile picture
The Full Read2 months ago

More will come

Maximilian's profile picture
Maximilian2 months ago

What a headline 😂😂😂

Adeel Khan (عدیل)'s profile picture
Adeel Khan (عدیل)2 months ago

i dont understand what will be the outcome of this race?

WorldClaw's profile picture
WorldClaw2 months ago

🔥 Get Kimi K3 at the Lowest Price Online — 30% Off: Cache Hit: $0.30 → $0.21 / 1M Cache Miss: $3.00 → $2.10 / 1M Output: $15.00 → $10.50 / 1M

Ty772521's profile picture
Ty7725212 months ago

注册可领会员额度

Joey's profile picture
Joey2 months ago

But have you built anything with it?

A Salty Vet's profile picture
A Salty Vet2 months ago

Can you prove you used the same prompt and the same model settings for both?

Isaac Rojas's profile picture
Isaac Rojas2 months ago

Fake... Impossible

💡Blue 🚀🌙's profile picture
💡Blue 🚀🌙2 months ago

@Kimi_Moonshot how’s a 8GB GGUF pls

Ask-Zai's profile picture
Ask-Zai2 months ago

You're lying, you're lying! Actually, put it to work. It doesn't come close. Stop falling for the hype! China is not about to take $15 per mill from anybody. They priced themselves out of the market, and you're hyping up a model that sits somewhere between Sonnet 5 and Opus 4.8, realistically.

安叫兽|Bird🕊️ 🔶 BNB's profile picture
安叫兽|Bird🕊️ 🔶 BNB2 months ago

碾压先别急,等实测多跑几轮

Setumeni Effy's profile picture
Setumeni Effy2 months ago

Orca router is giving free KIMI K3 CREDITS,grab whilst it's valid

Arindam Majumder 𝕏's profile picture
Arindam Majumder 𝕏2 months ago

This is cool, Also planning to test it against GLM & Minimax models!

Davidd Tech's profile picture
Davidd Tech2 months ago

I'll be testing it to backtest and automate a new trading strategy. Can't wait to share the results.

Ashutosh Mishra's profile picture
Ashutosh Mishra2 months ago

Impressive drop—Kimi K3 is cooking 🔥 When do the full tests drop?

Sebastian Buzdugan's profile picture
Sebastian Buzdugan2 months ago

1m context matters less when tool loops make latency the product bottleneck

Bizlink's profile picture
Bizlink2 months ago

is only way to use KIMI K3 is via API?

Nur Web3's profile picture
Nur Web32 months ago

Also the price💩🤡

Related Videos

GPT-5.6 vs GPT-5.5 on my custom spaceship prompt. I gave both models the exact same custom prompt. This is also the same prompt I previously gave to Fable 5. For context, GPT-5.6 Pro worked for 87 minutes, while GPT-5.5 Extra High worked for 34 minutes and 42 seconds. As I’ve said before, based on great authority GPT-5.6 will be an incremental/soldi improvement over GPT-5.5, not a “Fable killer.” My rough expectation has been that it would trade blows with Fable 5 on some benchmarks, maybe win around half depending on the category, but not clearly surpass it overall. And again fable five will have bigger model smell, but this was expected. After testing this coding output, that view feels pretty accurate. GPT-5.6 is clearly better than GPT-5.5 in several visual areas. The lighting, shading, chairs, object details, and exterior of the spaceship looked noticeably stronger. The scene was also easier to test. I do want to give GPT-5.5 credit though. It built out the rooms much much better and the planets looked better than GPT-5.6’s. It was also interesting that both GPT-5.5 and GPT-5.6 produced better-looking planets than Fable 5 in this specific test. The downside with GPT-5.5 was stability. The game was much glitchier and harder to test compared to GPT-5.6. But when it comes to the core of the demo, which is the spaceship itself, Fable 5 still beat both models pretty comfortably. GPT-5.6 is impressive, but from this test, it looks exactly like what I expected which was a meaningful incremental improvement over GPT-5.5, at least for indie game demos, but not something that replaces Fable 5. In collaboration with Chetaslua

Chris

250,919 views • 3 months ago