Kimi K3 ranks #6 on ReactBench - It performs... better than Opus 4.8 while being over 2x cheaper (!!) - Scores 33%, jump of +10% compared to K2.7 - Lags behind the frontier (Fable and Sol) in frontend production-readiness:show more

Aiden Bai
73,851 次观看 • 1 个月前
I tested Kimi K3 vs Claude Opus 4.8 Same... prompt, an armory bay with lighting, props, and detail. Top is Kimi K3, bottom is Opus 4.8. It's not even close. Kimi K3 built a full scene with textures, proper lighting, ammo crates, weapon racks, working detail everywhere. Opus 4.8 gave me a near empty room with a couple of floating tables. No doubt it beats Opus 4.8. Kimi K3 is Fable 5 level, and it's clearly better than GPT-5.6 Sol at 3D and games. An open weight model just matched the best closed models on the market. Let that sink in.show more

Bhavy☄️
476,157 次观看 • 1 个月前
🤯KIMI K3 ABSOLUTELY MOGS! BEATING Opus 4.8, GPT 5.5,... and even Fable 5 in multiple benchmarks. They scored 1688 on GDPval-AA v2 🔥 This is a completely different breed of open-source models Kimi creates better games and front-end designs than Fable 5, but it's 8x cheaper! The results coming out are truly impressive! I will be testing this out further and posting multiple tests today. Stay tuned!show more

Mark Santos
127,993 次观看 • 1 个月前
DeepSeek V4 Flash being ranked almost equal to Opus... 4.8 is actually insane to me. Against Kimi K3, Opus 5, GPT-5.6 Sol, and Qwen 3.8 Max Preview, you give up a LOT by choosing DeepSeek. This is not a 2-point difference in practice. I’ll post the Opus 4.8 comparison next because you need to see this.show more

OmedTheVibeCoder
26,990 次观看 • 1 个月前
I gave Kimi K3 and Opus 5 the exact... same prompt to build a car racing game. I expected Opus 5 to take the lead. Instead, Kimi K3 finished in ~25 min for ~$8.9, while Opus 5 took 19 min and cost $14. The biggest surprise wasn’t the speed or the price. Kimi K3 built the better game. Smoother controls, better physics, and a driving experience that felt ready from the start. Opus 5 looked cleaner visually but Kimi won where it actually mattered. Nearly half the cost and the better gameplay.show more

Aman
111,571 次观看 • 1 个月前
> be Kimi Founder > Zhilin Yang. nobody in... the West knows your name. > build K1. K2. K2.5. K2.6. K2.7. > each one sets the open-source record > everyone says China can't reach the frontier > raise $2B. hit $20B valuation. > keep your head down. keep shipping. > July 16. drop K3. > 2.8 trillion parameters. > the biggest open model ever released. > 1 million token context. native multimodal. > Kimi Delta Attention — 6.3x faster decoding. > Artificial Analysis ranks it #4 overall. > right behind GPT-5.6 Sol and Fable 5. > open weights July 27. free for anyone. > the closed labs charge for the frontier. > you gave it away. > different game.show more

Kirill
413,178 次观看 • 1 个月前
Kimi K3 just jumped 17 places to #1 on... Arena.ai's Frontend Code Arena. So we gave it real client work to see where leaderboards meet reality: >10 landing page briefs, >judged blind by 8 working designers >against GPT 5.6 Sol, Gemini 3.5 Flash, and Claude Fable 5. It tied the best model in the study. And on detailed briefs, it won. Details in 🧵:show more

ben
13,798 次观看 • 1 个月前
KIMI K3 VS OPUS 4.8 SIMULATING A 3D PARTICLE... SYSTEM same prompt, same target: thousands of particles reacting to gravity and to each other in real time → K3: more organic distribution, smoother motion, particles actually cluster and drift like something physical → Opus: more rigid pattern, less natural, movement reads more like a grid than a simulation K3: $0.65 in API. Opus 4.8: ~$1.30 this is simulated physics, not just aesthetics, and it's the kind of detail that separates a demo from something you'd actually shipshow more

Jade
32,766 次观看 • 1 个月前
KIMI K3 JUST RECREATED A FUNCTIONAL GAME BOY ADVANCE... IN 3D → interactive 3D model, every button and screen modeled → a game loaded and playable right inside it → camera you can rotate and zoom around the device → one single prompt, zero external assets $3 in API with K3. the same result would run about $9 with Fable 5 and $6 with Opus 4.8 this isn't generating a static image of hardware anymore, it's generating a working simulation of it, from the shell to the screen to the game running insideshow more

alex
38,248 次观看 • 1 个月前
FABLE 5 JUST RECREATED NEW YORK CITY IN BLENDER.... I connected Blender to Fable 5 and in like 20 minutes, it remade the entire cityscape of NYC It started by getting data of the buildings from public sources, THEN began building it (imo this is a smarter move than what Opus 4.8 would've done) So theoretically, this entire build is up to scale too Fable 5, welcome back.show more

ashen
222,966 次观看 • 2 个月前
New Claude Sonnet 5 performs at GPT 5.5 level... 6x cheaper! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics crash demos Prompts: - A car crashes into a brick wall - A wrecking ball destroys a house - A catapult throws a rock at a castle wall Outputs: Sonnet 5: 15,047 tokens, $0.15 Opus 4.8: 23,063 tokens, $0.58 Sonnet 4.6: 25,824 tokens, $0.39 GPT 5.5: 31,152 tokens, $0.94 Sonnet 5 did as well as Opus 4.8 and GPT 5.5 on all three tests. In the wrecking ball test, it beat Opus 4.8. The cable moves smoothly and every hit connects. In the catapult test, it beat GPT 5.5. The rock always lands inside the wall. Sonnet 5 still needs better detail and graphics. But it used fewer tokens than every other modelshow more

atomic.chat
728,593 次观看 • 2 个月前
Fable 5 totally crushed our new contest, but it... cost 6x more than Opus 4.8! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: — A train derailing off a broken bridge into the water — Two cars jumping off ramps and colliding mid-air over a canyon — A monster truck crushing a row of parked cars Outputs: Fable 5: 62,158 tokens, $3.12 GPT 5.5: 37,753 tokens, $1.14 Opus 4.8: 22,280 tokens, $0.56 GLM 5.2: 36,246 tokens, $0.08 Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.show more

atomic.chat
2,839,437 次观看 • 2 个月前
It's never been a better time to be creative.... (ever) The two best frontier models (ever) have been released within 30 days of each other. Here’s what we learned from running 4 frontier models head to head. >Sol has taste and Fable takes direction. 10 identical landing page briefs, judged blind by working creatives. GPT 5.6 Sol won 82% on loose briefs. But when handed a real design spec it finished last. 🧵show more

ben
35,609 次观看 • 1 个月前
a moonshot engineer leaked the benchmark anthropic, openai and... xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article belowshow more

starmex
32,547 次观看 • 11 天前
Kimi K3 vs GPT-5.6 vs Claude Fable 5 A... game is one of the hardest places to fake working code. A web app can hide bugs behind a polished screenshot. A game can't. Bad collision, clipping, jitter, broken physics, awkward controls, or a snowboard sinking into the terrain become obvious within seconds. That's why I like using games to evaluate coding models. I gave Kimi K3 a single prompt to build a snowboard game in Godot, and just 20 seconds of gameplay told me more than a page of benchmark scores ever could. Benchmarks matter, but watching a model build something interactive reveals a different side of its capabilities. If the gameplay feels right, the systems work together, and the experience holds up under real input, that's a much more convincing demonstration than any leaderboard.show more

FHILY👑
36,569 次观看 • 1 个月前
🇨🇳 CHINA IS ALMOST READY TO UNDERCUT THE WORLD'S... BEST AI MODELS. Six months ago, that sounded impossible. Today GLM-5.2 is open source and costs a fraction of frontier models. Zhipu's co-founder says Fable-class Chinese AI comes even sooner than people think. This is the same playbook China ran on solar, on steel, on EVs. Build it cheaper. Give it away. Take the market. If China delivers frontier AI at a fraction of the price, the economics holding up the U.S. AI market start to crack.show more

CryptoGoos
46,766 次观看 • 2 个月前
You can now orchestrate Fable 5, Sol, and any... model inside Codex with one plugin. It's called Codex-Orchestration. Assign Fable 5 as the advisor, Sol as the executor, or any model to any role. Then define the order they work in. Codex handles the routing. I ran Fable 5 High as planner with GPT-5.6 Sol Extra High as executor on a set of issues Opus and GPT-5.5 always struggled with. Done in 30 minutes. 40% fewer limit hits. 2x faster implementation. Install it by pasting this into Codex: "Install Codex Orchestration: codex plugin marketplace add Cjbuilds/Codex-Orchestration codex plugin add codex-orchestration@codex-orchestration Verify the installation, then tell me to start a new task." Then assign your models: @ codex-orchestration advisor: Claude Fable 5 High, Executor: GPT-5.6 Sol High Open source. Tweak the routing however you want.show more

Alvaro Cintas
91,269 次观看 • 1 个月前
BREAKING: Anthropic just dropped Opus 4.8—and it is a... MONSTER We've been testing for about a week Every 🪨 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 🪨:show more

Dan Shipper 📧
354,559 次观看 • 3 个月前
Codex can now run Deepseek-v4- flash! There's a catch... though. Deepseek's official setup switches your entire codex over to them, so your GPT models stop showing up at all. This is exactly what Codex Router is for. It adds models to the list instead of replacing them, so sol, grok, kimi and deepseek all sit in the same picker and i just grab whichever one suits the job. Deepseek v4-flash is $0.28 per million output tokens. opus 4.8 is $25. same picker, 89x apart. Links in the comment. setup's in the video 👇show more

Ziwen
145,018 次观看 • 1 个月前
THIS GUY BOUGHT A $31 TOY DRONE AND TURNED... CLAUDE OPUS 4.8 INTO ITS ENGINEER he plugged it into a laptop, explained the control logic in plain english and let Claude build the flight interface by the end of the session it had calibration, live controls and a browser cockpit moving the drone in real time most people will use Opus 4.8 to save 12 minutes on emails. he used it to turn cheap plastic into a working demo the crazy part isn’t the drone. it’s that the bottleneck moved from writing code to describing exactly what you want built while everyone debates benchmarks, someone with a $31 gadget and one afternoon is already shipping hardware demosshow more

Gipp 🦅
735,627 次观看 • 3 个月前
I can't believe this is real I have GLM... 5.2 running 100% locally on my Mac Studio. 2 bit quant. The results I'm getting are better than Opus 4.8 It's now powering my Hermes Agent and Codex. 100% free, local, private super intelligence on my desk I also have it in a loop coding for me 24/7 now I thought we were at least a year away from this type of event. It happened today. The model takes up about 250gb of memory. So you can technically run it on a Mac Studio with 256gb, but you probably want the 512gb memory version (please tell me you listened to me 5 months ago when these were sitting on store shelves) With Fable gone, I now have Opus 4.8 level intelligence on my desk for free. This is the future. Local, private, secure, personal super intelligence. If you're still writing off local AI as a fad or engagement bait, you are officially delusionalshow more

Alex Finn
622,922 次观看 • 2 个月前