KIMI K3 VS OPUS 4.8 SIMULATING A 3D PARTICLE... SYSTEM same prompt, same target: thousands of particles reacting to gravity and to each other in real time → K3: more organic distribution, smoother motion, particles actually cluster and drift like something physical → Opus: more rigid pattern, less natural, movement reads more like a grid than a simulation K3: $0.65 in API. Opus 4.8: ~$1.30 this is simulated physics, not just aesthetics, and it's the kind of detail that separates a demo from something you'd actually shipshow more

Jade
32,766 просмотров • 1 месяц назад
I gave Kimi K3 and Opus 5 the exact... same prompt to build a car racing game. I expected Opus 5 to take the lead. Instead, Kimi K3 finished in ~25 min for ~$8.9, while Opus 5 took 19 min and cost $14. The biggest surprise wasn’t the speed or the price. Kimi K3 built the better game. Smoother controls, better physics, and a driving experience that felt ready from the start. Opus 5 looked cleaner visually but Kimi won where it actually mattered. Nearly half the cost and the better gameplay.show more

Aman
111,571 просмотров • 1 месяц назад
DeepSeek V4 Flash being ranked almost equal to Opus... 4.8 is actually insane to me. Against Kimi K3, Opus 5, GPT-5.6 Sol, and Qwen 3.8 Max Preview, you give up a LOT by choosing DeepSeek. This is not a 2-point difference in practice. I’ll post the Opus 4.8 comparison next because you need to see this.show more

OmedTheVibeCoder
26,990 просмотров • 1 месяц назад
KIMI K3 JUST RECREATED A FUNCTIONAL GAME BOY ADVANCE... IN 3D → interactive 3D model, every button and screen modeled → a game loaded and playable right inside it → camera you can rotate and zoom around the device → one single prompt, zero external assets $3 in API with K3. the same result would run about $9 with Fable 5 and $6 with Opus 4.8 this isn't generating a static image of hardware anymore, it's generating a working simulation of it, from the shell to the screen to the game running insideshow more

alex
38,248 просмотров • 1 месяц назад
1-bit Kimi K3 performs at Opus 5 level on... 3D physics! We ran our Atomic Chat quant of Kimi K3 locally on 4x B200 against three cloud models and gave them all the same task, to build a giant anvil drop test as a single HTML file with real physics Outputs: K3 1bit (local): 15.8K tokens, $0 API cost Kimi K3 (API): 15.3K tokens, $0.30 API cost Opus 5: 22.8K tokens, $0.77 API cost GPT 5.6: 14.5K tokens, $0.72 API cost All four got the physics right. But only Kimi made a working winch. The drum turns and the chain drags the flat car off the pad. Opus 5 drew the most detail, road markings and sparks on the hit. And you can run a model at this level on your own box now. That still feels insane to usshow more

atomic.chat
53,845 просмотров • 1 месяц назад
I'M PROBABLY ONE OF THE PEOPLE CONTRIBUTING TO THE... GPUS DYING 😭 Built the same Crossy Road clone with both Kimi K3 and Opus 4.8 using their respective API keys inside Cursor, so neither model had a framework advantage despite the 2.7× difference in cost. Both produced fully playable games with solid physics, but Kimi's UI felt slightly more polished Kimi cost $7.11, Opus cost $19. Pretty impressive considering the price gapshow more

aditya
47,488 просмотров • 1 месяц назад
LongCat performed Opus 4.8 and GPT 5.5 level on... real physics tasks for $0! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics Prompts: - A cannon demolishing a brick wall - A bowling ball knocking down the pins - A tornado that sucks in random objects Outputs: LongCat: 18,015 tokens, $0.00 Opus 4.8: 18,872 tokens, $0.48 GPT 5.5: 32,588 tokens, $0.98 GLM 5.2: 31,062 tokens, $0.09 On the physics LongCat came out ahead of Opus 4.8 and GLM 5.2 - cleaner collisions, nothing clipping or falling through. On detail and rendering it matched GPT 5.5, the best looking of the four. Getting this quality for free is wild!show more

atomic.chat
105,523 просмотров • 2 месяцев назад
Fable 5 absolutely crushed the HTML5 physics contest, but... cost 6x more than Opus 4.8 and 39× more than GLM 5.2 in that test. Test was done on atomic[.]chat, a desktop app that runs LLMs locally. The test asked 4 models to generate self-contained canvas demos with believable motion and collisions. The scenes were not simple animations because every crash needed gravity, force, timing, and contact handling. Outputs: - Fable 5: 62,158 tokens, $3.12 - GPT 5.5: 37,753 tokens, $1.14 - Opus 4.8: 22,280 tokens, $0.56 - GLM 5.2: 36,246 tokens, $0.08show more

Rohan Paul
205,771 просмотров • 2 месяцев назад
a moonshot engineer leaked the benchmark anthropic, openai and... xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article belowshow more

starmex
32,547 просмотров • 20 дней назад
GPT-6 Astra vs Fable 5.1 vs Kimi K3 vs... Sol Astra: Took 11 minutes. This is from a follow-up after the first attempt had rendering and UI bugs. The smoke feels a little off. I expected more detail, but it did a good job understanding the intent. Cost: $12.85. Fable 5.1: Took 18 minutes to complete the same test with high detailing. Cost: $12.60. Kimi K3: Took 10 minutes with a similar level of output and cost only $6.95. SOL: Took 7 minutes, but the output seemed messy. Cost: $5.78. Personally, I like Kimi K3 for delivering similar quality with better token efficiency, followed by fable 5.1 for realism. What about you? I’ll do a couple more tests before coming to a conclusion.show more

Bhavy☄️
77,404 просмотров • 6 дней назад
Kimi K3 vs GPT-5.6 vs Claude Fable 5 A... game is one of the hardest places to fake working code. A web app can hide bugs behind a polished screenshot. A game can't. Bad collision, clipping, jitter, broken physics, awkward controls, or a snowboard sinking into the terrain become obvious within seconds. That's why I like using games to evaluate coding models. I gave Kimi K3 a single prompt to build a snowboard game in Godot, and just 20 seconds of gameplay told me more than a page of benchmark scores ever could. Benchmarks matter, but watching a model build something interactive reveals a different side of its capabilities. If the gameplay feels right, the systems work together, and the experience holds up under real input, that's a much more convincing demonstration than any leaderboard.show more

FHILY👑
36,766 просмотров • 1 месяц назад
🤯KIMI K3 ABSOLUTELY MOGS! BEATING Opus 4.8, GPT 5.5,... and even Fable 5 in multiple benchmarks. They scored 1688 on GDPval-AA v2 🔥 This is a completely different breed of open-source models Kimi creates better games and front-end designs than Fable 5, but it's 8x cheaper! The results coming out are truly impressive! I will be testing this out further and posting multiple tests today. Stay tuned!show more

Mark Santos
128,094 просмотров • 1 месяц назад
BREAKING: Anthropic just dropped Opus 4.8—and it is a... MONSTER We've been testing for about a week Every 🪨 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 🪨:show more

Dan Shipper
354,649 просмотров • 3 месяцев назад
$35 OF FREE CLAUDE OPUS 5 CREDITS FOR A... WEEK, NO CARD ANYWHERE • the credit > Sign up with Google and the trial activates on its own, nothing to claim manually: > Opus 5 sits alongside Opus 4.8, Sonnet 4.6 and Haiku on the same key. > The window runs a week from activation. • wiring it up > base url: > model id: claude-opus-5 > In OpenCode Desktop hit Ctrl+, and add a custom provider with the id aerolink. > The same key drops into Codex, Cursor and Claude Code with nothing else changed. One key, not locked to one client -> every tool that takes a custom OpenAI-compatible endpoint reads it the same way. It routes through a third-party proxy rather than an official Anthropic channel, so keep client work off it ↓show more

slash1s
20,882 просмотров • 25 дней назад
We are getting absurdly close to the point where... “learning Blender” means learning how to direct an AI. A shot like this looks soft and playful on the surface, but under it is the usual 3D pain: modeling, layout, materials, lighting, atmosphere, animation, and endless tiny fixes until the frame stops looking dead. That is why Kimi K3 matters. With Blender MCP, you can describe a scene like a robotic goat walking through a dreamy field and let the model help build the environment, place the camera, shape the materials, script the motion, and iterate inside the real Blender project. The real shift is not text-to-image. It is text-to-workflow. Kimi K3 does not just give you a pretty output and disappear. It can help move the actual scene from rough setup to something that looks art-directed. Soon the hardest part of 3D will not be the software. It will be whether your imagination is good enough to deserve tools like this.show more

Rina
45,651 просмотров • 1 месяц назад
This guy used Opus 5 to build a tiny... device that replaced scrolling for him. Imagine a small screen, barely bigger than a watch, with hundreds of glowing blue particles floating inside. When you tilt it, they slide. When you shake it, they splash. When you hold it still, they settle like water at the bottom of a glass. It looks like actual liquid is trapped behind the screen. But it's all code running on a $15 chip. Opus 5 wrote the entire thing in C. The physics, the lighting, the way each particle reacts to your hand movements. 100 fps on a device smaller than your palm. He says he picks it up instead of reaching for his phone now. Just sits there tilting it, watching particles move.show more

Vaibhav Sisinty
120,554 просмотров • 1 месяц назад
Opus 4.6 vs GPT 5.4 (High) (1/9) prompt: Build... a single-file HTML/CSS/JS (no libs) demo that uses SVG to simulate a plant growing: stem extends, leaves sprout + unfurl with springy/windy “physics”, then seamlessly loops forever. For the initial impressions I'm really impressed by GPT, for speed they both felt about the same, but gpt was still half cheaper than opus. I also much prefer the design and animation that gpt produced, physics on the leaves are super cool and it also loops pretty nicely whilst opus just fades out the plant. Still got a bunch of tests to run but this is really exciting.show more

Dev Ed
660,154 просмотров • 6 месяцев назад
Fable 5 totally crushed our new contest, but it... cost 6x more than Opus 4.8! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: — A train derailing off a broken bridge into the water — Two cars jumping off ramps and colliding mid-air over a canyon — A monster truck crushing a row of parked cars Outputs: Fable 5: 62,158 tokens, $3.12 GPT 5.5: 37,753 tokens, $1.14 Opus 4.8: 22,280 tokens, $0.56 GLM 5.2: 36,246 tokens, $0.08 Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.show more

atomic.chat
2,840,649 просмотров • 2 месяцев назад
Blender just became a prompt box. Kimi K3 +... Blender MCP can take a simple text prompt and build an entire 3D scene for you. terrain, buildings, lighting, materials, camera movement, animation, even the Python scripts holding everything together. but the crazy part isn’t just that it can build the scene. Kimi can inspect what it created, understand what looks wrong, then go back into Blender and fix the actual scene before rendering again. the camera can be clipping through a tree, the lighting can feel off, the city can look too artificial, and instead of starting over, Kimi can make the changes directly inside Blender. and because everything is happening in Blender, you’re not left with a locked AI-generated video. you get the actual 3D scene. every object, material, light, camera and keyframe stays editable. a few years ago, creating something like this could take weeks of modeling, scripting, lighting and animation. now you can describe the world, let Kimi build the first 90%, and spend your time making the final 10% actually look good. that’s what makes Kimi K3 + Blender MCP so interesting. it’s getting dangerously close to text-to-3D, except the result is a real Blender project you can keep editing.show more

MIKE
10,375 просмотров • 29 дней назад
After a few more hours, I think I've figured... out Opus 5. Opus 5 is trained to be more agentic than anything I've used. All Claude 5 models are like that. So what changes? The way to interact with Opus 5 or contextualize it won't work the same way as with other models. It loves exploring, so it doesn't need much guidance for it. Unique preferences, artifacts, and references compliment it well and enable cleaner and more effective exploration and execution. Now that it can explore more effectively on its own and understand intent better, the best thing to do is to get out of its way (e.g., it doesn't need examples of your preferences; a clear high-level description of it works best). It's truly agentic in that sense. A good first step to provide better context for Opus 5 is to distinguish between what's situational and what needs persistence. Regardless, persistent system prompts and CLAUDE.MD needs to stay lightweight. Remove memories and tool descriptions from these. CLAUDE.MD is also a great place to tap into progressive disclosure by linking command/skills to it. On the situational side, agent skills and auto-memory can leverage progressive disclosure and the improved ability of the model to use its external context/knowledge. Conflicting and unnecessary instructions, which are common at this layer (mainly to ensure reliability), are going to throw off this model easily. That's the biggest change I had to make. Simple, clean, and clear prompts and skills work best. I had to clean a lot of my skills and system prompts. The way I prompt remains the same (usually clear and well-scoped). MCP tool descriptions are also more descriptive and have been deduped from the system prompt. Anthropic released a guide on the new rules for context engineering, which was helpful here. I started to test the recommendations and created a little artifact with the things that worked along the way. This might feel like a lot of work. Believe me, it has been frustrating. But I think we can expect future frontier models to become more agentic and smarter at figuring out the right context/gaps. The best thing to do is to prepare for that now. Boris Cherny mentioned that Opus 5 is their least prompt-injectable model yet. I am not sure if that was something they intentionally trained for or if it emerged based on how it was trained, which is to be extremely agentic in nature and more direct in execution.show more

elvis
37,757 просмотров • 1 месяц назад
HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE... 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5show more

YanXbt
16,744 просмотров • 1 месяц назад