1-bit Kimi K3 performs at Opus 5 level on... 3D physics! We ran our Atomic Chat quant of Kimi K3 locally on 4x B200 against three cloud models and gave them all the same task, to build a giant anvil drop test as a single HTML file with real physics Outputs: K3 1bit (local): 15.8K tokens, $0 API cost Kimi K3 (API): 15.3K tokens, $0.30 API cost Opus 5: 22.8K tokens, $0.77 API cost GPT 5.6: 14.5K tokens, $0.72 API cost All four got the physics right. But only Kimi made a working winch. The drum turns and the chain drags the flat car off the pad. Opus 5 drew the most detail, road markings and sparks on the hit. And you can run a model at this level on your own box now. That still feels insane to usshow more

atomic.chat
53,845 görüntüleme • 1 ay önce
I gave Kimi K3 and Opus 5 the exact... same prompt to build a car racing game. I expected Opus 5 to take the lead. Instead, Kimi K3 finished in ~25 min for ~$8.9, while Opus 5 took 19 min and cost $14. The biggest surprise wasn’t the speed or the price. Kimi K3 built the better game. Smoother controls, better physics, and a driving experience that felt ready from the start. Opus 5 looked cleaner visually but Kimi won where it actually mattered. Nearly half the cost and the better gameplay.show more

Aman
111,571 görüntüleme • 1 ay önce
I tested Kimi K3 vs Claude Opus 4.8 Same... prompt, an armory bay with lighting, props, and detail. Top is Kimi K3, bottom is Opus 4.8. It's not even close. Kimi K3 built a full scene with textures, proper lighting, ammo crates, weapon racks, working detail everywhere. Opus 4.8 gave me a near empty room with a couple of floating tables. No doubt it beats Opus 4.8. Kimi K3 is Fable 5 level, and it's clearly better than GPT-5.6 Sol at 3D and games. An open weight model just matched the best closed models on the market. Let that sink in.show more

Bhavy☄️
476,936 görüntüleme • 1 ay önce
Kimi K3 beats GPT 5.6 Sol on both speed... and in cost in backend bug fixing ! Planted 4 defects in the same Python order service: a wrong-total calculation, a broken access-control check (IDOR), a crash on invalid input, and an inventory oversell. Both models got identical code and the same prompt, with an automated test suite verifying every fix over the live API. Both went 4/4, but by the time Kimi K3 finished all 4, GPT 5.6 had only fixed 2. Kimi wrapped in under half the time, at half the cost. Kimi K3: 1min 34s · $0.1 · 4/4 GPT 5.6 Sol: 4min 6s · $0.2 · 4/4show more

GMI Cloud
24,671 görüntüleme • 1 ay önce
a moonshot engineer leaked the benchmark anthropic, openai and... xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article belowshow more

starmex
32,547 görüntüleme • 14 gün önce
New Claude Sonnet 5 performs at GPT 5.5 level... 6x cheaper! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics crash demos Prompts: - A car crashes into a brick wall - A wrecking ball destroys a house - A catapult throws a rock at a castle wall Outputs: Sonnet 5: 15,047 tokens, $0.15 Opus 4.8: 23,063 tokens, $0.58 Sonnet 4.6: 25,824 tokens, $0.39 GPT 5.5: 31,152 tokens, $0.94 Sonnet 5 did as well as Opus 4.8 and GPT 5.5 on all three tests. In the wrecking ball test, it beat Opus 4.8. The cable moves smoothly and every hit connects. In the catapult test, it beat GPT 5.5. The rock always lands inside the wall. Sonnet 5 still needs better detail and graphics. But it used fewer tokens than every other modelshow more

atomic.chat
728,593 görüntüleme • 2 ay önce
LongCat performed Opus 4.8 and GPT 5.5 level on... real physics tasks for $0! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics Prompts: - A cannon demolishing a brick wall - A bowling ball knocking down the pins - A tornado that sucks in random objects Outputs: LongCat: 18,015 tokens, $0.00 Opus 4.8: 18,872 tokens, $0.48 GPT 5.5: 32,588 tokens, $0.98 GLM 5.2: 31,062 tokens, $0.09 On the physics LongCat came out ahead of Opus 4.8 and GLM 5.2 - cleaner collisions, nothing clipping or falling through. On detail and rendering it matched GPT 5.5, the best looking of the four. Getting this quality for free is wild!show more

atomic.chat
105,523 görüntüleme • 2 ay önce
KIMI K3 VS OPUS 4.8 SIMULATING A 3D PARTICLE... SYSTEM same prompt, same target: thousands of particles reacting to gravity and to each other in real time → K3: more organic distribution, smoother motion, particles actually cluster and drift like something physical → Opus: more rigid pattern, less natural, movement reads more like a grid than a simulation K3: $0.65 in API. Opus 4.8: ~$1.30 this is simulated physics, not just aesthetics, and it's the kind of detail that separates a demo from something you'd actually shipshow more

Jade
32,766 görüntüleme • 1 ay önce
Fable 5 absolutely crushed the HTML5 physics contest, but... cost 6x more than Opus 4.8 and 39× more than GLM 5.2 in that test. Test was done on atomic[.]chat, a desktop app that runs LLMs locally. The test asked 4 models to generate self-contained canvas demos with believable motion and collisions. The scenes were not simple animations because every crash needed gravity, force, timing, and contact handling. Outputs: - Fable 5: 62,158 tokens, $3.12 - GPT 5.5: 37,753 tokens, $1.14 - Opus 4.8: 22,280 tokens, $0.56 - GLM 5.2: 36,246 tokens, $0.08show more

Rohan Paul
205,771 görüntüleme • 2 ay önce
Qwen 3.8 Max is actually impressive Sent a sky... pic to it along with Claude Opus 5, Kimi K3, and GPT 5.6 Sol, and asked them to draw an animal silhouette based on the cloud shape Qwen 3.8 Max is honestly on par with all the other modelsshow more

Ann Nguyen
471,057 görüntüleme • 1 ay önce
KIMI K3 JUST RECREATED A FUNCTIONAL GAME BOY ADVANCE... IN 3D → interactive 3D model, every button and screen modeled → a game loaded and playable right inside it → camera you can rotate and zoom around the device → one single prompt, zero external assets $3 in API with K3. the same result would run about $9 with Fable 5 and $6 with Opus 4.8 this isn't generating a static image of hardware anymore, it's generating a working simulation of it, from the shell to the screen to the game running insideshow more

alex
38,248 görüntüleme • 1 ay önce
Fable 5 totally crushed our new contest, but it... cost 6x more than Opus 4.8! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: — A train derailing off a broken bridge into the water — Two cars jumping off ramps and colliding mid-air over a canyon — A monster truck crushing a row of parked cars Outputs: Fable 5: 62,158 tokens, $3.12 GPT 5.5: 37,753 tokens, $1.14 Opus 4.8: 22,280 tokens, $0.56 GLM 5.2: 36,246 tokens, $0.08 Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.show more

atomic.chat
2,839,437 görüntüleme • 2 ay önce
Grok 4.5 performed GPT Sol level for free! We... gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: -robot deathmatch, Tombstone vs Minotaur -a hydraulic press flattening stuff on a conveyor -a semi truck jumping a canyon Outputs: GPT-5.6 Sol: 12.9K tokens, $0.51 (~7 min) Grok 4.5: 10.8K tokens, $0 (~5 min) Muse Spark 1.1: 26.8K tokens, $0.12 (~7.5 min) GLM 5.2: 10.9K tokens, $0.02 (~12 min) Grok 4.5 handled all three scenes genuinely well and got surprisingly close to GPT-5.6 this round. On top of that, it ran on the free tier. GPT-5.6 Sol, the frontier model, put out solid but not standout work. GLM 5.2 rendered all three scenes for pennies, but it came out the roughest of the four. Meta's new Muse Spark burned the most tokens yet still stayed cheap, delivering an average result.show more

atomic.chat
70,490 görüntüleme • 1 ay önce
Open-weight LongCat 2.0 matched GPT-5.5 level on agentic game... dev for $0! We ran Meituan's LongCat 2.0 against cloud frontier GPT-5.5 in Kilo CLI with their agent. Same task for both - build a retro Duck Hunt game in one game.html, improved over 3 agent iterations with duck waves, ammo and physics Outputs: LongCat 2.0: 70.3K tokens, $0.00 GPT-5.5: 64.9K tokens, $0.65 LongCat kept up on graphics, physics and game logic. Ducks fly and fall when hit, the dog fetches them, ammo counts down, the waves keep coming. Both ran clean and nothing clipped. The only difference was the bill - GPT cost $0.65, LongCat ran local for $0show more

atomic.chat
40,046 görüntüleme • 1 ay önce
You can now use GPT 5.5, Gemini 3.7 Flash,... Kimi K3 and 47 other AI models completely free😱 No subscription. No credit card. Even the API usage costs $0. AIHubMix just opened a free catalog with 50 AI models. Some of the available models: • Ox Alpha • Gemini 3.7 Flash • GLM 5.2 • Kimi K3 • MiniMax M3 • GPT 5.5 • 40+ more And you don’t need separate API keys for each model. Setup takes 2 minutes: > Step 1: Go to > Create an account using your email or OAuth. No card needed. Step 2: Create one API key > The same key works with every free and paid model. Step 3: Add it to any OpenAI-compatible tool Base URL: Then choose any model ending in -free, such as: coding-glm-5.2-free gpt-5.5-free That’s it. One API key. 50 AI models. $0 for both input and output. Save this. You might need a free multi-model setup later.show more

CDG
15,278 görüntüleme • 11 gün önce
HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE... 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5show more

YanXbt
16,744 görüntüleme • 1 ay önce
Laguna S 2.1 performs at GLM-5.2 level on building... popular games with 6x fewer params! We gave three local models the same task: build three popular arcade games that play themselves. Each game is one self-contained HTML file with a bot that plays it. Prompts: – Geometry Dash – Doodle Jump – Air Hockey Outputs: Laguna S 2.1: 10.3K tokens GLM-5.2: 26.4K tokens Hy3: 10.4K tokens Laguna held its quality against a 753B model. We think Laguna's Geometry Dash looked the best of the three, the cube clears every spike and block. GLM won Air Hockey. Its table looked the most detailed of all. Hy3 was the only model that added shooting to its Doodle Jump. But Laguna is the only model in our benchmark that runs on a MacBook with 128GB!show more

atomic.chat
66,917 görüntüleme • 1 ay önce
KIMI K3 + OBSIDIAN + LOOP ENGINEERING = A... VAULT THAT RUNS ITSELF the core idea: the vault is the loop's state, not the chat window everything Kimi K3 knows lives in a .md file the loop: > capture - a thought lands in 00-inbox > context - K3 pulls links, tags, and neighbouring notes > draft - edits happen inside a git worktree, never the live vault > review - a critic agent checks the diff before anything ships > commit - appended to the vault, nothing gets rewritten the key insight: frontmatter fields like supports, contradicts, and supersedes are graph edges, not metadata - the note format is the write API start with a plain loop, it runs about 2-4x the cost of one direct call > only move to a full graph once state has to outlive the session, several agents need to coordinate, or you have to explain what changed - that jump can run 10-50x one review assistant climbed from 55% to 72% to 84% just by moving through these shapes in ordershow more

Mr. Buzzoni
10,680 görüntüleme • 15 gün önce
Kimi K3 vs GPT-5.6 vs Claude Fable 5 A... game is one of the hardest places to fake working code. A web app can hide bugs behind a polished screenshot. A game can't. Bad collision, clipping, jitter, broken physics, awkward controls, or a snowboard sinking into the terrain become obvious within seconds. That's why I like using games to evaluate coding models. I gave Kimi K3 a single prompt to build a snowboard game in Godot, and just 20 seconds of gameplay told me more than a page of benchmark scores ever could. Benchmarks matter, but watching a model build something interactive reveals a different side of its capabilities. If the gameplay feels right, the systems work together, and the experience holds up under real input, that's a much more convincing demonstration than any leaderboard.show more

FHILY👑
36,569 görüntüleme • 1 ay önce
OpenCode Go is now wired into Codex!! The pricing... is insane. $10 gets you 10,000 DeepSeek requests every 5 hours (no weekly limits). I converted that into DeepSeek API dollars because I thought I was reading it wrong, and the same 5 hours of usage would run somewhere between $10 and $30 depending on how big your context gets. So one afternoon of use already covers the whole sub. It comes with Kimi K3 as well. Both sit in the picker next to my other models now. Next time we hit a limit in the middle of a loop, we can grab it and keep going.show more

Ziwen
431,895 görüntüleme • 1 ay önce
Opus 4.6 vs GPT 5.4 (High) (1/9) prompt: Build... a single-file HTML/CSS/JS (no libs) demo that uses SVG to simulate a plant growing: stem extends, leaves sprout + unfurl with springy/windy “physics”, then seamlessly loops forever. For the initial impressions I'm really impressed by GPT, for speed they both felt about the same, but gpt was still half cheaper than opus. I also much prefer the design and animation that gpt produced, physics on the leaves are super cool and it also loops pretty nicely whilst opus just fades out the plant. Still got a bunch of tests to run but this is really exciting.show more

Dev Ed
660,154 görüntüleme • 6 ay önce