LongCat performed Opus 4.8 and GPT 5.5 level on... real physics tasks for $0! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics Prompts: - A cannon demolishing a brick wall - A bowling ball knocking down the pins - A tornado that sucks in random objects Outputs: LongCat: 18,015 tokens, $0.00 Opus 4.8: 18,872 tokens, $0.48 GPT 5.5: 32,588 tokens, $0.98 GLM 5.2: 31,062 tokens, $0.09 On the physics LongCat came out ahead of Opus 4.8 and GLM 5.2 - cleaner collisions, nothing clipping or falling through. On detail and rendering it matched GPT 5.5, the best looking of the four. Getting this quality for free is wild!show more

atomic.chat
105,523 görüntüleme • 2 ay önce
New Claude Sonnet 5 performs at GPT 5.5 level... 6x cheaper! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics crash demos Prompts: - A car crashes into a brick wall - A wrecking ball destroys a house - A catapult throws a rock at a castle wall Outputs: Sonnet 5: 15,047 tokens, $0.15 Opus 4.8: 23,063 tokens, $0.58 Sonnet 4.6: 25,824 tokens, $0.39 GPT 5.5: 31,152 tokens, $0.94 Sonnet 5 did as well as Opus 4.8 and GPT 5.5 on all three tests. In the wrecking ball test, it beat Opus 4.8. The cable moves smoothly and every hit connects. In the catapult test, it beat GPT 5.5. The rock always lands inside the wall. Sonnet 5 still needs better detail and graphics. But it used fewer tokens than every other modelshow more

atomic.chat
728,593 görüntüleme • 2 ay önce
Open-weight LongCat 2.0 matched GPT-5.5 level on agentic game... dev for $0! We ran Meituan's LongCat 2.0 against cloud frontier GPT-5.5 in Kilo CLI with their agent. Same task for both - build a retro Duck Hunt game in one game.html, improved over 3 agent iterations with duck waves, ammo and physics Outputs: LongCat 2.0: 70.3K tokens, $0.00 GPT-5.5: 64.9K tokens, $0.65 LongCat kept up on graphics, physics and game logic. Ducks fly and fall when hit, the dog fetches them, ammo counts down, the waves keep coming. Both ran clean and nothing clipped. The only difference was the bill - GPT cost $0.65, LongCat ran local for $0show more

atomic.chat
40,046 görüntüleme • 1 ay önce
Grok 4.5 performed GPT Sol level for free! We... gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: -robot deathmatch, Tombstone vs Minotaur -a hydraulic press flattening stuff on a conveyor -a semi truck jumping a canyon Outputs: GPT-5.6 Sol: 12.9K tokens, $0.51 (~7 min) Grok 4.5: 10.8K tokens, $0 (~5 min) Muse Spark 1.1: 26.8K tokens, $0.12 (~7.5 min) GLM 5.2: 10.9K tokens, $0.02 (~12 min) Grok 4.5 handled all three scenes genuinely well and got surprisingly close to GPT-5.6 this round. On top of that, it ran on the free tier. GPT-5.6 Sol, the frontier model, put out solid but not standout work. GLM 5.2 rendered all three scenes for pennies, but it came out the roughest of the four. Meta's new Muse Spark burned the most tokens yet still stayed cheap, delivering an average result.show more

atomic.chat
70,490 görüntüleme • 1 ay önce
New Hunyuan Hy3 hits Gemini 3.5 quality on physics... for 35x cheaper! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: - A bowling ball knocking down the pins - An air hockey rally that ends in a goal - A pool break scattering the rack Outputs: Hunyuan Hy3: 29,757 tokens, $0.006 Gemini 3.5: 23,300 tokens, $0.21 GLM-5.2: 25,454 tokens, $0.07 DeepSeek-V4: 50,600 tokens, $0.009 Tencent's Hy3 matched Gemini across all three: clean collisions, the puck bounced true, the pins scattered like a real strike, the rack broke with real momentum, nothing clipped or floated. GLM is genuinely strong on pure coding tasks, but the moment the job steps outside clean code it gives way. DeepSeek was the letdown, it burned the most tokens of anyone (50k, almost 2x Hy3) and still turned in the weakest scenesshow more

atomic.chat
96,930 görüntüleme • 2 ay önce
Nemotron 3 Ultra performed GPT 5.5 level 10× cheaper... We gave three same prompts to build HTML5 canvas with real physics. At first scene we have water in a spinning drum. Galton board - balls through pegs into bins. And a block collision setup with extreme mass differences. Outputs: Nemotron 3 Ultra: 11.3k tokens, $0.051 GPT 5.5: 11.0k tokens, $0.57 Nemotron stays right on GPT 5.5's heels, but at 10× cheaper. The gap in quality is far smaller than the gap in price.show more

atomic.chat
701,686 görüntüleme • 3 ay önce
Announcing GLM Arena! A series of tests (infographics, svgs,... sites, ect..) ran on GLM 5.2 and Opus 4.8, with prompts included. On average, GLM 5.2 produced 2x the tokens but was still faster + 3x cheaper with similar quality.show more

Hassan
75,421 görüntüleme • 2 ay önce
Sakana Fugu surprisingly performed near GLM 5.2 level but... 17× more expensive! We gave the same prompt to 4 models: build a complete live Trader Desk with both frontend and backend components, real-time market data fetched from external APIs for 8 symbols, and a custom dark-theme UI. Outputs: Fugu Ultra — 22,225 t, $0.51 Opus 4.8 — 15,802 t, $0.31 GPT-5.5 — 11,474 t, $0.26 GLM 5.2 — 13,677 t, $0.03 Fugu created the most polished and feature-rich trading desk in the run. GLM 5.2 was very close behind, with a similarly complete multi-panel interface and live data, but at a much lower cost. Opus and GPT also performed well, delivering solid results with a better balance between quality and costshow more

atomic.chat
744,642 görüntüleme • 2 ay önce
1-bit Kimi K3 performs at Opus 5 level on... 3D physics! We ran our Atomic Chat quant of Kimi K3 locally on 4x B200 against three cloud models and gave them all the same task, to build a giant anvil drop test as a single HTML file with real physics Outputs: K3 1bit (local): 15.8K tokens, $0 API cost Kimi K3 (API): 15.3K tokens, $0.30 API cost Opus 5: 22.8K tokens, $0.77 API cost GPT 5.6: 14.5K tokens, $0.72 API cost All four got the physics right. But only Kimi made a working winch. The drum turns and the chain drags the flat car off the pad. Opus 5 drew the most detail, road markings and sparks on the hit. And you can run a model at this level on your own box now. That still feels insane to usshow more

atomic.chat
53,845 görüntüleme • 1 ay önce
I tested Kimi K3 vs Claude Opus 4.8 Same... prompt, an armory bay with lighting, props, and detail. Top is Kimi K3, bottom is Opus 4.8. It's not even close. Kimi K3 built a full scene with textures, proper lighting, ammo crates, weapon racks, working detail everywhere. Opus 4.8 gave me a near empty room with a couple of floating tables. No doubt it beats Opus 4.8. Kimi K3 is Fable 5 level, and it's clearly better than GPT-5.6 Sol at 3D and games. An open weight model just matched the best closed models on the market. Let that sink in.show more

Bhavy☄️
476,936 görüntüleme • 1 ay önce
Laguna S 2.1 performs at GLM-5.2 level on building... popular games with 6x fewer params! We gave three local models the same task: build three popular arcade games that play themselves. Each game is one self-contained HTML file with a bot that plays it. Prompts: – Geometry Dash – Doodle Jump – Air Hockey Outputs: Laguna S 2.1: 10.3K tokens GLM-5.2: 26.4K tokens Hy3: 10.4K tokens Laguna held its quality against a 753B model. We think Laguna's Geometry Dash looked the best of the three, the cube clears every spike and block. GLM won Air Hockey. Its table looked the most detailed of all. Hy3 was the only model that added shooting to its Doodle Jump. But Laguna is the only model in our benchmark that runs on a MacBook with 128GB!show more

atomic.chat
66,917 görüntüleme • 1 ay önce
BREAKING: Anthropic just dropped Opus 4.8—and it is a... MONSTER We've been testing for about a week Every 🪨 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 🪨:show more

Dan Shipper 📧
354,559 görüntüleme • 3 ay önce
HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE... 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5show more

YanXbt
16,744 görüntüleme • 1 ay önce
Opus 4.6 vs GPT 5.4 (High) (1/9) prompt: Build... a single-file HTML/CSS/JS (no libs) demo that uses SVG to simulate a plant growing: stem extends, leaves sprout + unfurl with springy/windy “physics”, then seamlessly loops forever. For the initial impressions I'm really impressed by GPT, for speed they both felt about the same, but gpt was still half cheaper than opus. I also much prefer the design and animation that gpt produced, physics on the leaves are super cool and it also loops pretty nicely whilst opus just fades out the plant. Still got a bunch of tests to run but this is really exciting.show more

Dev Ed
660,154 görüntüleme • 6 ay önce
Salute to the Qwen team 🫡 We tested Qwen... 3.7-Max, Gemini 3.5 Flash, GPT-5.5, and Claude Opus 4.7. The biggest shock came from Qwen. In less than a month (3.6 Max dropped April 20), Qwen went from the worst multimodal output on our sakura tree test, barely keeping up with Gemini, GPT, and Claude , to matching Gemini 3.5 Flash frame for frame on this soccer test, and outperforming GPT-5.5 and Claude Opus 4.7. It rendered a perfectly proportioned soccer player and the most lifelike ball in the entire test. Remarkable spatial reasoning. Also: Gemini 3.5 Flash is now faster than GPT-5.5, which used to be the fastest in our past tests.show more

GMI Cloud
75,785 görüntüleme • 3 ay önce
You can now use GPT 5.5, Gemini 3.7 Flash,... Kimi K3 and 47 other AI models completely free😱 No subscription. No credit card. Even the API usage costs $0. AIHubMix just opened a free catalog with 50 AI models. Some of the available models: • Ox Alpha • Gemini 3.7 Flash • GLM 5.2 • Kimi K3 • MiniMax M3 • GPT 5.5 • 40+ more And you don’t need separate API keys for each model. Setup takes 2 minutes: > Step 1: Go to > Create an account using your email or OAuth. No card needed. Step 2: Create one API key > The same key works with every free and paid model. Step 3: Add it to any OpenAI-compatible tool Base URL: Then choose any model ending in -free, such as: coding-glm-5.2-free gpt-5.5-free That’s it. One API key. 50 AI models. $0 for both input and output. Save this. You might need a free multi-model setup later.show more

CDG
15,278 görüntüleme • 11 gün önce
🤯KIMI K3 ABSOLUTELY MOGS! BEATING Opus 4.8, GPT 5.5,... and even Fable 5 in multiple benchmarks. They scored 1688 on GDPval-AA v2 🔥 This is a completely different breed of open-source models Kimi creates better games and front-end designs than Fable 5, but it's 8x cheaper! The results coming out are truly impressive! I will be testing this out further and posting multiple tests today. Stay tuned!show more

Mark Santos
127,993 görüntüleme • 1 ay önce
I can't believe this is real I have GLM... 5.2 running 100% locally on my Mac Studio. 2 bit quant. The results I'm getting are better than Opus 4.8 It's now powering my Hermes Agent and Codex. 100% free, local, private super intelligence on my desk I also have it in a loop coding for me 24/7 now I thought we were at least a year away from this type of event. It happened today. The model takes up about 250gb of memory. So you can technically run it on a Mac Studio with 256gb, but you probably want the 512gb memory version (please tell me you listened to me 5 months ago when these were sitting on store shelves) With Fable gone, I now have Opus 4.8 level intelligence on my desk for free. This is the future. Local, private, secure, personal super intelligence. If you're still writing off local AI as a fad or engagement bait, you are officially delusionalshow more

Alex Finn
622,922 görüntüleme • 2 ay önce
Opus 4.6 vs GPT-5.4 (4/9) prompt: Build a production-quality... 3D flight-tracking web app using React + Vite + Three.js (react-three-fiber + drei) that visualizes live OpenSky aircraft data on a rotatable 3D Earth, with real-time plane motion, smooth interpolation, altitude-accurate positioning, and polished lighting/post-processing. Both models did really well on this one and honestly I’m impressed with both. GPT-5.4 had the nicer post-processing out of the box. I really liked the subtle light shimmer on the airplanes when rotating the planet, and the camera work when clicking a plane felt better overall. Opus 4.6 though had a few details I liked more. It automatically went and found a much nicer Earth texture on GitHub, while with GPT-5.4 I had to reprompt it to go look for a better one. I also preferred Opus’s plane model overall, it just looked more polished, whereas GPT-5.4’s plane asset looked a bit funny. One thing I noticed with GPT-5.4 is that when you click into the plane, the camera sometimes clips through the planet, which breaks the effect a bit. Opus handled that part more cleanly. Overall this felt like a strong result from both, just with different strengths. GPT-5.4 felt better on presentation and post-processing, while Opus had better asset choices and a more premium-looking Earth/plane combo.show more

Dev Ed
279,565 görüntüleme • 6 ay önce
Codex can now run Deepseek-v4- flash! There's a catch... though. Deepseek's official setup switches your entire codex over to them, so your GPT models stop showing up at all. This is exactly what Codex Router is for. It adds models to the list instead of replacing them, so sol, grok, kimi and deepseek all sit in the same picker and i just grab whichever one suits the job. Deepseek v4-flash is $0.28 per million output tokens. opus 4.8 is $25. same picker, 89x apart. Links in the comment. setup's in the video 👇show more

Ziwen
145,018 görüntüleme • 1 ay önce
a moonshot engineer leaked the benchmark anthropic, openai and... xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article belowshow more

starmex
32,547 görüntüleme • 14 gün önce