MiniMax M3: Opus-level coding at DeepSeek pricing. On Terminal-Bench... 2.1, it scores 66.0, only 0.1 behind Opus 4.7. I gave it a quick try on a few frontend tasks, and the output quality genuinely feels close to Opus 4.7.show more

Kai
25,730 просмотров • 2 месяцев назад
My first time building with Opus 4.7 and I... love it! I quickly built this Base AI Ecosystem map with names, links logos in the form of a galaxy that you can navigate through 🛰️ With some rigorous prompting, Opus 4.7 did everything for me including descriptions and logos!show more

Youssef
17,112 просмотров • 3 месяцев назад
we compared Gemini 3.1 Pro, Opus 4.7, and GPT... 5.5 to Kimi K2.6, Xiaomi Mimo v2.5, and Qwen 3.6 Max in average, closed source are faster. GPT 5.5 and Opus 4.7 the fastest. Kimi K2.6 comes after. It keeps up thanks to its native INT4 quantization. MiMo took the longest, but the refinement and aesthetic only ranks after Gemini 3.1. as if a mini Gemini. Its slow cuz it's trained for long-horizon agentic work. Claude Opus 4.7 is blooming in a very different style, probably because Anthropic trains for taste, not just accuracy.show more

GMI Cloud
37,446 просмотров • 2 месяцев назад
We’re excited to add support for Opus 4.6. With... adaptive thinking, Opus 4.6 will decide the right amount of reasoning for your query, reducing latency and improving quality. Internally, we've already resolved bugs that Opus 4.5 couldn't tackle. Try it out on the latest!show more

Warp
12,404 просмотров • 6 месяцев назад
Kimi K3 ranks #6 on ReactBench - It performs... better than Opus 4.8 while being over 2x cheaper (!!) - Scores 33%, jump of +10% compared to K2.7 - Lags behind the frontier (Fable and Sol) in frontend production-readiness:show more

Aiden Bai
73,851 просмотров • 16 дней назад
Opus 5 appears to be a top performing model... at agentic CAD design! It even outperforms Fable on a number on a number of mechanical design tasksshow more

adam
102,255 просмотров • 12 дней назад
HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE... 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5show more

YanXbt
16,744 просмотров • 11 дней назад
BREAKING: Anthropic just dropped Opus 4.8—and it is a... MONSTER We've been testing for about a week Every 📧 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 📧:show more

Dan Shipper 📧
354,163 просмотров • 2 месяцев назад
An OpenAI engineer stopped me at a hackathon in... Hayes Valley I had my terminal open on a table. Three panels. Live trades scrolling. He was walking past and froze. "That's not a demo. That's a live scoring engine. What model is that" I told him. Claude Opus 4.7. Four repos. $25 a month. He pulled up a chair without asking. "We benchmarked Opus 4.7 internally. It beat o3 on structured reasoning across every eval we ran. And you're telling me you're using it to trade" I told him it does more than trade. It reads 86 million trades and finds who wins and why. No fine-tuning. No prompting chains. Just raw context. He leaned back. "Show me the data source" I opened one link. 86 million trades. Every wallet. Every entry. Every exit. "You point Opus 4.7 at this and it reverse-engineers the strategy. It finds the wallets that win. Then it finds why they win. Then it copies the pattern" His team spent 14 months building something similar. 10 engineers. Custom infra. Still in staging. "The part that killed us was exit timing. Every model we trained nailed entries. But the best traders exit before the crowd. We never figured out the threshold" I told him my bot cuts at 85% of expected move. Or on a 3x volume spike. Whichever comes first. He stopped talking. "How did you find that" Opus 4.7 found it in poly_data. Top wallets exit before resolution 86% of the time. Losers hold to 58%. Exits are the entire game. I opened another tab. "Three commands. 500 markets. Opus scores them in 20 minutes" "That's our internal eval pipeline. Except it took us a year and you did it in a weekend with our competitor's model" My setup: Claude Opus 4.7 - $20/mo VPS - $5/mo poly_data - free polymarket-cli - free 214 trades. 74% win rate. +$9,400 in 19 days. Copytrade here: I showed him the article where I broke down every repo and every command. He read it twice. Then looked up. "You just published what we've been trying to ship for six months. Using the other team's model" He texted me the next day. "My manager found your thread. Delete it" Too late.show more

Lunar
136,549 просмотров • 3 месяцев назад
DeepSeek V4 Flash 0731 is now 90% off on... Nous Portal for the next 7 days, in partnership with Novita AI. At this discounted price, it is over 1000x cheaper than Fable 5 on comparable tasks while still beating it on Terminal-Bench 2.1. Try it today atshow more

Nous Research
834,021 просмотров • 3 дней назад
I tested Kimi K3 vs Claude Opus 4.8 Same... prompt, an armory bay with lighting, props, and detail. Top is Kimi K3, bottom is Opus 4.8. It's not even close. Kimi K3 built a full scene with textures, proper lighting, ammo crates, weapon racks, working detail everywhere. Opus 4.8 gave me a near empty room with a couple of floating tables. No doubt it beats Opus 4.8. Kimi K3 is Fable 5 level, and it's clearly better than GPT-5.6 Sol at 3D and games. An open weight model just matched the best closed models on the market. Let that sink in.show more

Bhavy☄️
468,355 просмотров • 20 дней назад
the real reason why SF is 6 months ahead... on AI is just physics: shipping a SOTA model from the bay area to NYC is a logistical nightmare imagine what it took to bring Opus 4.7 to Europe (we don't have trains here) 😢show more

fabian
1,850,659 просмотров • 2 месяцев назад
Opus 4.7 can build Lottie Animations. One prompt via... Lottie Creator MCP → 500 particles, each with its own path, easing, and arrival frame. I didn't touch a keyframe. What should I ask it to build next? Best reply, I'll make it.show more

Nattu
186,829 просмотров • 3 месяцев назад
Salute to the Qwen team 🫡 We tested Qwen... 3.7-Max, Gemini 3.5 Flash, GPT-5.5, and Claude Opus 4.7. The biggest shock came from Qwen. In less than a month (3.6 Max dropped April 20), Qwen went from the worst multimodal output on our sakura tree test, barely keeping up with Gemini, GPT, and Claude , to matching Gemini 3.5 Flash frame for frame on this soccer test, and outperforming GPT-5.5 and Claude Opus 4.7. It rendered a perfectly proportioned soccer player and the most lifelike ball in the entire test. Remarkable spatial reasoning. Also: Gemini 3.5 Flash is now faster than GPT-5.5, which used to be the fastest in our past tests.show more

GMI Cloud
75,785 просмотров • 2 месяцев назад
LongCat performed Opus 4.8 and GPT 5.5 level on... real physics tasks for $0! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics Prompts: - A cannon demolishing a brick wall - A bowling ball knocking down the pins - A tornado that sucks in random objects Outputs: LongCat: 18,015 tokens, $0.00 Opus 4.8: 18,872 tokens, $0.48 GPT 5.5: 32,588 tokens, $0.98 GLM 5.2: 31,062 tokens, $0.09 On the physics LongCat came out ahead of Opus 4.8 and GLM 5.2 - cleaner collisions, nothing clipping or falling through. On detail and rendering it matched GPT 5.5, the best looking of the four. Getting this quality for free is wild!show more

atomic.chat
105,523 просмотров • 1 месяц назад
I gave Kimi K3 and Opus 5 the exact... same prompt to build a car racing game. I expected Opus 5 to take the lead. Instead, Kimi K3 finished in ~25 min for ~$8.9, while Opus 5 took 19 min and cost $14. The biggest surprise wasn’t the speed or the price. Kimi K3 built the better game. Smoother controls, better physics, and a driving experience that felt ready from the start. Opus 5 looked cleaner visually but Kimi won where it actually mattered. Nearly half the cost and the better gameplay.show more

Aman
111,571 просмотров • 10 дней назад
We put grok 4.5 and opus 4.8 head to... head on the same browser task, on Hyperbrowser Sandboxes. We then asked it to open a page in a real sandboxed browser and pull the title. Grok build on Grok 4.5, Claude code on Opus 4.8, identical setup. Grok 4.5 came out ahead ↓show more

Hyperbrowser
386,220 просмотров • 28 дней назад
I had access to Opus 5 before release and... found it to be a good model if a quirky one. On shorter tasks, it could match or beat Fable levels of performance, at longer tasks it seemed less ambitious & would not deliver as complete a set of work. Here is its neo-gothic shader.show more

Ethan Mollick
101,297 просмотров • 12 дней назад
Fable 5, first Mythos-class model is live on AI/ML... API! We ran a fun test: Opus 4.8 vs Fable 5 are generating a 3D Pokemon. Verdict? Fable 5 is brilliant, fast, and rare as Mew… but Opus is still that nice little guy who does great stuff. 💛 Fable 5 Important bits: SOTA on nearly every benchmark, and the lead only grows on longer, complex tasks. • 1M Context • $10 / 1M input • $50 / 1M output Built for: long-horizon agentic coding, big migrations, vision-to-code, deep research.show more

AI/ML API
560,634 просмотров • 1 месяц назад
Grok 4.6 drops around August 7 with a 1.5... trillion parameter model and improved fine-tuning, per Elon Musk. Grok 4.7 follows a few weeks later at 2.1T, better in every benchmark but slightly slower to serve. xAI is stacking releases while Grok 4.5 already beats Opus 5 on price-performance in cybersecurity benchmarks.show more

Digg
35,731 просмотров • 8 дней назад
I don't think there's a single terminal ux that... handles agent swarms well With slate, you can literally use Opus 4.6 and GPT 5.4 at the exact same time But making it intuitive took a ton of work So heres a thread on how it works and how to actually use it 🧵show more

akira
139,114 просмотров • 4 месяцев назад