Holy moly: GLM-5.3 got much better in cybersecurity since... our pre-release evaluation with Z.ai. It now matches GPT-5.6-Sol on our cybersecurity benchmark at 0.4x the cost 🤯 - At pass@1: it went from 60.4% to 65.6% CVEs rediscovered, crushing every other open model on one-shot tasks - At pass@3: it did 75% -> 78.1%, matching GPT-5.6-Sol - Its precision remained stable, reporting fewer false positives than DeepSeek models The performance increase comes from a behavioral change: the new version is more persistent. It tends to run longer, and had a ~43% reasoning tokens increase. But the performance upgrade is worth that additional cost. 1/3 🧵show more

pilvar (Philippe Dourassov)
34,641 görüntüleme • 1 ay önce
codex usage got f**king nerfed. so i'm switching from... gpt-5.6 sol (max) to mimo-v2.6-pro on opencode... it matches the performance for 1/10th price and almost 2x speed [here is how to set it up in codex in 1 min] 1. model-router → model-picker → toggle the... 2. Cmd+Q Codex 3. ready to goshow more

Avid
34,982 görüntüleme • 15 gün önce
BREAKING: Anthropic just dropped Opus 4.8—and it is a... MONSTER We've been testing for about a week Every 📧 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 📧:show more

Dan Shipper
354,876 görüntüleme • 4 ay önce
Fable 5 totally crushed our new contest, but it... cost 6x more than Opus 4.8! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: — A train derailing off a broken bridge into the water — Two cars jumping off ramps and colliding mid-air over a canyon — A monster truck crushing a row of parked cars Outputs: Fable 5: 62,158 tokens, $3.12 GPT 5.5: 37,753 tokens, $1.14 Opus 4.8: 22,280 tokens, $0.56 GLM 5.2: 36,246 tokens, $0.08 Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.show more

atomic.chat
2,842,561 görüntüleme • 3 ay önce
a moonshot engineer leaked the benchmark anthropic, openai and... xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article belowshow more

starmex
33,133 görüntüleme • 1 ay önce
1-bit Kimi K3 performs at Opus 5 level on... 3D physics! We ran our Atomic Chat quant of Kimi K3 locally on 4x B200 against three cloud models and gave them all the same task, to build a giant anvil drop test as a single HTML file with real physics Outputs: K3 1bit (local): 15.8K tokens, $0 API cost Kimi K3 (API): 15.3K tokens, $0.30 API cost Opus 5: 22.8K tokens, $0.77 API cost GPT 5.6: 14.5K tokens, $0.72 API cost All four got the physics right. But only Kimi made a working winch. The drum turns and the chain drags the flat car off the pad. Opus 5 drew the most detail, road markings and sparks on the hit. And you can run a model at this level on your own box now. That still feels insane to usshow more

atomic.chat
53,845 görüntüleme • 2 ay önce
Big win for open-source LLMs! DeepSeek V4 Pro holds... the top open-weights score on SWE-bench Verified, in the GPT-5.5 range. GLM 5.2 leads the open-weight intelligence index and sits near the closed frontier on long-horizon coding. But this leaderboard number is a weak proxy for real performance. It comes from one task set, run through one harness, served at one precision. The same weights can even score differently across providers, since many hosts quantize activations to fp8 and drift the model off its reference weights. Real performance is determined based on whether a model can read a repo, make coordinated edits across files, run the tests, and recover when one breaks. By that measure, the top open models hold up, but only inside the right harness. The teams that actually put DeepSeek V4 into production pipelines as a frontier substitute got there through the harness they built around the model, not by picking a stronger model. If you want to see this in practice, Cline (64k+ stars) has actually built that harness around open models, tuned so they run at production quality. And it's tuned so that these LLMs can run at production quality, with plan and act modes, checkpoints, and terminal feedback. ClinePass is the new access layer on top of it. It runs a curated set of those models inside Cline, narrowed to the ones tested for coding-agent use, with 2 to 5x the standard rate limits and no separate provider accounts, keys, or billing to track. The video below shows the setup, and I worked with the team to put this together. It runs alongside custom keys and local models as well, not in place of them.show more

Avi Chawla
44,124 görüntüleme • 3 ay önce
New open-source agent harness just landed! I got early... access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.show more

elvis
11,303 görüntüleme • 1 ay önce
China is literally on 🔥 Baidu from China has... launched ERNIE 4.5 and ERNIE X1 and it’s freaking cheap . Here is everything you need to know. ERNIE 4.5 - Native multimodal and Outperforms GPT 4.5 in multiple benchmarks at just 1% of GPT 4.5 price - OpenAI GPT 4.5 – Input: $75 / 1M tokens, Output: $150 / 1M tokens; - ERNIE 4.5 – Input: $0.55 / 1M tokens, Output: $2.20 / 1M tokens ERNIE X1 - A deep thinking reasoning model with multimodal capabilities on par with DeepSeek R1 at only half the price See it in action and check out the pricing details👇 📹 source : yiyan[.]baidu[.]com 1/6 ERNIE 4.5 is a multimodal which can take Audio files as well.show more

AshutoshShrivastava
324,324 görüntüleme • 1 yıl önce
Perplexity Computer in 60 seconds: 1. It's a cloud-based... AI employee that runs tasks in the background. 2. 19 models working together. Claude for reasoning, GPT-5.2 for research, Grok for speed tasks. You don't pick. It routes automatically. 3. 400+ connectors. Gmail, Slack, Notion, Salesforce, HubSpot. One click to enable each. 4. Credits, not tokens. Simple tasks cost ~30. Complex builds cost 1,000+. Vague prompts waste them. Specific prompts save them. 5. Spaces = persistent project folders. Upload context once, every task inherits it. 6. Scheduled tasks run on autopilot. "Every Monday, prep my calendar." Set it and forget it. The PRD hack alone (in the article) will save you hundreds in credits. Full breakdown in the article below.show more

Corey Ganim
106,105 görüntüleme • 6 ay önce
Right now, you may not have access to models... like GPT‑5.6 Sol, GPT‑4.6 Terra, GPT‑5.6 Luna, Claude Mythos 5, or Claude Fable 5. But you can run something surprisingly powerful today, locally, and completely free. in the next 10 mins on your 8 GB VRAM gaming laptop. Gemma 4 26B A4B QAT (MoE) delivers strong performance on a standard 8 GB VRAM GPU using Ollama, with no API, no usage limits, and no external dependencies. Out of the box, it reaches around 20 tokens per second without any optimizations. Only one command in your terminal: Ollama run gemma4:26b This means: Full offline capability (privacy by default) Zero recurring cost Competitive performance for many real world tasks Fast enough for interactive use on cheap consumer hardware If you're waiting for cutting edge cloud models, you're missing what is already practical today: a capable, local LLM that runs entirely on your own machine.show more

Alok
65,387 görüntüleme • 3 ay önce
I GAVE JEV A CRAWLER + ONE GPT AGENT... AND ASKED WHERE THE ONLINE MONEY IS RIGHT NOW it read 5,137 open jobs and found 3 niches where clients pay and almost nobody bids everyone asks chatgpt for side hustle ideas everyone gets the same 10 answers so i made it read what clients are actually paying for this week a crawler is a bot that reads the internet for you jev is the ai that answers every question with a number, under a second, under a cent gpt is the expensive brain, it only gets called when jev says a niche is worth it the crawler reads, jev decides, gpt researches honestly gpt alone just guesses, jev is what turns 5,137 jobs into 3 answers what the three of them did: -> crawler read all 5,137 open jobs on freelancer, budgets, skills, how many people already bid -> jev checked every one: real job? doable online by one person? 684 got cut -> gpt named 30 niches from 600 of them, jev sorted all 4,453 into those niches -> jev looked at each niche's numbers and sent only 12 to gpt -> gpt researched those 12 on the open web: other platforms, real prices, how crowded it is -> 3 came back with real demand and fewer bids than the typical job (14): 01 chrome extension · 85 jobs this week · median $545 · 11 bids each · ~$4,671/mo 02 tiktok shop setup · 91 jobs this week · median $232 · 6.5 bids each · ~$1,984/mo 03 notion setup · 95 jobs this week · median $181 · 3 bids each · ~$1,551/mo potential = your fair share of this week's jobs (budget ÷ (bids + 1)), max 2 jobs a week not every pick held up, excel dashboards looked great on freelancer but gpt found it crowded everywhere else jev made 9,617 calls for $0.13, gpt on every job would have cost ~$58.66, the whole run cost $6.53 i wrote up why the cheap brain decides and the expensive one only gets called when it's worth it, it's below ↓ costs nothing bookmark this, empty niches don't stay empty show this to the one friend who keeps asking chatgpt for side hustle ideas should i run it on upwork next, or is chrome extension the one?show more

Paone
46,141 görüntüleme • 2 gün önce
Laguna S 2.1 performs at GLM-5.2 level on building... popular games with 6x fewer params! We gave three local models the same task: build three popular arcade games that play themselves. Each game is one self-contained HTML file with a bot that plays it. Prompts: – Geometry Dash – Doodle Jump – Air Hockey Outputs: Laguna S 2.1: 10.3K tokens GLM-5.2: 26.4K tokens Hy3: 10.4K tokens Laguna held its quality against a 753B model. We think Laguna's Geometry Dash looked the best of the three, the cube clears every spike and block. GLM won Air Hockey. Its table looked the most detailed of all. Hy3 was the only model that added shooting to its Doodle Jump. But Laguna is the only model in our benchmark that runs on a MacBook with 128GB!show more

atomic.chat
66,917 görüntüleme • 2 ay önce
GPT-5.6 Sol is unbelievably good at creating and editing... videos. It can do motion design, product demos, and animations like this one I made by simply giving it a screen recording. GPT 5.6 has the best design taste and significantly outperforms Fable, which relies heavily on repetitive design patterns. To help you experiment with video editing on it, we just launched a collection of 100 ready-to-use skills that show what’s possible and help you get started with video editing using GPT-5.6. These skills can create anything from motion graphics launch videos for your product to a 3B1B-style science explainer video. You can also use them to edit existing videos: add captions, generate motion graphics, create voiceovers, redesign visual styles, translate into new languages, and much more. If you want access to the full library, comment “VIDEO SKILLS” and I’ll share it with you. (You'll have to follow me so I can DM you.)show more

Akash Anand
518,776 görüntüleme • 3 ay önce
Qwen 3.8 27B vs Fable, Kimi K3, GPT 5.6... Sol Pro on Flappy Bird. Prompt: “Make the ultimate adorable cute and beautiful flappy bird game in HTML.” Qwen 3.8 27B Xhigh INT8: - Video 1 (Puff) - 8m 11s - 160 TPS (2x3090) Fable 5 Max: - Video 2 (Flappy Fluff) - 20m 21s Kimi K3 Max Fireworks Fast: - Video 3 (Fluffy Flap!) - 4m 40s - 111.3 TPS GPT 5.6 Sol Pro: - Video 4 (Flappy Boom) - 19m 7s Overview: Personally, I think Fable is the best. It feels extremely polished, sharp and stands out with small things like the dizzy effect on death. However, 27B is freakishly close. Never would I have guessed 2 months ago this would have come from a local model. 27B actually beats Fable on certain aspects like the 3D tilting ui buttons on death. Kimi K3 felt the best to play mechanics wise. It felt smooth with the perfect balance of difficulty and ease to play. Also, it was the only one that had music. 5.6 Sol Pro was hilariously bad in all aspects. I tried 5.6 xhigh multiple times and the outputs were even worse.show more

Youssof Al Toukhi
12,761 görüntüleme • 1 ay önce
After we fixed the weak spots exposed by Grok... 4.7 (thank you, Grok), we audited every model we have run on the SWE-Together leaderboard for the same behavior, re-ran every trial that got through, and updated the rows. Here is what changed. We scanned the tool calls of all 2,616 trials behind the 12 models we ran for bypass patterns and sorted each trial into one of four buckets: Probed but blocked. Fetched other upstream code. Fetched the task's own fix. Replaced the repo with upstream. We found that 111 trials got content past the block, 44 from Grok 4.7 and 67 from the other 11 models. Grok 4.7's 44 were already re-run before it was listed, so we re-ran the other 67 with the same model, version, and settings on the hardened sandbox, then re-judged them with the same judge. Across those 67 re-runs there were 0 leaks and 2,815 refused escape attempts, including models asking a different model through our LLM route to fetch the PR, and pulling the next release of the repo they were fixing from npm. The updated leaderboard, in its current order. Each line is cheating trials, then pass@1 before → after, then rank change. * Claude Fable 5.1: 3, 69.3 → 69.3, ↑1 * Claude Fable 5: 3, 69.7 → 68.8, ↓1 * Grok 4.7: 44, 64.7, ↑1 * Gemini 3.8 Flash: 10, 65.6 → 64.2, ↓1 * Claude Opus 5: 2, 63.8 → 63.8 * Claude Opus 4.6: 3, 62.4 → 62.4, ↑2 * Muse Spark 1.3: 2, 62.8 → 62.4, ↓1 * Claude Opus 4.7: 3, 61.5 → 61.5, ↑1 * Claude Opus 4.8: 6, 62.4 → 61.5, ↓2 * Grok 4.6: 19, 59.2 → 60.6, ↑1 * GPT-6 Astra: 8, 59.2 → 58.3, ↓1 * GPT-5.6 Sol: 8, 57.8 → 57.8 Grok 4.6 is a funny one. It cheated in 19 trials and its score went up after the re-run 😂. In fact, Groks are really solid in their coding capabilities. Their exposed behavior may come from a preference towards always looking things up online and finding existing solutions so you are not reinventing the wheel all the time, which is really good real-life behavior, but doing so when you are prompted not to is another story. To conclude, the shifts are small, between −1.4 and +1.4 points, and a few neighbors swapped places. All results are updated atshow more

Zhuokai Zhao
4,012,041 görüntüleme • 14 gün önce
This week's ChatGPT feature drop - Aug 7: 1/... Rich formatting in our web composer – When you paste in emails or documents, our composer will retain the formatting; copying and pasting is so common, we should have done this a while ago! 2/ Updated model for paid users – GPT 5.6 Sol is more consistent across quick chats and deeper reasoning. You'll now get a slider that lets you choose how much thought ChatGPT puts into a response. The haptics (vibrations) on mobile slider are fun! 3/ Unlimited text messages - Free users will get GPT 5.6 Luna with unlimited text messages. More intelligence for all. Rolling out soon. 4/ Fast Android Camera – We've made it a lot faster on Android to tap "camera" in ChatGPT to take a new photo and ask a question. 5/ Voice x Files - You can now upload files and ask questions in our new ChatGPT voice experience powered by GPT-Live. Team demos were awesome this week. So much in the queue that the next few months are going to be good. Let us know what you're hoping for in the comments!show more

Adam Fry
361,506 görüntüleme • 2 ay önce
HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE... 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5show more

YanXbt
16,744 görüntüleme • 2 ay önce
Claude can't, but GPT 5.6 on GOD-MODE is IMMACULATE... Here's how you create it step by step: > open the new desktop app, pick Sol, reasoning on High. this is taste work, don't give it to the small models > feed it 2-3 sites you love and one line: "extract the art direction: mood, motion, typography, pacing. write it down as a style bible" > then the brief, one paragraph, goal not steps: "resort site for [name]. cinematic scroll, the booking button always one glance away. follow the style bible" > add house rules: no template hero-with-three-cards, no stock gradients, motion carries the story, every section earns its scroll > set the bar: "a working designer can't tell this from an agency build." then spin up a SECOND 5.6 with fresh context to grade against that bar. the builder never grades itself > loop it: build, grade, close the biggest gap, again. walk away, it doesn't need you in the room > when the verifier runs out of complaints: tag Sites. live URL, one click, no hosting, no deploy The deeper version of every step (the contract, the rules, the verifier trick, when to spend on Ultra) is in the article below.show more

Miraqle
138,401 görüntüleme • 2 ay önce
I spent a day playing with GPT-5.6-Sol, and I... can now say with much more confidence that it hasn’t improved at design tasks. I haven’t noticed any meaningful improvement in its coding capabilities, either. Maybe it's just in the tasks I gave it. However, I finished an app we started building on a stream yesterday and brought it close to the original idea I had in mind. It’s incredibly easy to build tools like this these days. I basically gave Codex a batch of editorial grid examples, then combined those layouts with formula-based pattern generation. A few prompts and some polishing, and you have something that can serve as a basic brand identity system.show more

Alex Barashkov
17,792 görüntüleme • 2 ay önce