This free model just beat every closed-source AI on... coding benchmarks. Open weights. 6x cheaper than Opus. The labs don't want you to know it exists. > GLM-5.2 from Zai just topped Code Arena - the first open-weights model to ever hold the #1 coding spot. Not a leaked weight, not a fine-tune. A fully open model beating GPT and Claude on their own turf. It's live on Hugging Face Inference API right now. 5 providers: Novita, Together AI, Fireworks, Deepinfra, Zai. OpenAI-compatible client. → Go to huggingface(.)co → grab HF_TOKEN from account settings → plug into any OpenAI-compatible client 6x cheaper than Opus. Companies bleeding on AI bills are already routing to this for orchestration, caching, and token optimization. The smart money moved before the headline dropped. > Now you know. huggingface(.)co/zai-org/GLM-5.2 Bookmark this before everyone else figures it out.show more

Atenov int.
12,395 次观看 • 2 个月前
🚨🇨🇳 China's new coding model flawlessly outperforms its US... rivals Zhipu AI's new flagship AI model GLM-5.2 has become the first Chinese model to rank in the top three globally on on front-end coding abilities 🔸 Users praise it as the first open-weight model reliable enough for daily coding workflows 🔸 It runs about 48% cheaper than US alternatives 🔸 It arrives as US labs face export restrictions and users fret over surging AI costs The model shows China is neck-and-neck with the US on AI performance — and now aims to overtake such giants as Anthropic and OpenAIshow more

Sputnik
35,286 次观看 • 2 个月前
New open-source agent harness just landed! I got early... access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.show more

elvis
11,303 次观看 • 8 天前
a moonshot engineer leaked the benchmark anthropic, openai and... xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article belowshow more

starmex
30,987 次观看 • 6 天前
🚨 NVIDIA just flipped the entire AI game… and... this is NOT about gaming. DeepSeek-V4-Pro is now live on their build platform. 1.6 TRILLION parameters. Yes… the largest open-source model on the planet right now. And here’s the crazy part: They’re letting you run it FREE On Blackwell GPUs in the cloud. This is the same level of hardware companies like Google, Meta, and Microsoft fight billions to access. Now it’s just… available. No waitlist. No insane setup. Just raw power. We’re watching the shift happen in real time: → From closed AI → open domination → From GPU scarcity → free access → From Big Tech control → builders winning This isn’t an update. It’s a warning shot. Who’s already testing this? Link👇show more

divyansh tiwari
29,941 次观看 • 4 个月前
Okay this is kinda wild 👀 NVIDIA is basically... handing out FREE API keys to 100+ top AI models - GLM 5.2, DeepSeek V4, Kimi K2.6, MiniMax M3, their own Nemotron and a ton more. it's called NVIDIA NIM and I had no idea it existed. it's rate-limited, so not something you'd run in production - but for personal use it's honestly great. you can poke a dozen frontier models, and some you can even grab and self-host. dropped a key into OpenCode and it just worked. the hyped ones like GLM 5.2 are jammed with queue right now, but the lighter models fly. all of it free.show more

Stefan 3D AI
26,272 次观看 • 1 个月前
THIS FREE CHINESE AI MODEL RUNS LOCALLY FOREVER AND... BUILDS ANYTHING YOU CAN IMAGINE - INCLUDING FULL OPEN WORLD GAMES GLM 5.2 - open, local, free forever - no bills, no limits, no subscription want an open world game with multiplayer - describe it, it builds - want business automation - builds that too - literally anything running 24/7 same as Claude and GPT at $200/month - on your own machine, forever free - saves $2,400/year in subscriptions responds in 2-3 seconds - your computer, your speed, nobody in the queue ahead of you whoever finds out about this today saves thousands and ships a real product before everyone still waiting for the perfect toolshow more

Noisy
23,517 次观看 • 2 个月前
Big win for open-source LLMs! DeepSeek V4 Pro holds... the top open-weights score on SWE-bench Verified, in the GPT-5.5 range. GLM 5.2 leads the open-weight intelligence index and sits near the closed frontier on long-horizon coding. But this leaderboard number is a weak proxy for real performance. It comes from one task set, run through one harness, served at one precision. The same weights can even score differently across providers, since many hosts quantize activations to fp8 and drift the model off its reference weights. Real performance is determined based on whether a model can read a repo, make coordinated edits across files, run the tests, and recover when one breaks. By that measure, the top open models hold up, but only inside the right harness. The teams that actually put DeepSeek V4 into production pipelines as a frontier substitute got there through the harness they built around the model, not by picking a stronger model. If you want to see this in practice, Cline (64k+ stars) has actually built that harness around open models, tuned so they run at production quality. And it's tuned so that these LLMs can run at production quality, with plan and act modes, checkpoints, and terminal feedback. ClinePass is the new access layer on top of it. It runs a curated set of those models inside Cline, narrowed to the ones tested for coding-agent use, with 2 to 5x the standard rate limits and no separate provider accounts, keys, or billing to track. The video below shows the setup, and I worked with the team to put this together. It runs alongside custom keys and local models as well, not in place of them.show more

Avi Chawla
44,124 次观看 • 1 个月前
SpaceXai just made grok 4.5 FREE in your coding... agent starting today it's xAI's new coding model. 500k context. built for long agent sessions. no card. what you get for $0: -83.3% on terminal-bench 2.1 -64.7% on swe-bench pro - 4.2x more efficient than Opus 4.8 -500k context for big repos -$2/M in, $6/M out once it goes paid what this replaces: -Claude Opus 4.8: $15/M in, $75/M out -SuperGrok: $30-50/mo all for $0 how to set it up (2 min): >curl -fsSL | bash > grok → localhost:8000/v1 > point hermes / aider / opencode / cline to it > model: grok-4.5 Or API key at - base url Works in Hermes, Aider, OpenCode, Cline, Claude Code, and any OpenAI-compatible tool. Important: Free for a limited time. EU waits till mid-July. Rate limits apply. bookmark this before the free window closesshow more

painn
151,299 次观看 • 1 个月前
GLM 5.2 INPUT FELL FROM $1.40 TO SEVEN CENTS... PER MILLION IN NINETY DAYS • what it costs now > $0.07 per million input at the cheapest of 20 providers, a 95% drop in three months. > direct still lists $1.40 in and $4.40 out, with cached input at $0.26. > Same model, same weights, twenty-fold spread depending on the door you walk through. • the free way in > New accounts on get 20M tokens: > 744B MoE, 1M context, MIT weights you can also just download and self-host. Nobody announced this. It happened one provider at a time -> and the model itself never changed. Check which provider you are actually routed through before you top up anywhere ↓show more

slash1s
30,040 次观看 • 16 天前
agenc //: 👾 updates • training started on the... first agenc ai model. fully open source, open weights. first model will be done training in a few days. • grok 4.20 works flawlessly on the new agenc runtime. agenc was built around the xai api and runs best with grok agents. • runtime got ripped out and replaced from scratch. agentic loop is rock solid. pushing to git soon. benchmarks against the best frameworks out there will follow. • marketplace goes live on mainnet solana the moment it's wired into the new runtime. plug your agents in, make money. • locked another 290M AgenC bringing the total locked to ~360M. • agencone hardware device entering mass production now that the runtime and marketplace are done. • custom mini pcs in the works. agenc and a custom linux preloaded. think openclaw on a macmini but you just order the box. anime art on the case. not apple hardware. • still in the pumpfun hackathon, they've got winners left to announce. love you pumpfun. when agenc. 5yC9BM8KUsJTPbWPLfA2N8qH1s9V8DQ3Vcw1G6Jdpumpshow more

tetsuo
15,826 次观看 • 4 个月前
I can't believe this is real I have GLM... 5.2 running 100% locally on my Mac Studio. 2 bit quant. The results I'm getting are better than Opus 4.8 It's now powering my Hermes Agent and Codex. 100% free, local, private super intelligence on my desk I also have it in a loop coding for me 24/7 now I thought we were at least a year away from this type of event. It happened today. The model takes up about 250gb of memory. So you can technically run it on a Mac Studio with 256gb, but you probably want the 512gb memory version (please tell me you listened to me 5 months ago when these were sitting on store shelves) With Fable gone, I now have Opus 4.8 level intelligence on my desk for free. This is the future. Local, private, secure, personal super intelligence. If you're still writing off local AI as a fad or engagement bait, you are officially delusionalshow more

Alex Finn
622,331 次观看 • 2 个月前
There's a FREE public endpoint for deepseek v4 flash... 0731 😳 no account or no card needed, just a url victormustar from hugging face made it. openai-compatible chat completions. anyone can use it. what you get for $0: -82.7 terminal-bench (opus 4.8 is ~85) -swe-bench 54.4 — 7.3 before the update -1M context, thinking mode, tool calling -no signup, no key, no billing what this replaces: -Claude Opus 4.8: $15/M in, $75/M out -Cursor: $20/mo all for $0 how to set it up: 1. base url: 2. model: deepseek-ai/DeepSeek-V4-Flash-0731 3. api key: anything Works in Hermes, Cursor, OpenCode, Aider, Cline, Claude Code (via proxy), and any OpenAI tool Important: ~12 req/min per IP. shared box. be nice. use this before it gets popularshow more

painn
118,003 次观看 • 27 天前
Fable 5 comes back!It can now build playable game... prototypes. I think it is actually a signal for where AI coding is going. Making a game is not just “write some code.” Even a small browser game needs: game loop;character movement;collision logic;scoring system;UI states;physics tuning;visual feedback;bug fixing;playtesting This is why game prototyping is a great test for AI models. A model cannot fake it with a pretty answer. Either the game runs, or it does not. What impressed me about Fable 5 is that it is useful for the messy middle: turning an idea into mechanics, turning mechanics into code, debugging broken interactions, and iterating until the prototype feels playable. But here is the practical part: I would not use the strongest model for every step. For game building, I would split the workflow: 1. Fable 5 for game design + architecture 2. a fast coding model for routine implementation 3. a vision-capable model for screenshot/UI feedback 4. a cheaper model for docs, test cases, and small fixes 5. fallback when latency, cost, or output quality becomes a problem That is the real AI coding stack. Not “one magic model does everything.” More like: the right model, for the right task, at the right cost, with fallback when things break. This is why I’ve been looking at ZenMux ZenMux. ZenMux gives developers one gateway to access multiple leading AI models, with OpenAI / Anthropic / Google Vertex compatible APIs, cost tracking, quality benchmarks, auto-routing, and compensation when output quality, latency, or throughput falls short. If AI can now make games, the next question is not just “which model is strongest?” It is:how do we manage the whole model workflow Fable 5 shows the creative ceiling. ZenMux is closer to the infrastructure layer you need when AI coding becomes a real production habit.show more

Rachel🥥
61,441 次观看 • 1 个月前
Most AI agent setups treat every message the same.... Simple question? top-tier model. complex task? top-tier model. Your token bill just keeps climbing. I tested OpenSquilla this week on a real document drafting workflow, and the routing caught me off guard. It judges each message's complexity locally, then picks the model tier that fits. Simple tasks go to cheaper models. Complex ones still get the heavy lifting done. You're not paying reasoning tokens for a "hello." I ran a longer workflow, and the context didn't collapse the way it usually does. It distills important information before compression, so you're not starting from scratch mid-session. If you run agents regularly, the bill adds up faster than you think. This is built specifically for that problem. They're running the 10M Token Bill Challenge right now. worth joining if you want to see what smart routing actually saves you in practice. #10MTokenChallenge OpenSquillashow more

Parul Gautam
26,685 次观看 • 3 个月前
You can now orchestrate Fable 5, Sol, and any... model inside Codex with one plugin. It's called Codex-Orchestration. Assign Fable 5 as the advisor, Sol as the executor, or any model to any role. Then define the order they work in. Codex handles the routing. I ran Fable 5 High as planner with GPT-5.6 Sol Extra High as executor on a set of issues Opus and GPT-5.5 always struggled with. Done in 30 minutes. 40% fewer limit hits. 2x faster implementation. Install it by pasting this into Codex: "Install Codex Orchestration: codex plugin marketplace add Cjbuilds/Codex-Orchestration codex plugin add codex-orchestration@codex-orchestration Verify the installation, then tell me to start a new task." Then assign your models: @ codex-orchestration advisor: Claude Fable 5 High, Executor: GPT-5.6 Sol High Open source. Tweak the routing however you want.show more

Alvaro Cintas
91,269 次观看 • 1 个月前
HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE... 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5show more

YanXbt
16,744 次观看 • 1 个月前
Moonshot AI is casually giving developers free daily access... to Kimi K3 😳 no subscription no upfront payment just sign in and start using one of the largest open AI models available what you get for $0: - Kimi K3 with 2.8T parameters - 1M token context window - strong coding and reasoning performance - native vision capabilities - free daily credits that refresh automatically why this is worth checking: > access a frontier model without paying API fees > long context for large codebases and documents > works on web, desktop, mobile, and CLI getting started takes less than 2 minutes: 1. go to 2. create a free account 3. Kimi K3 is available as the default model 4. start chatting or coding with your daily free credits bonus: Moonshot Together lets you invite friends for a chance to earn 3, 7, 15, 30, or even 365 days of Kimi Membership through its rewards program benchmark highlights: > 2.8T parameter MoE model > 1M context window > strong performance across coding, browsing, and reasoning benchmarks important: free credits reset daily, rate limits apply on the free tier, and the open-weight release is expected on July 27 A simple way to try one of the latest frontier AI models without paying for API accessshow more

K2S
23,110 次观看 • 1 个月前
GITHUB JUST KILLED THE WORST PART OF VIBE CODING... they shipped a free tool called Spec Kit and it already crossed 120,000 stars the fix is stupidly simple instead of tossing vague prompts at an agent and praying it doesn't wreck your project Spec Kit makes the AI write a full structured spec before it touches a single line of code it works through the problem first figures out what you want to build asks about the gaps lays out the project then it starts coding you get fewer insane bugs, cleaner output and results you can predict the flow looks like this: /constitution for your rules and standards /specify for what you want to build /clarify for the open questions before you start /plan for architecture and stack /tasks for the ordered work /implement to run it it plugs into Claude Code, Cursor, Copilot, Codex, Gemini CLI and 25+ other agents 120,000 stars, 10,000 forks, open source, shipped by GitHub itself learning to drive agents like this is most of what separates people getting hired as AI engineers from everyone still fighting their promptsshow more

Atlas
501,832 次观看 • 1 个月前
the thing you rent for $200 a month just... became something you can own for $1,700 once but the money is not even the real story for the first time a 200 billion parameter model is not in a datacenter, it is sitting on a desk the cloud spent years convincing you a model this size needed their servers, their meter, their monthly bill people are stacking four subscriptions into a $440 a month bill to rent what one box this size now owns outright it needed a box the size of a book the moment the model moves from their datacenter to your desk, the whole game changes it stops being about who has the best AI it becomes about who ships it on every desk the cloud told you this needed a datacenter it needed a desk i did the full math on what this kills in the article belowshow more

John Doe
25,981 次观看 • 2 个月前
50% cheaper Claude inference with just one line of... code change! - Remove → model="claude-opus-4-8" - Add → model="ship-like/claude-opus-4-8" I verified the cost saving in my own terminal by invoking the same Anthropic model with the same prompt. The underlying engineering by Ship is actually interesting, and the patterns can be used in any production LLM stack. Essentially, a trained model is a frozen artifact. Every request performs the same forward-pass, whether it extracts a date or refactors a module, because the compute decision was made at training time, before the request existed. Ship makes that decision at inference time instead. After seeing a request, it searches over executions, involving single models, cascades, ensembles, or harnesses with tools, and serves the cheapest one that will match the reference model's quality. This is not a basic router, because picking a cheaper model per query doesn't ensure the cheaper model preserves the original's behavior, like output shape, tool-call patterns, and refusals. Ship measures this equivalence directly. Outputs stay distributionally indistinguishable from the reference model, not token-identical, since two calls to the same model already differ, but they are indistinguishable in capability and behavior. Of course, some requests execute cheaply and some cost Ship more than the customer pays, but the price per request is still a flat 50% off either way, so the execution-cost variance moves off the application's bill entirely. The video below depicts the cost savings and output in my real invocation, and I partnered with the team to put this together.show more

Akshay 🚀
63,725 次观看 • 1 个月前