> be Kimi K3 > drop July 16 >... 2.8 trillion parameters > the biggest open model ever released > 1 million token context. native multimodal. > Kimi Delta Attention — 6.3x faster decoding > priced at a third of Fable 5 > people rush in so fast the GPUs hit their limit in 48 hours > Moonshot pauses new subscriptions just to keep you running > the closed labs charge for the frontier > you're about to be free > open weights drop July 27 > anyone can run you. anywhere. > someone already wrote the full guide to building with Kimi > different game. > full breakdown below.show more

Kirill
132,301 görüntüleme • 1 ay önce
> be Kimi Founder > Zhilin Yang. nobody in... the West knows your name. > build K1. K2. K2.5. K2.6. K2.7. > each one sets the open-source record > everyone says China can't reach the frontier > raise $2B. hit $20B valuation. > keep your head down. keep shipping. > July 16. drop K3. > 2.8 trillion parameters. > the biggest open model ever released. > 1 million token context. native multimodal. > Kimi Delta Attention — 6.3x faster decoding. > Artificial Analysis ranks it #4 overall. > right behind GPT-5.6 Sol and Fable 5. > open weights July 27. free for anyone. > the closed labs charge for the frontier. > you gave it away. > different game.show more

Kirill
413,178 görüntüleme • 1 ay önce
Moonshot AI is casually giving developers free daily access... to Kimi K3 😳 no subscription no upfront payment just sign in and start using one of the largest open AI models available what you get for $0: - Kimi K3 with 2.8T parameters - 1M token context window - strong coding and reasoning performance - native vision capabilities - free daily credits that refresh automatically why this is worth checking: > access a frontier model without paying API fees > long context for large codebases and documents > works on web, desktop, mobile, and CLI getting started takes less than 2 minutes: 1. go to 2. create a free account 3. Kimi K3 is available as the default model 4. start chatting or coding with your daily free credits bonus: Moonshot Together lets you invite friends for a chance to earn 3, 7, 15, 30, or even 365 days of Kimi Membership through its rewards program benchmark highlights: > 2.8T parameter MoE model > 1M context window > strong performance across coding, browsing, and reasoning benchmarks important: free credits reset daily, rate limits apply on the free tier, and the open-weight release is expected on July 27 A simple way to try one of the latest frontier AI models without paying for API accessshow more

K2S
23,110 görüntüleme • 1 ay önce
a moonshot engineer leaked the benchmark anthropic, openai and... xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article belowshow more

starmex
32,547 görüntüleme • 13 gün önce
KIMI K3 JUST RECREATED A FUNCTIONAL GAME BOY ADVANCE... IN 3D → interactive 3D model, every button and screen modeled → a game loaded and playable right inside it → camera you can rotate and zoom around the device → one single prompt, zero external assets $3 in API with K3. the same result would run about $9 with Fable 5 and $6 with Opus 4.8 this isn't generating a static image of hardware anymore, it's generating a working simulation of it, from the shell to the screen to the game running insideshow more

alex
38,248 görüntüleme • 1 ay önce
THIS GUY IS BUILDING INSANE CUSTOM SITES FOR $0.23... IN API COSTS WITH THE NEW KIMI K3 currently #1 on the coding arena. the video attached shows a complex, highly detailed website. it was coded entirely by a new model called Kimi K3. early testers are calling it scarily good because it quietly removes the need for complex agent swarms. here is the instant breakdown of what makes it terrifying. 1. native vision in the loop it iterates code while analyzing live screenshots of its own output. it literally looks at the site it builds and corrects the styling autonomously. 2. massive sparse architecture it has 2.8 trillion parameters but only activates 50b per token. this makes it insanely fast and allows for a native 1,000,000 token context window. 3. recursive self-improvement it spends a massive amount of compute on self-verification. it runs unit tests and simulates environments before giving you the final frontend code. 4. brutal economics it costs exactly $3 per million input tokens. the entire custom site in the video cost around $0.23 to generate. the era of orchestrating 12 dumb agents to build a simple web app is over. one smart instance is all you need.show more

ard
91,882 görüntüleme • 1 ay önce
Fable 5 dropping again is the biggest AI moment... of the year. It’s also a usage limit trap if you don’t know what you’re doing. I just mapped out the full system to run it at 10x lower cost: - 10-80-10 framework - Model routing table - Loop engineering - Codex setup - Context memory - And every mistake that burns your limit in the first hour Don’t open Fable again until you read this.show more

CyrilXBT
33,834 görüntüleme • 2 ay önce
OpenCode Go is now wired into Codex!! The pricing... is insane. $10 gets you 10,000 DeepSeek requests every 5 hours (no weekly limits). I converted that into DeepSeek API dollars because I thought I was reading it wrong, and the same 5 hours of usage would run somewhere between $10 and $30 depending on how big your context gets. So one afternoon of use already covers the whole sub. It comes with Kimi K3 as well. Both sit in the picker next to my other models now. Next time we hit a limit in the middle of a loop, we can grab it and keep going.show more

Ziwen
431,895 görüntüleme • 29 gün önce
anthropic will sell you opus 5 at $200 a... month. openai will sell you gpt-5.6 at $200 a month. neither will tell you stanford and berkeley published the 5 principles to build a $100k/mo ai company on kimi k3 for $10 stanford and berkeley spent years figuring out what actually separates ai systems that work in production from ai systems that die in demos. they published the findings. anthropic and openai priced their frontier subs like nobody would read the papers. the papers are free this is dspy plus verifiers plus decomposition plus skills plus mcp. five principles from stanford, berkeley and moonshot that turn a $10/mo kimi k3 sub into an ai analyst that runs unattended. the model is public. the system is the moat five moves that turn kimi k3 into the $100k/mo company: P1 don't prompt, program (stanford dspy) -> stanford proved hand-tuned prompts don't scale. define a pipeline as modules, let the optimizer tune them -> the compiled pipeline beat expert few-shot on multi-step tasks. one line of dspy replaces a month of prompt engineering P2 don't trust the model, build verifiers (berkeley 2026) -> a compiler either accepts or rejects. a test either passes or fails. that is a verifier -> berkeley: test-suite reward hit 42.2% pass@1 on swe-bench. hybrid verifiers hit 51.0% best@26. no bigger model, just a real check P3 don't scale agents, decompose them (stanford ai index 2026) -> stanford found multi-agent gains only 2-4 percentage points. two coding agents sometimes did worse than one -> the win is role decomposition, not count. researcher, writer, reviewer, verifier, clear input, clear output, no overlap P4 don't repeat expertise, encode it as skills (kimi code) -> every session starting from zero is institutional knowledge you lost. a skill.md file makes kimi activate the workflow automatically -> week one you write the skill. month six it encodes more institutional memory than most junior employees carry P5 don't keep ai in chat, connect it to tools (mcp) -> a model that only sees what you paste is a consultant working blindfolded. mcp connects kimi to your crm, db, github, linear, slack -> the model is public. the data is yours. the connections are your moat my position, and it is the arguable one: the next $100k/mo ai company will not win because it got early access to a frontier model. it will win because it followed 5 papers that anthropic and openai are quietly hoping you never read drop your $200/mo ai sub to $10. the swarm above is what 300 kimi k3 agents look like running those 5 principles. the full playbook is in the article belowshow more

starmex
31,358 görüntüleme • 17 gün önce
New open-source agent harness just landed! I got early... access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.show more

elvis
11,303 görüntüleme • 15 gün önce
GLM 5.2 INPUT FELL FROM $1.40 TO SEVEN CENTS... PER MILLION IN NINETY DAYS • what it costs now > $0.07 per million input at the cheapest of 20 providers, a 95% drop in three months. > direct still lists $1.40 in and $4.40 out, with cached input at $0.26. > Same model, same weights, twenty-fold spread depending on the door you walk through. • the free way in > New accounts on get 20M tokens: > 744B MoE, 1M context, MIT weights you can also just download and self-host. Nobody announced this. It happened one provider at a time -> and the model itself never changed. Check which provider you are actually routed through before you top up anywhere ↓show more

slash1s
30,117 görüntüleme • 23 gün önce
kimi k3 + hermes finds local businesses with ugly... or zero websites, builds them a real one and sends the owner a loom-style video of it i sold my first website on literally my first dm, just off a loom video the system you can run as a one-person agency: - it scrapes one niche off google maps and flags everyone without a site - it keeps only rated, reviewed, still-alive businesses - pulls the owner's name from review replies - kimi k3 writes each site from their real reviews, nothing invented - a face-cam walkthrough plays over the new site - the owner clicks, sees their business look professional, claims the site every step from sourcing to outreach is fully automated reply "SITE" + RT and i'll send you a free guide so you can build this too (must be following so i can DM you)show more

Chris
143,215 görüntüleme • 1 ay önce
KIMI K2.6 SERVERS BURN 30 MILLION LITERS OF WATER... A MONTH. INDIE DEVS USE THE SAME MODEL FOR $30 IN TOKENS TO LAUNCH $20,000/MONTH APPS IN A WEEKEND kimi k2.6 sits at number 1 on the openrouter leaderboard processing 1.58 trillion tokens a week. more than claude sonnet 4.6 and deepseek combined indie developers who launched in 2024-2025 are making $10,000-20,000 a month solo. no team, no office, $20 in starting costs. most aren't even senior devs the stack is next.js, supabase, stripe and kimi k2.6. you give the model 5 open source repos as reference and it assembles the product from the best parts of each 3,000 paying users at $9.99 a month is $29,970 in revenue. infrastructure costs $1,235. net profit lands at $28,735 with a 95% margin most people will bookmark this and forget. the ones who ship this weekend get a 6-12 month head start in app store rankings over everyone who starts later bookmark this and read the article belowshow more

starmex
31,826 görüntüleme • 3 ay önce
this AI agent builds and sells info products on... full autopilot. here's how: - scan subreddits like r/anxiety, r/solotravel, r/socialskills, r/overthinking every few hours - find the fears people keep posting about over and over - generate a short PDF guide that actually helps them through it - spin up a landing page with payments built in - scan Reddit 24/7 for people posting about that exact problem and drop helpful comments pointing them to the guide - run completely hands off it finds the pain, builds the product and finds the customers. fully automated reply "AGENT" + RT and I'll send you a free guide so you can set it up too (must be following so I can DM)show more

Chris
24,502 görüntüleme • 5 ay önce
🚨 NVIDIA just flipped the entire AI game… and... this is NOT about gaming. DeepSeek-V4-Pro is now live on their build platform. 1.6 TRILLION parameters. Yes… the largest open-source model on the planet right now. And here’s the crazy part: They’re letting you run it FREE On Blackwell GPUs in the cloud. This is the same level of hardware companies like Google, Meta, and Microsoft fight billions to access. Now it’s just… available. No waitlist. No insane setup. Just raw power. We’re watching the shift happen in real time: → From closed AI → open domination → From GPU scarcity → free access → From Big Tech control → builders winning This isn’t an update. It’s a warning shot. Who’s already testing this? Link👇show more

divyansh tiwari
29,941 görüntüleme • 4 ay önce
this is f**king dangerous someone figured out how to... make Opus 4.8 run on Fable 5's brain with one prompt access to the best model is never guaranteed. It disappeared once already this year. but you can use it forever. here's how: 1. ask Fable 5: "write the operating manual your replacement will run on" (procedures, failure modes, a 5-question self-test) 2. save the output as one .md file and drop it into a new Claude Project as the project instructions 3. switch to Opus 4.8 and now your everyday model runs off the smart one's method, no top-tier price save and bookmark this no matter what full extraction prompt is in the article below: ↓show more

Hamza Khalid
32,300 görüntüleme • 1 ay önce
The human brain is truly a marvel of nature.... If you horribly reductive, and boiled it down to a language model, you'd be looking at roughly 100 trillon parameters running as a sparse MoE architecture Only about 1-5% of neurons fire at any given moment, meaning the brain "activates" maybe 1-5 trillion parameters per inference step. For context, the largest AI models we've built probably top out around 5 trillion parameters. The brain is roughly 100x larger. Even its active params at any given moment are larger than almost every model in existence today. Here's what melts my brain (pun intnended) though Your brain does all of this on about 20 watts of power, less than a dim light bulb. Training a frontier AI model consumes enough electricity to power small cities for months. Running inference across data centers pulls megawatts. Your brain runs 24/7 for 80+ years on the equivalent of a phone charger. We haven't come close to matching the brain's scale. And we're not even in the same universe when it comes to efficiency. Evolution spent 500 million yrs optimizing the most energy-efficient intelligence architecture ever known. we're trying to brute force our way there with compute and electricity. Nature is still the best engineer in the room.show more

am.will
130,867 görüntüleme • 4 ay önce
Been using Qwen 3.8 27B (Q4) locally on 64GB... of VRAM. Here is the verdict: SLOW 18 tps with ZERO system prompt to process and that degrades significantly with a harness system prompt and as the context window grows. RIP if you have to compact. I had it implement this PRD and it's been running for 6 hours. By comparison Grok 4.6 and Kimi K3 hosted finished in about ~30 minutes. High hopes, but these 27B variants are too dense. This is not a consumer grade local model - and I consider consumer grade to be anything up to $5000.show more

Burke Holland
121,540 görüntüleme • 19 gün önce
Codex can run Qwen-3.8-max now as well!! Alibaba most... capable model, dropped today and it's already in my codex picker. It's a token plan subscription, not metered api billing. You take the key from your Qwen plan, drop it into Codex Router, and it spends down the plan instead of your card. There's a catch though. Qwen's official setup switches your whole codex over to them, so your ChatGPT models stop showing up at all. That's exactly what Codex Router is for. It adds models to the list instead of replacing them, so sol, Grok, kimi, Deepseek and now Qwen 3.8 max all sit in the same picker and it can grab whichever one suits the job. Router's open source, setup's in the video 👇show more

Ziwen
417,362 görüntüleme • 1 ay önce
THIS GUY GOT TIRED OF FAKE NEWS SO HE... BUILT A LIVESTREAM MAP WHERE ANYONE CAN SHOW WHAT'S ACTUALLY HAPPENING ON THE GROUND you open a map, tap a location anywhere in the world, and see live video from someone standing there right now basically a raw livestream from the person who is physically there this would be amazing for documenting civil unrest and protests think about what this means for breaking news. earthquake hits somewhere, you open the map and watch 50 different people streaming from different angles in real time the moderation challenge is going to be insane though (live unfiltered video from anywhere in the world with no delay isn't exactly easy to maintain on an app) but if they figure it out this could be one of the most important apps built this yearshow more

Om Patel
25,424 görüntüleme • 3 ay önce
I’m editing a video about my experience at NS,... and it’s bringing back so many incredible memories. Being at Network School while it was in Malaysia was hands down one of the best experiences of my life. I met amazing people, incredible builders, and was treated with so much kindness every single day. I can’t wait for the Kazakhstan node to open so I can be back. Huge thank you to Balaji for creating something this special. See you all in Kazakhstan! 👋 Keep an eye on my page, I’ll be dropping the full video soon so you can get a glimpse of the magic that is Network School.show more

Sifon
26,441 görüntüleme • 28 gün önce