正在加载视频...

视频加载失败

Kimi K3 beats GPT 5.6 Sol on both speed and in cost in backend bug fixing ! Planted 4 defects in the same Python order service: a wrong-total calculation, a broken access-control check (IDOR), a crash on invalid input, and an inventory oversell. Both models got identical code and...

24,715 次观看 • 2 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

a moonshot engineer leaked the benchmark anthropic, openai and xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article below

starmex

33,133 次观看 • 1 个月前

Kimi K3 + GPT-6 Astra can become something bigger than two agents: a Two-Brain AI Operating System the formula: Two-Brain OS = Research Brain + Execution Brain + Shared State + Router + Tools + Verification not two models doing the same job. two specialized brains connected through one persistent system step 1 -> Kimi K3 becomes the research brain. complex research, source comparison, long-context synthesis and strategic planning happen here. its job is to explore the problem and compress the findings into a clear plan. step 2 -> GPT-6 Astra becomes the execution brain. it receives the plan, writes code, operates tools, creates artifacts and turns decisions into finished work. step 3 -> build a shared state layer. store the objective, evidence, decisions, constraints, failed attempts and next action outside the chat. both brains always know what happened and where the system should continue. step 4 -> add the router. uncertainty and open-ended questions move to Kimi K3. execution, tool use and deterministic tasks move to GPT-6 Astra. every job reaches the brain designed to handle it. step 5 -> create the handoff loop: research -> plan -> execute -> inspect -> update state -> continue. if execution reveals missing information, the task returns to Kimi. if the plan is ready, Astra takes control again. step 6 -> verify before completion. tests, source checks, constraints and explicit success criteria decide whether the result ships or returns to the correct brain with a clear failure signal. that is the difference between using two AI models and building a Two-Brain AI Operating System. Kimi K3 expands the search space. GPT-6 Astra converts it into action. shared state preserves progress, the router controls every handoff, tools execute the work and verification decides when the system is actually finished. one model can generate an answer. two specialized brains connected through memory, routing and verification can run an entire workflow. the full Kimi K3 + GPT-6 Astra Two-Brain OS breakdown is below ↓

Alex

23,962 次观看 • 22 天前