Kimi-K3 VS Opus-4.8 GTA IV test. Left : Kimi,... Right : Opus-4.8 Kimi-K3 got impressively better than its previous model. + Muse Spark 1.1 generated something that can’t be played at all. + Gemini 3.5 Flash called itself from Anthropic and the game can’t be played too.show more

Jun Song
26,335 görüntüleme • 1 ay önce
I tested Kimi K3 vs Claude Opus 4.8 Same... prompt, an armory bay with lighting, props, and detail. Top is Kimi K3, bottom is Opus 4.8. It's not even close. Kimi K3 built a full scene with textures, proper lighting, ammo crates, weapon racks, working detail everywhere. Opus 4.8 gave me a near empty room with a couple of floating tables. No doubt it beats Opus 4.8. Kimi K3 is Fable 5 level, and it's clearly better than GPT-5.6 Sol at 3D and games. An open weight model just matched the best closed models on the market. Let that sink in.show more

Bhavy☄️
476,157 görüntüleme • 1 ay önce
🤯KIMI K3 ABSOLUTELY MOGS! BEATING Opus 4.8, GPT 5.5,... and even Fable 5 in multiple benchmarks. They scored 1688 on GDPval-AA v2 🔥 This is a completely different breed of open-source models Kimi creates better games and front-end designs than Fable 5, but it's 8x cheaper! The results coming out are truly impressive! I will be testing this out further and posting multiple tests today. Stay tuned!show more

Mark Santos
127,993 görüntüleme • 1 ay önce
a moonshot engineer leaked the benchmark anthropic, openai and... xai all buried the same week: kimi k3 beat opus 5, gpt-5.6 and grok 4.6 at $0.94 a task. stop paying anthropic $200 a month for opus 5 and openai $200 for gpt-5.6 when kimi does the same work for $8 the leak showed kimi k3 winning 9 of 12 categories against opus 5, gpt-5.6 and grok 4.6. within 48 hours all three labs quietly pushed pricing pages and one very specific comparison chart off their sites. nobody announced anything. they just deleted, which tells you everything the four numbers they scrubbed: cost per task · $0.94 vs $1.80 -> opus 5 charges $1.80 to finish one task. gpt-5.6 $1.04. grok 4.6 $0.61. kimi k3 $0.94 and it landed 487 of 500 clean -> anthropic is billing you double for a model that lost the benchmark it paid to promote the weights · free, sitting on huggingface right now -> the entire model is a public download. pull it, keep it, run it forever, nobody can switch it off -> a model you can hold cannot be rented at $200 a month. that single fact is what three labs deleted a chart over the switch · one line of bash -> moonshot ships an anthropic-compatible endpoint. one env variable and claude code points at kimi -> same cli, same keybindings, same /model. you change a url, opus 5 never knows it lost the seat the bill · $400 down to $8 -> opus 5 max plus gpt-5.6 pro is $400 a month. kimi runs the same daily work for $8 metered -> that is a 98% cut for output that beat both of them 9 categories to 3 here is the part they will fight me on: the frontier tax died the week this leaked and all three labs know it. once the weights are public the price has a ceiling, because anyone can serve the same model. anthropic, openai and xai are charging 2025 prices on a lead that ended in a benchmark they deleted instead of answered drop your $400/mo ai stack to $8. the run above is kimi k3 finishing the task opus 5 bills $1.80 for. the full breakdown is in the article belowshow more

starmex
32,547 görüntüleme • 13 gün önce
KIMI K3 JUST RECREATED A FUNCTIONAL GAME BOY ADVANCE... IN 3D → interactive 3D model, every button and screen modeled → a game loaded and playable right inside it → camera you can rotate and zoom around the device → one single prompt, zero external assets $3 in API with K3. the same result would run about $9 with Fable 5 and $6 with Opus 4.8 this isn't generating a static image of hardware anymore, it's generating a working simulation of it, from the shell to the screen to the game running insideshow more

alex
38,248 görüntüleme • 1 ay önce
WOW KIMI K3 ON MY MACBOOK PRO IS MIND... BLOWING! I am getting side by side comparisons that I can’t believe. By far the most powerful model I have ever run on a laptop. It is easily better than ChatGPT 3.5 at just a 1 bit quantized and close to ChatGPT 5. It beats most Claude models. More soon!show more

Brian Roemmele
107,434 görüntüleme • 1 ay önce
Opus 5 is soo good at 3D this has... to be the craziest 3D output i’ve seen from any model looks like anthropic saw Kimi leading in frontend and 3D and went all in only problem is it’s still token hungry and slow as ass what do you think?show more

J A Z I I
32,419 görüntüleme • 1 ay önce
HERMES AGENT NOW RUNS CLAUDE OPUS 5. NEAR FABLE... 5 INTELLIGENCE. HALF THE PRICE. SELF-VERIFIES ITS OWN WORK. AVAILABLE TODAY VIA NOUS PORTAL (20% OFF ALL MODELS). Anthropic shipped Opus 5 on July 24, 2026. same $5/$25 per million tokens as Opus 4.8. but the benchmarks tell a different story. WHAT CHANGED FROM OPUS 4.8: FrontierBench v0.1: Opus 5: 43.3%. Opus 4.8: 18.7%. 2.3x jump on the same test. ARC-AGI-3: Opus 5: 30.2%. 3x better than the next closest model. beat Fable 5 on 8 out of 13 benchmarks. at half the cost ($5/$25 vs $10/$50). same price as Opus 4.8. twice the intelligence. no reason to stay on 4.8. THE SPECS: model ID: claude-opus-5 context: 1M tokens (default and maximum) max output: 128K tokens thinking: on by default effort toggle: low / medium / high per request fast mode: $10/$50, 2.5x faster knowledge cutoff: May 2026 minimum cacheable prompt: 512 tokens (was 1,024) SELF-VERIFICATION (the biggest change): Opus 5 checks its own work automatically. Anthropic says: delete your verification prompts. "include a final verification step" now causes OVER-verification because the model already does it. for Hermes /goal tasks this is a direct upgrade. the judge checks evidence. the model also checks evidence. double layer of verification without extra tokens. EFFORT TOGGLE: low: fast, cheap, routine work. medium: balanced, daily tasks. high: full reasoning, complex problems. set per request. not a global switch. matches Hermes /reasoning command: /reasoning low (routine) /reasoning high (complex) Opus 5 effort toggle + Hermes reasoning control = precise cost management per turn. WHERE OPUS 5 FITS IN HERMES: DAILY DRIVER (replaces Opus 4.8): same price. 2.3x better benchmarks. set as your main model: Desktop app / Dashboard: Models → claude-opus-5 CHIEF OF STAFF: synthesis across multiple agents. reads Kanban, prioritizes, routes tasks. self-verification catches routing errors before they cascade. COMPLEX CODING: SOTA on agentic coding benchmarks. FrontierBench 43.3% = best public model for coding. set as coder profile model. /GOAL TASKS: self-verification + completion contracts = the model proves its work AND double-checks the proof. long-horizon goals finish correctly more often. MoA AGGREGATOR: strongest synthesis model at $5/$25. pair with GPT-5.6 and Grok 4.5 as references. Opus 5 aggregates. best quality at mid-range price. presets: max-quality: reference_models: - provider: openai-codex model: gpt-5.6-sol - provider: xai model: grok-4.5 aggregator: provider: anthropic model: claude-opus-5 COMPUTER USE: near-Fable 5 quality for browser automation. at half the token cost per session. computer_use tasks burn lots of vision tokens. Opus 5 halves that bill vs Fable 5. WHAT TO KEEP OPUS 5 AWAY FROM: cron monitoring: too expensive. use DeepSeek or no_agent mode. sub-agent grunt work: use GPT-5.6 Luna ($1/$6) or DeepSeek. auxiliary tasks: use Gemini Flash. routine web extraction: use a cheap model. Opus 5 is for the turns where quality compounds. planning, synthesis, verification, complex reasoning. budget models handle everything else. NOUS PORTAL: 20% OFF ALL MODELS Nous Portal currently runs a 20% discount on all models including Opus 5. $5/$25 official → $4/$20 through Nous Portal. the cheapest way to run Opus 5 right now. hermes setup --portal select claude-opus-5 as your model. discount applies automatically. Opus 5 replaces Opus 4.8 everywhere. same price. better at everything. no tradeoff. straight upgrade. hermes update /model claude-opus-5show more

YanXbt
16,744 görüntüleme • 1 ay önce
Kimi K3 vs GPT-5.6 vs Claude Fable 5 A... game is one of the hardest places to fake working code. A web app can hide bugs behind a polished screenshot. A game can't. Bad collision, clipping, jitter, broken physics, awkward controls, or a snowboard sinking into the terrain become obvious within seconds. That's why I like using games to evaluate coding models. I gave Kimi K3 a single prompt to build a snowboard game in Godot, and just 20 seconds of gameplay told me more than a page of benchmark scores ever could. Benchmarks matter, but watching a model build something interactive reveals a different side of its capabilities. If the gameplay feels right, the systems work together, and the experience holds up under real input, that's a much more convincing demonstration than any leaderboard.show more

FHILY👑
36,569 görüntüleme • 1 ay önce
anthropic will sell you opus 5 at $200 a... month. openai will sell you gpt-5.6 at $200 a month. neither will tell you stanford and berkeley published the 5 principles to build a $100k/mo ai company on kimi k3 for $10 stanford and berkeley spent years figuring out what actually separates ai systems that work in production from ai systems that die in demos. they published the findings. anthropic and openai priced their frontier subs like nobody would read the papers. the papers are free this is dspy plus verifiers plus decomposition plus skills plus mcp. five principles from stanford, berkeley and moonshot that turn a $10/mo kimi k3 sub into an ai analyst that runs unattended. the model is public. the system is the moat five moves that turn kimi k3 into the $100k/mo company: P1 don't prompt, program (stanford dspy) -> stanford proved hand-tuned prompts don't scale. define a pipeline as modules, let the optimizer tune them -> the compiled pipeline beat expert few-shot on multi-step tasks. one line of dspy replaces a month of prompt engineering P2 don't trust the model, build verifiers (berkeley 2026) -> a compiler either accepts or rejects. a test either passes or fails. that is a verifier -> berkeley: test-suite reward hit 42.2% pass@1 on swe-bench. hybrid verifiers hit 51.0% best@26. no bigger model, just a real check P3 don't scale agents, decompose them (stanford ai index 2026) -> stanford found multi-agent gains only 2-4 percentage points. two coding agents sometimes did worse than one -> the win is role decomposition, not count. researcher, writer, reviewer, verifier, clear input, clear output, no overlap P4 don't repeat expertise, encode it as skills (kimi code) -> every session starting from zero is institutional knowledge you lost. a skill.md file makes kimi activate the workflow automatically -> week one you write the skill. month six it encodes more institutional memory than most junior employees carry P5 don't keep ai in chat, connect it to tools (mcp) -> a model that only sees what you paste is a consultant working blindfolded. mcp connects kimi to your crm, db, github, linear, slack -> the model is public. the data is yours. the connections are your moat my position, and it is the arguable one: the next $100k/mo ai company will not win because it got early access to a frontier model. it will win because it followed 5 papers that anthropic and openai are quietly hoping you never read drop your $200/mo ai sub to $10. the swarm above is what 300 kimi k3 agents look like running those 5 principles. the full playbook is in the article belowshow more

starmex
31,358 görüntüleme • 17 gün önce
BREAKING: Anthropic just dropped Opus 4.8—and it is a... MONSTER We've been testing for about a week Every 🪨 and our verdict is they could've just called it Opus 5, it's that good. Here's our vibe check: - Beats GPT-5.5 on Senior Engineer bench. On our toughest benchmark Opus 4.8 scores a 63—a hair higher than GPT-5.5's score of 62, and a full 30 points higher than Opus 4.7. It tackled a ground-up rewrite of a production codebase, and actually built something that works. HOWEVER: Coding performance varied a lot at different reasoning levels. We recommend using it on xhigh for best results. - Incredibly good writer. Opus 4.8 scored a 79.6 on our writing benchmark—measuring models on real-world writing tasks we do all of the time like essay writing, promo email writing, and more. It beats GPT-5.5 by 6 points. It produces well-written prose with fewer "AI-isms". It's also very good at writing in your voice given the right context. HOWEVER: Writing performance also varied with reasoning levels. Medium reasoning had higher incidence of AI-isms—we found best results with high. - Beast at knowledge work. Opus 4.8 is very good at general knowledge work tasks like report creation, research and more. It produced the best PowerPoint one-shot we've ever seen on our deck generation benchmark. - Emotionally intelligent, willing to question the frame. I've also found it to be quite good at talking through psychological or interpersonal issues. It has a high EQ, and it's also good at not glazing and helping to expand your perspective. Its thought process feels extremely rich and dynamic. THE BAD: These days a model is only as good as its harness, and Codex is still a far superior harness to the Claude Desktop app. This has kept me using Codex + GPT-5.5 as my daily driver, but I am flipping back and forth a lot more between Codex and Claude. Anthropic is back baby! Read the rest on Every 🪨:show more

Dan Shipper 📧
354,559 görüntüleme • 3 ay önce
Claude Opus 5 + Kimi K3 find businesses with... unused land, render what fits, mail the owner, and take a commission on their pick. here's what it does: - pulls every luxury restaurant, hotel and club off google maps - joins each address to the county record for the owner - drops anyone leasing. only owners can approve a build - measures the open ground they own, land minus building - four builds gate off that number: seating, shade, planting, lighting - renders each onto a real photo of their own building - kimi writes every venue its own landing page with the numbers - prints a postcard, mailed to the owner on the parcel record - they scan, pick one, the lead goes to a contractor fully automated, these are six figure jobs and you just collect the referral reply "SETUP" + RT and i'll send you a free guide so you can build this too (must be following so i can DM you)show more

Chris
14,612 görüntüleme • 1 ay önce
After a few more hours, I think I've figured... out Opus 5. Opus 5 is trained to be more agentic than anything I've used. All Claude 5 models are like that. So what changes? The way to interact with Opus 5 or contextualize it won't work the same way as with other models. It loves exploring, so it doesn't need much guidance for it. Unique preferences, artifacts, and references compliment it well and enable cleaner and more effective exploration and execution. Now that it can explore more effectively on its own and understand intent better, the best thing to do is to get out of its way (e.g., it doesn't need examples of your preferences; a clear high-level description of it works best). It's truly agentic in that sense. A good first step to provide better context for Opus 5 is to distinguish between what's situational and what needs persistence. Regardless, persistent system prompts and CLAUDE.MD needs to stay lightweight. Remove memories and tool descriptions from these. CLAUDE.MD is also a great place to tap into progressive disclosure by linking command/skills to it. On the situational side, agent skills and auto-memory can leverage progressive disclosure and the improved ability of the model to use its external context/knowledge. Conflicting and unnecessary instructions, which are common at this layer (mainly to ensure reliability), are going to throw off this model easily. That's the biggest change I had to make. Simple, clean, and clear prompts and skills work best. I had to clean a lot of my skills and system prompts. The way I prompt remains the same (usually clear and well-scoped). MCP tool descriptions are also more descriptive and have been deduped from the system prompt. Anthropic released a guide on the new rules for context engineering, which was helpful here. I started to test the recommendations and created a little artifact with the things that worked along the way. This might feel like a lot of work. Believe me, it has been frustrating. But I think we can expect future frontier models to become more agentic and smarter at figuring out the right context/gaps. The best thing to do is to prepare for that now. Boris Cherny mentioned that Opus 5 is their least prompt-injectable model yet. I am not sure if that was something they intentionally trained for or if it emerged based on how it was trained, which is to be extremely agentic in nature and more direct in execution.show more

elvis
37,685 görüntüleme • 1 ay önce
New open-source agent harness just landed! I got early... access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.show more

elvis
11,303 görüntüleme • 14 gün önce
not sure why nobody is talking about this but... Google Omni is insane at video editing Original Video (left) vs Omni Edited Video (right) everyone is comparing it to Seedance and missing the point completely. Seedance is for generating videos from scratch. Google Omni is for editing videos that already exist. which are two completely different use cases this is like when Nano Banana 1 first came out and nobody realized how big it was going to be. this is the first AI that can actually properly edit videos.. i've generated a few hundred videos with this model and it can do literally any type of edit you can think of. changing voices, swapping characters, removing watermarks, adding captions, transitions, pop ups, whatever. if you can describe the edit you want it can do it this completely crushes every other model on the market when it comes to video editing. nothing else even comes close right now and this is just the flash model. imagine what the pro version is going to be able to do when it drops in a couple months this should have way more hype than it's getting..show more

Miko
29,783 görüntüleme • 3 ay önce
23 year old David Puig is having another great... season, finishing inside the points in every event on the LIV Golf League to be 10th on the Standings and finishing T11 or better in all 3 DP World Tour starts. He’s been a model of consistency but is without a win and he says it’s something he’s looking forward to the challenge of trying to overcome: “I’ve been having a really good season and playing very consistent. I think I’ve been doing well in every part of the game if you look at my strokes gained. Lots of birdies, some eagles, but too many mistakes to be winning events and it’s something I’m really looking forward to, the challenge of minimising the mistakes and trying to get it done this year. “There’s times where I’ve made the right decision, it was just poorly executed and a few unlucky breaks here and there that stopped me from winning. But I’m looking forward to the challenge and really happy with how I’ve been playing.” David is an electric player to watch who is not scared to take on tough shots and can get hot at any time running off strings of birdies. His speed is his biggest asset but it’s something he says he needs to learn how to utilise: “I might be a little different than the average player. I hit the ball pretty far and that makes it a bit more challenging when I’m out of position in the rough. We’ve caught a couple of fliers where the ball reacts different than expected out of the different lies and that makes it more challenging. But I’ve just got to learn from those mistakes and keep getting better. “My game is getting into good shape. I’m positive with the way I’ve been working on and off the course to get to that point that I can get the win. I’m excited for it and I’m learning. I just need to keep the mistakes off the card.” David is playing this week at LIV Golf UK by JCB and it’s a course which he says he enjoys. He also played well here last season coming T11. Hopefully he can continue his great form this season and contend again this weekend at the JCB Golf and Country Club.show more

Flushing It
34,470 görüntüleme • 1 yıl önce
We are in an insane run of open-weight drops.... Every modality, open source is winning. This is what an open source AI summer ☀️ looks like: 🧠 LLMs & Reasoning → DeepSeek-V4-Flash-0731 (my king 👑): 304B MoE refresh, Terminal-Bench 2.1 jumps 61.8→82.7 over the preview, DeepSWE 7.3→54.4. Closes in on Opus-4.8 on Agents' Last Exam (25.2 vs 25.7). MIT. → Muse-Glimmer-30B, from Meta (they are back!!): their first open agentic model. ~29.6B dense + perception encoder, 131k+ context, built to run fully local, no cloud. Apache 2.0. → Liquid AI LFM2.5-2.6B: 2.69B params, 131k context, 220 tok/s on an M5 Max in under 2.5GB RAM. Competitive with models 4x larger on agentic tasks. → inclusionAI Ling-3.0-flash: 124B total, only 5.1B active, ~12% the size of their old 1T flagship Ring-2.6, matches it on key benchmarks. MIT. → inclusionAI Ling-3.0-tiny: 7.9B total, 1.3B active, 86-90 tok/s on an M4 Pro MacBook at ~8GB peak memory. MIT. → NVIDIA Nemotron-3.5-Lightning-30B-A3B: hybrid Mamba-2+MoE+Attention, up to 1M context, runs on a single H100 or DGX Spark, SWE-bench Verified 52.8. → deepgrove maple-preview: 20B-A1B ternary-weight reasoner, 218 tok/s on a Mac mini M4, 5.3GB checkpoint. MIT. → BigBang-v1 (endless-frontier): fine-tuned from Qwen3.6-35B-A3B via a self-evolving generator/critic synthetic-data loop. Lands aggregate performance between DeepSeek V4 Flash (284B) and V4 Pro (1.6T), at 35B. Apache 2.0. 🎬 Video → MiniMax-H3: 33B dense omni model, native stereo audio, up to 2K/15s. 3.6k+ likes already. → Minimax-H3-Turbo (lightx2v): Apache-2.0 turbo distillation of H3 for fast inference. → Lightricks LTX-2.5: image-to-video update, custom Gemma-4-12B text encoder, a markedly stronger distilled model. 🔊 Voice → NVIDIA NemotronLabs VoiceChat-11B: full-duplex speech-to-speech, ~450ms turn-taking, #2 on open VoiceBench, and the first open full-duplex model with live tool-calling mid-conversation. 🛡️ Safety → Mistral Shieldstral-1.0-3B: 3B multimodal guardrail that takes your safety policy as plain text instead of fixed categories. Beats LlamaGuard-4-12B and ShieldGemma-9B on HarmBench (99.4) and ToxicChat (84.1) at a fraction of the size. Apache 2.0.show more

Victor M
54,264 görüntüleme • 22 gün önce
this is the worst local ai will ever be.... it only gets better from here. if you are not expanding your mind with these small models you are missing what's happening right now 99 percent tool call success rate. when steered well with the right skills and a framework like hermes agent the node becomes a cognition layer. not a chatbot. not a toy. an extension of how you think. i was cranking this node at 35 to 50 tok/s all day on personal experiments and now after all the work is done qwen 3.5 9B is iterating on its own code. the game it created. fixing its own bugs autonomously. and the part you should probably not miss is that all of this is happening on a RTX 3060. not an H100. not an A100. the card most of you have sitting in a drawer right now. if you just open that drawer and put that intelligence to work every tensor core on that card should be running for you. your work. your experiments. your thinking. you all have it but because nobody told you what this hardware can actually do in 2026 you never tried. the day it unlocks is the day you test your workload, understand the tradeoffs, debug the loops, and then decide if you need to scale the hardware. there is no point buying 3 mac studios when things done well you can squeeze a similar level of intelligence from 9B compared to 70B. but only when you create the right environment for your model through the right harness. and let me tell you i have tried claude code as a local harness. i have tried opencode. i have tried various others. somehow i landed on hermes agent and never left. there is something magical going on at Nous Research. the tool call parsers, the skills system, the way it handles small models natively. nothing else comes close for local inference. own your cognition. your AI. your agent. your prompts. your experiments. why give them away for free. those are who you are and they don't belong on someone else's servers being monitored. just give it a shot with your existing hardware. you run into a problem the community will help you. and if you are migrating from openclaw to hermes i will personally help you make the switch.show more

Sudo su
58,717 görüntüleme • 5 ay önce
I went a little overboard with Codex last week... and burned through my entire weekly allowance in two days. Luckily, my quota reset today. Otherwise, I’m not sure what I would’ve done. It got me thinking: instead of asking one large model to handle everything from start to finish, why not let a stronger model plan the project and review the work, while a model built for execution handles the day-to-day implementation? So I tried it. The result was better than I expected. I used GPT-5.6 Sol in Codex as the decision-maker, then ran Ling-3.0-flash from Ant Ling inside OpenCode as the execution engine. Together, they built a small 3D farming game. Before writing any code, I had Codex create four documents: SPEC.md defined the product scope and the lines we couldn’t cross. ARCHITECTURE.md laid out the isometric coordinate system, state machine, and module boundaries. TASKS.md broke the project into small jobs Ling could tackle one at a time. ACCEPTANCE.md explained how each step would be tested and what “done” actually meant. Then I gave Ling a very straightforward role: You are the execution model for this project. Read all four documents before you begin. Work only on the task assigned for this round. When you’re done, run typecheck, test, and build. If anything fails, read the error, fix it, and run the checks again. Do not move on to the next task early. Ling handled dependency installation, project structure, strict TypeScript configuration, test setup, and a production build in 6 minutes and 3 seconds. It ran into issues with the Vite test config, a TS6310 error, and a missing jsdom dependency along the way. Instead of stopping at the first error, it kept reading the logs and fixing the problems until all three checks passed. The speed was honestly hard to believe. If you exclude the time spent waiting on tools, it was producing more than 100 tokens per second. That made the whole development loop feel noticeably faster. After this experiment, I’m planning to keep using the same workflow. If the task is small, there’s no reason to call an expensive planning model for every single step. If the task is large, handing the entire project to a Flash model in one prompt isn’t a great idea either. The setup that makes more sense to me is: Use a more capable model such as Codex to explore the project, make architectural decisions, and break the work down. Put the constraints into specs, schemas, types, and tests instead of leaving them buried in chat history. Give Ling-3.0-flash a steady stream of clear, verifiable implementation tasks. Report bugs with structured context and actual error logs, rather than saying, “It still doesn’t work.” Bring Codex back in for architecture reviews, visual checks, and changes that affect multiple parts of the project. The point of this setup isn’t to give AI a big “build the whole project” button. It’s to turn software development into a pipeline with a much more sensible cost structure: Codex figures out the plan, sets the boundaries, and catches problems. Ling-3.0-flash moves quickly, calls tools reliably, and works through well-defined tasks at scale. For agent workflows that involve lots of repetitive edits, production tasks, and tool calls, this may be a more practical answer than simply using the biggest model for everything.show more

雪踏乌云
23,107 görüntüleme • 1 ay önce
America is witnessing history at Super Bowl LIX as... President Donald Trump and his daughter Ivanka Trump stand proudly, applauding the U.S. flag and the players on the field, embraced by a roaring crowd. It’s a moment of pure patriotism, a striking contrast to the division that has plagued the nation for years. This is more than just a sports event; it’s a reminder of what a united America looks like—something that has felt out of reach for far too long. As Trump works tirelessly to clean up Washington, where establishment forces are desperate to cling to power, the energy in the South tells a different story. Here, the people are standing together, celebrating the values that truly matter. The stadium is electric, the atmosphere charged with a sense of renewal and national pride. This is not about left or right, but about an America that refuses to be torn apart by those who profit from division. While the political elite bicker and scheme behind closed doors, real Americans are here, standing shoulder to shoulder in a moment of unfiltered unity. The contrast could not be more obvious. Washington may still be tangled in its own corruption, but the heart of the country beats strong and undivided. Trump’s presence at this historic event is more than symbolic—it’s a testament to his unbreakable bond with the American people. While his enemies try to destroy him with endless lawfare and media attacks, the people show where they truly stand. His reception tonight is proof that the connection between Trump and the nation he fights for has never been stronger. The cheers, the flags, the roaring applause—it’s a reminder that despite all the attempts to silence and control, the spirit of America cannot be extinguished. This moment is a glimpse into what the country can and should be—unapologetic in its patriotism, fearless in its unity, and unwavering in its belief in American greatness. The establishment may try to dictate the narrative, but the people have already spoken. Tonight, under the bright stadium lights, America is reminded of its true strength: its people, standing together, honoring their flag, and refusing to back down. It’s a sight to behold, and one that will not be forgotten.show more

Torsten Prochnow
33,608 görüntüleme • 1 yıl önce
“CONSTITUTION HILL should retire” 🫠 He’s only had ONE... bad run in his career. 🙄 Yesterday Punchestown was that. He’s had two falls prior to that, but travelled well in both races until falling. The Champion Hurdle was too far from home to know how it would have played out, but at Aintree we knew he had some spark. What the falls have done to him physically (nothing shown by vets) or mentally more so I’m unsure, but he was just flat yesterday. We all have off days, all lose our confidence at times, but to write him off already is just wrong, IMO. If something comes to light and the call it quits then they are doing right by him, but let’s not forget he’s only 8 years old.. He was 10 / 10 prior to the Cheltenham Festival, with 8 x Grade 1 wins. 🔥 Bt JONBON 22 lengths in Supreme. Bt STATE MAN 9 lengths Champion. Bt LOSSIMOUTH 2 1/2 at Christmas. Bt EPATANTE 12 lengths Fighting 5th. All multiple Grade 1 winners! 🏇🏻 Let’s stop writing horses off so easily, and give this lad a break! ☹️ Yes, he won’t get the chance to be one of THE greatest… but he’s still great! 🖤🤍show more

RACING LEE
65,819 görüntüleme • 1 yıl önce