I tested GLM 5.3 Flash vs Kimi K3. Same... prompts, both capped at $2 per scene. > Same cost, but GLM Flash couldn't build realistic scenes with proper physics. > Kimi K3 nailed it. It actually seemed to understand how objects should move and interact. What are your thoughts?show more

Bhavy☄️
43,472 Aufrufe • vor 16 Tagen
I tested Gemini 3.7 Flash vs Qwen 3.8 vs... Grok 4.6 vs GLM 5.3 Same prompt, 2 scenes: a harvester and farmers working. honestly? Qwen 3.8 impressed me. it just did what I asked. Grok 4.6 failed on coloring. Qwen>GLM> Grok> Gemini which one did better for you?show more

Bhavy☄️
110,316 Aufrufe • vor 27 Tagen
GOOGLE HAS BEEN FAILING PRETTY BADLY TO PUT OUT... A MODEL THAT CAN ACTUALLY HANG WITH THE BEST LATELY tested Gemini 3.8 Flash against Kimi K3 on the same Sticky Ball game Flash was insanely fast and used way fewer tokens, but the actual game was nowhere close. K3 had much better mechanics, movement, and overall game design Flash clearly has the speed and efficiency part down, but the gap in what it can actually build is pretty bigshow more

aditya
135,305 Aufrufe • vor 9 Tagen
I gave Kimi K3 and Opus 5 the exact... same prompt to build a car racing game. I expected Opus 5 to take the lead. Instead, Kimi K3 finished in ~25 min for ~$8.9, while Opus 5 took 19 min and cost $14. The biggest surprise wasn’t the speed or the price. Kimi K3 built the better game. Smoother controls, better physics, and a driving experience that felt ready from the start. Opus 5 looked cleaner visually but Kimi won where it actually mattered. Nearly half the cost and the better gameplay.show more

Aman
111,571 Aufrufe • vor 1 Monat
I tested Gemini 3.7 Flash vs Qwen 3.8 vs... DeepSeek V4 Pro vs GLM 5.3 same prompt: a busy train interior, then leaving the station. Qwen 3.8 and GLM 5.3 look best to me. but Gemini did it in 1 min 45 secs lol, others took more than 7 mins. Which one wins for you?show more

Bhavy☄️
24,968 Aufrufe • vor 25 Tagen
I tested Kimi K3 vs Claude Opus 4.8 Same... prompt, an armory bay with lighting, props, and detail. Top is Kimi K3, bottom is Opus 4.8. It's not even close. Kimi K3 built a full scene with textures, proper lighting, ammo crates, weapon racks, working detail everywhere. Opus 4.8 gave me a near empty room with a couple of floating tables. No doubt it beats Opus 4.8. Kimi K3 is Fable 5 level, and it's clearly better than GPT-5.6 Sol at 3D and games. An open weight model just matched the best closed models on the market. Let that sink in.show more

Bhavy☄️
478,350 Aufrufe • vor 1 Monat
I tested Qwen 3.8 Max vs Kimi K3 Same... scene, lighting, props, detail requirements. Qwen's output is decent, but it's not close to Kimi K3. Qwen's output came out messy, text handling was off, and the composition felt unorganized next to Kimi's version. In my opinion, Qwen isn't second to Fable 5 either, and it's also below GPT-5.6 Sol. It's a solid model, but not at that tier. Kimi K3 is still clearly ahead here.show more

Bhavy☄️
34,045 Aufrufe • vor 1 Monat
GPT-6 Astra vs Fable 5.1 vs Kimi K3 vs... Sol Astra: Took 11 minutes. This is from a follow-up after the first attempt had rendering and UI bugs. The smoke feels a little off. I expected more detail, but it did a good job understanding the intent. Cost: $12.85. Fable 5.1: Took 18 minutes to complete the same test with high detailing. Cost: $12.60. Kimi K3: Took 10 minutes with a similar level of output and cost only $6.95. SOL: Took 7 minutes, but the output seemed messy. Cost: $5.78. Personally, I like Kimi K3 for delivering similar quality with better token efficiency, followed by fable 5.1 for realism. What about you? I’ll do a couple more tests before coming to a conclusion.show more

Bhavy☄️
77,499 Aufrufe • vor 7 Tagen
1-bit Kimi K3 performs at Opus 5 level on... 3D physics! We ran our Atomic Chat quant of Kimi K3 locally on 4x B200 against three cloud models and gave them all the same task, to build a giant anvil drop test as a single HTML file with real physics Outputs: K3 1bit (local): 15.8K tokens, $0 API cost Kimi K3 (API): 15.3K tokens, $0.30 API cost Opus 5: 22.8K tokens, $0.77 API cost GPT 5.6: 14.5K tokens, $0.72 API cost All four got the physics right. But only Kimi made a working winch. The drum turns and the chain drags the flat car off the pad. Opus 5 drew the most detail, road markings and sparks on the hit. And you can run a model at this level on your own box now. That still feels insane to usshow more

atomic.chat
53,845 Aufrufe • vor 1 Monat
You can now use GPT 5.5, Gemini 3.7 Flash,... Kimi K3 and 47 other AI models completely free😱 No subscription. No credit card. Even the API usage costs $0. AIHubMix just opened a free catalog with 50 AI models. Some of the available models: • Ox Alpha • Gemini 3.7 Flash • GLM 5.2 • Kimi K3 • MiniMax M3 • GPT 5.5 • 40+ more And you don’t need separate API keys for each model. Setup takes 2 minutes: > Step 1: Go to > Create an account using your email or OAuth. No card needed. Step 2: Create one API key > The same key works with every free and paid model. Step 3: Add it to any OpenAI-compatible tool Base URL: Then choose any model ending in -free, such as: coding-glm-5.2-free gpt-5.5-free That’s it. One API key. 50 AI models. $0 for both input and output. Save this. You might need a free multi-model setup later.show more

CDG
15,388 Aufrufe • vor 19 Tagen
Blender just became a prompt box. Kimi K3 +... Blender MCP can take a simple text prompt and build an entire 3D scene for you. terrain, buildings, lighting, materials, camera movement, animation, even the Python scripts holding everything together. but the crazy part isn’t just that it can build the scene. Kimi can inspect what it created, understand what looks wrong, then go back into Blender and fix the actual scene before rendering again. the camera can be clipping through a tree, the lighting can feel off, the city can look too artificial, and instead of starting over, Kimi can make the changes directly inside Blender. and because everything is happening in Blender, you’re not left with a locked AI-generated video. you get the actual 3D scene. every object, material, light, camera and keyframe stays editable. a few years ago, creating something like this could take weeks of modeling, scripting, lighting and animation. now you can describe the world, let Kimi build the first 90%, and spend your time making the final 10% actually look good. that’s what makes Kimi K3 + Blender MCP so interesting. it’s getting dangerously close to text-to-3D, except the result is a real Blender project you can keep editing.show more

MIKE
10,375 Aufrufe • vor 1 Monat
Got DeepSeek V4 Flash running on 2x H200s at... 160–200 tok/s on JarvisLabsAI this speed is perfect for in the loop things I do with these agents. It feels like a GLM-5.2-class model with much lower hardware requirements. I’m going to daily-drive it for a bit and see how it performs, but first impressions are really good.show more

Atharva Ingle
21,913 Aufrufe • vor 1 Monat
holy sh*t this is f**king dangerous i just figured... out how to run Opencode in Codex It has Deepseek V4 Flash which replaces Opus 5 at 1/4th price You can get 10,000 request in only $10/month along with Kimi K3 and Qwen 3.5 Max [here is how you set it up] 1. install the 'codex-router' 2. Put in the Opencode Go API key 3. Done that's it Save this no matter what. This will be the best thing you do this weekshow more

Avid
117,270 Aufrufe • vor 1 Monat
We are getting absurdly close to the point where... “learning Blender” means learning how to direct an AI. A shot like this looks soft and playful on the surface, but under it is the usual 3D pain: modeling, layout, materials, lighting, atmosphere, animation, and endless tiny fixes until the frame stops looking dead. That is why Kimi K3 matters. With Blender MCP, you can describe a scene like a robotic goat walking through a dreamy field and let the model help build the environment, place the camera, shape the materials, script the motion, and iterate inside the real Blender project. The real shift is not text-to-image. It is text-to-workflow. Kimi K3 does not just give you a pretty output and disappear. It can help move the actual scene from rough setup to something that looks art-directed. Soon the hardest part of 3D will not be the software. It will be whether your imagination is good enough to deserve tools like this.show more

Rina
45,651 Aufrufe • vor 1 Monat
Codex can now run Deepseek-v4- flash! There's a catch... though. Deepseek's official setup switches your entire codex over to them, so your GPT models stop showing up at all. This is exactly what Codex Router is for. It adds models to the list instead of replacing them, so sol, grok, kimi and deepseek all sit in the same picker and i just grab whichever one suits the job. Deepseek v4-flash is $0.28 per million output tokens. opus 4.8 is $25. same picker, 89x apart. Links in the comment. setup's in the video 👇show more

Ziwen
145,018 Aufrufe • vor 1 Monat
Fable 5 totally crushed our new contest, but it... cost 6x more than Opus 4.8! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: — A train derailing off a broken bridge into the water — Two cars jumping off ramps and colliding mid-air over a canyon — A monster truck crushing a row of parked cars Outputs: Fable 5: 62,158 tokens, $3.12 GPT 5.5: 37,753 tokens, $1.14 Opus 4.8: 22,280 tokens, $0.56 GLM 5.2: 36,246 tokens, $0.08 Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.show more

atomic.chat
2,840,649 Aufrufe • vor 2 Monaten
THIS GUY IS BUILDING INSANE CUSTOM SITES FOR $0.23... IN API COSTS WITH THE NEW KIMI K3 currently #1 on the coding arena. the video attached shows a complex, highly detailed website. it was coded entirely by a new model called Kimi K3. early testers are calling it scarily good because it quietly removes the need for complex agent swarms. here is the instant breakdown of what makes it terrifying. 1. native vision in the loop it iterates code while analyzing live screenshots of its own output. it literally looks at the site it builds and corrects the styling autonomously. 2. massive sparse architecture it has 2.8 trillion parameters but only activates 50b per token. this makes it insanely fast and allows for a native 1,000,000 token context window. 3. recursive self-improvement it spends a massive amount of compute on self-verification. it runs unit tests and simulates environments before giving you the final frontend code. 4. brutal economics it costs exactly $3 per million input tokens. the entire custom site in the video cost around $0.23 to generate. the era of orchestrating 12 dumb agents to build a simple web app is over. one smart instance is all you need.show more

ard
91,882 Aufrufe • vor 1 Monat
Qwen3.8-Flash-Next is starting to feel like the local model... Opus fans have been waiting for. Someone ran the NVFP4 176B-class Flash-Next on 2× DGX Sparks, and the results are wild. Real measured scaling → C1: 44.2 tok/s → C2: 64.6 tok/s → C4: 86.8 tok/s aggregate The per-stream speed drops with concurrency, but total throughput keeps climbing. Long-context behavior was even more impressive: → 5K: needle retrieved → 21K: needle retrieved → 84K: needle retrieved → 167K: needle retrieved → 262K: prefill succeeded, but the window was saturated That 167K retrieval test is the one I care about. Long agent runs are where models usually start losing the plot. Flash-Next didn’t. It also held up surprisingly well on physics-heavy reasoning, artifact generation, research workflows, evidence checking, and long-horizon planning. The personality is interesting too. DeepSeek V4 Flash feels like the dependable workhorse. GLM-5.2 feels like the problem-solving machine. Qwen3.8-Flash-Next feels more insightful. It has that rare ability to understand what you’re actually asking rather than just following the surface pattern. The main weakness I’ve noticed is instruction following. It can occasionally drift between prose turns where DeepSeek and GLM stay tighter. And this is why the 256GB M5 Ultra conversation gets interesting. If Apple can pair that huge unified-memory pool with enough bandwidth, this model class becomes genuinely practical for long-running local agents. We’re talking about frontier-class reasoning on hardware sitting on a desk.show more

FHILY👑
20,217 Aufrufe • vor 15 Tagen
OpenCode Go is now wired into Codex!! The pricing... is insane. $10 gets you 10,000 DeepSeek requests every 5 hours (no weekly limits). I converted that into DeepSeek API dollars because I thought I was reading it wrong, and the same 5 hours of usage would run somewhere between $10 and $30 depending on how big your context gets. So one afternoon of use already covers the whole sub. It comes with Kimi K3 as well. Both sit in the picker next to my other models now. Next time we hit a limit in the middle of a loop, we can grab it and keep going.show more

Ziwen
432,614 Aufrufe • vor 1 Monat
Grok 4.5 performed GPT Sol level for free! We... gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: -robot deathmatch, Tombstone vs Minotaur -a hydraulic press flattening stuff on a conveyor -a semi truck jumping a canyon Outputs: GPT-5.6 Sol: 12.9K tokens, $0.51 (~7 min) Grok 4.5: 10.8K tokens, $0 (~5 min) Muse Spark 1.1: 26.8K tokens, $0.12 (~7.5 min) GLM 5.2: 10.9K tokens, $0.02 (~12 min) Grok 4.5 handled all three scenes genuinely well and got surprisingly close to GPT-5.6 this round. On top of that, it ran on the free tier. GPT-5.6 Sol, the frontier model, put out solid but not standout work. GLM 5.2 rendered all three scenes for pennies, but it came out the roughest of the four. Meta's new Muse Spark burned the most tokens yet still stayed cheap, delivering an average result.show more

atomic.chat
70,490 Aufrufe • vor 2 Monaten
New open-source agent harness just landed! I got early... access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.show more

elvis
11,303 Aufrufe • vor 23 Tagen