atomic.chat's banner
atomic.chat's profile picture

atomic.chat

@atomic_chat_hq11,982 subscribers

Local AI chat and Inference Engine. Enhanced by TurboQuant. Team: @gladkos @skinbagwbones @AlexFromAtomic @danyurkin

Shorts

Fable 5 totally crushed our new contest, but it cost 6x more than Opus 4.8! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: — A train derailing off a broken bridge into the water — Two cars jumping off ramps and colliding mid-air over a canyon — A monster truck crushing a row of parked cars Outputs: Fable 5: 62,158 tokens, $3.12 GPT 5.5: 37,753 tokens, $1.14 Opus 4.8: 22,280 tokens, $0.56 GLM 5.2: 36,246 tokens, $0.08 Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.

Fable 5 totally crushed our new contest, but it cost 6x more than Opus 4.8! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: — A train derailing off a broken bridge into the water — Two cars jumping off ramps and colliding mid-air over a canyon — A monster truck crushing a row of parked cars Outputs: Fable 5: 62,158 tokens, $3.12 GPT 5.5: 37,753 tokens, $1.14 Opus 4.8: 22,280 tokens, $0.56 GLM 5.2: 36,246 tokens, $0.08 Fable 5 did all three scenes at A+. The crashes looked real, things fell and broke the right way, and nothing went through the ground or floated. GPT 5.5 was the closest to Fable. In the Bigfoot show, we think GPT was even a little better. GLM 5.2 did not win any scene, but it was the cheapest by far. Fable is the best pick for quality, but you pay more for it.

2,833,992 views

New Claude Sonnet 5 performs at GPT 5.5 level 6x cheaper! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics crash demos Prompts: - A car crashes into a brick wall - A wrecking ball destroys a house - A catapult throws a rock at a castle wall Outputs: Sonnet 5: 15,047 tokens, $0.15 Opus 4.8: 23,063 tokens, $0.58 Sonnet 4.6: 25,824 tokens, $0.39 GPT 5.5: 31,152 tokens, $0.94 Sonnet 5 did as well as Opus 4.8 and GPT 5.5 on all three tests. In the wrecking ball test, it beat Opus 4.8. The cable moves smoothly and every hit connects. In the catapult test, it beat GPT 5.5. The rock always lands inside the wall. Sonnet 5 still needs better detail and graphics. But it used fewer tokens than every other model

New Claude Sonnet 5 performs at GPT 5.5 level 6x cheaper! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics crash demos Prompts: - A car crashes into a brick wall - A wrecking ball destroys a house - A catapult throws a rock at a castle wall Outputs: Sonnet 5: 15,047 tokens, $0.15 Opus 4.8: 23,063 tokens, $0.58 Sonnet 4.6: 25,824 tokens, $0.39 GPT 5.5: 31,152 tokens, $0.94 Sonnet 5 did as well as Opus 4.8 and GPT 5.5 on all three tests. In the wrecking ball test, it beat Opus 4.8. The cable moves smoothly and every hit connects. In the catapult test, it beat GPT 5.5. The rock always lands inside the wall. Sonnet 5 still needs better detail and graphics. But it used fewer tokens than every other model

725,764 views

Sakana Fugu surprisingly performed near GLM 5.2 level but 17× more expensive! We gave the same prompt to 4 models: build a complete live Trader Desk with both frontend and backend components, real-time market data fetched from external APIs for 8 symbols, and a custom dark-theme UI. Outputs: Fugu Ultra — 22,225 t, $0.51 Opus 4.8 — 15,802 t, $0.31 GPT-5.5 — 11,474 t, $0.26 GLM 5.2 — 13,677 t, $0.03 Fugu created the most polished and feature-rich trading desk in the run. GLM 5.2 was very close behind, with a similarly complete multi-panel interface and live data, but at a much lower cost. Opus and GPT also performed well, delivering solid results with a better balance between quality and cost

Sakana Fugu surprisingly performed near GLM 5.2 level but 17× more expensive! We gave the same prompt to 4 models: build a complete live Trader Desk with both frontend and backend components, real-time market data fetched from external APIs for 8 symbols, and a custom dark-theme UI. Outputs: Fugu Ultra — 22,225 t, $0.51 Opus 4.8 — 15,802 t, $0.31 GPT-5.5 — 11,474 t, $0.26 GLM 5.2 — 13,677 t, $0.03 Fugu created the most polished and feature-rich trading desk in the run. GLM 5.2 was very close behind, with a similarly complete multi-panel interface and live data, but at a much lower cost. Opus and GPT also performed well, delivering solid results with a better balance between quality and cost

742,420 views

Mistral OCR 4 turned a handwritten calculus exam into clean LaTeX! We gave it a photo of a hand-written exam page. The model read the handwriting and rebuilt every formula into structured digital text Output: Time: 5.1s · Cost: $0.09 Formulas came through exactly right - the hard part was nailed. The graph, unfortunately, it didn’t redraw. But that’s the telling part: most OCR tools just dump the text and quietly drop the figure. OCR 4 caught the plot, boxed it, and tagged it as a chart. It doesn’t get redrawn, but it gets read and accounted for

Mistral OCR 4 turned a handwritten calculus exam into clean LaTeX! We gave it a photo of a hand-written exam page. The model read the handwriting and rebuilt every formula into structured digital text Output: Time: 5.1s · Cost: $0.09 Formulas came through exactly right - the hard part was nailed. The graph, unfortunately, it didn’t redraw. But that’s the telling part: most OCR tools just dump the text and quietly drop the figure. OCR 4 caught the plot, boxed it, and tagged it as a chart. It doesn’t get redrawn, but it gets read and accounted for

416,647 views

Nemotron 3 Ultra performed GPT 5.5 level 10× cheaper We gave three same prompts to build HTML5 canvas with real physics. At first scene we have water in a spinning drum. Galton board - balls through pegs into bins. And a block collision setup with extreme mass differences. Outputs: Nemotron 3 Ultra: 11.3k tokens, $0.051 GPT 5.5: 11.0k tokens, $0.57 Nemotron stays right on GPT 5.5's heels, but at 10× cheaper. The gap in quality is far smaller than the gap in price.

Nemotron 3 Ultra performed GPT 5.5 level 10× cheaper We gave three same prompts to build HTML5 canvas with real physics. At first scene we have water in a spinning drum. Galton board - balls through pegs into bins. And a block collision setup with extreme mass differences. Outputs: Nemotron 3 Ultra: 11.3k tokens, $0.051 GPT 5.5: 11.0k tokens, $0.57 Nemotron stays right on GPT 5.5's heels, but at 10× cheaper. The gap in quality is far smaller than the gap in price.

700,327 views

Qwen 3.7-max beats Opus 4.7 and GPT-5.5 We tested three frontier models on a real agentic task: write a Tetris bot that plays the game and trains itself. Each model could read its own code, run benchmarks, and rewrite itself across 10 iterations. Then we compared the final bots head to head. Qwen 3.7-Max: training cost $1.32, bot improvement +56% Claude Opus 4.7: training cost $12.15, bot improvement +28% GPT-5.5: training cost $2.85, bot improvement +7% Qwen won on every dimension - biggest jump, 9× cheaper than Claude, 2× cheaper than GPT. Long agentic loops is where Qwen Max actually delivers.

Qwen 3.7-max beats Opus 4.7 and GPT-5.5 We tested three frontier models on a real agentic task: write a Tetris bot that plays the game and trains itself. Each model could read its own code, run benchmarks, and rewrite itself across 10 iterations. Then we compared the final bots head to head. Qwen 3.7-Max: training cost $1.32, bot improvement +56% Claude Opus 4.7: training cost $12.15, bot improvement +28% GPT-5.5: training cost $2.85, bot improvement +7% Qwen won on every dimension - biggest jump, 9× cheaper than Claude, 2× cheaper than GPT. Long agentic loops is where Qwen Max actually delivers.

865,484 views

Bonsai 27B running locally on an iPhone in Atomic Chat! Bonsai is the first 27B-class model that fits on a phone. PrismML built it on Qwen3.6 27B with 1-bit weights. It takes 3.9GB instead of 54GB and keeps ~90% of the benchmark scores. Available now on iPhone and Android

Bonsai 27B running locally on an iPhone in Atomic Chat! Bonsai is the first 27B-class model that fits on a phone. PrismML built it on Qwen3.6 27B with 1-bit weights. It takes 3.9GB instead of 54GB and keeps ~90% of the benchmark scores. Available now on iPhone and Android

51,580 views

Grok 4.5 performed GPT Sol level for free! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: -robot deathmatch, Tombstone vs Minotaur -a hydraulic press flattening stuff on a conveyor -a semi truck jumping a canyon Outputs: GPT-5.6 Sol: 12.9K tokens, $0.51 (~7 min) Grok 4.5: 10.8K tokens, $0 (~5 min) Muse Spark 1.1: 26.8K tokens, $0.12 (~7.5 min) GLM 5.2: 10.9K tokens, $0.02 (~12 min) Grok 4.5 handled all three scenes genuinely well and got surprisingly close to GPT-5.6 this round. On top of that, it ran on the free tier. GPT-5.6 Sol, the frontier model, put out solid but not standout work. GLM 5.2 rendered all three scenes for pennies, but it came out the roughest of the four. Meta's new Muse Spark burned the most tokens yet still stayed cheap, delivering an average result.

Grok 4.5 performed GPT Sol level for free! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: -robot deathmatch, Tombstone vs Minotaur -a hydraulic press flattening stuff on a conveyor -a semi truck jumping a canyon Outputs: GPT-5.6 Sol: 12.9K tokens, $0.51 (~7 min) Grok 4.5: 10.8K tokens, $0 (~5 min) Muse Spark 1.1: 26.8K tokens, $0.12 (~7.5 min) GLM 5.2: 10.9K tokens, $0.02 (~12 min) Grok 4.5 handled all three scenes genuinely well and got surprisingly close to GPT-5.6 this round. On top of that, it ran on the free tier. GPT-5.6 Sol, the frontier model, put out solid but not standout work. GLM 5.2 rendered all three scenes for pennies, but it came out the roughest of the four. Meta's new Muse Spark burned the most tokens yet still stayed cheap, delivering an average result.

70,064 views

New Hunyuan Hy3 hits Gemini 3.5 quality on physics for 35x cheaper! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: - A bowling ball knocking down the pins - An air hockey rally that ends in a goal - A pool break scattering the rack Outputs: Hunyuan Hy3: 29,757 tokens, $0.006 Gemini 3.5: 23,300 tokens, $0.21 GLM-5.2: 25,454 tokens, $0.07 DeepSeek-V4: 50,600 tokens, $0.009 Tencent's Hy3 matched Gemini across all three: clean collisions, the puck bounced true, the pins scattered like a real strike, the rack broke with real momentum, nothing clipped or floated. GLM is genuinely strong on pure coding tasks, but the moment the job steps outside clean code it gives way. DeepSeek was the letdown, it burned the most tokens of anyone (50k, almost 2x Hy3) and still turned in the weakest scenes

New Hunyuan Hy3 hits Gemini 3.5 quality on physics for 35x cheaper! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: - A bowling ball knocking down the pins - An air hockey rally that ends in a goal - A pool break scattering the rack Outputs: Hunyuan Hy3: 29,757 tokens, $0.006 Gemini 3.5: 23,300 tokens, $0.21 GLM-5.2: 25,454 tokens, $0.07 DeepSeek-V4: 50,600 tokens, $0.009 Tencent's Hy3 matched Gemini across all three: clean collisions, the puck bounced true, the pins scattered like a real strike, the rack broke with real momentum, nothing clipped or floated. GLM is genuinely strong on pure coding tasks, but the moment the job steps outside clean code it gives way. DeepSeek was the letdown, it burned the most tokens of anyone (50k, almost 2x Hy3) and still turned in the weakest scenes

96,930 views

LongCat performed Opus 4.8 and GPT 5.5 level on real physics tasks for $0! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics Prompts: - A cannon demolishing a brick wall - A bowling ball knocking down the pins - A tornado that sucks in random objects Outputs: LongCat: 18,015 tokens, $0.00 Opus 4.8: 18,872 tokens, $0.48 GPT 5.5: 32,588 tokens, $0.98 GLM 5.2: 31,062 tokens, $0.09 On the physics LongCat came out ahead of Opus 4.8 and GLM 5.2 - cleaner collisions, nothing clipping or falling through. On detail and rendering it matched GPT 5.5, the best looking of the four. Getting this quality for free is wild!

LongCat performed Opus 4.8 and GPT 5.5 level on real physics tasks for $0! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics Prompts: - A cannon demolishing a brick wall - A bowling ball knocking down the pins - A tornado that sucks in random objects Outputs: LongCat: 18,015 tokens, $0.00 Opus 4.8: 18,872 tokens, $0.48 GPT 5.5: 32,588 tokens, $0.98 GLM 5.2: 31,062 tokens, $0.09 On the physics LongCat came out ahead of Opus 4.8 and GLM 5.2 - cleaner collisions, nothing clipping or falling through. On detail and rendering it matched GPT 5.5, the best looking of the four. Getting this quality for free is wild!

104,522 views

MTP speedup Qwen by 2.5x in Atomic Chat Dense vs MoE models on 2x RTX 5090 Qwen3.6 27B: 51 → 117 tps +137% Qwen3.6 35B-A3B: 218 → 267 tps +25% MTP drafts several tokens ahead and verifies them in one pass. The speedup depends on memory moved per pass. Dense 27B reads all 27B params per token, MoE 35B-A3B only reads 3B active. Dense had way more to save by batching. The baseline tps also differ (218 vs 51) for the same reason from the other side. Token generation is memory-bandwidth bound, and MoE moves ~8x less memory per token, so its baseline is already 4x ahead. ~80% draft acceptance. Zero accuracy loss. ~1 GB extra VRAM. Open-source code and local AI app – in the comments 👇

MTP speedup Qwen by 2.5x in Atomic Chat Dense vs MoE models on 2x RTX 5090 Qwen3.6 27B: 51 → 117 tps +137% Qwen3.6 35B-A3B: 218 → 267 tps +25% MTP drafts several tokens ahead and verifies them in one pass. The speedup depends on memory moved per pass. Dense 27B reads all 27B params per token, MoE 35B-A3B only reads 3B active. Dense had way more to save by batching. The baseline tps also differ (218 vs 51) for the same reason from the other side. Token generation is memory-bandwidth bound, and MoE moves ~8x less memory per token, so its baseline is already 4x ahead. ~80% draft acceptance. Zero accuracy loss. ~1 GB extra VRAM. Open-source code and local AI app – in the comments 👇

170,338 views

Open-weight LongCat 2.0 matched GPT-5.5 level on agentic game dev for $0! We ran Meituan's LongCat 2.0 against cloud frontier GPT-5.5 in Kilo CLI with their agent. Same task for both - build a retro Duck Hunt game in one game.html, improved over 3 agent iterations with duck waves, ammo and physics Outputs: LongCat 2.0: 70.3K tokens, $0.00 GPT-5.5: 64.9K tokens, $0.65 LongCat kept up on graphics, physics and game logic. Ducks fly and fall when hit, the dog fetches them, ammo counts down, the waves keep coming. Both ran clean and nothing clipped. The only difference was the bill - GPT cost $0.65, LongCat ran local for $0

Open-weight LongCat 2.0 matched GPT-5.5 level on agentic game dev for $0! We ran Meituan's LongCat 2.0 against cloud frontier GPT-5.5 in Kilo CLI with their agent. Same task for both - build a retro Duck Hunt game in one game.html, improved over 3 agent iterations with duck waves, ammo and physics Outputs: LongCat 2.0: 70.3K tokens, $0.00 GPT-5.5: 64.9K tokens, $0.65 LongCat kept up on graphics, physics and game logic. Ducks fly and fall when hit, the dog fetches them, ammo counts down, the waves keep coming. Both ran clean and nothing clipped. The only difference was the bill - GPT cost $0.65, LongCat ran local for $0

39,771 views

Liquid's LFM2.5-8B-A1B smashed OpenAI's gpt-oss-20b on tool calling We ran both locally on a MacBook Pro M5 Max, 64GB, and gave each the same trip-planning request that only completes if the model fires all 7 tool calls - weather for 3 cities, two currency conversions, an email and a reminder Outputs: LFM2.5-8B-A1B: 4.8 GB RAM usage, 7/7 tool-calls, 266 tok/s, 6.9s OpenAI gpt-oss-20b: 11 GB RAM usage, 3/7 tool-calls, 146 tok/s, 15.0s The 8B used less than half the RAM and still fired all 7 calls, while the 20B silently dropped more than half of its own. It also ran ~2x faster, wrapping the full agentic request in 6.9s against 15s. That's what 38T training tokens buy: a 1B-active MoE that nails the agentic tool calls a model 2.5x its active size keeps dropping

Liquid's LFM2.5-8B-A1B smashed OpenAI's gpt-oss-20b on tool calling We ran both locally on a MacBook Pro M5 Max, 64GB, and gave each the same trip-planning request that only completes if the model fires all 7 tool calls - weather for 3 cities, two currency conversions, an email and a reminder Outputs: LFM2.5-8B-A1B: 4.8 GB RAM usage, 7/7 tool-calls, 266 tok/s, 6.9s OpenAI gpt-oss-20b: 11 GB RAM usage, 3/7 tool-calls, 146 tok/s, 15.0s The 8B used less than half the RAM and still fired all 7 calls, while the 20B silently dropped more than half of its own. It also ran ~2x faster, wrapping the full agentic request in 6.9s against 15s. That's what 38T training tokens buy: a 1B-active MoE that nails the agentic tool calls a model 2.5x its active size keeps dropping

90,063 views

Laguna XS 2.1 performed on Qwen 3.6 35B's level in Tetris building and ran 2x faster We tested two open models on a single RTX 3090 in the Poolside coding agent. The task was building a playable retro Tetris as one self-contained html file. Each model wrote and rewrote the game across 3 iterations Outputs: Laguna XS 2.1: 45K tokens, 158 tok/s Qwen 3.6 35B: 39K tokens, 81 tok/s The two Tetris builds are near identical. Poolside's Laguna has a couple of small visual bugs that Qwen 3.6 35B doesn't, but it built the same game twice as fast by its built-in DFlash speculative decoding

Laguna XS 2.1 performed on Qwen 3.6 35B's level in Tetris building and ran 2x faster We tested two open models on a single RTX 3090 in the Poolside coding agent. The task was building a playable retro Tetris as one self-contained html file. Each model wrote and rewrote the game across 3 iterations Outputs: Laguna XS 2.1: 45K tokens, 158 tok/s Qwen 3.6 35B: 39K tokens, 81 tok/s The two Tetris builds are near identical. Poolside's Laguna has a couple of small visual bugs that Qwen 3.6 35B doesn't, but it built the same game twice as fast by its built-in DFlash speculative decoding

23,256 views

Compared Qwen3.6 35B and 27B in the same conditions with Google TurboQuant Device: MacBook Pro M5Max 64GB RAM Outputs characteristics: Qwen3.6 35B: 6672 tokens, 2m 10s, 65 tok/s Qwen3.6 27B: 7344 tokens, 5m 22s, 24 tok/s Conclusion: Both models were asked to draw waves using HTML, 35B responded quickly but the result feels weak and messy, while 27B took more time and delivered a much cleaner and more consistent result, because it is built for thinking and planning, so it works better on tasks that need structure, overall 27B is a better choice for tasks where planning matters, while 35B is more suitable for everyday use when you just need a fast response

Compared Qwen3.6 35B and 27B in the same conditions with Google TurboQuant Device: MacBook Pro M5Max 64GB RAM Outputs characteristics: Qwen3.6 35B: 6672 tokens, 2m 10s, 65 tok/s Qwen3.6 27B: 7344 tokens, 5m 22s, 24 tok/s Conclusion: Both models were asked to draw waves using HTML, 35B responded quickly but the result feels weak and messy, while 27B took more time and delivered a much cleaner and more consistent result, because it is built for thinking and planning, so it works better on tasks that need structure, overall 27B is a better choice for tasks where planning matters, while 35B is more suitable for everyday use when you just need a fast response

55,540 views

Videos

No more content to load