
atomic.chat
@atomic_chat_hq • 11,982 subscribers
Local AI chat and Inference Engine. Enhanced by TurboQuant. Team: @gladkos @skinbagwbones @AlexFromAtomic @danyurkin
Shorts
Videos

GPT-5.6 Sol Ultra lost to GPT-5.5 on physics at 3× the cost! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: - A monster truck backflipping onto a parked car - A stunt car jumping six buses into a brick wall - A train derailing off a broken bridge into the water Outputs: GPT-5.6 Sol Ultra: 32.9K tokens, $0.33 Opus 4.8: 9.2K tokens, $0.24 GPT-5.5: 12.4K tokens, $0.11 Grok 4.5: 7.0K tokens, $0.08 Sol Ultra draws just like GPT-5.5, only with more detail and it shows clearest on the bus jump where the two look almost identical. In the other tests Sol Ultra came out worse than its predecessor and the physics really got weaker. We think GPT-5.5 took the truck flip and the train outright. With Sol Ultra you basically get GPT-5.5 with weaker physics and a nicer picture for 3x the price. Newborn Grok 4.5 failed two of the three tests and only came good on the train
atomic.chat895,609 views • 11 days ago

New Fable 5 beats Opus 4.8 on real world physics simulations We gave both models the same three prompts and asked them to build self contained HTML5 sims with real physics and no libraries: 1. Chaotic double pendulum 2. Galton board 3. Water in a spinning drum (WCSPH) Generation cost Fable 5: $3.35 on 68.7k tokens, time 14m 47s Opus 4.8: $0.93 on 38.9k tokens, time 8m 10s Fable clearly did better on the water simulation, producing a much more solid and continuous body of water. Opus left larger gaps near the walls, scattered particles around the scene, and struggled to keep the fluid stable.
atomic.chat1,585,909 views • 1 month ago

1-bit Hy3 running locally is 2.2x faster than its API at the same quality! We gave both models the same task and compared one-shot outputs. 1-bit Hy3 295B GGUF (92GB) ran locally on 4x RTX 5090 with 128GB VRAM against the same Hy3 over cloud API Tasks: - Flappy Bird - Arkanoid - Snake Outputs: Hy3 1-bit local: 76.9K tokens, 15.5 min Hy3 cloud API: 75.1K tokens, 34.3 min The 1-bit games look the same as the API ones. Birds fly through the pipes, bricks break, the snake eats and grows. Nothing froze or crashed. Both models even made the same slip: the snake can cross itself and the game does not end. Getting this quality from 1 bit running locally is wild! Run Hy3 GGUF yourself in Atomic Chat in 2 clicks
atomic.chat88,121 views • 6 days ago

New Opus 4.8 crashed Opus 4.7 at physics on canvas! We gave both models the same three prompts: simulate a real physics phenomenon on raw HTML5 canvas. Prompt 1: "A triple pendulum swings into chaos and paints glowing trails with its tip" Prompt 2: "A 1 kg block bounces between a wall and a 100.000 kg block. The collisions count out the digits of pi" Prompt 3: "Balls fall through a grid of pegs and pile into a bell curve"
atomic.chat637,426 views • 1 month ago

DFlash makes Qwen 2.2x faster with no quality loss! We ran the same Qwen3.6-27B locally three ways on one RTX 6000: baseline, MTP, DFlash. The tasks only differ in one thing - how predictable the next word is: quicksort, describe a file in JSON, a logic puzzle, a sci-fi story. Outputs: Baseline: 44 tok/s · 1.00x MTP: 65 tok/s · 1.45x · 71% accepted DFlash: 98 tok/s · 2.20x · 30% accepted Baseline writes one token per step. MTP works inside the model itself and guesses 3 tokens ahead. DFlash is a separate small model that writes 15 tokens at once, and the big model only checks them. In JSON the same words repeat all the time, so most guesses were right: 152 tok/s, 3.4x speedup. In the story 9 guesses out of 10 were wrong. DFlash did all that extra work for nothing and became slower than baseline: 42 vs 44 tok/s. MTP guesses only 3 tokens, so a wrong guess costs very little: 46 tok/s and the win in that round. The output is identical in all three modes - DFlash is the pick for tasks with predictable output, like coding, and for chat and creative writing MTP works better. DFlash is now natively integrated into Atomic Chat on llama.cpp - speed up your Qwen models!
atomic.chat75,006 views • 7 days ago

New Kimi K2.7 Code performs at GPT-5.5 level 3x cheaper! We gave both models the same three prompts: build a self-contained HTML5 canvas sim with real physics, no libraries. A spring pendulum on a stretching coil, a 1 kg block trading collisions with a 100,000 kg block, and 22 balls churning in a spinning hexagon Outputs: Kimi K2.7 Code: $0.28 on 52.4k tokens GPT-5.5: $0.93 on 23.4k tokens Spring pendulums and blocks came out even. The balls Kimi did better: its pile spins with the drum when GPT's bounce around in pure chaos. On price to quality, K2.7 Code is the clear pick
atomic.chat139,760 views • 1 month ago

New Google Gemma 4 12B claims near-26B performance - we tested both! We ran both models locally on one RTX 4090 and gave each the same task: write a self-contained HTML5 canvas animation with real physics in one file without libraries. Three scenes - a Galton board, two blocks colliding off a wall, and a chaotic triple pendulum Outputs: Gemma 4 26B-A4B: 15 GB VRAM usage, 6.9k tokens, 138 tok/s Gemma 4 12B: 9 GB VRAM usage, 8.9k tokens, 80 tok/s Same Gemma 4 family, but the 26B-A4B won every scene and ran ~1.7x faster - on just 4B active params. The 12B stayed very close though, on almost half the VRAM - which makes it the ideal model for a 16 GB laptop
atomic.chat151,786 views • 1 month ago

Open-weight MiniMax M3 filled out a US customs form from a driver's license photo For this test we deployed MiniMax M3 Q4 using MLX-VLM on a Mac Studio M3 Ultra 512GB RAM. The model was tasked with reading a scanned document and an ID card photo, then completing a declaration form Output: 736 tokens · Input: 1,847 tokens · Time: ~31s The model analyzed both inputs, streamed its reasoning, and then called three tools: write_field for text fields, mark for Yes/No checkboxes, and sign for the signature and date. It extracted the required information, mapped it to the correct fields and completed the form without any manual input
atomic.chat109,369 views • 1 month ago

Diffusion Gemma is 4x faster, but makes 6x more mistakes! We benchmarked the new diffusion LLM against its autoregressive twin on a single H100 (FP8). We gave each the same three tasks: write a Steve Jobs biography, the history of Tetris, and the story of BeOS - every next topic less popular than the previous one. Then we fact-checked every claim in every answer. Gemma4 got 45 facts right, 5 wrong. DiffusionGemma got 33 right, 28 wrong. The less popular the topic, the worse it got: 4 mistakes on Jobs, 12 on Tetris, 12 on BeOS. It named Clara Clley as Steve Jobs' mother, invented a colleague for Pajitnov named Geri Gulovik and priced the BeBox at $9,999. The real one cost $1,600. Outputs: Gemma4 26B A4B: 218 tok/s · 15.1s total · 45 facts · 5 mistakes DiffusionGemma 26B A4B: 763 tok/s · 3.7s total · 33 facts · 28 mistakes The reason is simple. DiffusionGemma throws 256 tokens on the screen at once and polishes them pass after pass until the text sounds smooth. Smooth is all it cares about: a fake name, date or number sounds just as smooth as a real one, so it stays. Regular Gemma4 meanwhile writes one word at a time and checks every new word against everything before it. Google says it themselves in the launch post: quality is lower, use regular Gemma 4 when facts matter.
atomic.chat75,760 views • 1 month ago

New Z.ai GLM-5.2 beats Kimi K2.7 Code on physics contest! We gave both models the same three prompts and asked them to build self contained HTML5 sims with real physics and no libraries: 1. Pool break 2. Block on a bed of springs 3. Galton board Outputs: GLM-5.2: 12,640 tokens Kimi K2.7 Code: 7,420 tokens GLM 5.2 nailed all three, and it did it with way more detail and polish. The break conserved momentum, the block bounced off the springs and the Galton beads spread into a clean bell curve. Kimi struggled on every scene: its block fell straight through the springs, its break didn't look realistic with the balls colliding all wrong, and on the Galton board its balls overlapped and piled into each other instead of spreading out
atomic.chat60,899 views • 1 month ago

MiniMax M3 turned a napkin sketch into a playable game We handed MiniMax M3 a hand-drawn draft of a Doodle Jump style platformer. It read the elements off the draft, wrote the logic, drew the interface and shipped it as one self-contained HTML game Input: 6,920 tokens Output: 9,933 tokens Cost: $0.028 MiniMax (official) drops M3 on Hugging Face next week
atomic.chat64,640 views • 1 month ago
No more content to load