
atomic.chat
@atomic_chat_hq • 15,309 subscribers
Local AI chat and Inference Engine. Enhanced by TurboQuant. Team: @gladkos @skinbagwbones @AlexFromAtomic @danyurkin @quantizedden
Shorts
Videos

Opus 5 crushed Fable 5 at 3D destruction physics for 2x cheaper! We gave four models the same task: build three self-contained HTML scenes with real physics Prompts: - A tornado that sucks in a whole field - A wrecking ball taking down an apartment block - An overloaded truck collapsing a truss bridge Outputs: - Opus 5: 55.9K tokens, $1.40 - Fable 5: 55.1K tokens, $2.82 - Kimi K3: 35.7K tokens, $0.55 - GPT 5.6: 20.1K tokens, $0.31 Opus got all three right unlike the other models. Houses fly up the funnel and out the top, the wall breaks where the ball hits and the rubble piles up, the bridge drops the truck into the river. Fable had almost nothing on the ground for the tornado to pick up, its building collapsed on its own before the ball even touched it and its bridge blew into sticks all at once. GPT is the cheapest here but its ball never reached the building at all and its bridge fell apart in a way nothing falls apart in real life. Kimi K3, the new Chinese frontier model, ended up in the same place as GPT
atomic.chat1,481,334 Aufrufe • vor 1 Monat

New Fable 5 beats Opus 4.8 on real world physics simulations We gave both models the same three prompts and asked them to build self contained HTML5 sims with real physics and no libraries: 1. Chaotic double pendulum 2. Galton board 3. Water in a spinning drum (WCSPH) Generation cost Fable 5: $3.35 on 68.7k tokens, time 14m 47s Opus 4.8: $0.93 on 38.9k tokens, time 8m 10s Fable clearly did better on the water simulation, producing a much more solid and continuous body of water. Opus left larger gaps near the walls, scattered particles around the scene, and struggled to keep the fluid stable.
atomic.chat1,587,969 Aufrufe • vor 2 Monaten

GPT-5.6 Sol Ultra lost to GPT-5.5 on physics at 3× the cost! We gave 4 models the same prompt: build three self-contained HTML5 canvas scenes with real physics demos Prompts: - A monster truck backflipping onto a parked car - A stunt car jumping six buses into a brick wall - A train derailing off a broken bridge into the water Outputs: GPT-5.6 Sol Ultra: 32.9K tokens, $0.33 Opus 4.8: 9.2K tokens, $0.24 GPT-5.5: 12.4K tokens, $0.11 Grok 4.5: 7.0K tokens, $0.08 Sol Ultra draws just like GPT-5.5, only with more detail and it shows clearest on the bus jump where the two look almost identical. In the other tests Sol Ultra came out worse than its predecessor and the physics really got weaker. We think GPT-5.5 took the truck flip and the train outright. With Sol Ultra you basically get GPT-5.5 with weaker physics and a nicer picture for 3x the price. Newborn Grok 4.5 failed two of the three tests and only came good on the train
atomic.chat898,293 Aufrufe • vor 1 Monat

Run OpenClaw 2.0 through Atomic Chat local LLM server on a 16GB MacBook 💻🦞 We gave OpenClaw a task to edit a video using Gemma 4 E4B, it successfully created a welcome video in two minutes by writing custom React code Load the model in start the LLM server and OpenClaw 2.0 picks it right up
atomic.chat43,275 Aufrufe • vor 3 Tagen

Agent Zero crushed Codex by 2.6× on token efficiency! We gave both agents the same task on the same local model (Qwen3.6 27B): build a retro Q*bert arcade game. Each agent wrote and rewrote it across iterations, adding enemies and a HUD. Outputs: • Agent Zero: 239K tokens, 53 min • Codex: 627K tokens, 51 min Both agents built a working Q*bert, but GitLawb 's Agent Zero is the cleaner build, with crisper textures and hops that land where they should. Codex refills the context window with its whole history every turn, so the model loses the thread and hallucinates code it never wrote. Agent Zero sends only what's new, so it keeps more of the window open and makes fewer mistakes. Run agents on local AI models in Atomic Chat!
atomic.chat320,187 Aufrufe • vor 1 Monat
0:30
Sensitive content
This media may contain sensitive content.

Qwen 3.8 Max beat Fable 5 at building 3D physics scenes for 7x cheaper! We gave two models the same task. Build three self-contained 3D scenes, each one HTML file with real physics that runs itself. Prompts: - A marble machine that lifts marbles up a wheel and drops them on a loop - A car factory assembly line - A sawmill cutting logs into planks Outputs: Qwen 3.8 Max: 46.6K tokens, $0.28 Fable 5: 38.7K tokens, $1.93 Qwen worked much harder on the details. Its textures and shadows look real. Fable 5 built a fast prototype that works and left it there. In the marble machine, Fable's marbles show up from nowhere at the top of the wheel. In Qwen's, you can see them go up. On the car line, Qwen's cars look much better and more real. Fable's sawmill is very strange. The saw looks wrong, and the log goes right through a beam before the saw cuts it. Atomic Chat will have day zero support to run Qwen 3.8 Max locally!
atomic.chat210,569 Aufrufe • vor 1 Monat

New Opus 4.8 crashed Opus 4.7 at physics on canvas! We gave both models the same three prompts: simulate a real physics phenomenon on raw HTML5 canvas. Prompt 1: "A triple pendulum swings into chaos and paints glowing trails with its tip" Prompt 2: "A 1 kg block bounces between a wall and a 100.000 kg block. The collisions count out the digits of pi" Prompt 3: "Balls fall through a grid of pegs and pile into a bell curve"
atomic.chat638,293 Aufrufe • vor 3 Monaten

New LFM2.5-2.6B hits DeepSeek-V4 level on tool calling and runs 3.7x faster! We ran Liquid AI 's new LFM2.5-2.6B against DeepSeek-V4-Flash on one box with 4x RTX 5090. Both got the same three jobs, and each one only completes if the model fires every tool call Topics: -weather and local time in six cities -one budget into six currencies -four hotels checked and booked for one date Outputs: LFM2.5-2.6B: 35/35 tool-calls, 366 tok/s, 19s DeepSeek-V4-Flash: 35/35 tool-calls, 77 tok/s, 70s Both models made all 35 calls, and both fired twelve of them in one turn. LFM was trained for agents, and tool calling is where that shows. Strong result for a model you can run right on your phone
atomic.chat112,816 Aufrufe • vor 28 Tagen

Open weight Kimi K3 crushed cloud frontier GPT 5.6 at 3D destruction physics! Moonshot AI became the first lab to open-source frontier model weights. We tested them on 8x B300 We gave four models the same task: build three self-contained HTML scenes with real physics Prompts: – A monster truck crushing a row of cars – Two cars jumping a canyon and colliding head-on mid-air – A giant anvil drop test flattening cars one by one Outputs: – Kimi K3 (local): 32.7K tokens, $0 – GPT 5.6: 19.2K tokens, $0.30 – Grok 4.5: 55.2K tokens, $0.45 – GLM 5.2: 66.8K tokens, $0.15 Kimi K3 made the most realistic scenes and won all three against every other model. Its truck rides over the cars and crushes them one by one, and the wrecks stay crushed. The canyon scene shows the gap best: Kimi's cars smash into each other and fall into the gorge, GPT's cars fly past the crash and land with no damage. Grok 4.5 and GLM 5.2 surprised us with the collision animation in the second scene: the impact moment is the most epic of the whole run, bumpers fold, wheels fly off, debris flies everywhere. Their other scenes are nowhere near that level
atomic.chat150,343 Aufrufe • vor 1 Monat

Open-weight Qwen 3.8 2.4T built a Call of Duty clone in one prompt 🪖 Qwen released the 2.4T Max weights, so we rented a B200 cluster and asked the model to make a Call of Duty clone Output: ~1.1M tokens · 5 hours · one prompt Almost no one can run 2.4T at home, luckily there is a 27B version that runs on a 16GB MacBook Air with our Atomic Dynamic quants
atomic.chat44,720 Aufrufe • vor 13 Tagen

Meta's Muse Spark 1.1 beat Gemini 3.6 at making 3D arcade games for 3x cheaper! We gave two models the same task: build three 3D tabletop games that play themselves Prompts: - An 8-ball pool break scattering the rack - An air hockey rally that ends in a goal - A foosball table where the rods play the ball Outputs: Muse Spark 1.1: 31.3K tokens, $0.026 Gemini 3.6: 28.8K tokens, $0.073 Muse Spark played all three tables clean with far more detail than Gemini, which barely rendered its backgrounds. Its air hockey glitched and its foosball lagged into a buggy still picture. Muse Spark 1.1 ran the full set perfectly for 3x less
atomic.chat125,857 Aufrufe • vor 1 Monat

New Kimi K3 beat Opus 4.8 at making retro games for 2x cheaper! We gave three models the same task: make three old arcade games that play by themselves. Each game is one HTML file with a bot that plays it Prompts: – Road Fighter – Battle City – Q*bert Outputs: Kimi K3: 18.4K tokens, $0.28 GPT-5.6: 18.1K tokens, $0.28 Opus 4.8: 21.3K tokens, $0.54 Kimi K3 made the best games. Its Q*bert is the best of all three. The player jumps on the cubes, paints them and runs from the purple snake. Its Battle City looks just like the real game and plays right. The tanks move, aim, shoot and break the walls. GPT-5.6 could not handle the cars. Its Battle City broke too. The tank died on its own and the base got hit. But GPT had the nicest textures of the three
atomic.chat100,082 Aufrufe • vor 1 Monat

1-bit Hy3 running locally is 2.2x faster than its API at the same quality! We gave both models the same task and compared one-shot outputs. 1-bit Hy3 295B GGUF (92GB) ran locally on 4x RTX 5090 with 128GB VRAM against the same Hy3 over cloud API Tasks: - Flappy Bird - Arkanoid - Snake Outputs: Hy3 1-bit local: 76.9K tokens, 15.5 min Hy3 cloud API: 75.1K tokens, 34.3 min The 1-bit games look the same as the API ones. Birds fly through the pipes, bricks break, the snake eats and grows. Nothing froze or crashed. Both models even made the same slip: the snake can cross itself and the game does not end. Getting this quality from 1 bit running locally is wild! Run Hy3 GGUF yourself in Atomic Chat in 2 clicks
atomic.chat91,182 Aufrufe • vor 1 Monat

New Kimi K2.7 Code performs at GPT-5.5 level 3x cheaper! We gave both models the same three prompts: build a self-contained HTML5 canvas sim with real physics, no libraries. A spring pendulum on a stretching coil, a 1 kg block trading collisions with a 100,000 kg block, and 22 balls churning in a spinning hexagon Outputs: Kimi K2.7 Code: $0.28 on 52.4k tokens GPT-5.5: $0.93 on 23.4k tokens Spring pendulums and blocks came out even. The balls Kimi did better: its pile spins with the drum when GPT's bounce around in pure chaos. On price to quality, K2.7 Code is the clear pick
atomic.chat139,760 Aufrufe • vor 2 Monaten

New Google Gemma 4 12B claims near-26B performance - we tested both! We ran both models locally on one RTX 4090 and gave each the same task: write a self-contained HTML5 canvas animation with real physics in one file without libraries. Three scenes - a Galton board, two blocks colliding off a wall, and a chaotic triple pendulum Outputs: Gemma 4 26B-A4B: 15 GB VRAM usage, 6.9k tokens, 138 tok/s Gemma 4 12B: 9 GB VRAM usage, 8.9k tokens, 80 tok/s Same Gemma 4 family, but the 26B-A4B won every scene and ran ~1.7x faster - on just 4B active params. The 12B stayed very close though, on almost half the VRAM - which makes it the ideal model for a 16 GB laptop
atomic.chat152,526 Aufrufe • vor 3 Monaten

DFlash2 makes Qwen3.8 27B up to 4x faster⚡️ We rented one RTX 6000 and ran the same Qwen3.8-27B four ways: baseline, MTP, DFlash, DFlash2. We picked 4 different tasks for the test: a music library as JSON, a Python rate limiter, a seating puzzle, a heist story. Outputs: Baseline: 47.4 tok/s MTP: 114.7 tok/s · 52% accepted DFlash: 99.3 tok/s · 34% accepted DFlash2: 140.6 tok/s · 56% accepted The diff between DFlash and DFlash2 is essentially this: DFlash only looks at the current token, makes one guess and sends it, which often results in poor acceptance rates. DFlash2 on the other hand looks at the previous and current tokens, creates 16 candidates and ranks each guess on two criteria: 1. how well it fits as the next token 2. does it connect with the previous token. After ranking, it picks the best candidate and submits it as the guess. This is how DFlash2 achieves much higher acceptance rates and as a result is often much faster than DFlash
atomic.chat27,777 Aufrufe • vor 14 Tagen