
Atomic Agent
@atomicagent_io • 2,379 subscribers
Local-First AI Agent. Team: @skinbagwbones @sosidudku @schwarzerrttr
Videos

Qwen3.8-Flash-Next from Qwen has day-0 support in Atomic Agent! 125B main model, 51B N-gram embeddings, 6B activated per token and 62.5 on SWE-bench Pro. We gave it a folder and it checked every file to sort them by content. Total: 25 steps + 230K tokens ≈ $0.07
Atomic Agent43,968 Aufrufe • vor 3 Tagen

Open weight Kimi K3 performs at Claude Fable 5 level on 3D arcade games! We gave both models the same prompts: ✦ Pac-Man with four hunting ghosts ✦ Nokia Snake remade in 3D ✦ Retro space pinball table with real physics Outputs (thinking included): ✦ Kimi K3 (local, 8x B300): ~237K tokens, $0 ✦ Claude Fable 5: 92K tokens, $4.50 Every one of Kimi's tokens was free on our GPUs. Even via Moonshot's API the same Kimi run would be $3.60, still under Fable. Kimi held its quality against the strongest cloud model. Its snake is the liveliest scene of all six: soft shadows, floating dust, a flicking tongue, and a bot that checks it can still reach its own tail before every move. Fable's snake plays clean but looks flat next to it. Pac-Man is the closest call of the three. Kimi builds a fresh labyrinth on every restart and draws a live minimap in the corner while the red ghost tracks the player turn by turn through the corridors, and it even caught our bot on camera. Fable's maze answers with pearl pellets, +10 popups bursting over the floor and a wide-eyed cartoon ghost patrol. Fable won the pinball table: painted playfield art, a chrome ball that mirrors the neon lights, even little synth sound effects.
Atomic Agent309,981 Aufrufe • vor 1 Monat

First local agent to beat Hermes on benchmarks! ✦ runs Qwen, Gemma, Llama via llama.cpp ✦ stable-prefix caching keeps sessions cheap ✦ TurboQuant cuts the KV-cache 6.4× smaller ✦ 37 tasks solved vs Hermes' 31 on GAIA Level 1 Open source on macOS, Windows & Linux 👇
Atomic Agent299,521 Aufrufe • vor 1 Monat

Qwen 3.8 27B on Atomic Agent was 2x faster than on Hermes and Prime agents! Outputs: Atomic: 5,535 lines, 19 modules, 2h35min Hermes: 10,424 lines, 109 modules, 4h15min Prime: 12,226 lines, 129 modules, 4h42min We gave three agents the same tasks, rebuild these games: ✦ Hole Io ✦ Tower Bloxx ✦ Subway Surfers Results: ✦ Atomic Agent gave us a real city going down the hole, skyscrapers tipping in, cars on the crossings, shops eaten one by one. In the tower every block swings in on its cable and lands where you drop it. Its runner is the cleanest of the three, though the world behind the boy stands still while he runs. ✦ Hermes dropped its hole onto a bare brown field with a few crates, and the hole grew while nothing ever fell in. Its tower buildings hang in the air with no cable at all, the logs reporting 12/12 tests passing on a game it had just broken. Its runner is an empty corridor with almost nothing in it. ✦ Prime built the prettiest town of the three, and five seconds in the whole thing tore off the ground into a tornado spinning around the hole, cars and houses circling forever and never falling in. Its tower lags on every drop, colours flickering on the buildings, and the lag is what brings the whole thing down by the fifth floor. In its runner the coins move at exactly the boy's speed, so the gap never closes and you can never pick one up. Run Qwen 3.8 27B in Atomic Agent!
Atomic Agent93,468 Aufrufe • vor 15 Tagen

Qwen3.8-Max became the brain of Atomic Agent, Hermes and OpenClaw. We gave the same task: Turn a photo of a hand-drawn floor plan into an interactive 3D walkthrough of that apartment and open it in the browser. Outputs: – Atomic Agent: 66 min, 557K tokens, $2.01 – OpenClaw: 32 min, 1.2M tokens, $1.12 – Hermes: 2 h 14 min, 4.2M tokens, $6.42 Before the start we leveled the field: one model endpoint, equal step and token budgets, equal timeouts, full autonomy, memory wiped on all three. Atomic Agent reads images through its vision tool, so it interrogated the sketch 14 times until every room, door and window turned into data. Then it drafted the whole scene in its head six times, threw away five drafts, and wrote the finished 19.8 KB file in one single write. After that it opened Chrome, checked its own render, and only then replied. The only agent of the three that verified its work, and the only one that stopped on its own. OpenClaw was twice as fast and the cheapest of the three, but its image tool kept timing out mid-run, and it shipped the palest apartment of the day: white rooms, no floor colors, one texture visibly glitched, and furniture you have to squint to find. It read the full plan three times, cut 11 room crops, wrote the scene in chunks, and landed the fastest and cheapest apartment of the day in 32 minutes. Then it kept polishing the finished file until we pulled the plug. Hermes worked the longest: two hours, 97 model calls, 4.2M tokens, and the apartment came out wrong anyway: doors standing loose in the middle of rooms, a 2 by 1.8 bath sprawled across a quarter of the flat, furniture drifting away from the plan. It measured everything twice and still built the least accurate apartment. Atomic Agent will run Qwen3.8-27B locally on day zero, next week!
Atomic Agent122,569 Aufrufe • vor 26 Tagen

Atomic Agent beat Hermes on GAIA: 69.8% vs 58.5%, and it was 1.6x faster! We ran both agents through the full GAIA Level 1 benchmark, 53 real-world tasks, same 4-bit qwen-3.6-35b on the same Apple M4 Max. Results: ✦ Atomic Agent: 37 of 53 solved, done in 3h 12m ✦ Hermes Agent: 31 of 53 solved, took 5h 10m Atomic solved 6 more tasks and finished nearly 2 hours sooner. Hermes ran into the 900s timeout on 7 tasks; Atomic on just 2. Hermes burned 71% of its total time on tasks it still failed, Atomic, 48%. Where it showed: ✦ Audre Lorde poem, which stanza is indented: Atomic pushed through a dead source, switched tools, and answered in 7.6 min. Hermes ran the full clock and returned a blank. ✦ Vietnamese specimens, which city they ended up in: Atomic pulled it from the first source and normalized the answer in 33s. Hermes spent 7.3 min and never answered. ✦ The dinosaur featured-article nominator: Atomic walked the Wikipedia chain to "FunkMonk" in 57s. Hermes guessed a wrong name after 11 min. Atomic keeps a byte-stable prompt prefix, so llama-server reuses the KV-cache instead of re-encoding the whole context every turn, and it emits one JSON array of tool calls per inference, then compresses results back instead of pasting them in full, so the context never balloons and a small model stays sharp deep into a task. On top of that a no-progress guard vetoes repeated identical tool calls (warn at 3, hard veto at 5) and forces a reply, so Atomic never sinks 15 minutes into re-scanning one page the way Hermes did. Both agents missed some of the same questions, and on a few Hermes got there and Atomic did not, usually format slips where Atomic computed the right number but printed the working instead of the bare value. But on identical hardware and identical weights, the runtime that reuses its cache and refuses to spin came out ahead on accuracy and speed. Getting this from the runtime alone is wild. Run the same 53 GAIA tasks on Atomic Agent!
Atomic Agent111,357 Aufrufe • vor 1 Monat
Keine weiteren Inhalte verfügbar