Загрузка видео...
Не удалось загрузить видео
Meet Wally. Our inference stack for open frontier models, built to be the fastest place to run them. Performance snapshot: GLM-5.3 Flash: 380 tok/s. GLM-5.3 Max: 790 tok/s Qwen3.8-27B: 485 tok/s DeepSeek-V4.1 Flash: 615 tok/s 1/7
61,190 просмотров • 4 дней назад •via X (Twitter)
Комментарии: 60

Benchmarks don't always tell the reality. So we ran a real prompt, same OpenCode harness, GLM-5.3 Flash on Wally, @nebiusai , @FireworksAI_HQ and @Zai_org , all building at once. Wally delivers 43% more throughput than Nebius, 3.1× Fireworks, and 4.5× Z ai. 2/7

Drop Wally into the harness you already use. Live right now for: ->Claude Code ->Claude Desktop ->OpenCode ->DeepSeek Harness ->Hermes ->OpenClaw Blazing fast open-model inference, first token in under 300 ms. $5 free, Drop this in your favorite coding agent to get started: "set up 3/7

We raced Wally, Nebius, Fireworks and Z ai 3 times on a real prompt. Here are the results: 4/7

GLM-5.3 Flash across every provider on Artificial Analysis, Sep 2026. Nebius 313 tok/s. Fireworks 185 tok/s. The model maker's own API: 72 tok/s. Wally: 380 tok/s, measured the same way. 5/7

Every model on wally runs at state-of-the-art speed. That matters the most for realtime apps:Voice pipelines with an LLM in the loop, coding agents, browser automation, live copilots. On our GPUs, or yours(BYOC/on-prem) 6/7

Get Started: Drop this into your favorite coding agent: "set up or Curl : curl -fsSL | sh Windows: irm | iex wally login wally claude-code -m glm-5.3-flash $5 free on signup, no card. Standard per-token after. 7/7

insane numbers

I couldn't use your console, nothing works. Cannot create keys, cannot run playground, console crashes most of the time and have to refresh. Seems sloppy launch

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Great thanks, its working now

Congrats on the launch!!

🫱🏻🫲🏽🙌🏽

installing!

🫱🏻🫲🏽🫱🏻🫲🏽

Those numbers are fast enough to make "Where's Wally?" a rhetorical question.

lfg!

AWESOME!

Wally 🫱🏻🫲🏽 @tryrevyl ? 👀

Congrats on the launch 🚀

🫱🏻🫲🏽

Those inference speeds are seriously impressive.

⚡️⚡️

Lfg team!! 🔥🔥

🫱🏻🫲🏽

Super cool stuff!

🙏🫱🏻🫲🏽

so cool

The stop button deserves some love too. Wally's docs say Esc in Claude Code cancels the hosted request; during prefill it waits for the first token's request ID. Handy when an agent confidently starts solving the wrong problem.

huge congrats on the launch!

🫱🏻🫲🏽🫱🏻🫲🏽

where are intelligence numbers? there's gotta be insane fall off...

nice video

Made by GLM 5.3 flash using Wally :)

LFG 🚀🚀

🫱🏻🫲🏽🫱🏻🫲🏽

Where are the prices? It says it's in the console, but the console doesn't show models? Making a key also doesn't work; too much traffic?😅

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Doesn't seem to be working yet. In EU (Netherlands) I'm getting infinite load on AJAX calls to "

Seems to be working now 😄

LFG

These numbers are absurd—getting nearly 500 tok/s on a 27B parameter model like Qwen completely changes the architecture for real-time applications. Are these benchmarks on Apple Silicon unified memory or a cloud GPU cluster? Need to know before I hook it into My App

No pricing and so broken

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

But did you find... You know?

those deepseek speeds are getting ridiculous

Seems pretty decent but I got rate limited just doing some really basic stuff. Did resume a previous session but still.

Thank you for the feedback! We’re working hard to support the demand

Seems like a cool service but just hit that again. Likely the agent got stale and just sending its cache + 1 tiny command put me over the limit.

flash isn't a speed tier, it's a compression trick

Will you offer plans, or just pay-per-use?

More info soon. Stay tuned!

guys when are you adding deepseek v4.1 flash? really waiting for it

The real-prompt comparison is more useful than a synthetic leaderboard. Same harness and workload make the throughput gap much easier to understand.

Congrats this is awesome!!🔥

@yoheinakajima bro that’s wall-e

@yoheinakajima We like Wally!

On what hardware

Whats the quantization? No point in a lobotomized model thats faster

at 790 tok/s, a 100-token answer takes about 127 ms to decode. throughput alone won't show whether time-to-first-token or queueing dominates the inference stack.

790 tok/s on GLM-5.3 Max is insane, how much of that speedup comes from custom kernels versus speculative decoding?
