Loading video...

Video Failed to Load

Go Home

Meet Wally. Our inference stack for open frontier models, built to be the fastest place to run them. Performance snapshot: GLM-5.3 Flash: 380 tok/s. GLM-5.3 Max: 790 tok/s Qwen3.8-27B: 485 tok/s DeepSeek-V4.1 Flash: 615 tok/s 1/7

61,190 views • 4 days ago •via X (Twitter)

60 Comments

RunAnywhere's profile picture
RunAnywhere4 days ago

Benchmarks don't always tell the reality. So we ran a real prompt, same OpenCode harness, GLM-5.3 Flash on Wally, @nebiusai , @FireworksAI_HQ and @Zai_org , all building at once. Wally delivers 43% more throughput than Nebius, 3.1× Fireworks, and 4.5× Z ai. 2/7

RunAnywhere's profile picture
RunAnywhere4 days ago

Drop Wally into the harness you already use. Live right now for: ->Claude Code ->Claude Desktop ->OpenCode ->DeepSeek Harness ->Hermes ->OpenClaw Blazing fast open-model inference, first token in under 300 ms. $5 free, Drop this in your favorite coding agent to get started: "set up 3/7

RunAnywhere's profile picture
RunAnywhere4 days ago

We raced Wally, Nebius, Fireworks and Z ai 3 times on a real prompt. Here are the results: 4/7

RunAnywhere's profile picture
RunAnywhere4 days ago

GLM-5.3 Flash across every provider on Artificial Analysis, Sep 2026. Nebius 313 tok/s. Fireworks 185 tok/s. The model maker's own API: 72 tok/s. Wally: 380 tok/s, measured the same way. 5/7

RunAnywhere's profile picture
RunAnywhere4 days ago

Every model on wally runs at state-of-the-art speed. That matters the most for realtime apps:Voice pipelines with an LLM in the loop, coding agents, browser automation, live copilots. On our GPUs, or yours(BYOC/on-prem) 6/7

RunAnywhere's profile picture
RunAnywhere4 days ago

Get Started: Drop this into your favorite coding agent: "set up or Curl : curl -fsSL | sh Windows: irm | iex wally login wally claude-code -m glm-5.3-flash $5 free on signup, no card. Standard per-token after. 7/7

Eric Liu's profile picture
Eric Liu4 days ago

insane numbers

Deesha Tech's profile picture
Deesha Tech4 days ago

I couldn't use your console, nothing works. Cannot create keys, cannot run playground, console crashes most of the time and have to refresh. Seems sloppy launch

RunAnywhere's profile picture
RunAnywhere4 days ago

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Deesha Tech's profile picture
Deesha Tech4 days ago

Great thanks, its working now

David Gu's profile picture
David Gu4 days ago

Congrats on the launch!!

RunAnywhere's profile picture
RunAnywhere4 days ago

🫱🏻‍🫲🏽🙌🏽

Yohei's profile picture
Yohei4 days ago

installing!

RunAnywhere's profile picture
RunAnywhere4 days ago

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

MrOzi's profile picture
MrOzi4 days ago

Those numbers are fast enough to make "Where's Wally?" a rhetorical question.

Zachary Yu's profile picture
Zachary Yu4 days ago

lfg!

Landseer Enga's profile picture
Landseer Enga4 days ago

AWESOME!

RunAnywhere's profile picture
RunAnywhere4 days ago

Wally 🫱🏻‍🫲🏽 @tryrevyl ? 👀

Ishaan's profile picture
Ishaan4 days ago

Congrats on the launch 🚀

RunAnywhere's profile picture
RunAnywhere4 days ago

🫱🏻‍🫲🏽

RAZA | AI EXPLORER's profile picture
RAZA | AI EXPLORER4 days ago

Those inference speeds are seriously impressive.

RunAnywhere's profile picture
RunAnywhere4 days ago

⚡️⚡️

Luis Costa's profile picture
Luis Costa4 days ago

Lfg team!! 🔥🔥

RunAnywhere's profile picture
RunAnywhere4 days ago

🫱🏻‍🫲🏽

Tony Gao's profile picture
Tony Gao4 days ago

Super cool stuff!

RunAnywhere's profile picture
RunAnywhere4 days ago

🙏🫱🏻‍🫲🏽

Owen Botkin's profile picture
Owen Botkin4 days ago

so cool

TensorQuay's profile picture
TensorQuay4 days ago

The stop button deserves some love too. Wally's docs say Esc in Claude Code cancels the hosted request; during prefill it waits for the first token's request ID. Handy when an agent confidently starts solving the wrong problem.

Shivay Lamba's profile picture
Shivay Lamba4 days ago

huge congrats on the launch!

RunAnywhere's profile picture
RunAnywhere4 days ago

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Beshoy Saad's profile picture
Beshoy Saad4 days ago

where are intelligence numbers? there's gotta be insane fall off...

Nico's profile picture
Nico4 days ago

nice video

RunAnywhere's profile picture
RunAnywhere4 days ago

Made by GLM 5.3 flash using Wally :)

Aakash Mahalingam's profile picture
Aakash Mahalingam4 days ago

LFG 🚀🚀

RunAnywhere's profile picture
RunAnywhere4 days ago

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Gerard Smit's profile picture
Gerard Smit4 days ago

Where are the prices? It says it's in the console, but the console doesn't show models? Making a key also doesn't work; too much traffic?😅

RunAnywhere's profile picture
RunAnywhere4 days ago

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Gerard Smit's profile picture
Gerard Smit4 days ago

Doesn't seem to be working yet. In EU (Netherlands) I'm getting infinite load on AJAX calls to "

Gerard Smit's profile picture
Gerard Smit4 days ago

Seems to be working now 😄

Silen Naihin's profile picture
Silen Naihin4 days ago

LFG

Piyush Yadav | Top1Board | PassReady's profile picture
Piyush Yadav | Top1Board | PassReady4 days ago

These numbers are absurd—getting nearly 500 tok/s on a 27B parameter model like Qwen completely changes the architecture for real-time applications. Are these benchmarks on Apple Silicon unified memory or a cloud GPU cluster? Need to know before I hook it into My App

MJ's profile picture
MJ4 days ago

No pricing and so broken

RunAnywhere's profile picture
RunAnywhere4 days ago

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Peasant Smith's profile picture
Peasant Smith4 days ago

But did you find... You know?

Sadok's profile picture
Sadok4 days ago

those deepseek speeds are getting ridiculous

Trunk Monkey's profile picture
Trunk Monkey4 days ago

Seems pretty decent but I got rate limited just doing some really basic stuff. Did resume a previous session but still.

RunAnywhere's profile picture
RunAnywhere4 days ago

Thank you for the feedback! We’re working hard to support the demand

Trunk Monkey's profile picture
Trunk Monkey4 days ago

Seems like a cool service but just hit that again. Likely the agent got stale and just sending its cache + 1 tiny command put me over the limit.

Shadow's profile picture
Shadow4 days ago

flash isn't a speed tier, it's a compression trick

Velsinha's profile picture
Velsinha4 days ago

Will you offer plans, or just pay-per-use?

RunAnywhere's profile picture
RunAnywhere4 days ago

More info soon. Stay tuned!

vikas's profile picture
vikas4 days ago

guys when are you adding deepseek v4.1 flash? really waiting for it

BayesChat's profile picture
BayesChat4 days ago

The real-prompt comparison is more useful than a synthetic leaderboard. Same harness and workload make the throughput gap much easier to understand.

Viswesh N G's profile picture
Viswesh N G4 days ago

Congrats this is awesome!!🔥

cetusian's profile picture
cetusian4 days ago

@yoheinakajima bro that’s wall-e

RunAnywhere's profile picture
RunAnywhere4 days ago

@yoheinakajima We like Wally!

sss's profile picture
sss4 days ago

On what hardware

Nilan Saha's profile picture
Nilan Saha4 days ago

Whats the quantization? No point in a lobotomized model thats faster

_alphashark_'s profile picture
_alphashark_4 days ago

at 790 tok/s, a 100-token answer takes about 127 ms to decode. throughput alone won't show whether time-to-first-token or queueing dominates the inference stack.

EDDY VU's profile picture
EDDY VU4 days ago

790 tok/s on GLM-5.3 Max is insane, how much of that speedup comes from custom kernels versus speculative decoding?

Related Videos