正在加载视频...

视频加载失败

Meet Wally. Our inference stack for open frontier models, built to be the fastest place to run them. Performance snapshot: GLM-5.3 Flash: 380 tok/s. GLM-5.3 Max: 790 tok/s Qwen3.8-27B: 485 tok/s DeepSeek-V4.1 Flash: 615 tok/s 1/7

61,190 次观看 • 4 天前 •via X (Twitter)

60 条评论

RunAnywhere 的头像
RunAnywhere4 天前

Benchmarks don't always tell the reality. So we ran a real prompt, same OpenCode harness, GLM-5.3 Flash on Wally, @nebiusai , @FireworksAI_HQ and @Zai_org , all building at once. Wally delivers 43% more throughput than Nebius, 3.1× Fireworks, and 4.5× Z ai. 2/7

RunAnywhere 的头像
RunAnywhere4 天前

Drop Wally into the harness you already use. Live right now for: ->Claude Code ->Claude Desktop ->OpenCode ->DeepSeek Harness ->Hermes ->OpenClaw Blazing fast open-model inference, first token in under 300 ms. $5 free, Drop this in your favorite coding agent to get started: "set up 3/7

RunAnywhere 的头像
RunAnywhere4 天前

We raced Wally, Nebius, Fireworks and Z ai 3 times on a real prompt. Here are the results: 4/7

RunAnywhere 的头像
RunAnywhere4 天前

GLM-5.3 Flash across every provider on Artificial Analysis, Sep 2026. Nebius 313 tok/s. Fireworks 185 tok/s. The model maker's own API: 72 tok/s. Wally: 380 tok/s, measured the same way. 5/7

RunAnywhere 的头像
RunAnywhere4 天前

Every model on wally runs at state-of-the-art speed. That matters the most for realtime apps:Voice pipelines with an LLM in the loop, coding agents, browser automation, live copilots. On our GPUs, or yours(BYOC/on-prem) 6/7

RunAnywhere 的头像
RunAnywhere4 天前

Get Started: Drop this into your favorite coding agent: "set up or Curl : curl -fsSL | sh Windows: irm | iex wally login wally claude-code -m glm-5.3-flash $5 free on signup, no card. Standard per-token after. 7/7

Eric Liu 的头像
Eric Liu4 天前

insane numbers

Deesha Tech 的头像
Deesha Tech4 天前

I couldn't use your console, nothing works. Cannot create keys, cannot run playground, console crashes most of the time and have to refresh. Seems sloppy launch

RunAnywhere 的头像
RunAnywhere4 天前

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Deesha Tech 的头像
Deesha Tech4 天前

Great thanks, its working now

David Gu 的头像
David Gu4 天前

Congrats on the launch!!

RunAnywhere 的头像
RunAnywhere4 天前

🫱🏻‍🫲🏽🙌🏽

Yohei 的头像
Yohei4 天前

installing!

RunAnywhere 的头像
RunAnywhere4 天前

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

MrOzi 的头像
MrOzi4 天前

Those numbers are fast enough to make "Where's Wally?" a rhetorical question.

Zachary Yu 的头像
Zachary Yu4 天前

lfg!

Landseer Enga 的头像
Landseer Enga4 天前

AWESOME!

RunAnywhere 的头像
RunAnywhere4 天前

Wally 🫱🏻‍🫲🏽 @tryrevyl ? 👀

Ishaan 的头像
Ishaan4 天前

Congrats on the launch 🚀

RunAnywhere 的头像
RunAnywhere4 天前

🫱🏻‍🫲🏽

RAZA | AI EXPLORER 的头像
RAZA | AI EXPLORER4 天前

Those inference speeds are seriously impressive.

RunAnywhere 的头像
RunAnywhere4 天前

⚡️⚡️

Luis Costa 的头像
Luis Costa4 天前

Lfg team!! 🔥🔥

RunAnywhere 的头像
RunAnywhere4 天前

🫱🏻‍🫲🏽

Tony Gao 的头像
Tony Gao4 天前

Super cool stuff!

RunAnywhere 的头像
RunAnywhere4 天前

🙏🫱🏻‍🫲🏽

Owen Botkin 的头像
Owen Botkin4 天前

so cool

TensorQuay 的头像
TensorQuay4 天前

The stop button deserves some love too. Wally's docs say Esc in Claude Code cancels the hosted request; during prefill it waits for the first token's request ID. Handy when an agent confidently starts solving the wrong problem.

Shivay Lamba 的头像
Shivay Lamba4 天前

huge congrats on the launch!

RunAnywhere 的头像
RunAnywhere4 天前

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Beshoy Saad 的头像
Beshoy Saad4 天前

where are intelligence numbers? there's gotta be insane fall off...

Nico 的头像
Nico4 天前

nice video

RunAnywhere 的头像
RunAnywhere4 天前

Made by GLM 5.3 flash using Wally :)

Aakash Mahalingam 的头像
Aakash Mahalingam4 天前

LFG 🚀🚀

RunAnywhere 的头像
RunAnywhere4 天前

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Gerard Smit 的头像
Gerard Smit4 天前

Where are the prices? It says it's in the console, but the console doesn't show models? Making a key also doesn't work; too much traffic?😅

RunAnywhere 的头像
RunAnywhere4 天前

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Gerard Smit 的头像
Gerard Smit4 天前

Doesn't seem to be working yet. In EU (Netherlands) I'm getting infinite load on AJAX calls to "

Gerard Smit 的头像
Gerard Smit4 天前

Seems to be working now 😄

Silen Naihin 的头像
Silen Naihin4 天前

LFG

Piyush Yadav | Top1Board | PassReady 的头像
Piyush Yadav | Top1Board | PassReady4 天前

These numbers are absurd—getting nearly 500 tok/s on a 27B parameter model like Qwen completely changes the architecture for real-time applications. Are these benchmarks on Apple Silicon unified memory or a cloud GPU cluster? Need to know before I hook it into My App

MJ 的头像
MJ4 天前

No pricing and so broken

RunAnywhere 的头像
RunAnywhere4 天前

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Peasant Smith 的头像
Peasant Smith4 天前

But did you find... You know?

Sadok 的头像
Sadok4 天前

those deepseek speeds are getting ridiculous

Trunk Monkey 的头像
Trunk Monkey4 天前

Seems pretty decent but I got rate limited just doing some really basic stuff. Did resume a previous session but still.

RunAnywhere 的头像
RunAnywhere4 天前

Thank you for the feedback! We’re working hard to support the demand

Trunk Monkey 的头像
Trunk Monkey4 天前

Seems like a cool service but just hit that again. Likely the agent got stale and just sending its cache + 1 tiny command put me over the limit.

Shadow 的头像
Shadow4 天前

flash isn't a speed tier, it's a compression trick

Velsinha 的头像
Velsinha4 天前

Will you offer plans, or just pay-per-use?

RunAnywhere 的头像
RunAnywhere4 天前

More info soon. Stay tuned!

vikas 的头像
vikas4 天前

guys when are you adding deepseek v4.1 flash? really waiting for it

BayesChat 的头像
BayesChat4 天前

The real-prompt comparison is more useful than a synthetic leaderboard. Same harness and workload make the throughput gap much easier to understand.

Viswesh N G 的头像
Viswesh N G4 天前

Congrats this is awesome!!🔥

cetusian 的头像
cetusian4 天前

@yoheinakajima bro that’s wall-e

RunAnywhere 的头像
RunAnywhere4 天前

@yoheinakajima We like Wally!

sss 的头像
sss4 天前

On what hardware

Nilan Saha 的头像
Nilan Saha4 天前

Whats the quantization? No point in a lobotomized model thats faster

_alphashark_ 的头像
_alphashark_4 天前

at 790 tok/s, a 100-token answer takes about 127 ms to decode. throughput alone won't show whether time-to-first-token or queueing dominates the inference stack.

EDDY VU 的头像
EDDY VU4 天前

790 tok/s on GLM-5.3 Max is insane, how much of that speedup comes from custom kernels versus speculative decoding?

相关视频