Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Meet Wally. Our inference stack for open frontier models, built to be the fastest place to run them. Performance snapshot: GLM-5.3 Flash: 380 tok/s. GLM-5.3 Max: 790 tok/s Qwen3.8-27B: 485 tok/s DeepSeek-V4.1 Flash: 615 tok/s 1/7

61,190 Aufrufe • vor 4 Tagen •via X (Twitter)

60 Kommentare

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

Benchmarks don't always tell the reality. So we ran a real prompt, same OpenCode harness, GLM-5.3 Flash on Wally, @nebiusai , @FireworksAI_HQ and @Zai_org , all building at once. Wally delivers 43% more throughput than Nebius, 3.1× Fireworks, and 4.5× Z ai. 2/7

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

Drop Wally into the harness you already use. Live right now for: ->Claude Code ->Claude Desktop ->OpenCode ->DeepSeek Harness ->Hermes ->OpenClaw Blazing fast open-model inference, first token in under 300 ms. $5 free, Drop this in your favorite coding agent to get started: "set up 3/7

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

We raced Wally, Nebius, Fireworks and Z ai 3 times on a real prompt. Here are the results: 4/7

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

GLM-5.3 Flash across every provider on Artificial Analysis, Sep 2026. Nebius 313 tok/s. Fireworks 185 tok/s. The model maker's own API: 72 tok/s. Wally: 380 tok/s, measured the same way. 5/7

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

Every model on wally runs at state-of-the-art speed. That matters the most for realtime apps:Voice pipelines with an LLM in the loop, coding agents, browser automation, live copilots. On our GPUs, or yours(BYOC/on-prem) 6/7

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

Get Started: Drop this into your favorite coding agent: "set up or Curl : curl -fsSL | sh Windows: irm | iex wally login wally claude-code -m glm-5.3-flash $5 free on signup, no card. Standard per-token after. 7/7

Profilbild von Eric Liu
Eric Liuvor 4 Tagen

insane numbers

Profilbild von Deesha Tech
Deesha Techvor 4 Tagen

I couldn't use your console, nothing works. Cannot create keys, cannot run playground, console crashes most of the time and have to refresh. Seems sloppy launch

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Profilbild von Deesha Tech
Deesha Techvor 4 Tagen

Great thanks, its working now

Profilbild von David Gu
David Guvor 4 Tagen

Congrats on the launch!!

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

🫱🏻‍🫲🏽🙌🏽

Profilbild von Yohei
Yoheivor 4 Tagen

installing!

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Profilbild von MrOzi
MrOzivor 4 Tagen

Those numbers are fast enough to make "Where's Wally?" a rhetorical question.

Profilbild von Zachary Yu
Zachary Yuvor 4 Tagen

lfg!

Profilbild von Landseer Enga
Landseer Engavor 4 Tagen

AWESOME!

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

Wally 🫱🏻‍🫲🏽 @tryrevyl ? 👀

Profilbild von Ishaan
Ishaanvor 4 Tagen

Congrats on the launch 🚀

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

🫱🏻‍🫲🏽

Profilbild von RAZA | AI EXPLORER
RAZA | AI EXPLORERvor 4 Tagen

Those inference speeds are seriously impressive.

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

⚡️⚡️

Profilbild von Luis Costa
Luis Costavor 4 Tagen

Lfg team!! 🔥🔥

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

🫱🏻‍🫲🏽

Profilbild von Tony Gao
Tony Gaovor 4 Tagen

Super cool stuff!

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

🙏🫱🏻‍🫲🏽

Profilbild von Owen Botkin
Owen Botkinvor 4 Tagen

so cool

Profilbild von TensorQuay
TensorQuayvor 4 Tagen

The stop button deserves some love too. Wally's docs say Esc in Claude Code cancels the hosted request; during prefill it waits for the first token's request ID. Handy when an agent confidently starts solving the wrong problem.

Profilbild von Shivay Lamba
Shivay Lambavor 4 Tagen

huge congrats on the launch!

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Profilbild von Beshoy Saad
Beshoy Saadvor 4 Tagen

where are intelligence numbers? there's gotta be insane fall off...

Profilbild von Nico
Nicovor 4 Tagen

nice video

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

Made by GLM 5.3 flash using Wally :)

Profilbild von Aakash Mahalingam
Aakash Mahalingamvor 4 Tagen

LFG 🚀🚀

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Profilbild von Gerard Smit
Gerard Smitvor 4 Tagen

Where are the prices? It says it's in the console, but the console doesn't show models? Making a key also doesn't work; too much traffic?😅

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Profilbild von Gerard Smit
Gerard Smitvor 4 Tagen

Doesn't seem to be working yet. In EU (Netherlands) I'm getting infinite load on AJAX calls to "

Profilbild von Gerard Smit
Gerard Smitvor 4 Tagen

Seems to be working now 😄

Profilbild von Silen Naihin
Silen Naihinvor 4 Tagen

LFG

Profilbild von Piyush Yadav | Top1Board | PassReady
Piyush Yadav | Top1Board | PassReadyvor 4 Tagen

These numbers are absurd—getting nearly 500 tok/s on a 27B parameter model like Qwen completely changes the architecture for real-time applications. Are these benchmarks on Apple Silicon unified memory or a cloud GPU cluster? Need to know before I hook it into My App

Profilbild von MJ
MJvor 4 Tagen

No pricing and so broken

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Profilbild von Peasant Smith
Peasant Smithvor 4 Tagen

But did you find... You know?

Profilbild von Sadok
Sadokvor 4 Tagen

those deepseek speeds are getting ridiculous

Profilbild von Trunk Monkey
Trunk Monkeyvor 4 Tagen

Seems pretty decent but I got rate limited just doing some really basic stuff. Did resume a previous session but still.

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

Thank you for the feedback! We’re working hard to support the demand

Profilbild von Trunk Monkey
Trunk Monkeyvor 4 Tagen

Seems like a cool service but just hit that again. Likely the agent got stale and just sending its cache + 1 tiny command put me over the limit.

Profilbild von Shadow
Shadowvor 4 Tagen

flash isn't a speed tier, it's a compression trick

Profilbild von Velsinha
Velsinhavor 4 Tagen

Will you offer plans, or just pay-per-use?

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

More info soon. Stay tuned!

Profilbild von vikas
vikasvor 4 Tagen

guys when are you adding deepseek v4.1 flash? really waiting for it

Profilbild von BayesChat
BayesChatvor 4 Tagen

The real-prompt comparison is more useful than a synthetic leaderboard. Same harness and workload make the throughput gap much easier to understand.

Profilbild von Viswesh N G
Viswesh N Gvor 4 Tagen

Congrats this is awesome!!🔥

Profilbild von cetusian
cetusianvor 4 Tagen

@yoheinakajima bro that’s wall-e

Profilbild von RunAnywhere
RunAnywherevor 4 Tagen

@yoheinakajima We like Wally!

Profilbild von sss
sssvor 4 Tagen

On what hardware

Profilbild von Nilan Saha
Nilan Sahavor 4 Tagen

Whats the quantization? No point in a lobotomized model thats faster

Profilbild von _alphashark_
_alphashark_vor 4 Tagen

at 790 tok/s, a 100-token answer takes about 127 ms to decode. throughput alone won't show whether time-to-first-token or queueing dominates the inference stack.

Profilbild von EDDY VU
EDDY VUvor 4 Tagen

790 tok/s on GLM-5.3 Max is insane, how much of that speedup comes from custom kernels versus speculative decoding?

Ähnliche Videos