Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Meet Wally. Our inference stack for open frontier models, built to be the fastest place to run them. Performance snapshot: GLM-5.3 Flash: 380 tok/s. GLM-5.3 Max: 790 tok/s Qwen3.8-27B: 485 tok/s DeepSeek-V4.1 Flash: 615 tok/s 1/7

61,190 görüntüleme • 4 gün önce •via X (Twitter)

60 Yorum

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

Benchmarks don't always tell the reality. So we ran a real prompt, same OpenCode harness, GLM-5.3 Flash on Wally, @nebiusai , @FireworksAI_HQ and @Zai_org , all building at once. Wally delivers 43% more throughput than Nebius, 3.1× Fireworks, and 4.5× Z ai. 2/7

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

Drop Wally into the harness you already use. Live right now for: ->Claude Code ->Claude Desktop ->OpenCode ->DeepSeek Harness ->Hermes ->OpenClaw Blazing fast open-model inference, first token in under 300 ms. $5 free, Drop this in your favorite coding agent to get started: "set up 3/7

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

We raced Wally, Nebius, Fireworks and Z ai 3 times on a real prompt. Here are the results: 4/7

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

GLM-5.3 Flash across every provider on Artificial Analysis, Sep 2026. Nebius 313 tok/s. Fireworks 185 tok/s. The model maker's own API: 72 tok/s. Wally: 380 tok/s, measured the same way. 5/7

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

Every model on wally runs at state-of-the-art speed. That matters the most for realtime apps:Voice pipelines with an LLM in the loop, coding agents, browser automation, live copilots. On our GPUs, or yours(BYOC/on-prem) 6/7

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

Get Started: Drop this into your favorite coding agent: "set up or Curl : curl -fsSL | sh Windows: irm | iex wally login wally claude-code -m glm-5.3-flash $5 free on signup, no card. Standard per-token after. 7/7

Eric Liu profil fotoğrafı
Eric Liu4 gün önce

insane numbers

Deesha Tech profil fotoğrafı
Deesha Tech4 gün önce

I couldn't use your console, nothing works. Cannot create keys, cannot run playground, console crashes most of the time and have to refresh. Seems sloppy launch

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Deesha Tech profil fotoğrafı
Deesha Tech4 gün önce

Great thanks, its working now

David Gu profil fotoğrafı
David Gu4 gün önce

Congrats on the launch!!

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

🫱🏻‍🫲🏽🙌🏽

Yohei profil fotoğrafı
Yohei4 gün önce

installing!

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

MrOzi profil fotoğrafı
MrOzi4 gün önce

Those numbers are fast enough to make "Where's Wally?" a rhetorical question.

Zachary Yu profil fotoğrafı
Zachary Yu4 gün önce

lfg!

Landseer Enga profil fotoğrafı
Landseer Enga4 gün önce

AWESOME!

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

Wally 🫱🏻‍🫲🏽 @tryrevyl ? 👀

Ishaan profil fotoğrafı
Ishaan4 gün önce

Congrats on the launch 🚀

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

🫱🏻‍🫲🏽

RAZA | AI EXPLORER profil fotoğrafı
RAZA | AI EXPLORER4 gün önce

Those inference speeds are seriously impressive.

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

⚡️⚡️

Luis Costa profil fotoğrafı
Luis Costa4 gün önce

Lfg team!! 🔥🔥

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

🫱🏻‍🫲🏽

Tony Gao profil fotoğrafı
Tony Gao4 gün önce

Super cool stuff!

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

🙏🫱🏻‍🫲🏽

Owen Botkin profil fotoğrafı
Owen Botkin4 gün önce

so cool

TensorQuay profil fotoğrafı
TensorQuay4 gün önce

The stop button deserves some love too. Wally's docs say Esc in Claude Code cancels the hosted request; during prefill it waits for the first token's request ID. Handy when an agent confidently starts solving the wrong problem.

Shivay Lamba profil fotoğrafı
Shivay Lamba4 gün önce

huge congrats on the launch!

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Beshoy Saad profil fotoğrafı
Beshoy Saad4 gün önce

where are intelligence numbers? there's gotta be insane fall off...

Nico profil fotoğrafı
Nico4 gün önce

nice video

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

Made by GLM 5.3 flash using Wally :)

Aakash Mahalingam profil fotoğrafı
Aakash Mahalingam4 gün önce

LFG 🚀🚀

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Gerard Smit profil fotoğrafı
Gerard Smit4 gün önce

Where are the prices? It says it's in the console, but the console doesn't show models? Making a key also doesn't work; too much traffic?😅

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Gerard Smit profil fotoğrafı
Gerard Smit4 gün önce

Doesn't seem to be working yet. In EU (Netherlands) I'm getting infinite load on AJAX calls to "

Gerard Smit profil fotoğrafı
Gerard Smit4 gün önce

Seems to be working now 😄

Silen Naihin profil fotoğrafı
Silen Naihin4 gün önce

LFG

Piyush Yadav | Top1Board | PassReady profil fotoğrafı
Piyush Yadav | Top1Board | PassReady4 gün önce

These numbers are absurd—getting nearly 500 tok/s on a 27B parameter model like Qwen completely changes the architecture for real-time applications. Are these benchmarks on Apple Silicon unified memory or a cloud GPU cluster? Need to know before I hook it into My App

MJ profil fotoğrafı
MJ4 gün önce

No pricing and so broken

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Peasant Smith profil fotoğrafı
Peasant Smith4 gün önce

But did you find... You know?

Sadok profil fotoğrafı
Sadok4 gün önce

those deepseek speeds are getting ridiculous

Trunk Monkey profil fotoğrafı
Trunk Monkey4 gün önce

Seems pretty decent but I got rate limited just doing some really basic stuff. Did resume a previous session but still.

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

Thank you for the feedback! We’re working hard to support the demand

Trunk Monkey profil fotoğrafı
Trunk Monkey4 gün önce

Seems like a cool service but just hit that again. Likely the agent got stale and just sending its cache + 1 tiny command put me over the limit.

Shadow profil fotoğrafı
Shadow4 gün önce

flash isn't a speed tier, it's a compression trick

Velsinha profil fotoğrafı
Velsinha4 gün önce

Will you offer plans, or just pay-per-use?

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

More info soon. Stay tuned!

vikas profil fotoğrafı
vikas4 gün önce

guys when are you adding deepseek v4.1 flash? really waiting for it

BayesChat profil fotoğrafı
BayesChat4 gün önce

The real-prompt comparison is more useful than a synthetic leaderboard. Same harness and workload make the throughput gap much easier to understand.

Viswesh N G profil fotoğrafı
Viswesh N G4 gün önce

Congrats this is awesome!!🔥

cetusian profil fotoğrafı
cetusian4 gün önce

@yoheinakajima bro that’s wall-e

RunAnywhere profil fotoğrafı
RunAnywhere4 gün önce

@yoheinakajima We like Wally!

sss profil fotoğrafı
sss4 gün önce

On what hardware

Nilan Saha profil fotoğrafı
Nilan Saha4 gün önce

Whats the quantization? No point in a lobotomized model thats faster

_alphashark_ profil fotoğrafı
_alphashark_4 gün önce

at 790 tok/s, a 100-token answer takes about 127 ms to decode. throughput alone won't show whether time-to-first-token or queueing dominates the inference stack.

EDDY VU profil fotoğrafı
EDDY VU4 gün önce

790 tok/s on GLM-5.3 Max is insane, how much of that speedup comes from custom kernels versus speculative decoding?

Benzer Videolar