Загрузка видео...

Не удалось загрузить видео

На главную

Meet Wally. Our inference stack for open frontier models, built to be the fastest place to run them. Performance snapshot: GLM-5.3 Flash: 380 tok/s. GLM-5.3 Max: 790 tok/s Qwen3.8-27B: 485 tok/s DeepSeek-V4.1 Flash: 615 tok/s 1/7

61,190 просмотров • 4 дней назад •via X (Twitter)

Комментарии: 60

Фото профиля RunAnywhere
RunAnywhere4 дней назад

Benchmarks don't always tell the reality. So we ran a real prompt, same OpenCode harness, GLM-5.3 Flash on Wally, @nebiusai , @FireworksAI_HQ and @Zai_org , all building at once. Wally delivers 43% more throughput than Nebius, 3.1× Fireworks, and 4.5× Z ai. 2/7

Фото профиля RunAnywhere
RunAnywhere4 дней назад

Drop Wally into the harness you already use. Live right now for: ->Claude Code ->Claude Desktop ->OpenCode ->DeepSeek Harness ->Hermes ->OpenClaw Blazing fast open-model inference, first token in under 300 ms. $5 free, Drop this in your favorite coding agent to get started: "set up 3/7

Фото профиля RunAnywhere
RunAnywhere4 дней назад

We raced Wally, Nebius, Fireworks and Z ai 3 times on a real prompt. Here are the results: 4/7

Фото профиля RunAnywhere
RunAnywhere4 дней назад

GLM-5.3 Flash across every provider on Artificial Analysis, Sep 2026. Nebius 313 tok/s. Fireworks 185 tok/s. The model maker's own API: 72 tok/s. Wally: 380 tok/s, measured the same way. 5/7

Фото профиля RunAnywhere
RunAnywhere4 дней назад

Every model on wally runs at state-of-the-art speed. That matters the most for realtime apps:Voice pipelines with an LLM in the loop, coding agents, browser automation, live copilots. On our GPUs, or yours(BYOC/on-prem) 6/7

Фото профиля RunAnywhere
RunAnywhere4 дней назад

Get Started: Drop this into your favorite coding agent: "set up or Curl : curl -fsSL | sh Windows: irm | iex wally login wally claude-code -m glm-5.3-flash $5 free on signup, no card. Standard per-token after. 7/7

Фото профиля Eric Liu
Eric Liu4 дней назад

insane numbers

Фото профиля Deesha Tech
Deesha Tech4 дней назад

I couldn't use your console, nothing works. Cannot create keys, cannot run playground, console crashes most of the time and have to refresh. Seems sloppy launch

Фото профиля RunAnywhere
RunAnywhere4 дней назад

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Фото профиля Deesha Tech
Deesha Tech4 дней назад

Great thanks, its working now

Фото профиля David Gu
David Gu4 дней назад

Congrats on the launch!!

Фото профиля RunAnywhere
RunAnywhere4 дней назад

🫱🏻‍🫲🏽🙌🏽

Фото профиля Yohei
Yohei4 дней назад

installing!

Фото профиля RunAnywhere
RunAnywhere4 дней назад

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Фото профиля MrOzi
MrOzi4 дней назад

Those numbers are fast enough to make "Where's Wally?" a rhetorical question.

Фото профиля Zachary Yu
Zachary Yu4 дней назад

lfg!

Фото профиля Landseer Enga
Landseer Enga4 дней назад

AWESOME!

Фото профиля RunAnywhere
RunAnywhere4 дней назад

Wally 🫱🏻‍🫲🏽 @tryrevyl ? 👀

Фото профиля Ishaan
Ishaan4 дней назад

Congrats on the launch 🚀

Фото профиля RunAnywhere
RunAnywhere4 дней назад

🫱🏻‍🫲🏽

Фото профиля RAZA | AI EXPLORER
RAZA | AI EXPLORER4 дней назад

Those inference speeds are seriously impressive.

Фото профиля RunAnywhere
RunAnywhere4 дней назад

⚡️⚡️

Фото профиля Luis Costa
Luis Costa4 дней назад

Lfg team!! 🔥🔥

Фото профиля RunAnywhere
RunAnywhere4 дней назад

🫱🏻‍🫲🏽

Фото профиля Tony Gao
Tony Gao4 дней назад

Super cool stuff!

Фото профиля RunAnywhere
RunAnywhere4 дней назад

🙏🫱🏻‍🫲🏽

Фото профиля Owen Botkin
Owen Botkin4 дней назад

so cool

Фото профиля TensorQuay
TensorQuay4 дней назад

The stop button deserves some love too. Wally's docs say Esc in Claude Code cancels the hosted request; during prefill it waits for the first token's request ID. Handy when an agent confidently starts solving the wrong problem.

Фото профиля Shivay Lamba
Shivay Lamba4 дней назад

huge congrats on the launch!

Фото профиля RunAnywhere
RunAnywhere4 дней назад

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Фото профиля Beshoy Saad
Beshoy Saad4 дней назад

where are intelligence numbers? there's gotta be insane fall off...

Фото профиля Nico
Nico4 дней назад

nice video

Фото профиля RunAnywhere
RunAnywhere4 дней назад

Made by GLM 5.3 flash using Wally :)

Фото профиля Aakash Mahalingam
Aakash Mahalingam4 дней назад

LFG 🚀🚀

Фото профиля RunAnywhere
RunAnywhere4 дней назад

🫱🏻‍🫲🏽🫱🏻‍🫲🏽

Фото профиля Gerard Smit
Gerard Smit4 дней назад

Where are the prices? It says it's in the console, but the console doesn't show models? Making a key also doesn't work; too much traffic?😅

Фото профиля RunAnywhere
RunAnywhere4 дней назад

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Фото профиля Gerard Smit
Gerard Smit4 дней назад

Doesn't seem to be working yet. In EU (Netherlands) I'm getting infinite load on AJAX calls to "

Фото профиля Gerard Smit
Gerard Smit4 дней назад

Seems to be working now 😄

Фото профиля Silen Naihin
Silen Naihin4 дней назад

LFG

Фото профиля Piyush Yadav | Top1Board | PassReady
Piyush Yadav | Top1Board | PassReady4 дней назад

These numbers are absurd—getting nearly 500 tok/s on a 27B parameter model like Qwen completely changes the architecture for real-time applications. Are these benchmarks on Apple Silicon unified memory or a cloud GPU cluster? Need to know before I hook it into My App

Фото профиля MJ
MJ4 дней назад

No pricing and so broken

Фото профиля RunAnywhere
RunAnywhere4 дней назад

Hey, sorry for the inconvenience. The console may have been unavailable for a short period due to an unexpected surge in traffic. Everything is back up and running now.

Фото профиля Peasant Smith
Peasant Smith4 дней назад

But did you find... You know?

Фото профиля Sadok
Sadok4 дней назад

those deepseek speeds are getting ridiculous

Фото профиля Trunk Monkey
Trunk Monkey4 дней назад

Seems pretty decent but I got rate limited just doing some really basic stuff. Did resume a previous session but still.

Фото профиля RunAnywhere
RunAnywhere4 дней назад

Thank you for the feedback! We’re working hard to support the demand

Фото профиля Trunk Monkey
Trunk Monkey4 дней назад

Seems like a cool service but just hit that again. Likely the agent got stale and just sending its cache + 1 tiny command put me over the limit.

Фото профиля Shadow
Shadow4 дней назад

flash isn't a speed tier, it's a compression trick

Фото профиля Velsinha
Velsinha4 дней назад

Will you offer plans, or just pay-per-use?

Фото профиля RunAnywhere
RunAnywhere4 дней назад

More info soon. Stay tuned!

Фото профиля vikas
vikas4 дней назад

guys when are you adding deepseek v4.1 flash? really waiting for it

Фото профиля BayesChat
BayesChat4 дней назад

The real-prompt comparison is more useful than a synthetic leaderboard. Same harness and workload make the throughput gap much easier to understand.

Фото профиля Viswesh N G
Viswesh N G4 дней назад

Congrats this is awesome!!🔥

Фото профиля cetusian
cetusian4 дней назад

@yoheinakajima bro that’s wall-e

Фото профиля RunAnywhere
RunAnywhere4 дней назад

@yoheinakajima We like Wally!

Фото профиля sss
sss4 дней назад

On what hardware

Фото профиля Nilan Saha
Nilan Saha4 дней назад

Whats the quantization? No point in a lobotomized model thats faster

Фото профиля _alphashark_
_alphashark_4 дней назад

at 790 tok/s, a 100-token answer takes about 127 ms to decode. throughput alone won't show whether time-to-first-token or queueing dominates the inference stack.

Фото профиля EDDY VU
EDDY VU4 дней назад

790 tok/s on GLM-5.3 Max is insane, how much of that speedup comes from custom kernels versus speculative decoding?

Похожие видео