Loading video...
Video Failed to Load
ALERT🚨: AMD INCREASED vLLM PERFORMANCE BY 11x IN LESS THAN 19 DAYS ON MI355X AGENTIC WORKLOADS on the modern MiniMax M3 model! This was done entirely through software optimizations, mainly by optimizing the long-context attention op, along with other optimizations! This is the power of the ROCm stack: buy... show more
69,214 views • 11 days ago •via X (Twitter)
31 Comments

@vllm_project So you mean it was so underoptimized that it was 11x slower before the fix?

@vllm_project An 11x gain in 19 days without changing the chip is the real $AMD bull case: ROCm is finally turning software progress into free hardware upgrades.

@vllm_project AMD $1200

@vllm_project Is this a particular feature of ROCm? Or was it just so unoptimized before? I’m not sure I’d call this a feature…

@vllm_project 11x from software alone means the MI355X launch numbers were never a hardware ceiling, they were a kernel backlog. Every "AMD is behind Nvidia" benchmark from launch week is now measuring vLLM's TODO list, not the silicon.

@vllm_project Software-only gains on long-context attention usually mean they fixed memory layout bottlenecks or kernel dispatch overhead.

@vllm_project Didn’t you say this exact thing for the DSv4 CUDA increase? Like, word for word, fund-and-replace CUDA by Rocm?

@vllm_project 11x from the attention kernel and people still talk about ROCm like a side project. the card was sitting there, the kernel was not

@vllm_project How shit were the $AMD engineers before this release? Like if it could be improved 11x, that means what was shipped s'cked no?

@vllm_project I would expect a vendor to play finally fixing their software issues as some kind of benefit but would not expect a 3rd party to describe praying they decide to finally pay attention to the model you run as their “power”

@vllm_project Hardware's a purchase. Software keeps getting faster for free.

@vllm_project @grok can you vet this information and also when was this posted for this increase?

@vllm_project results?

@vllm_project $AMD

@vllm_project An 11x in 19 days is less a story about how good the software got and more about how much was being left on the table at day zero — nobody finds 11x in a mature stack. Still the right trade for AMD though: silicon ships once, the software gap is the part you can actually close.

@vllm_project next bottleneck is the rocm scheduler

@vllm_project 謝嘍 請謝謝我的肝

@vllm_project just as astra to write better kernel

@vllm_project I wish every GPU I buy came with free performance gains. Oh wait... 11x in 19 days on ROCm.

@vllm_project An eleven fold gain in nineteen days measures how much of that silicon sat idle in July. The updates are free because the first version was not finished.

@vllm_project This is a fantastic example of performance work compounding quickly when the optimization target is chosen well. The 11x result also makes a strong case for treating kernels and serving paths as product surfaces, not invisible plumbing.

@vllm_project Not too long ago you were unbiased. I don't know why ?? Currently very slanted reporting in your post. Particulars with AMD. Every post on AMD is overhyped to the extreme. You've come from the top of the mountain. Now containually rolling down. Very sad to see.

@vllm_project 11× in 19 days on agentic/long-context is a software ceiling unlock, not a silicon step-function. need the baseline (vllm + batch/ctx), tokens/$ and W/token, and whether it generalizes past minimax m3 — otherwise it's one-workload tuning on mi355x

@vllm_project $AMD going to skyrocket Tomorrow as this huge. Called out $AMD and $INTC Thursday along with Memory chips . Your Welcome

@vllm_project Insane, this tech is unreal. The progress is too fast. So cool.

@vllm_project Finally ! They could have done that years ago !

@vllm_project You know it just smells like curry

@vllm_project Huge jump for long context agent runs. Curious how much of that 11x came from prefill speedups versus decode throughput.

@vllm_project the 11x is on one attention op for one model. thats a missing kernel getting written rather than the whole stack getting faster, so the next architecture with a different attention shape starts the clock again. real gain, just doesnt carry forward

@vllm_project what does this Indian have to do with ROCm?

@vllm_project 11x is impressive, but I wonder how much of that was low-hanging fruit in the attention kernel. Sustained gains usually require deeper architectural changes, not just software tweaks. Curious to see long-term benchmarks.


![[Punishing: Gray Raven | Babylonia Dev Comms VOL.1] The sun's past slowly unfolds, as the prelude to homecoming awaits its cue. The first edition of Babylonia Dev Comms is now live. In this issue, we’ll be sharing with Commandants our future story plans, upcoming events, game optimizations, and more—along with a sneak peek at the new content and activities in version "The Dying Sun". Through this Dev Comms, we hope to offer Commandants a new perspective on Kamui.](https://image.24vids.com/tw-2038813544611971229/media/HEqXIVzawAA4EcD.jpg)