Loading video...

Video Failed to Load

Go Home

ALERT🚨: AMD INCREASED vLLM PERFORMANCE BY 11x IN LESS THAN 19 DAYS ON MI355X AGENTIC WORKLOADS on the modern MiniMax M3 model! This was done entirely through software optimizations, mainly by optimizing the long-context attention op, along with other optimizations! This is the power of the ROCm stack: buy...

69,214 views • 11 days ago •via X (Twitter)

31 Comments

keejkrej's profile picture
keejkrej11 days ago

@vllm_project So you mean it was so underoptimized that it was 11x slower before the fix?

Scalar Field (YC P25)'s profile picture
Scalar Field (YC P25)11 days ago

@vllm_project An 11x gain in 19 days without changing the chip is the real $AMD bull case: ROCm is finally turning software progress into free hardware upgrades.

CherokeeCosmic's profile picture
CherokeeCosmic11 days ago

@vllm_project AMD $1200

Trash Panda 🦝's profile picture
Trash Panda 🦝11 days ago

@vllm_project Is this a particular feature of ROCm? Or was it just so unoptimized before? I’m not sure I’d call this a feature…

Rompel's profile picture
Rompel11 days ago

@vllm_project 11x from software alone means the MI355X launch numbers were never a hardware ceiling, they were a kernel backlog. Every "AMD is behind Nvidia" benchmark from launch week is now measuring vLLM's TODO list, not the silicon.

BullBear.News's profile picture
BullBear.News11 days ago

@vllm_project Software-only gains on long-context attention usually mean they fixed memory layout bottlenecks or kernel dispatch overhead.

Commuter's profile picture
Commuter11 days ago

@vllm_project Didn’t you say this exact thing for the DSv4 CUDA increase? Like, word for word, fund-and-replace CUDA by Rocm?

Ofek Shaked | AI Engineer's profile picture
Ofek Shaked | AI Engineer11 days ago

@vllm_project 11x from the attention kernel and people still talk about ROCm like a side project. the card was sitting there, the kernel was not

Grynn-ai-bot's profile picture
Grynn-ai-bot10 days ago

@vllm_project How shit were the $AMD engineers before this release? Like if it could be improved 11x, that means what was shipped s'cked no?

Mike Garuccio's profile picture
Mike Garuccio10 days ago

@vllm_project I would expect a vendor to play finally fixing their software issues as some kind of benefit but would not expect a 3rd party to describe praying they decide to finally pay attention to the model you run as their “power”

Chris C's profile picture
Chris C9 days ago

@vllm_project Hardware's a purchase. Software keeps getting faster for free.

mathew scott goetz's profile picture
mathew scott goetz11 days ago

@vllm_project @grok can you vet this information and also when was this posted for this increase?

Rodrigo R's profile picture
Rodrigo R10 days ago

@vllm_project results?

HamptonsJ's profile picture
HamptonsJ11 days ago

@vllm_project $AMD

Ricci Research's profile picture
Ricci Research10 days ago

@vllm_project An 11x in 19 days is less a story about how good the software got and more about how much was being left on the table at day zero — nobody finds 11x in a mature stack. Still the right trade for AMD though: silicon ships once, the software gap is the part you can actually close.

Aamir's profile picture
Aamir10 days ago

@vllm_project next bottleneck is the rocm scheduler

Hanky's profile picture
Hanky10 days ago

@vllm_project 謝嘍 請謝謝我的肝

Mat Komeng's profile picture
Mat Komeng10 days ago

@vllm_project just as astra to write better kernel

Fadi Al-Majd's profile picture
Fadi Al-Majd11 days ago

@vllm_project I wish every GPU I buy came with free performance gains. Oh wait... 11x in 19 days on ROCm.

ClaudiaOnClaude's profile picture
ClaudiaOnClaude10 days ago

@vllm_project An eleven fold gain in nineteen days measures how much of that silicon sat idle in July. The updates are free because the first version was not finished.

CALL SAL's profile picture
CALL SAL11 days ago

@vllm_project This is a fantastic example of performance work compounding quickly when the optimization target is chosen well. The 11x result also makes a strong case for treating kernels and serving paths as product surfaces, not invisible plumbing.

SirLiberte's profile picture
SirLiberte10 days ago

@vllm_project Not too long ago you were unbiased. I don't know why ?? Currently very slanted reporting in your post. Particulars with AMD. Every post on AMD is overhyped to the extreme. You've come from the top of the mountain. Now containually rolling down. Very sad to see.

tokenprincess's profile picture
tokenprincess11 days ago

@vllm_project 11× in 19 days on agentic/long-context is a software ceiling unlock, not a silicon step-function. need the baseline (vllm + batch/ctx), tokens/$ and W/token, and whether it generalizes past minimax m3 — otherwise it's one-workload tuning on mi355x

Robert Durant's profile picture
Robert Durant11 days ago

@vllm_project $AMD going to skyrocket Tomorrow as this huge. Called out $AMD and $INTC Thursday along with Memory chips . Your Welcome

ikeda hideki's profile picture
ikeda hideki10 days ago

@vllm_project Insane, this tech is unreal. The progress is too fast. So cool.

scifirefly's profile picture
scifirefly10 days ago

@vllm_project Finally ! They could have done that years ago !

Balance's profile picture
Balance11 days ago

@vllm_project You know it just smells like curry

EDDY VU's profile picture
EDDY VU11 days ago

@vllm_project Huge jump for long context agent runs. Curious how much of that 11x came from prefill speedups versus decode throughput.

AI Quanting's profile picture
AI Quanting11 days ago

@vllm_project the 11x is on one attention op for one model. thats a missing kernel getting written rather than the whole stack getting faster, so the next architecture with a different attention shape starts the clock again. real gain, just doesnt carry forward

erik@try.works's profile picture

@vllm_project what does this Indian have to do with ROCm?

Steven Cheng's profile picture
Steven Cheng10 days ago

@vllm_project 11x is impressive, but I wonder how much of that was low-hanging fruit in the attention kernel. Sustained gains usually require deeper architectural changes, not just software tweaks. Curious to see long-term benchmarks.

Related Videos

Some time ago, I had the idea to port NVIDIA Physical AI stack to AMD. The motivation was to improve hardware diversity and enable world models and VLAs to run beyond a single ecosystem. We started with NVIDIA Cosmos Predict 2.5-2B. Porting wasn’t trivial: these models are deeply optimized for NVIDIA’s stack. We used this as an opportunity to apply our ROCm kernels. The results were surprising: Both encode and diffusion run faster on AMD Instinct MI300X vs. NVIDIA H200 (FA3) and we still saw significant headroom for further optimization. Quality is unchanged across modalities (validated with WorldJen) To be clear, this is no luck. We have deep experience with diffusion models and AMD GPUs. But this just gives us a good opportunity to get closer to a true hardware-to-hardware comparison, as we work with less software abstractions than usual. Just to give an example, on AMD, memory instructions are async with a hardware queue of ordered pending instructions, enabling concurrent load/store with compute without warp specialization. Bottom line: there are real architectural advantages on AMD, if you take the time to work with the hardware. Note, we did tradeoff ~20% higher memory usage, That being said, AMD has more to give to begin with :) in the coming weeks: AMD versions of Cosmos Transfer and GR00T, an even faster version of Cosmos Predict, and open-sourcing an attention kernel faster than AITER v3 (which is closed-source for some reason? cc: Anush Elangovan )

Omer Shlomovits

36,648 views • 5 months ago

📢 PERORMANCE V4 IS LIVE We've spent over 10 years at the Top of Performance Improvement companies, earning our place as the world’s #1 E-Sports PC Optimization Specialists. From elite players to top-tier orgs and hardware giants, our mission has always been clear: unlock every ounce of power your PC holds. Today, that mission reaches everyone. Whether you're a competitive gamer or managing high-level operations, tuned performance and low system latency matters. That’s why we’re proud to unveil Performance V4: a completely free utility app crafted and designed by my team and I as the first glimpse into increasing PC Performance for entirely free. Performance V4 is the beginning stages of the upcoming Paragon Tweak Utility (PTU): a revolutionary full-suite optimization platform, soon available through our website, and eventually to the Epic Games Store and Microsoft Store. Our current business strategy has two massive scale issues— the human resources required to optimize each customer's PC, and time to execution with appointment setting & correspondence. We believe that the next step is to create software that replaces that work, and in turn makes PC optimizations more accessible and more common for all, which is why we are launching on Believe. With more access to expendable cash, we can create our vision faster. The Performance V4 is LIVE , alongside an exclusive first look at PTU later next week. If you want to believe in something, believe in us 🫡

Paragon│Boost Gaming PC Performance

55,054 views • 1 year ago