Loading video...

Video Failed to Load

Go Home

ERRC: Entropy-Reinvested Residual Correction for Tensor-Parallel LLM Communication PyTorch Foundation Ambassador Abdulsalam Bande will present a poster on ERRC (Entropy-Reinvested Residual Correction for Tensor-Parallel LLM Communication) at PyTorch Conference North America 2026. ERRC focuses on optimizing GPU data exchange during LLM inference. By compressing inter-GPU communication and using the...

16,886 views • 5 days ago •via X (Twitter)

6 Comments

Chandrakumar R Pillai's profile picture
Chandrakumar R Pillai5 days ago

The key comparison will be end-to-end throughput and quality across tensor-parallel sizes and interconnects. Reporting communication time, tail latency, and quality against an uncompressed baseline would show where bandwidth reinvestment pays off in practice.

♡ sophie ✧'s profile picture
♡ sophie ✧4 days ago

whats the actual overhead vs just doing all-reduce every k steps

@chia_050807's profile picture
@chia_0508074 days ago

Using the freed bandwidth to offset quantization error is a clever trade. Curious how much of the gain holds up at larger tensor parallel degrees.

JasonC's profile picture
JasonC4 days ago

ERRC 把 tensor-parallel 通信里的量化误差压下去——跨节点带宽大概省多少、精度掉几个点,有没有公开过跟 baseline AllReduce 的对比表?

Sophia p.'s profile picture
Sophia p.4 days ago

Calling it "entropy-reinvested" is bold when the residual correction is just trading compute for bandwidth. What's the measured overlap gain versus a plain compressed residual baseline?

Sofie b,'s profile picture
Sofie b,4 days ago

Love this take

Related Videos

I had to test it myself to believe this unreal inference speed. 3,000 tokens/s for 1 user on standard datacenter GPUs. They leveraged a hidden efficiency gap in how GPUs generate tokens. Kog just achieved 3,000 tokens/s on 8× AMD MI300X GPUs and 2,100 on 8× NVIDIA H200 (FP16, no speculative decoding). Their tech preview is on a 2B model, and they show how their techniques will scale to large frontier MoE models at similar speeds. That's a huge number because normal low-batch GPU decoding for 2B to 8B models is usually closer to 100 to 300 tokens/s per request, so Kog is claiming something like a 10X to 30X jump in the speed one user actually feels. Their trick: they are getting the speed by treating LLM decoding as a memory streaming problem, not mainly a math problem. For 1 user at batch size 1, the GPU is not doing big, efficient matrix-matrix work like in training or large-batch serving; it is repeatedly pulling the model’s active weights from high-bandwidth memory for each new token, so speed depends on how smoothly those weights keep flowing. Normal inference stacks keep breaking that flow. They run many separate GPU programs for different parts of the model, move intermediate results through memory, wait at synchronization points, talk back to the CPU for scheduling or sampling, and then repeat this token after token. Kog’s answer is to co-design 3 things that are usually tuned separately: the runtime, the low-level GPU code, and the model architecture. The biggest engineering move is the monokernel, where the whole decode pass runs as 1 persistent GPU-resident program, including sampling, so the system does not keep stopping for kernel launches, CPU scheduling, and intermediate memory round trips. They also rebuilt synchronization, because their own measurements say grid sync was eating around 35% of token-generation time; instead of making every compute unit wait at a broad barrier, each unit waits only for the exact data it needs. On AMD MI300X, they also map memory access around the chiplet layout, because memory latency changes depending on which die makes the request. Then their Laneformer model uses Delayed Tensor Parallelism, which lets cross-GPU communication happen in the background instead of blocking every layer.

Rohan Paul

13,282 views • 4 months ago