Загрузка видео...

Не удалось загрузить видео

На главную

I tested MTPLX v2 with QWEN 3.6 27B and compared it with oMLX without cache on M5 Max and DGX Spark on vllm using nvfp4 model version. More details in 🧵 I've reached 82.8 tps of max decoding speed! 🔥 Custom Metal Kernel design specifically for this model and...

15,885 просмотров • 2 месяцев назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Day 11/90 of Inference Engineering How does vLLM work and how is it used in production? Before we discuss how vLLM works internally, it helps to understand what vLLM is. At a high level, vLLM is an inference engine that is designed to serve LLMs to thousands of concurrent users efficiently while managing scarce compute and memory. The goal for vLLM is to maximize throughput and minimize latency; optimizing for the best inference economics and experience for end users. With every request from the end user, it eventually ends up in the engine core, gets scheduled alongside other requests from other concurrent users, executes on the GPU, and updates the KV cache with the new key and value vectors, and streams the tokens back to the user. The Scheduler decides what requests should execute next while continuously batching requests together to maximize GPU utilization. Continuous batching is an inference optimization that allows new requests to join a running batch as other requests finish generating tokens. This helps with keeping the GPU utilization high instead of letting it sit idle waiting for an entire batch to complete generating. After the scheduler dispatches the selected batch to the Model Executor, the Model Executor prepares the tensors and metadata required for inference, retrieves each request’s block table from KV Cache Manager, launches the optimized transformer forward pass on the GPU, computes the logits, updates the KV cache with the new key and value vectors, and finally returns the results for sampling and streaming. The KV Cache Manager uses the PagedAttention memory layout to allocate fixed-size cache blocks on demand and maintains a Free Block Queue on the CPU that tracks which blocks in the GPU’s Paged KV Cache are currently free. When a request needs additional KV cache space, the KV Cache manager takes a free block from the queue and assigns it to that request, thus avoiding an expensive search through GPU memory for available cache blocks. All of these components form the core of vLLM’s inference engine. The Scheduler determines what requests are executed, the Model Executor determines how those requests are executed, the KV Cache Manager determines where each request’s KV cache lives using the PagedAttention Memory Layout. This architecture enables vLLM to serve thousands of concurrent requests with high throughput, low latency, and efficient GPU memory utilization. Heres a little animation that visualizes everything! - I've also completed the forward pass for my mnist.c project. I had a nice chat with shrey birmiwal, such a knowledgeable guy. Excited to learn more about vLLM and implement a tiny-vLLM one day.

max fu

70,797 просмотров • 2 месяцев назад

Astra (GPT-6) is here!!! I've had early access and tested it like crazy with things like games, code, writing, browser control, presentations and general knowledge work. This is the best model I've ever used. Period. (Incredible demos below in this thread ⬇️) Here's my take on Astra: > It's insanely capable. This feels like a massive improvement, not just an incremental change. This is especially true with zero-shot prompts. > It's all about knowledge work. Slide creation, analysis, writing, and browser control. And oh my...it's so good at browser control. GPT-5.6 was already fantastic at doing things in the browser, Astra is another level and significantly faster. > We're closer than ever (arrived?) at prompt-to-playable game. And I don't just mean only playable, these are actually fun games. I bet if someone with a great eye for games used Astra, they could create a viral game within 1-2 weeks. > Astra is better at writing but not perfect. It removes much of the "AI Smell" we're all familiar with but some stink still survived. > It has a tendency to use the same design colors and look/feel as GPT-5.6 (forrest green anyone?) but it is more steerable in design than previous models. > It's highly steerable in general. A little nudge goes a long way. When I first started using Astra, almost every task I gave it would go for ~30 minutes. I wanted it to keep working. Adding more specifics to a prompt helped greatly with it's ability to work for a long time. > Astra's 3D understanding is unmatched. 3D asset creation was consistent and easy and its spacial awareness while building complex 3D worlds blew me away. I'm still getting familiar with Astra but this will now be my go-to model for any difficult work I have. Check out the demos below: 👇

Matthew Berman

1,907,752 просмотров • 1 месяц назад