正在加载视频...

视频加载失败

Introducing SubQ - a major breakthrough in LLM intelligence. It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA), And the first frontier model with a 12 million token context window which is: - 52x faster than FlashAttention at 1MM tokens - Less than 5% the...

13,483,053 次观看 • 4 个月前 •via X (Twitter)

70 条评论

Alexander Whedon 的头像
Alexander Whedon4 个月前

SubQ is available for early access today, alongside our coding agent, SubQ Code Get access today ↓

Alexander Whedon 的头像
Alexander Whedon4 个月前

We were a little slow on this, but we just got a technical blog post up with more details. Please take a look! We have a model card coming next week, and we are happy to take requests for any specific details there. I am happy to answer any questions here!

Martin Shkreli 的头像
Martin Shkreli4 个月前

congrats!! why only 150tok/sec if 52x faster, etc.?

Alexander Whedon 的头像
Alexander Whedon4 个月前

52x is for prefill speed. Decoding speedups are coming too! This is the first research announcement, with many to follow.

Chris B. Ward (e/bored) 🇬🇧🇮🇪🤐 的头像
Chris B. Ward (e/bored) 🇬🇧🇮🇪🤐2 个月前

@MartinShkreli absolute grift, didn't deliver

Alexander Whedon 的头像
Alexander Whedon2 个月前

@MartinShkreli Model card shared!

Vincent 的头像
Vincent4 个月前

Any papers? Seems too good to be true

Alexander Whedon 的头像
Alexander Whedon4 个月前

The model card is coming next week! We are releasing a technical blog post with more details later today.

Vincent 的头像
Vincent4 个月前

Looking forward to reading it

shirish 的头像
shirish4 个月前

i still can't believe this.. 12M context window with 98% accuracy 💀 @alex_whedon @subquadratic

Alexander Whedon 的头像
Alexander Whedon4 个月前

@subquadratic And at a fraction of the cost...

• nanou • 的头像
• nanou •4 个月前

If this works as described, it basically changes how we build with LLMs. A lot of current pipelines exist just to work around context limits.

Alexander Whedon 的头像
Alexander Whedon4 个月前

And now that constraint goes away, what people choose to build is going to be very different.

Linus ✦ Ekenstam 的头像
Linus ✦ Ekenstam4 个月前

I remember when Gemini first made the claim of 1M in lab situations. We've come quite a bit from there. I'm extremely intrigued to see the real-world implications of this in bigger and more complex AI powered software. There is the zero shot environment, but there is also the qualitative repetitive work that is prone to drift and degradation over time. I really hope you will share case-studies, since I think this will be one of the most impressive ways to highlight the capabilities. Amazing work Alex and team. Kudos.

Alexander Whedon 的头像
Alexander Whedon4 个月前

This is the framing we'd love to test against. If you've got a workload that historically drifts, would genuinely want you to throw it at SubQ and see what happens.

LADI ⟡⭐ 的头像
LADI ⟡⭐4 个月前

If SubQ only attends to a small subset of tokens, how does it guarantee recall of critical long range dependencies in worst case inputs (e.g adversarial prompts or tasks where the important tokens aren’t locally obvious)

Alexander Whedon 的头像
Alexander Whedon4 个月前

We dynamically select the token relationships that matter and train it against retrieval problems with distractor docs, etc.

LADI ⟡⭐ 的头像
LADI ⟡⭐4 个月前

What happens when the selector misses the one token that matters? Any fallback

LADI ⟡⭐ 的头像
LADI ⟡⭐4 个月前

I agree, it’s about frequency, not perfection.

Veselin Stoyanov 的头像
Veselin Stoyanov4 个月前

A few weeks ago I was in a discussion under an X post about LLMs, their adoption and cost. Most people argued it doesn't scale and brought up the usual economic concerns. My take was that it's only a matter of time before the next breakthrough. It's always been this way with tech. Congrats on this one 👏

Alexander Whedon 的头像
Alexander Whedon4 个月前

This is only the first breakthrough we have announced! More coming.

Rohan Paul 的头像
Rohan Paul4 个月前

What 🤯 So Opus 4.7 is at $15 per million tokens and you guyes are giving it for under $1.50.

Cynical Optimist 的头像
Cynical Optimist4 个月前

I hope you consider releasing some variants as open weights. This would change the game for private on-premise work.

Alexander Whedon 的头像
Alexander Whedon4 个月前

This first model won't be open-source, but we do want to contribute to the open-source community!

Comrade Banana (ML/UT) 的头像
Comrade Banana (ML/UT)4 个月前

@ChemPhysMajor Atleast sharing how you did it?

Cynical Optimist 的头像
Cynical Optimist4 个月前

@CommiePat1776 @alex_whedon A white paper would go a long way.

anpaure 的头像
anpaure4 个月前

close enough, welcome back brampton

Carl Vellotti 🥞 的头像
Carl Vellotti 🥞4 个月前

The transformer was the first workable answer to long context. Everyone scaled it so hard nobody wanted to admit it was a local maximum. You finally found a way out.

Alexander Whedon 的头像
Alexander Whedon4 个月前

Thanks @carlvellotti. Try it out at the team will appreciate your insights.

Poonam Soni 的头像
Poonam Soni4 个月前

5% of Anthropic’s price at 98% accuracy at 12M tokens. The math on every agent product being built right now just changed completely.

Alexander Whedon 的头像
Alexander Whedon4 个月前

The unit economics of agentic products were quietly upside down. Well, not anymore.

Zephyr 的头像
Zephyr4 个月前

spent 2 years engineering around AI that couldn't read long docs. that limit just got removed.

Alexander Whedon 的头像
Alexander Whedon4 个月前

Throw your hardest doc workload at it, and let me know how it performs. Check out

AshutoshShrivastava 的头像
AshutoshShrivastava4 个月前

Long-context at under 10% of Anthropic's price damn…..

Alexander Whedon 的头像
Alexander Whedon4 个月前

I mean, someone had to do it

Muratcan Koylan 的头像
Muratcan Koylan4 个月前

Release the API or it never happened

Allie K. Miller 的头像
Allie K. Miller4 个月前

As someone who has worked with you for years, I am excited to see you finally releasing this. ✨ may your waitlist be buzzing and your tokens be cheap ✨

Alexander Whedon 的头像
Alexander Whedon4 个月前

Thank you!

Danny Limanseta 的头像
Danny Limanseta4 个月前

Is this real? Wow. Is there any evals where we see it with other forontier models?

Rezoan Ferdose 的头像
Rezoan Ferdose4 个月前

Congrats to the @subquadratic team! It’s awesome to see someone finally breaking out of the standard transformer box

🅿️ 的头像
🅿️4 个月前

Being a SaaS investor in 2026 sounds like Dante’s Inferno

vas 的头像
vas4 个月前

Wtf

Vincent Koc 的头像
Vincent Koc4 个月前

I love not wasting tokens, superb

Alexander Whedon 的头像
Alexander Whedon4 个月前

Much of standard attention's compute is models talking to themselves about words that don't matter.

kanav 的头像
kanav4 个月前

- Less than 5% the cost of Opus TAKE MY MONEY

Morgan 的头像
Morgan4 个月前

Super interesting Alexander, and 12m token context window, holy moly 🤯

Alexander Whedon 的头像
Alexander Whedon4 个月前

52x faster than FlashAttention, <5% of Opus pricing

Morgan 的头像
Morgan4 个月前

Totally insane.

Burny - Effective Curiosity 的头像
Burny - Effective Curiosity4 个月前

"It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA)" This doesn't seem right? That already exists?

Leon Lin 的头像
Leon Lin4 个月前

this is literally epic wtf. ik another lab would make it

Ray Fernando 的头像
Ray Fernando4 个月前

Congrats on the announcement and I'm looking forward to giving this a try.

vxnuaj 的头像
vxnuaj4 个月前

I’m sorry, wtf. Paper?

Alexander Whedon 的头像
Alexander Whedon4 个月前

Next week!

vxnuaj 的头像
vxnuaj3 个月前

paper?

Alexander Whedon 的头像
Alexander Whedon3 个月前

This week!

Mahy 的头像
Mahy3 个月前

@vxnuaj This paper is never gonna get released lol.

vxnuaj 的头像
vxnuaj3 个月前

@alex_whedon patience young padawan.

vxnuaj 的头像
vxnuaj3 个月前

@alex_whedon learn to have patience, you must

Dan McAteer 的头像
Dan McAteer4 个月前

ummm...this is a "BIG F*CKING DEAL"?

Alexander Whedon 的头像
Alexander Whedon4 个月前

and it's available for early access at

Dan McAteer 的头像
Dan McAteer4 个月前

just submitted a request! would love to try it out and write a post about it.

Subah Wadhwani 的头像
Subah Wadhwani4 个月前

Was a pleasure collaborating w you guys over the last few weeks!

Devansh Tripathi 的头像
Devansh Tripathi4 个月前

congratulations it'll be interesting to see how it manages to avoid the quality cliff on a long context window

Alexander Whedon 的头像
Alexander Whedon4 个月前

@NotTheCh05en1 Thanks Devansh, would love for you to try it out.

Paul Couvert 的头像
Paul Couvert4 个月前

Congrats on the launch. Can't believe we have a model with a context window this large with this accuracy!

Alexander Whedon 的头像
Alexander Whedon4 个月前

Thanks @itsPaulAi! Would love for you try out the product & share any feedback

Csaba Kissi 的头像
Csaba Kissi4 个月前

Finally, an LLM thats fast, powerful, and cheap at the same time. Opus is extremely expensive.

Alexander Whedon 的头像
Alexander Whedon4 个月前

The tradeoff people learned to accept was "pick two of fast/cheap/smart." We changed that.

Nikhil N 的头像
Nikhil N4 个月前

If someone said that enterprises are burning cash with AI subscriptions, that ends today. Great work @alex_whedon

Alexander Whedon 的头像
Alexander Whedon4 个月前

No subquadratic tax!

相关视频

Announcing How Transformer LLMs Work, created with Jay Alammar and Maarten Grootendorst, co-authors of the beautifully illustrated book, “Hands-On Large Language Models.” This course offers a deep dive into the inner workings of the transformer architecture that powers large language models (LLMs). The transformer architecture revolutionized generative AI; in fact, the "GPT" in ChatGPT stands for "Generative Pre-Trained Transformer." Originally introduced in the Google Brain team's groundbreaking 2017 paper "Attention Is All You Need," by Vaswani and others, transformers were a highly scalable model for machine translation tasks. Variants of this architecture now power today’s LLMs such as those from OpenAI, Google, Meta, Cohere, Anthropic and DeepSeek. In this course, you’ll learn in detail how LLMs process text. You'll also work through code examples that illustrate that transformer's individual components. In details, you’ll learn: - How the representation of language has evolved, from Bag-of-Words to Word2Vec embeddings to the transformer architecture that captures a word's meanings taking into account the context of other words in the input. - How inputs are broken down into tokens before they are sent to the language model. - The details of a transformer's main stages: Tokenization and embedding, the stack of transformer blocks, and the language model head. - The inner workings of the transformer block, including attention, which calculates relevance scores, and the feedforward layer, which incorporates stored information learned in training. - How cached calculations make transformers faster. - Some of the most recent ideas in the latest models such as Mixture-of-Experts (MoE) which uses multiple sub-models and a router on each layer to improve the quality of LLMs. By the end of this course, you’ll have a deep understanding of how LLMs actually process text and be able to read through papers describing the latest models and understand the details. Gaining this intuition will improve your approach to building LLM applications. Please sign up here:

Andrew Ng

259,920 次观看 • 1 年前

A tricky LLM interview question: You're serving a reasoning model on vLLM, and it keeps running out of GPU memory on long traces. So you add KV cache compression and evict 90% of the cached tokens. VRAM usage stays as is and GPU still runs out of memory. Why? (answer below) Evicting 90% of the KV cache can free almost none of the memory it was using. This sounds counterintuitive, but it follows directly from how production servers store the cache today. The KV cache grows with every token a model generates. Each token appends its key and value vectors across every layer, and nothing is freed while generation continues. This is the dominant memory cost for reasoning models. If a 32K-token CoT caches ~32K tokens of KV vectors, a Qwen3-32B with 4-bit weights will run out-of-memory around 24K tokens on a 24GB GPU. One obvious solution is to keep the important tokens and drop the rest, since attention is sparse enough to allow it. But this does not solve the memory problem yet. The reason is paged attention, which is the memory manager behind vLLM and most production servers. Under the hood, it splits GPU memory into fixed physical blocks, each one holds the KV for about 16 tokens. This block returns to the allocator only when every slot inside it is empty. Since the eviction logic selects tokens by importance, and such tokens are scattered across blocks... ...so despite eviction, almost every block is left with at least some survivor tokens. For instance, if the logic evicts 14k of 16k tokens across 1,000 blocks, most likely every block will still have a token. This means the allocator frees almost nothing. Placing the new tokens into those freed slots is not ideal because it breaks the cache's layout. Say token 16,001 arrives, and it's placed in the slot the 40th token used to hold. The cache now reads position 38, then 16,001, then 41, so the cache is no longer in token order. Attention can still compute the right answer from that, but only if every slot now carries a separate note recording which position it actually holds. This introduces another bookkeeping cost that an in-order layout inherently avoids. So the cache is logically 90% smaller and still physically the same size. Many compression results miss this because they measure on pre-allocated contiguous tensors rather than a paged server. There's another problem. Eviction methods pick which tokens to keep by looking at the attention scores themselves (as expected). But fast attention kernels used in production, like FlashAttention, never save those scores. They compute attention in small pieces and throw the full score grid away as they go, which is also why they're fast. So the exact signal eviction methods need isn't available in memory. The workaround is to fall back to eager attention and build the full matrix, which gives up the speed FlashAttention was there to provide. NVIDIA published a method called TriAttention to solve both these problems. It never needs attention scores. Instead, it scores tokens from the geometry of the model's key and query vectors before RoPE is applied, where those vectors sit in stable clusters. For the memory problem, it runs a compaction pass every 128 decoded tokens. The surviving tokens slide forward to close the holes eviction creates, so whole blocks empty out and return to the allocator while the cache stays in token order. On long reasoning traces, the approach matches full-attention accuracy while decoding 2.5x faster and using 10.7x less KV memory. KV cache compression is a big infrastructure problem. The number that decides whether it works is the count of freed blocks, not the count of evicted tokens. You can find the NVIDIA write-up here: I wrote a first-principles breakdown of how the KV cache works. It walks through why the model stores keys and values at all, why the cache grows with every token, and a comparison of LLM generation speed with and without KV caching. Read it below.

Avi Chawla

271,839 次观看 • 2 个月前

Researchers found a way to make LLMs 8.5x faster! (without compromising accuracy) Speculative decoding is quite an effective way to address the single-token bottleneck in traditional LLM inference. A small "draft" model first generates the next several tokens, then the large model verifies all of them at once in a single forward pass. If a token at any position is wrong, you keep everything before it and restart from there. This never does worse than normal decoding. But current drafters in Speculative decoding still guess one token at a time. That makes the drafting step itself a bottleneck, capping real-world speedups at 2-3x. DFlash is a new technique that swaps the autoregressive drafter with a lightweight block diffusion model that guesses all tokens in one parallel shot. Drafting cost stays flat no matter how many tokens you speculate. On top of that, the drafter is conditioned on hidden features pulled from multiple layers of the target model and injected into every draft layer, so it makes significantly better guesses than a drafter working from scratch. In the side-by-side demo below, vanilla decoding runs at 48.5 tokens/sec. DFlash hits 415 tokens/sec on the same model, with zero quality loss. It's already integrated with vLLM, SGLang, and Transformers, with draft models on HuggingFace for several models like Qwen3, Qwen3.5, Llama 3.1, Kimi-K2.5, gpt-oss, and many more. I have shared the GitHub repo in the replies! KV caching is another must-know technique to boost LLM inference. I recently wrote an article about it. Read it below. 👉 Over to you: What use case are you working on that can benefit from this new technique?

Avi Chawla

157,390 次观看 • 4 个月前

New short course: Attention in Transformers: Concepts and Code in PyTorch. Last week we released a course on how LLM transformers work. This week, go deeper and learn about the technical ideas behind the attention mechanism, and see how to code it in PyTorch. This course is built with Joshua Starmer, Founder and CEO of StatQuest. The attention mechanism was a breakthrough that led to transformers, the architecture powering large language models like ChatGPT. Transformers, introduced in the 2017 paper: "Attention is All You Need" by Viswani and others, took off because of its highly scalable design. In this course, you’ll learn how the attention mechanism, a key element of transformer-based LLMs, works and implement it in PyTorch. You'll develop deep intuition about building reliable, functional, and scalable AI applications. What you will do: - Understand the evolution of the attention mechanism, a key breakthrough that led to transformers. - Learn the relationships between word embeddings, positional embeddings, and attention. - Learn about the Query, Key, and Value matrices, and how to produce and use them in attention. - Walk through the math required to calculate self-attention and masked self-attention to learn why and how they work. - Understand the difference between self-attention and masked self-attention and how one is used in the encoder to build context-aware embeddings and the other is used in the decoder for generative outputs. - Learn the details of the encoder-decoder architecture, cross-attention, and multi-head attention and how they are all incorporated into a transformer. - Use PyTorch to code a class that implements self-attention, masked self-attention, and multi-head attention. There're lots of exciting technical details in this course. Please sign up here:

Andrew Ng

132,400 次观看 • 1 年前

Jonathan Ross just revealed why AI companies aren’t growing faster. Not demand. Not competition. Physics. Ross: “The demand for compute is insatiable.” There isn’t enough compute in the world. Not a temporary shortage. A fundamental gap between what the market wants and what the infrastructure can deliver. Ross: “Right now, one of the biggest complaints of Anthropic is the rate limits. People can’t get enough tokens.” Rate limits aren’t product decisions. They’re rationing. Companies forced to regulate access because infrastructure cannot meet demand. Slower services. Token caps. The only things standing between these companies and a revenue surge they can’t access. Every token cap is a revenue cap. Every slowdown is a sale that didn’t happen. Ross: “If Anthropic was given twice the inference compute, within one month their revenue would almost double.” Read that again. Double the compute. Double the revenue. Within thirty days. That’s not a growth projection. That’s a measurement of how deep the backlog already is. The demand exists right now. It’s sitting in a queue. The only thing between these companies and that revenue is physical hardware they don’t have. This breaks every assumption about how tech companies scale. Usually you scale by finding customers. AI companies have infinite customers. They scale by finding hardware. The constraint isn’t market fit. It isn’t distribution. It isn’t competition. It’s processing power. This is why Jensen Huang is the most important person in the world right now. NVIDIA doesn’t just make chips. It makes the thing every government, every AI lab, and every company racing for this future needs more of and can’t get enough of. The compute bottleneck isn’t a tech industry problem. It’s a civilizational one. The winner of this era isn’t determined by who builds the smartest model. Every major lab has a frontier model. The winner is whoever secures the most compute fastest while everyone else rations what’s left. The race isn’t for intelligence. It’s for infrastructure. And right now there isn’t enough to go around.

Dustin

28,395 次观看 • 6 个月前