Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing SubQ - a major breakthrough in LLM intelligence. It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA), And the first frontier model with a 12 million token context window which is: - 52x faster than FlashAttention at 1MM tokens - Less than 5% the...

13,483,177 Aufrufe • vor 4 Monaten •via X (Twitter)

70 Kommentare

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

SubQ is available for early access today, alongside our coding agent, SubQ Code Get access today ↓

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

We were a little slow on this, but we just got a technical blog post up with more details. Please take a look! We have a model card coming next week, and we are happy to take requests for any specific details there. I am happy to answer any questions here!

Profilbild von Martin Shkreli
Martin Shkrelivor 4 Monaten

congrats!! why only 150tok/sec if 52x faster, etc.?

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

52x is for prefill speed. Decoding speedups are coming too! This is the first research announcement, with many to follow.

Profilbild von Chris B. Ward (e/bored) 🇬🇧🇮🇪🤐
Chris B. Ward (e/bored) 🇬🇧🇮🇪🤐vor 2 Monaten

@MartinShkreli absolute grift, didn't deliver

Profilbild von Alexander Whedon
Alexander Whedonvor 2 Monaten

@MartinShkreli Model card shared!

Profilbild von Vincent
Vincentvor 4 Monaten

Any papers? Seems too good to be true

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

The model card is coming next week! We are releasing a technical blog post with more details later today.

Profilbild von Vincent
Vincentvor 4 Monaten

Looking forward to reading it

Profilbild von shirish
shirishvor 4 Monaten

i still can't believe this.. 12M context window with 98% accuracy 💀 @alex_whedon @subquadratic

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

@subquadratic And at a fraction of the cost...

Profilbild von • nanou •
• nanou •vor 4 Monaten

If this works as described, it basically changes how we build with LLMs. A lot of current pipelines exist just to work around context limits.

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

And now that constraint goes away, what people choose to build is going to be very different.

Profilbild von Linus ✦ Ekenstam
Linus ✦ Ekenstamvor 4 Monaten

I remember when Gemini first made the claim of 1M in lab situations. We've come quite a bit from there. I'm extremely intrigued to see the real-world implications of this in bigger and more complex AI powered software. There is the zero shot environment, but there is also the qualitative repetitive work that is prone to drift and degradation over time. I really hope you will share case-studies, since I think this will be one of the most impressive ways to highlight the capabilities. Amazing work Alex and team. Kudos.

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

This is the framing we'd love to test against. If you've got a workload that historically drifts, would genuinely want you to throw it at SubQ and see what happens.

Profilbild von LADI ⟡⭐
LADI ⟡⭐vor 4 Monaten

If SubQ only attends to a small subset of tokens, how does it guarantee recall of critical long range dependencies in worst case inputs (e.g adversarial prompts or tasks where the important tokens aren’t locally obvious)

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

We dynamically select the token relationships that matter and train it against retrieval problems with distractor docs, etc.

Profilbild von LADI ⟡⭐
LADI ⟡⭐vor 4 Monaten

What happens when the selector misses the one token that matters? Any fallback

Profilbild von LADI ⟡⭐
LADI ⟡⭐vor 4 Monaten

I agree, it’s about frequency, not perfection.

Profilbild von Veselin Stoyanov
Veselin Stoyanovvor 4 Monaten

A few weeks ago I was in a discussion under an X post about LLMs, their adoption and cost. Most people argued it doesn't scale and brought up the usual economic concerns. My take was that it's only a matter of time before the next breakthrough. It's always been this way with tech. Congrats on this one 👏

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

This is only the first breakthrough we have announced! More coming.

Profilbild von Rohan Paul
Rohan Paulvor 4 Monaten

What 🤯 So Opus 4.7 is at $15 per million tokens and you guyes are giving it for under $1.50.

Profilbild von Cynical Optimist
Cynical Optimistvor 4 Monaten

I hope you consider releasing some variants as open weights. This would change the game for private on-premise work.

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

This first model won't be open-source, but we do want to contribute to the open-source community!

Profilbild von Comrade Banana (ML/UT)
Comrade Banana (ML/UT)vor 4 Monaten

@ChemPhysMajor Atleast sharing how you did it?

Profilbild von Cynical Optimist
Cynical Optimistvor 4 Monaten

@CommiePat1776 @alex_whedon A white paper would go a long way.

Profilbild von anpaure
anpaurevor 4 Monaten

close enough, welcome back brampton

Profilbild von Carl Vellotti 🥞
Carl Vellotti 🥞vor 4 Monaten

The transformer was the first workable answer to long context. Everyone scaled it so hard nobody wanted to admit it was a local maximum. You finally found a way out.

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

Thanks @carlvellotti. Try it out at the team will appreciate your insights.

Profilbild von Poonam Soni
Poonam Sonivor 4 Monaten

5% of Anthropic’s price at 98% accuracy at 12M tokens. The math on every agent product being built right now just changed completely.

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

The unit economics of agentic products were quietly upside down. Well, not anymore.

Profilbild von Zephyr
Zephyrvor 4 Monaten

spent 2 years engineering around AI that couldn't read long docs. that limit just got removed.

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

Throw your hardest doc workload at it, and let me know how it performs. Check out

Profilbild von AshutoshShrivastava
AshutoshShrivastavavor 4 Monaten

Long-context at under 10% of Anthropic's price damn…..

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

I mean, someone had to do it

Profilbild von Muratcan Koylan
Muratcan Koylanvor 4 Monaten

Release the API or it never happened

Profilbild von Allie K. Miller
Allie K. Millervor 4 Monaten

As someone who has worked with you for years, I am excited to see you finally releasing this. ✨ may your waitlist be buzzing and your tokens be cheap ✨

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

Thank you!

Profilbild von Danny Limanseta
Danny Limansetavor 4 Monaten

Is this real? Wow. Is there any evals where we see it with other forontier models?

Profilbild von Rezoan Ferdose
Rezoan Ferdosevor 4 Monaten

Congrats to the @subquadratic team! It’s awesome to see someone finally breaking out of the standard transformer box

Profilbild von 🅿️
🅿️vor 4 Monaten

Being a SaaS investor in 2026 sounds like Dante’s Inferno

Profilbild von vas
vasvor 4 Monaten

Wtf

Profilbild von Vincent Koc
Vincent Kocvor 4 Monaten

I love not wasting tokens, superb

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

Much of standard attention's compute is models talking to themselves about words that don't matter.

Profilbild von kanav
kanavvor 4 Monaten

- Less than 5% the cost of Opus TAKE MY MONEY

Profilbild von Morgan
Morganvor 4 Monaten

Super interesting Alexander, and 12m token context window, holy moly 🤯

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

52x faster than FlashAttention, <5% of Opus pricing

Profilbild von Morgan
Morganvor 4 Monaten

Totally insane.

Profilbild von Burny - Effective Curiosity
Burny - Effective Curiosityvor 4 Monaten

"It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA)" This doesn't seem right? That already exists?

Profilbild von Leon Lin
Leon Linvor 4 Monaten

this is literally epic wtf. ik another lab would make it

Profilbild von Ray Fernando
Ray Fernandovor 4 Monaten

Congrats on the announcement and I'm looking forward to giving this a try.

Profilbild von vxnuaj
vxnuajvor 4 Monaten

I’m sorry, wtf. Paper?

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

Next week!

Profilbild von vxnuaj
vxnuajvor 3 Monaten

paper?

Profilbild von Alexander Whedon
Alexander Whedonvor 3 Monaten

This week!

Profilbild von Mahy
Mahyvor 3 Monaten

@vxnuaj This paper is never gonna get released lol.

Profilbild von vxnuaj
vxnuajvor 3 Monaten

@alex_whedon patience young padawan.

Profilbild von vxnuaj
vxnuajvor 3 Monaten

@alex_whedon learn to have patience, you must

Profilbild von Dan McAteer
Dan McAteervor 4 Monaten

ummm...this is a "BIG F*CKING DEAL"?

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

and it's available for early access at

Profilbild von Dan McAteer
Dan McAteervor 4 Monaten

just submitted a request! would love to try it out and write a post about it.

Profilbild von Subah Wadhwani
Subah Wadhwanivor 4 Monaten

Was a pleasure collaborating w you guys over the last few weeks!

Profilbild von Devansh Tripathi
Devansh Tripathivor 4 Monaten

congratulations it'll be interesting to see how it manages to avoid the quality cliff on a long context window

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

@NotTheCh05en1 Thanks Devansh, would love for you to try it out.

Profilbild von Paul Couvert
Paul Couvertvor 4 Monaten

Congrats on the launch. Can't believe we have a model with a context window this large with this accuracy!

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

Thanks @itsPaulAi! Would love for you try out the product & share any feedback

Profilbild von Csaba Kissi
Csaba Kissivor 4 Monaten

Finally, an LLM thats fast, powerful, and cheap at the same time. Opus is extremely expensive.

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

The tradeoff people learned to accept was "pick two of fast/cheap/smart." We changed that.

Profilbild von Nikhil N
Nikhil Nvor 4 Monaten

If someone said that enterprises are burning cash with AI subscriptions, that ends today. Great work @alex_whedon

Profilbild von Alexander Whedon
Alexander Whedonvor 4 Monaten

No subquadratic tax!

Ähnliche Videos

Announcing How Transformer LLMs Work, created with Jay Alammar and Maarten Grootendorst, co-authors of the beautifully illustrated book, “Hands-On Large Language Models.” This course offers a deep dive into the inner workings of the transformer architecture that powers large language models (LLMs). The transformer architecture revolutionized generative AI; in fact, the "GPT" in ChatGPT stands for "Generative Pre-Trained Transformer." Originally introduced in the Google Brain team's groundbreaking 2017 paper "Attention Is All You Need," by Vaswani and others, transformers were a highly scalable model for machine translation tasks. Variants of this architecture now power today’s LLMs such as those from OpenAI, Google, Meta, Cohere, Anthropic and DeepSeek. In this course, you’ll learn in detail how LLMs process text. You'll also work through code examples that illustrate that transformer's individual components. In details, you’ll learn: - How the representation of language has evolved, from Bag-of-Words to Word2Vec embeddings to the transformer architecture that captures a word's meanings taking into account the context of other words in the input. - How inputs are broken down into tokens before they are sent to the language model. - The details of a transformer's main stages: Tokenization and embedding, the stack of transformer blocks, and the language model head. - The inner workings of the transformer block, including attention, which calculates relevance scores, and the feedforward layer, which incorporates stored information learned in training. - How cached calculations make transformers faster. - Some of the most recent ideas in the latest models such as Mixture-of-Experts (MoE) which uses multiple sub-models and a router on each layer to improve the quality of LLMs. By the end of this course, you’ll have a deep understanding of how LLMs actually process text and be able to read through papers describing the latest models and understand the details. Gaining this intuition will improve your approach to building LLM applications. Please sign up here:

Andrew Ng

259,920 Aufrufe • vor 1 Jahr

A tricky LLM interview question: You're serving a reasoning model on vLLM, and it keeps running out of GPU memory on long traces. So you add KV cache compression and evict 90% of the cached tokens. VRAM usage stays as is and GPU still runs out of memory. Why? (answer below) Evicting 90% of the KV cache can free almost none of the memory it was using. This sounds counterintuitive, but it follows directly from how production servers store the cache today. The KV cache grows with every token a model generates. Each token appends its key and value vectors across every layer, and nothing is freed while generation continues. This is the dominant memory cost for reasoning models. If a 32K-token CoT caches ~32K tokens of KV vectors, a Qwen3-32B with 4-bit weights will run out-of-memory around 24K tokens on a 24GB GPU. One obvious solution is to keep the important tokens and drop the rest, since attention is sparse enough to allow it. But this does not solve the memory problem yet. The reason is paged attention, which is the memory manager behind vLLM and most production servers. Under the hood, it splits GPU memory into fixed physical blocks, each one holds the KV for about 16 tokens. This block returns to the allocator only when every slot inside it is empty. Since the eviction logic selects tokens by importance, and such tokens are scattered across blocks... ...so despite eviction, almost every block is left with at least some survivor tokens. For instance, if the logic evicts 14k of 16k tokens across 1,000 blocks, most likely every block will still have a token. This means the allocator frees almost nothing. Placing the new tokens into those freed slots is not ideal because it breaks the cache's layout. Say token 16,001 arrives, and it's placed in the slot the 40th token used to hold. The cache now reads position 38, then 16,001, then 41, so the cache is no longer in token order. Attention can still compute the right answer from that, but only if every slot now carries a separate note recording which position it actually holds. This introduces another bookkeeping cost that an in-order layout inherently avoids. So the cache is logically 90% smaller and still physically the same size. Many compression results miss this because they measure on pre-allocated contiguous tensors rather than a paged server. There's another problem. Eviction methods pick which tokens to keep by looking at the attention scores themselves (as expected). But fast attention kernels used in production, like FlashAttention, never save those scores. They compute attention in small pieces and throw the full score grid away as they go, which is also why they're fast. So the exact signal eviction methods need isn't available in memory. The workaround is to fall back to eager attention and build the full matrix, which gives up the speed FlashAttention was there to provide. NVIDIA published a method called TriAttention to solve both these problems. It never needs attention scores. Instead, it scores tokens from the geometry of the model's key and query vectors before RoPE is applied, where those vectors sit in stable clusters. For the memory problem, it runs a compaction pass every 128 decoded tokens. The surviving tokens slide forward to close the holes eviction creates, so whole blocks empty out and return to the allocator while the cache stays in token order. On long reasoning traces, the approach matches full-attention accuracy while decoding 2.5x faster and using 10.7x less KV memory. KV cache compression is a big infrastructure problem. The number that decides whether it works is the count of freed blocks, not the count of evicted tokens. You can find the NVIDIA write-up here: I wrote a first-principles breakdown of how the KV cache works. It walks through why the model stores keys and values at all, why the cache grows with every token, and a comparison of LLM generation speed with and without KV caching. Read it below.

Avi Chawla

271,839 Aufrufe • vor 2 Monaten

Researchers found a way to make LLMs 8.5x faster! (without compromising accuracy) Speculative decoding is quite an effective way to address the single-token bottleneck in traditional LLM inference. A small "draft" model first generates the next several tokens, then the large model verifies all of them at once in a single forward pass. If a token at any position is wrong, you keep everything before it and restart from there. This never does worse than normal decoding. But current drafters in Speculative decoding still guess one token at a time. That makes the drafting step itself a bottleneck, capping real-world speedups at 2-3x. DFlash is a new technique that swaps the autoregressive drafter with a lightweight block diffusion model that guesses all tokens in one parallel shot. Drafting cost stays flat no matter how many tokens you speculate. On top of that, the drafter is conditioned on hidden features pulled from multiple layers of the target model and injected into every draft layer, so it makes significantly better guesses than a drafter working from scratch. In the side-by-side demo below, vanilla decoding runs at 48.5 tokens/sec. DFlash hits 415 tokens/sec on the same model, with zero quality loss. It's already integrated with vLLM, SGLang, and Transformers, with draft models on HuggingFace for several models like Qwen3, Qwen3.5, Llama 3.1, Kimi-K2.5, gpt-oss, and many more. I have shared the GitHub repo in the replies! KV caching is another must-know technique to boost LLM inference. I recently wrote an article about it. Read it below. 👉 Over to you: What use case are you working on that can benefit from this new technique?

Avi Chawla

157,390 Aufrufe • vor 4 Monaten

New short course: Attention in Transformers: Concepts and Code in PyTorch. Last week we released a course on how LLM transformers work. This week, go deeper and learn about the technical ideas behind the attention mechanism, and see how to code it in PyTorch. This course is built with Joshua Starmer, Founder and CEO of StatQuest. The attention mechanism was a breakthrough that led to transformers, the architecture powering large language models like ChatGPT. Transformers, introduced in the 2017 paper: "Attention is All You Need" by Viswani and others, took off because of its highly scalable design. In this course, you’ll learn how the attention mechanism, a key element of transformer-based LLMs, works and implement it in PyTorch. You'll develop deep intuition about building reliable, functional, and scalable AI applications. What you will do: - Understand the evolution of the attention mechanism, a key breakthrough that led to transformers. - Learn the relationships between word embeddings, positional embeddings, and attention. - Learn about the Query, Key, and Value matrices, and how to produce and use them in attention. - Walk through the math required to calculate self-attention and masked self-attention to learn why and how they work. - Understand the difference between self-attention and masked self-attention and how one is used in the encoder to build context-aware embeddings and the other is used in the decoder for generative outputs. - Learn the details of the encoder-decoder architecture, cross-attention, and multi-head attention and how they are all incorporated into a transformer. - Use PyTorch to code a class that implements self-attention, masked self-attention, and multi-head attention. There're lots of exciting technical details in this course. Please sign up here:

Andrew Ng

132,400 Aufrufe • vor 1 Jahr

Jonathan Ross just revealed why AI companies aren’t growing faster. Not demand. Not competition. Physics. Ross: “The demand for compute is insatiable.” There isn’t enough compute in the world. Not a temporary shortage. A fundamental gap between what the market wants and what the infrastructure can deliver. Ross: “Right now, one of the biggest complaints of Anthropic is the rate limits. People can’t get enough tokens.” Rate limits aren’t product decisions. They’re rationing. Companies forced to regulate access because infrastructure cannot meet demand. Slower services. Token caps. The only things standing between these companies and a revenue surge they can’t access. Every token cap is a revenue cap. Every slowdown is a sale that didn’t happen. Ross: “If Anthropic was given twice the inference compute, within one month their revenue would almost double.” Read that again. Double the compute. Double the revenue. Within thirty days. That’s not a growth projection. That’s a measurement of how deep the backlog already is. The demand exists right now. It’s sitting in a queue. The only thing between these companies and that revenue is physical hardware they don’t have. This breaks every assumption about how tech companies scale. Usually you scale by finding customers. AI companies have infinite customers. They scale by finding hardware. The constraint isn’t market fit. It isn’t distribution. It isn’t competition. It’s processing power. This is why Jensen Huang is the most important person in the world right now. NVIDIA doesn’t just make chips. It makes the thing every government, every AI lab, and every company racing for this future needs more of and can’t get enough of. The compute bottleneck isn’t a tech industry problem. It’s a civilizational one. The winner of this era isn’t determined by who builds the smartest model. Every major lab has a frontier model. The winner is whoever secures the most compute fastest while everyone else rations what’s left. The race isn’t for intelligence. It’s for infrastructure. And right now there isn’t enough to go around.

Dustin

28,395 Aufrufe • vor 6 Monaten