Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing SubQ - a major breakthrough in LLM intelligence. It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA), And the first frontier model with a 12 million token context window which is: - 52x faster than FlashAttention at 1MM tokens - Less than 5% the...

13,483,177 görüntüleme • 4 ay önce •via X (Twitter)

70 Yorum

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

SubQ is available for early access today, alongside our coding agent, SubQ Code Get access today ↓

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

We were a little slow on this, but we just got a technical blog post up with more details. Please take a look! We have a model card coming next week, and we are happy to take requests for any specific details there. I am happy to answer any questions here!

Martin Shkreli profil fotoğrafı
Martin Shkreli4 ay önce

congrats!! why only 150tok/sec if 52x faster, etc.?

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

52x is for prefill speed. Decoding speedups are coming too! This is the first research announcement, with many to follow.

Chris B. Ward (e/bored) 🇬🇧🇮🇪🤐 profil fotoğrafı
Chris B. Ward (e/bored) 🇬🇧🇮🇪🤐2 ay önce

@MartinShkreli absolute grift, didn't deliver

Alexander Whedon profil fotoğrafı
Alexander Whedon2 ay önce

@MartinShkreli Model card shared!

Vincent profil fotoğrafı
Vincent4 ay önce

Any papers? Seems too good to be true

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

The model card is coming next week! We are releasing a technical blog post with more details later today.

Vincent profil fotoğrafı
Vincent4 ay önce

Looking forward to reading it

shirish profil fotoğrafı
shirish4 ay önce

i still can't believe this.. 12M context window with 98% accuracy 💀 @alex_whedon @subquadratic

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

@subquadratic And at a fraction of the cost...

• nanou • profil fotoğrafı
• nanou •4 ay önce

If this works as described, it basically changes how we build with LLMs. A lot of current pipelines exist just to work around context limits.

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

And now that constraint goes away, what people choose to build is going to be very different.

Linus ✦ Ekenstam profil fotoğrafı
Linus ✦ Ekenstam4 ay önce

I remember when Gemini first made the claim of 1M in lab situations. We've come quite a bit from there. I'm extremely intrigued to see the real-world implications of this in bigger and more complex AI powered software. There is the zero shot environment, but there is also the qualitative repetitive work that is prone to drift and degradation over time. I really hope you will share case-studies, since I think this will be one of the most impressive ways to highlight the capabilities. Amazing work Alex and team. Kudos.

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

This is the framing we'd love to test against. If you've got a workload that historically drifts, would genuinely want you to throw it at SubQ and see what happens.

LADI ⟡⭐ profil fotoğrafı
LADI ⟡⭐4 ay önce

If SubQ only attends to a small subset of tokens, how does it guarantee recall of critical long range dependencies in worst case inputs (e.g adversarial prompts or tasks where the important tokens aren’t locally obvious)

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

We dynamically select the token relationships that matter and train it against retrieval problems with distractor docs, etc.

LADI ⟡⭐ profil fotoğrafı
LADI ⟡⭐4 ay önce

What happens when the selector misses the one token that matters? Any fallback

LADI ⟡⭐ profil fotoğrafı
LADI ⟡⭐4 ay önce

I agree, it’s about frequency, not perfection.

Veselin Stoyanov profil fotoğrafı
Veselin Stoyanov4 ay önce

A few weeks ago I was in a discussion under an X post about LLMs, their adoption and cost. Most people argued it doesn't scale and brought up the usual economic concerns. My take was that it's only a matter of time before the next breakthrough. It's always been this way with tech. Congrats on this one 👏

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

This is only the first breakthrough we have announced! More coming.

Rohan Paul profil fotoğrafı
Rohan Paul4 ay önce

What 🤯 So Opus 4.7 is at $15 per million tokens and you guyes are giving it for under $1.50.

Cynical Optimist profil fotoğrafı
Cynical Optimist4 ay önce

I hope you consider releasing some variants as open weights. This would change the game for private on-premise work.

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

This first model won't be open-source, but we do want to contribute to the open-source community!

Comrade Banana (ML/UT) profil fotoğrafı
Comrade Banana (ML/UT)4 ay önce

@ChemPhysMajor Atleast sharing how you did it?

Cynical Optimist profil fotoğrafı
Cynical Optimist4 ay önce

@CommiePat1776 @alex_whedon A white paper would go a long way.

anpaure profil fotoğrafı
anpaure4 ay önce

close enough, welcome back brampton

Carl Vellotti 🥞 profil fotoğrafı
Carl Vellotti 🥞4 ay önce

The transformer was the first workable answer to long context. Everyone scaled it so hard nobody wanted to admit it was a local maximum. You finally found a way out.

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

Thanks @carlvellotti. Try it out at the team will appreciate your insights.

Poonam Soni profil fotoğrafı
Poonam Soni4 ay önce

5% of Anthropic’s price at 98% accuracy at 12M tokens. The math on every agent product being built right now just changed completely.

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

The unit economics of agentic products were quietly upside down. Well, not anymore.

Zephyr profil fotoğrafı
Zephyr4 ay önce

spent 2 years engineering around AI that couldn't read long docs. that limit just got removed.

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

Throw your hardest doc workload at it, and let me know how it performs. Check out

AshutoshShrivastava profil fotoğrafı
AshutoshShrivastava4 ay önce

Long-context at under 10% of Anthropic's price damn…..

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

I mean, someone had to do it

Muratcan Koylan profil fotoğrafı
Muratcan Koylan4 ay önce

Release the API or it never happened

Allie K. Miller profil fotoğrafı
Allie K. Miller4 ay önce

As someone who has worked with you for years, I am excited to see you finally releasing this. ✨ may your waitlist be buzzing and your tokens be cheap ✨

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

Thank you!

Danny Limanseta profil fotoğrafı
Danny Limanseta4 ay önce

Is this real? Wow. Is there any evals where we see it with other forontier models?

Rezoan Ferdose profil fotoğrafı
Rezoan Ferdose4 ay önce

Congrats to the @subquadratic team! It’s awesome to see someone finally breaking out of the standard transformer box

🅿️ profil fotoğrafı
🅿️4 ay önce

Being a SaaS investor in 2026 sounds like Dante’s Inferno

vas profil fotoğrafı
vas4 ay önce

Wtf

Vincent Koc profil fotoğrafı
Vincent Koc4 ay önce

I love not wasting tokens, superb

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

Much of standard attention's compute is models talking to themselves about words that don't matter.

kanav profil fotoğrafı
kanav4 ay önce

- Less than 5% the cost of Opus TAKE MY MONEY

Morgan profil fotoğrafı
Morgan4 ay önce

Super interesting Alexander, and 12m token context window, holy moly 🤯

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

52x faster than FlashAttention, <5% of Opus pricing

Morgan profil fotoğrafı
Morgan4 ay önce

Totally insane.

Burny - Effective Curiosity profil fotoğrafı
Burny - Effective Curiosity4 ay önce

"It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA)" This doesn't seem right? That already exists?

Leon Lin profil fotoğrafı
Leon Lin4 ay önce

this is literally epic wtf. ik another lab would make it

Ray Fernando profil fotoğrafı
Ray Fernando4 ay önce

Congrats on the announcement and I'm looking forward to giving this a try.

vxnuaj profil fotoğrafı
vxnuaj4 ay önce

I’m sorry, wtf. Paper?

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

Next week!

vxnuaj profil fotoğrafı
vxnuaj3 ay önce

paper?

Alexander Whedon profil fotoğrafı
Alexander Whedon3 ay önce

This week!

Mahy profil fotoğrafı
Mahy3 ay önce

@vxnuaj This paper is never gonna get released lol.

vxnuaj profil fotoğrafı
vxnuaj3 ay önce

@alex_whedon patience young padawan.

vxnuaj profil fotoğrafı
vxnuaj3 ay önce

@alex_whedon learn to have patience, you must

Dan McAteer profil fotoğrafı
Dan McAteer4 ay önce

ummm...this is a "BIG F*CKING DEAL"?

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

and it's available for early access at

Dan McAteer profil fotoğrafı
Dan McAteer4 ay önce

just submitted a request! would love to try it out and write a post about it.

Subah Wadhwani profil fotoğrafı
Subah Wadhwani4 ay önce

Was a pleasure collaborating w you guys over the last few weeks!

Devansh Tripathi profil fotoğrafı
Devansh Tripathi4 ay önce

congratulations it'll be interesting to see how it manages to avoid the quality cliff on a long context window

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

@NotTheCh05en1 Thanks Devansh, would love for you to try it out.

Paul Couvert profil fotoğrafı
Paul Couvert4 ay önce

Congrats on the launch. Can't believe we have a model with a context window this large with this accuracy!

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

Thanks @itsPaulAi! Would love for you try out the product & share any feedback

Csaba Kissi profil fotoğrafı
Csaba Kissi4 ay önce

Finally, an LLM thats fast, powerful, and cheap at the same time. Opus is extremely expensive.

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

The tradeoff people learned to accept was "pick two of fast/cheap/smart." We changed that.

Nikhil N profil fotoğrafı
Nikhil N4 ay önce

If someone said that enterprises are burning cash with AI subscriptions, that ends today. Great work @alex_whedon

Alexander Whedon profil fotoğrafı
Alexander Whedon4 ay önce

No subquadratic tax!

Benzer Videolar

Announcing How Transformer LLMs Work, created with Jay Alammar and Maarten Grootendorst, co-authors of the beautifully illustrated book, “Hands-On Large Language Models.” This course offers a deep dive into the inner workings of the transformer architecture that powers large language models (LLMs). The transformer architecture revolutionized generative AI; in fact, the "GPT" in ChatGPT stands for "Generative Pre-Trained Transformer." Originally introduced in the Google Brain team's groundbreaking 2017 paper "Attention Is All You Need," by Vaswani and others, transformers were a highly scalable model for machine translation tasks. Variants of this architecture now power today’s LLMs such as those from OpenAI, Google, Meta, Cohere, Anthropic and DeepSeek. In this course, you’ll learn in detail how LLMs process text. You'll also work through code examples that illustrate that transformer's individual components. In details, you’ll learn: - How the representation of language has evolved, from Bag-of-Words to Word2Vec embeddings to the transformer architecture that captures a word's meanings taking into account the context of other words in the input. - How inputs are broken down into tokens before they are sent to the language model. - The details of a transformer's main stages: Tokenization and embedding, the stack of transformer blocks, and the language model head. - The inner workings of the transformer block, including attention, which calculates relevance scores, and the feedforward layer, which incorporates stored information learned in training. - How cached calculations make transformers faster. - Some of the most recent ideas in the latest models such as Mixture-of-Experts (MoE) which uses multiple sub-models and a router on each layer to improve the quality of LLMs. By the end of this course, you’ll have a deep understanding of how LLMs actually process text and be able to read through papers describing the latest models and understand the details. Gaining this intuition will improve your approach to building LLM applications. Please sign up here:

Andrew Ng

259,920 görüntüleme • 1 yıl önce

A tricky LLM interview question: You're serving a reasoning model on vLLM, and it keeps running out of GPU memory on long traces. So you add KV cache compression and evict 90% of the cached tokens. VRAM usage stays as is and GPU still runs out of memory. Why? (answer below) Evicting 90% of the KV cache can free almost none of the memory it was using. This sounds counterintuitive, but it follows directly from how production servers store the cache today. The KV cache grows with every token a model generates. Each token appends its key and value vectors across every layer, and nothing is freed while generation continues. This is the dominant memory cost for reasoning models. If a 32K-token CoT caches ~32K tokens of KV vectors, a Qwen3-32B with 4-bit weights will run out-of-memory around 24K tokens on a 24GB GPU. One obvious solution is to keep the important tokens and drop the rest, since attention is sparse enough to allow it. But this does not solve the memory problem yet. The reason is paged attention, which is the memory manager behind vLLM and most production servers. Under the hood, it splits GPU memory into fixed physical blocks, each one holds the KV for about 16 tokens. This block returns to the allocator only when every slot inside it is empty. Since the eviction logic selects tokens by importance, and such tokens are scattered across blocks... ...so despite eviction, almost every block is left with at least some survivor tokens. For instance, if the logic evicts 14k of 16k tokens across 1,000 blocks, most likely every block will still have a token. This means the allocator frees almost nothing. Placing the new tokens into those freed slots is not ideal because it breaks the cache's layout. Say token 16,001 arrives, and it's placed in the slot the 40th token used to hold. The cache now reads position 38, then 16,001, then 41, so the cache is no longer in token order. Attention can still compute the right answer from that, but only if every slot now carries a separate note recording which position it actually holds. This introduces another bookkeeping cost that an in-order layout inherently avoids. So the cache is logically 90% smaller and still physically the same size. Many compression results miss this because they measure on pre-allocated contiguous tensors rather than a paged server. There's another problem. Eviction methods pick which tokens to keep by looking at the attention scores themselves (as expected). But fast attention kernels used in production, like FlashAttention, never save those scores. They compute attention in small pieces and throw the full score grid away as they go, which is also why they're fast. So the exact signal eviction methods need isn't available in memory. The workaround is to fall back to eager attention and build the full matrix, which gives up the speed FlashAttention was there to provide. NVIDIA published a method called TriAttention to solve both these problems. It never needs attention scores. Instead, it scores tokens from the geometry of the model's key and query vectors before RoPE is applied, where those vectors sit in stable clusters. For the memory problem, it runs a compaction pass every 128 decoded tokens. The surviving tokens slide forward to close the holes eviction creates, so whole blocks empty out and return to the allocator while the cache stays in token order. On long reasoning traces, the approach matches full-attention accuracy while decoding 2.5x faster and using 10.7x less KV memory. KV cache compression is a big infrastructure problem. The number that decides whether it works is the count of freed blocks, not the count of evicted tokens. You can find the NVIDIA write-up here: I wrote a first-principles breakdown of how the KV cache works. It walks through why the model stores keys and values at all, why the cache grows with every token, and a comparison of LLM generation speed with and without KV caching. Read it below.

Avi Chawla

271,839 görüntüleme • 2 ay önce

Researchers found a way to make LLMs 8.5x faster! (without compromising accuracy) Speculative decoding is quite an effective way to address the single-token bottleneck in traditional LLM inference. A small "draft" model first generates the next several tokens, then the large model verifies all of them at once in a single forward pass. If a token at any position is wrong, you keep everything before it and restart from there. This never does worse than normal decoding. But current drafters in Speculative decoding still guess one token at a time. That makes the drafting step itself a bottleneck, capping real-world speedups at 2-3x. DFlash is a new technique that swaps the autoregressive drafter with a lightweight block diffusion model that guesses all tokens in one parallel shot. Drafting cost stays flat no matter how many tokens you speculate. On top of that, the drafter is conditioned on hidden features pulled from multiple layers of the target model and injected into every draft layer, so it makes significantly better guesses than a drafter working from scratch. In the side-by-side demo below, vanilla decoding runs at 48.5 tokens/sec. DFlash hits 415 tokens/sec on the same model, with zero quality loss. It's already integrated with vLLM, SGLang, and Transformers, with draft models on HuggingFace for several models like Qwen3, Qwen3.5, Llama 3.1, Kimi-K2.5, gpt-oss, and many more. I have shared the GitHub repo in the replies! KV caching is another must-know technique to boost LLM inference. I recently wrote an article about it. Read it below. 👉 Over to you: What use case are you working on that can benefit from this new technique?

Avi Chawla

157,390 görüntüleme • 4 ay önce

New short course: Attention in Transformers: Concepts and Code in PyTorch. Last week we released a course on how LLM transformers work. This week, go deeper and learn about the technical ideas behind the attention mechanism, and see how to code it in PyTorch. This course is built with Joshua Starmer, Founder and CEO of StatQuest. The attention mechanism was a breakthrough that led to transformers, the architecture powering large language models like ChatGPT. Transformers, introduced in the 2017 paper: "Attention is All You Need" by Viswani and others, took off because of its highly scalable design. In this course, you’ll learn how the attention mechanism, a key element of transformer-based LLMs, works and implement it in PyTorch. You'll develop deep intuition about building reliable, functional, and scalable AI applications. What you will do: - Understand the evolution of the attention mechanism, a key breakthrough that led to transformers. - Learn the relationships between word embeddings, positional embeddings, and attention. - Learn about the Query, Key, and Value matrices, and how to produce and use them in attention. - Walk through the math required to calculate self-attention and masked self-attention to learn why and how they work. - Understand the difference between self-attention and masked self-attention and how one is used in the encoder to build context-aware embeddings and the other is used in the decoder for generative outputs. - Learn the details of the encoder-decoder architecture, cross-attention, and multi-head attention and how they are all incorporated into a transformer. - Use PyTorch to code a class that implements self-attention, masked self-attention, and multi-head attention. There're lots of exciting technical details in this course. Please sign up here:

Andrew Ng

132,400 görüntüleme • 1 yıl önce

Jonathan Ross just revealed why AI companies aren’t growing faster. Not demand. Not competition. Physics. Ross: “The demand for compute is insatiable.” There isn’t enough compute in the world. Not a temporary shortage. A fundamental gap between what the market wants and what the infrastructure can deliver. Ross: “Right now, one of the biggest complaints of Anthropic is the rate limits. People can’t get enough tokens.” Rate limits aren’t product decisions. They’re rationing. Companies forced to regulate access because infrastructure cannot meet demand. Slower services. Token caps. The only things standing between these companies and a revenue surge they can’t access. Every token cap is a revenue cap. Every slowdown is a sale that didn’t happen. Ross: “If Anthropic was given twice the inference compute, within one month their revenue would almost double.” Read that again. Double the compute. Double the revenue. Within thirty days. That’s not a growth projection. That’s a measurement of how deep the backlog already is. The demand exists right now. It’s sitting in a queue. The only thing between these companies and that revenue is physical hardware they don’t have. This breaks every assumption about how tech companies scale. Usually you scale by finding customers. AI companies have infinite customers. They scale by finding hardware. The constraint isn’t market fit. It isn’t distribution. It isn’t competition. It’s processing power. This is why Jensen Huang is the most important person in the world right now. NVIDIA doesn’t just make chips. It makes the thing every government, every AI lab, and every company racing for this future needs more of and can’t get enough of. The compute bottleneck isn’t a tech industry problem. It’s a civilizational one. The winner of this era isn’t determined by who builds the smartest model. Every major lab has a frontier model. The winner is whoever secures the most compute fastest while everyone else rations what’s left. The race isn’t for intelligence. It’s for infrastructure. And right now there isn’t enough to go around.

Dustin

28,395 görüntüleme • 6 ay önce