正在加载视频...
视频加载失败
Introducing SubQ - a major breakthrough in LLM intelligence. It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA), And the first frontier model with a 12 million token context window which is: - 52x faster than FlashAttention at 1MM tokens - Less than 5% the... show more
70 条评论

SubQ is available for early access today, alongside our coding agent, SubQ Code Get access today ↓

We were a little slow on this, but we just got a technical blog post up with more details. Please take a look! We have a model card coming next week, and we are happy to take requests for any specific details there. I am happy to answer any questions here!

congrats!! why only 150tok/sec if 52x faster, etc.?

52x is for prefill speed. Decoding speedups are coming too! This is the first research announcement, with many to follow.

@MartinShkreli absolute grift, didn't deliver

@MartinShkreli Model card shared!

Any papers? Seems too good to be true

The model card is coming next week! We are releasing a technical blog post with more details later today.

Looking forward to reading it

i still can't believe this.. 12M context window with 98% accuracy 💀 @alex_whedon @subquadratic

@subquadratic And at a fraction of the cost...

If this works as described, it basically changes how we build with LLMs. A lot of current pipelines exist just to work around context limits.

And now that constraint goes away, what people choose to build is going to be very different.

I remember when Gemini first made the claim of 1M in lab situations. We've come quite a bit from there. I'm extremely intrigued to see the real-world implications of this in bigger and more complex AI powered software. There is the zero shot environment, but there is also the qualitative repetitive work that is prone to drift and degradation over time. I really hope you will share case-studies, since I think this will be one of the most impressive ways to highlight the capabilities. Amazing work Alex and team. Kudos.

This is the framing we'd love to test against. If you've got a workload that historically drifts, would genuinely want you to throw it at SubQ and see what happens.

If SubQ only attends to a small subset of tokens, how does it guarantee recall of critical long range dependencies in worst case inputs (e.g adversarial prompts or tasks where the important tokens aren’t locally obvious)

We dynamically select the token relationships that matter and train it against retrieval problems with distractor docs, etc.

What happens when the selector misses the one token that matters? Any fallback

I agree, it’s about frequency, not perfection.

A few weeks ago I was in a discussion under an X post about LLMs, their adoption and cost. Most people argued it doesn't scale and brought up the usual economic concerns. My take was that it's only a matter of time before the next breakthrough. It's always been this way with tech. Congrats on this one 👏

This is only the first breakthrough we have announced! More coming.

What 🤯 So Opus 4.7 is at $15 per million tokens and you guyes are giving it for under $1.50.

I hope you consider releasing some variants as open weights. This would change the game for private on-premise work.

This first model won't be open-source, but we do want to contribute to the open-source community!

@ChemPhysMajor Atleast sharing how you did it?

@CommiePat1776 @alex_whedon A white paper would go a long way.

close enough, welcome back brampton

The transformer was the first workable answer to long context. Everyone scaled it so hard nobody wanted to admit it was a local maximum. You finally found a way out.

Thanks @carlvellotti. Try it out at the team will appreciate your insights.

5% of Anthropic’s price at 98% accuracy at 12M tokens. The math on every agent product being built right now just changed completely.

The unit economics of agentic products were quietly upside down. Well, not anymore.

spent 2 years engineering around AI that couldn't read long docs. that limit just got removed.

Throw your hardest doc workload at it, and let me know how it performs. Check out

Long-context at under 10% of Anthropic's price damn…..

I mean, someone had to do it

Release the API or it never happened

As someone who has worked with you for years, I am excited to see you finally releasing this. ✨ may your waitlist be buzzing and your tokens be cheap ✨

Thank you!

Is this real? Wow. Is there any evals where we see it with other forontier models?

Congrats to the @subquadratic team! It’s awesome to see someone finally breaking out of the standard transformer box

Being a SaaS investor in 2026 sounds like Dante’s Inferno

Wtf

I love not wasting tokens, superb

Much of standard attention's compute is models talking to themselves about words that don't matter.

- Less than 5% the cost of Opus TAKE MY MONEY

Super interesting Alexander, and 12m token context window, holy moly 🤯

52x faster than FlashAttention, <5% of Opus pricing

Totally insane.

"It is the first model built on a fully sub-quadratic sparse-attention architecture (SSA)" This doesn't seem right? That already exists?

this is literally epic wtf. ik another lab would make it

Congrats on the announcement and I'm looking forward to giving this a try.

I’m sorry, wtf. Paper?

Next week!

paper?

This week!

@vxnuaj This paper is never gonna get released lol.

@alex_whedon patience young padawan.

@alex_whedon learn to have patience, you must

ummm...this is a "BIG F*CKING DEAL"?

and it's available for early access at

just submitted a request! would love to try it out and write a post about it.

Was a pleasure collaborating w you guys over the last few weeks!

congratulations it'll be interesting to see how it manages to avoid the quality cliff on a long context window

@NotTheCh05en1 Thanks Devansh, would love for you to try it out.

Congrats on the launch. Can't believe we have a model with a context window this large with this accuracy!

Thanks @itsPaulAi! Would love for you try out the product & share any feedback

Finally, an LLM thats fast, powerful, and cheap at the same time. Opus is extremely expensive.

The tradeoff people learned to accept was "pick two of fast/cheap/smart." We changed that.

If someone said that enterprises are burning cash with AI subscriptions, that ends today. Great work @alex_whedon

No subquadratic tax!
