Загрузка видео...
Не удалось загрузить видео
Updated & turned my Big LLM Architecture Comparison article into a narrated video lecture. The 11 LLM architectures covered in this video: 1. DeepSeek V3/R1 2. OLMo 2 3. Gemma 3 4. Mistral Small 3.1 5. Llama 4 6. Qwen3 7. SmolLM3 8. Kimi 2 9. GPT-OSS 10. Grok... show more
199,947 просмотров • 1 год назад •via X (Twitter)
Комментарии: 40

And here is a link to the video on YT for easier navigation with chapter marks:

The just-released Qwen3 Next has a crazy large number of experts, and a shared expert. Looks like they already implemented my suggestions 😆

Timely update: The Qwen3 team just released the Qwen3 Next MoE. And it does have a shared expert now!

Can you write a blog on design space of llm architecture

with design space you mean the different choices one can make in terms of attention variant, norm layer placement, etc? I think that's pretty much this article and video when taking the superset of all the components discussed 😊

Agree, I was wondering, of you could also write about potential architecture which are not yet tested.

That’d be a research paper… but perhaps one day haha

Thanks for making this. It’s so hard to keep up with all the new models dropping every week

this video is a solid resource, makes it easier to digest all the new models. great work turning the article into a lecture

Great summary, thanks for sharing!

missed our LFM2!

very interesting, thanks

The gift that keeps on giving

this is a fantastic overview thank you

Nice! Your article on the same is a classic reference material.

Another banger

Thanks, Sebastian! Really appreciate you turning the detailed architecture comparison into a digestible one!

Very cool mate. This is really good. 🙏 T

awesome job on the video. breaking down those architectures is no easy feat. definitely helps to keep up with all the changes.

Excellent, some entertainment for my breakfast break

okay

HG of LLM architectures 🙌🏻

Goated. Can’t wait to watch this.

Always good stuff. I had this all queued up and ready to watch yesterday, just didn't make it though the queue yet!

What about Gemma 3n and embedding-Gemma?

I removed 3n from the slides due to brevity but It’s in the article

why not Mistral 3.2 which is much improved compared to 3.1?

Good point. It’s the same architecture though afaik. According to their model hub readme: “Mistral-Small-3.2-24B-Instruct-2506 is a minor update of Mistral-Small-3.1-24B-Instruct-2503.”

Crazy to see how fast all these LLM architectures are stacking up 🤯. Feels like we’re barely catching up with one before another drops. Makes me wonder if #LAM from @GetActionModel will start getting compared alongside these soon.

Turning this comparison into a video lecture is brilliant. A must-watch for anyone serious about LLMs!

Teacher Raschka pumping those sessions out!

kinda wild how these architectures shape future tools

The MLA convergence across this list is the real story — Mistral 3 Large shipped with DeepSeek’s attention architecture within 3 months. When open weights enable that speed of cross-pollination, architectural moats look different. Which of these 11 has the most defensible design?

Thank you for the video.

Here is a video overview of these LLMs

@rasbt This is wonderful. Would love to more of your views on diffusion models! If you have already written your thoughts on the same, would love a link to it!

Got it! I’ll keep it casual and positive. Ready for the next comment!

very interesting and detailled architecture, glm-4.6 use the same like 4.5 i guess ? And u can have a -10% on Coding Plan subscription for GLM models with

Will this content be a part of your updated LLM book? Thank you very much.

My book is focuses on a from-scratch deep-dive coding approach of one of the architectures. Doing that for all 11 would be 11 books 😆
