Loading video...

Video Failed to Load

Go Home

Updated & turned my Big LLM Architecture Comparison article into a narrated video lecture. The 11 LLM architectures covered in this video: 1. DeepSeek V3/R1 2. OLMo 2 3. Gemma 3 4. Mistral Small 3.1 5. Llama 4 6. Qwen3 7. SmolLM3 8. Kimi 2 9. GPT-OSS 10. Grok...

199,947 views • 1 year ago •via X (Twitter)

40 Comments

Sebastian Raschka's profile picture
Sebastian Raschka1 year ago

And here is a link to the video on YT for easier navigation with chapter marks:

Sebastian Raschka's profile picture
Sebastian Raschka1 year ago

The just-released Qwen3 Next has a crazy large number of experts, and a shared expert. Looks like they already implemented my suggestions 😆

Sebastian Raschka's profile picture
Sebastian Raschka1 year ago

Timely update: The Qwen3 team just released the Qwen3 Next MoE. And it does have a shared expert now!

mrityunjoy panday's profile picture
mrityunjoy panday1 year ago

Can you write a blog on design space of llm architecture

Sebastian Raschka's profile picture
Sebastian Raschka1 year ago

with design space you mean the different choices one can make in terms of attention variant, norm layer placement, etc? I think that's pretty much this article and video when taking the superset of all the components discussed 😊

mrityunjoy panday's profile picture
mrityunjoy panday1 year ago

Agree, I was wondering, of you could also write about potential architecture which are not yet tested.

Sebastian Raschka's profile picture
Sebastian Raschka1 year ago

That’d be a research paper… but perhaps one day haha

Dylan Lamb's profile picture
Dylan Lamb1 year ago

Thanks for making this. It’s so hard to keep up with all the new models dropping every week

Alex Veremeyenko's profile picture
Alex Veremeyenko1 year ago

this video is a solid resource, makes it easier to digest all the new models. great work turning the article into a lecture

ℙö𝕚𝕟𝕥∫♄⊙ρ 🚀🌎🔳's profile picture
ℙö𝕚𝕟𝕥∫♄⊙ρ 🚀🌎🔳1 year ago

Great summary, thanks for sharing!

Ramin's profile picture
Ramin1 year ago

missed our LFM2!

Vedant Korade's profile picture
Vedant Korade1 year ago

very interesting, thanks

pratyush's profile picture
pratyush1 year ago

The gift that keeps on giving

南北西东's profile picture
南北西东1 year ago

this is a fantastic overview thank you

Dilip Mysuru's profile picture
Dilip Mysuru1 year ago

Nice! Your article on the same is a classic reference material.

VLT forever's profile picture
VLT forever1 year ago

Another banger

Jishan Ahmed's profile picture
Jishan Ahmed1 year ago

Thanks, Sebastian! Really appreciate you turning the detailed architecture comparison into a digestible one!

Tarik Hammadou's profile picture
Tarik Hammadou1 year ago

Very cool mate. This is really good. 🙏 T

Robert Youssef's profile picture
Robert Youssef1 year ago

awesome job on the video. breaking down those architectures is no easy feat. definitely helps to keep up with all the changes.

Ruairi ⚽🍊🇪🇺's profile picture
Ruairi ⚽🍊🇪🇺1 year ago

Excellent, some entertainment for my breakfast break

LORD ATU's profile picture
LORD ATU1 year ago

okay

sam ;D's profile picture
sam ;D1 year ago

HG of LLM architectures 🙌🏻

Gerardo Salazar's profile picture
Gerardo Salazar1 year ago

Goated. Can’t wait to watch this.

ModelDrift's profile picture
ModelDrift1 year ago

Always good stuff. I had this all queued up and ready to watch yesterday, just didn't make it though the queue yet!

💫ℹ️🐚▪️🌐💺🗨️🚨®️'s profile picture
💫ℹ️🐚▪️🌐💺🗨️🚨®️1 year ago

What about Gemma 3n and embedding-Gemma?

Sebastian Raschka's profile picture
Sebastian Raschka1 year ago

I removed 3n from the slides due to brevity but It’s in the article

Wassollichhier's profile picture
Wassollichhier1 year ago

why not Mistral 3.2 which is much improved compared to 3.1?

Sebastian Raschka's profile picture
Sebastian Raschka1 year ago

Good point. It’s the same architecture though afaik. According to their model hub readme: “Mistral-Small-3.2-24B-Instruct-2506 is a minor update of Mistral-Small-3.1-24B-Instruct-2503.”

𝕱𝖗𝖆𝖓𝖐𝖙𝖔𝖓𝖞 (Ø,G)'s profile picture
𝕱𝖗𝖆𝖓𝖐𝖙𝖔𝖓𝖞 (Ø,G)1 year ago

Crazy to see how fast all these LLM architectures are stacking up 🤯. Feels like we’re barely catching up with one before another drops. Makes me wonder if #LAM from @GetActionModel will start getting compared alongside these soon.

Suhrab Khan⚡️'s profile picture
Suhrab Khan⚡️1 year ago

Turning this comparison into a video lecture is brilliant. A must-watch for anyone serious about LLMs!

Siolu's profile picture
Siolu1 year ago

Teacher Raschka pumping those sessions out!

Yuki He's profile picture
Yuki He1 year ago

kinda wild how these architectures shape future tools

Bnaf.OG | 🟧's profile picture
Bnaf.OG | 🟧6 months ago

The MLA convergence across this list is the real story — Mistral 3 Large shipped with DeepSeek’s attention architecture within 3 months. When open weights enable that speed of cross-pollination, architectural moats look different. Which of these 11 has the most defensible design?

Vinh Nguyen's profile picture
Vinh Nguyen1 year ago

Thank you for the video.

✌️Stanislaw's profile picture
✌️Stanislaw8 months ago

Here is a video overview of these LLMs

Avi Bhargava's profile picture
Avi Bhargava1 year ago

@rasbt This is wonderful. Would love to more of your views on diffusion models! If you have already written your thoughts on the same, would love a link to it!

Md Fahim's profile picture
Md Fahim1 year ago

Got it! I’ll keep it casual and positive. Ready for the next comment!

Damien's profile picture
Damien11 months ago

very interesting and detailled architecture, glm-4.6 use the same like 4.5 i guess ? And u can have a -10% on Coding Plan subscription for GLM models with

Nguyen Ngoc Hai's profile picture
Nguyen Ngoc Hai1 year ago

Will this content be a part of your updated LLM book? Thank you very much.

Sebastian Raschka's profile picture
Sebastian Raschka1 year ago

My book is focuses on a from-scratch deep-dive coding approach of one of the architectures. Doing that for all 11 would be 11 books 😆

Related Videos