Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Updated & turned my Big LLM Architecture Comparison article into a narrated video lecture. The 11 LLM architectures covered in this video: 1. DeepSeek V3/R1 2. OLMo 2 3. Gemma 3 4. Mistral Small 3.1 5. Llama 4 6. Qwen3 7. SmolLM3 8. Kimi 2 9. GPT-OSS 10. Grok...

199,947 Aufrufe • vor 1 Jahr •via X (Twitter)

40 Kommentare

Profilbild von Sebastian Raschka
Sebastian Raschkavor 1 Jahr

And here is a link to the video on YT for easier navigation with chapter marks:

Profilbild von Sebastian Raschka
Sebastian Raschkavor 1 Jahr

The just-released Qwen3 Next has a crazy large number of experts, and a shared expert. Looks like they already implemented my suggestions 😆

Profilbild von Sebastian Raschka
Sebastian Raschkavor 1 Jahr

Timely update: The Qwen3 team just released the Qwen3 Next MoE. And it does have a shared expert now!

Profilbild von mrityunjoy panday
mrityunjoy pandayvor 1 Jahr

Can you write a blog on design space of llm architecture

Profilbild von Sebastian Raschka
Sebastian Raschkavor 1 Jahr

with design space you mean the different choices one can make in terms of attention variant, norm layer placement, etc? I think that's pretty much this article and video when taking the superset of all the components discussed 😊

Profilbild von mrityunjoy panday
mrityunjoy pandayvor 1 Jahr

Agree, I was wondering, of you could also write about potential architecture which are not yet tested.

Profilbild von Sebastian Raschka
Sebastian Raschkavor 1 Jahr

That’d be a research paper… but perhaps one day haha

Profilbild von Dylan Lamb
Dylan Lambvor 1 Jahr

Thanks for making this. It’s so hard to keep up with all the new models dropping every week

Profilbild von Alex Veremeyenko
Alex Veremeyenkovor 1 Jahr

this video is a solid resource, makes it easier to digest all the new models. great work turning the article into a lecture

Profilbild von ℙö𝕚𝕟𝕥∫♄⊙ρ 🚀🌎🔳
ℙö𝕚𝕟𝕥∫♄⊙ρ 🚀🌎🔳vor 1 Jahr

Great summary, thanks for sharing!

Profilbild von Ramin
Raminvor 1 Jahr

missed our LFM2!

Profilbild von Vedant Korade
Vedant Koradevor 1 Jahr

very interesting, thanks

Profilbild von pratyush
pratyushvor 1 Jahr

The gift that keeps on giving

Profilbild von 南北西东
南北西东vor 1 Jahr

this is a fantastic overview thank you

Profilbild von Dilip Mysuru
Dilip Mysuruvor 1 Jahr

Nice! Your article on the same is a classic reference material.

Profilbild von VLT forever
VLT forevervor 1 Jahr

Another banger

Profilbild von Jishan Ahmed
Jishan Ahmedvor 1 Jahr

Thanks, Sebastian! Really appreciate you turning the detailed architecture comparison into a digestible one!

Profilbild von Tarik Hammadou
Tarik Hammadouvor 1 Jahr

Very cool mate. This is really good. 🙏 T

Profilbild von Robert Youssef
Robert Youssefvor 1 Jahr

awesome job on the video. breaking down those architectures is no easy feat. definitely helps to keep up with all the changes.

Profilbild von Ruairi ⚽🍊🇪🇺
Ruairi ⚽🍊🇪🇺vor 1 Jahr

Excellent, some entertainment for my breakfast break

Profilbild von LORD ATU
LORD ATUvor 1 Jahr

okay

Profilbild von sam ;D
sam ;Dvor 1 Jahr

HG of LLM architectures 🙌🏻

Profilbild von Gerardo Salazar
Gerardo Salazarvor 1 Jahr

Goated. Can’t wait to watch this.

Profilbild von ModelDrift
ModelDriftvor 1 Jahr

Always good stuff. I had this all queued up and ready to watch yesterday, just didn't make it though the queue yet!

Profilbild von 💫ℹ️🐚▪️🌐💺🗨️🚨®️
💫ℹ️🐚▪️🌐💺🗨️🚨®️vor 1 Jahr

What about Gemma 3n and embedding-Gemma?

Profilbild von Sebastian Raschka
Sebastian Raschkavor 1 Jahr

I removed 3n from the slides due to brevity but It’s in the article

Profilbild von Wassollichhier
Wassollichhiervor 1 Jahr

why not Mistral 3.2 which is much improved compared to 3.1?

Profilbild von Sebastian Raschka
Sebastian Raschkavor 1 Jahr

Good point. It’s the same architecture though afaik. According to their model hub readme: “Mistral-Small-3.2-24B-Instruct-2506 is a minor update of Mistral-Small-3.1-24B-Instruct-2503.”

Profilbild von 𝕱𝖗𝖆𝖓𝖐𝖙𝖔𝖓𝖞 (Ø,G)
𝕱𝖗𝖆𝖓𝖐𝖙𝖔𝖓𝖞 (Ø,G)vor 1 Jahr

Crazy to see how fast all these LLM architectures are stacking up 🤯. Feels like we’re barely catching up with one before another drops. Makes me wonder if #LAM from @GetActionModel will start getting compared alongside these soon.

Profilbild von Suhrab Khan⚡️
Suhrab Khan⚡️vor 1 Jahr

Turning this comparison into a video lecture is brilliant. A must-watch for anyone serious about LLMs!

Profilbild von Siolu
Sioluvor 1 Jahr

Teacher Raschka pumping those sessions out!

Profilbild von Yuki He
Yuki Hevor 1 Jahr

kinda wild how these architectures shape future tools

Profilbild von Bnaf.OG | 🟧
Bnaf.OG | 🟧vor 6 Monaten

The MLA convergence across this list is the real story — Mistral 3 Large shipped with DeepSeek’s attention architecture within 3 months. When open weights enable that speed of cross-pollination, architectural moats look different. Which of these 11 has the most defensible design?

Profilbild von Vinh Nguyen
Vinh Nguyenvor 1 Jahr

Thank you for the video.

Profilbild von ✌️Stanislaw
✌️Stanislawvor 8 Monaten

Here is a video overview of these LLMs

Profilbild von Avi Bhargava
Avi Bhargavavor 1 Jahr

@rasbt This is wonderful. Would love to more of your views on diffusion models! If you have already written your thoughts on the same, would love a link to it!

Profilbild von Md Fahim
Md Fahimvor 1 Jahr

Got it! I’ll keep it casual and positive. Ready for the next comment!

Profilbild von Damien
Damienvor 11 Monaten

very interesting and detailled architecture, glm-4.6 use the same like 4.5 i guess ? And u can have a -10% on Coding Plan subscription for GLM models with

Profilbild von Nguyen Ngoc Hai
Nguyen Ngoc Haivor 1 Jahr

Will this content be a part of your updated LLM book? Thank you very much.

Profilbild von Sebastian Raschka
Sebastian Raschkavor 1 Jahr

My book is focuses on a from-scratch deep-dive coding approach of one of the architectures. Doing that for all 11 would be 11 books 😆

Ähnliche Videos