正在加载视频...

视频加载失败

Updated & turned my Big LLM Architecture Comparison article into a narrated video lecture. The 11 LLM architectures covered in this video: 1. DeepSeek V3/R1 2. OLMo 2 3. Gemma 3 4. Mistral Small 3.1 5. Llama 4 6. Qwen3 7. SmolLM3 8. Kimi 2 9. GPT-OSS 10. Grok...

199,947 次观看 • 1 年前 •via X (Twitter)

40 条评论

Sebastian Raschka 的头像
Sebastian Raschka1 年前

And here is a link to the video on YT for easier navigation with chapter marks:

Sebastian Raschka 的头像
Sebastian Raschka1 年前

The just-released Qwen3 Next has a crazy large number of experts, and a shared expert. Looks like they already implemented my suggestions 😆

Sebastian Raschka 的头像
Sebastian Raschka1 年前

Timely update: The Qwen3 team just released the Qwen3 Next MoE. And it does have a shared expert now!

mrityunjoy panday 的头像
mrityunjoy panday1 年前

Can you write a blog on design space of llm architecture

Sebastian Raschka 的头像
Sebastian Raschka1 年前

with design space you mean the different choices one can make in terms of attention variant, norm layer placement, etc? I think that's pretty much this article and video when taking the superset of all the components discussed 😊

mrityunjoy panday 的头像
mrityunjoy panday1 年前

Agree, I was wondering, of you could also write about potential architecture which are not yet tested.

Sebastian Raschka 的头像
Sebastian Raschka1 年前

That’d be a research paper… but perhaps one day haha

Dylan Lamb 的头像
Dylan Lamb1 年前

Thanks for making this. It’s so hard to keep up with all the new models dropping every week

Alex Veremeyenko 的头像
Alex Veremeyenko1 年前

this video is a solid resource, makes it easier to digest all the new models. great work turning the article into a lecture

ℙö𝕚𝕟𝕥∫♄⊙ρ 🚀🌎🔳 的头像
ℙö𝕚𝕟𝕥∫♄⊙ρ 🚀🌎🔳1 年前

Great summary, thanks for sharing!

Ramin 的头像
Ramin1 年前

missed our LFM2!

Vedant Korade 的头像
Vedant Korade1 年前

very interesting, thanks

pratyush 的头像
pratyush1 年前

The gift that keeps on giving

南北西东 的头像
南北西东1 年前

this is a fantastic overview thank you

Dilip Mysuru 的头像
Dilip Mysuru1 年前

Nice! Your article on the same is a classic reference material.

VLT forever 的头像
VLT forever1 年前

Another banger

Jishan Ahmed 的头像
Jishan Ahmed1 年前

Thanks, Sebastian! Really appreciate you turning the detailed architecture comparison into a digestible one!

Tarik Hammadou 的头像
Tarik Hammadou1 年前

Very cool mate. This is really good. 🙏 T

Robert Youssef 的头像
Robert Youssef1 年前

awesome job on the video. breaking down those architectures is no easy feat. definitely helps to keep up with all the changes.

Ruairi ⚽🍊🇪🇺 的头像
Ruairi ⚽🍊🇪🇺1 年前

Excellent, some entertainment for my breakfast break

LORD ATU 的头像
LORD ATU1 年前

okay

sam ;D 的头像
sam ;D1 年前

HG of LLM architectures 🙌🏻

Gerardo Salazar 的头像
Gerardo Salazar1 年前

Goated. Can’t wait to watch this.

ModelDrift 的头像
ModelDrift1 年前

Always good stuff. I had this all queued up and ready to watch yesterday, just didn't make it though the queue yet!

💫ℹ️🐚▪️🌐💺🗨️🚨®️ 的头像
💫ℹ️🐚▪️🌐💺🗨️🚨®️1 年前

What about Gemma 3n and embedding-Gemma?

Sebastian Raschka 的头像
Sebastian Raschka1 年前

I removed 3n from the slides due to brevity but It’s in the article

Wassollichhier 的头像
Wassollichhier1 年前

why not Mistral 3.2 which is much improved compared to 3.1?

Sebastian Raschka 的头像
Sebastian Raschka1 年前

Good point. It’s the same architecture though afaik. According to their model hub readme: “Mistral-Small-3.2-24B-Instruct-2506 is a minor update of Mistral-Small-3.1-24B-Instruct-2503.”

𝕱𝖗𝖆𝖓𝖐𝖙𝖔𝖓𝖞 (Ø,G) 的头像
𝕱𝖗𝖆𝖓𝖐𝖙𝖔𝖓𝖞 (Ø,G)1 年前

Crazy to see how fast all these LLM architectures are stacking up 🤯. Feels like we’re barely catching up with one before another drops. Makes me wonder if #LAM from @GetActionModel will start getting compared alongside these soon.

Suhrab Khan⚡️ 的头像
Suhrab Khan⚡️1 年前

Turning this comparison into a video lecture is brilliant. A must-watch for anyone serious about LLMs!

Siolu 的头像
Siolu1 年前

Teacher Raschka pumping those sessions out!

Yuki He 的头像
Yuki He1 年前

kinda wild how these architectures shape future tools

Bnaf.OG | 🟧 的头像
Bnaf.OG | 🟧6 个月前

The MLA convergence across this list is the real story — Mistral 3 Large shipped with DeepSeek’s attention architecture within 3 months. When open weights enable that speed of cross-pollination, architectural moats look different. Which of these 11 has the most defensible design?

Vinh Nguyen 的头像
Vinh Nguyen1 年前

Thank you for the video.

✌️Stanislaw 的头像
✌️Stanislaw8 个月前

Here is a video overview of these LLMs

Avi Bhargava 的头像
Avi Bhargava1 年前

@rasbt This is wonderful. Would love to more of your views on diffusion models! If you have already written your thoughts on the same, would love a link to it!

Md Fahim 的头像
Md Fahim1 年前

Got it! I’ll keep it casual and positive. Ready for the next comment!

Damien 的头像
Damien11 个月前

very interesting and detailled architecture, glm-4.6 use the same like 4.5 i guess ? And u can have a -10% on Coding Plan subscription for GLM models with

Nguyen Ngoc Hai 的头像
Nguyen Ngoc Hai1 年前

Will this content be a part of your updated LLM book? Thank you very much.

Sebastian Raschka 的头像
Sebastian Raschka1 年前

My book is focuses on a from-scratch deep-dive coding approach of one of the architectures. Doing that for all 11 would be 11 books 😆

相关视频