Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Updated & turned my Big LLM Architecture Comparison article into a narrated video lecture. The 11 LLM architectures covered in this video: 1. DeepSeek V3/R1 2. OLMo 2 3. Gemma 3 4. Mistral Small 3.1 5. Llama 4 6. Qwen3 7. SmolLM3 8. Kimi 2 9. GPT-OSS 10. Grok...

199,947 görüntüleme • 1 yıl önce •via X (Twitter)

40 Yorum

Sebastian Raschka profil fotoğrafı
Sebastian Raschka1 yıl önce

And here is a link to the video on YT for easier navigation with chapter marks:

Sebastian Raschka profil fotoğrafı
Sebastian Raschka1 yıl önce

The just-released Qwen3 Next has a crazy large number of experts, and a shared expert. Looks like they already implemented my suggestions 😆

Sebastian Raschka profil fotoğrafı
Sebastian Raschka1 yıl önce

Timely update: The Qwen3 team just released the Qwen3 Next MoE. And it does have a shared expert now!

mrityunjoy panday profil fotoğrafı
mrityunjoy panday1 yıl önce

Can you write a blog on design space of llm architecture

Sebastian Raschka profil fotoğrafı
Sebastian Raschka1 yıl önce

with design space you mean the different choices one can make in terms of attention variant, norm layer placement, etc? I think that's pretty much this article and video when taking the superset of all the components discussed 😊

mrityunjoy panday profil fotoğrafı
mrityunjoy panday1 yıl önce

Agree, I was wondering, of you could also write about potential architecture which are not yet tested.

Sebastian Raschka profil fotoğrafı
Sebastian Raschka1 yıl önce

That’d be a research paper… but perhaps one day haha

Dylan Lamb profil fotoğrafı
Dylan Lamb1 yıl önce

Thanks for making this. It’s so hard to keep up with all the new models dropping every week

Alex Veremeyenko profil fotoğrafı
Alex Veremeyenko1 yıl önce

this video is a solid resource, makes it easier to digest all the new models. great work turning the article into a lecture

ℙö𝕚𝕟𝕥∫♄⊙ρ 🚀🌎🔳 profil fotoğrafı
ℙö𝕚𝕟𝕥∫♄⊙ρ 🚀🌎🔳1 yıl önce

Great summary, thanks for sharing!

Ramin profil fotoğrafı
Ramin1 yıl önce

missed our LFM2!

Vedant Korade profil fotoğrafı
Vedant Korade1 yıl önce

very interesting, thanks

pratyush profil fotoğrafı
pratyush1 yıl önce

The gift that keeps on giving

南北西东 profil fotoğrafı
南北西东1 yıl önce

this is a fantastic overview thank you

Dilip Mysuru profil fotoğrafı
Dilip Mysuru1 yıl önce

Nice! Your article on the same is a classic reference material.

VLT forever profil fotoğrafı
VLT forever1 yıl önce

Another banger

Jishan Ahmed profil fotoğrafı
Jishan Ahmed1 yıl önce

Thanks, Sebastian! Really appreciate you turning the detailed architecture comparison into a digestible one!

Tarik Hammadou profil fotoğrafı
Tarik Hammadou1 yıl önce

Very cool mate. This is really good. 🙏 T

Robert Youssef profil fotoğrafı
Robert Youssef1 yıl önce

awesome job on the video. breaking down those architectures is no easy feat. definitely helps to keep up with all the changes.

Ruairi ⚽🍊🇪🇺 profil fotoğrafı
Ruairi ⚽🍊🇪🇺1 yıl önce

Excellent, some entertainment for my breakfast break

LORD ATU profil fotoğrafı
LORD ATU1 yıl önce

okay

sam ;D profil fotoğrafı
sam ;D1 yıl önce

HG of LLM architectures 🙌🏻

Gerardo Salazar profil fotoğrafı
Gerardo Salazar1 yıl önce

Goated. Can’t wait to watch this.

ModelDrift profil fotoğrafı
ModelDrift1 yıl önce

Always good stuff. I had this all queued up and ready to watch yesterday, just didn't make it though the queue yet!

💫ℹ️🐚▪️🌐💺🗨️🚨®️ profil fotoğrafı
💫ℹ️🐚▪️🌐💺🗨️🚨®️1 yıl önce

What about Gemma 3n and embedding-Gemma?

Sebastian Raschka profil fotoğrafı
Sebastian Raschka1 yıl önce

I removed 3n from the slides due to brevity but It’s in the article

Wassollichhier profil fotoğrafı
Wassollichhier1 yıl önce

why not Mistral 3.2 which is much improved compared to 3.1?

Sebastian Raschka profil fotoğrafı
Sebastian Raschka1 yıl önce

Good point. It’s the same architecture though afaik. According to their model hub readme: “Mistral-Small-3.2-24B-Instruct-2506 is a minor update of Mistral-Small-3.1-24B-Instruct-2503.”

𝕱𝖗𝖆𝖓𝖐𝖙𝖔𝖓𝖞 (Ø,G) profil fotoğrafı
𝕱𝖗𝖆𝖓𝖐𝖙𝖔𝖓𝖞 (Ø,G)1 yıl önce

Crazy to see how fast all these LLM architectures are stacking up 🤯. Feels like we’re barely catching up with one before another drops. Makes me wonder if #LAM from @GetActionModel will start getting compared alongside these soon.

Suhrab Khan⚡️ profil fotoğrafı
Suhrab Khan⚡️1 yıl önce

Turning this comparison into a video lecture is brilliant. A must-watch for anyone serious about LLMs!

Siolu profil fotoğrafı
Siolu1 yıl önce

Teacher Raschka pumping those sessions out!

Yuki He profil fotoğrafı
Yuki He1 yıl önce

kinda wild how these architectures shape future tools

Bnaf.OG | 🟧 profil fotoğrafı
Bnaf.OG | 🟧6 ay önce

The MLA convergence across this list is the real story — Mistral 3 Large shipped with DeepSeek’s attention architecture within 3 months. When open weights enable that speed of cross-pollination, architectural moats look different. Which of these 11 has the most defensible design?

Vinh Nguyen profil fotoğrafı
Vinh Nguyen1 yıl önce

Thank you for the video.

✌️Stanislaw profil fotoğrafı
✌️Stanislaw8 ay önce

Here is a video overview of these LLMs

Avi Bhargava profil fotoğrafı
Avi Bhargava1 yıl önce

@rasbt This is wonderful. Would love to more of your views on diffusion models! If you have already written your thoughts on the same, would love a link to it!

Md Fahim profil fotoğrafı
Md Fahim1 yıl önce

Got it! I’ll keep it casual and positive. Ready for the next comment!

Damien profil fotoğrafı
Damien11 ay önce

very interesting and detailled architecture, glm-4.6 use the same like 4.5 i guess ? And u can have a -10% on Coding Plan subscription for GLM models with

Nguyen Ngoc Hai profil fotoğrafı
Nguyen Ngoc Hai1 yıl önce

Will this content be a part of your updated LLM book? Thank you very much.

Sebastian Raschka profil fotoğrafı
Sebastian Raschka1 yıl önce

My book is focuses on a from-scratch deep-dive coding approach of one of the architectures. Doing that for all 11 would be 11 books 😆

Benzer Videolar