Загрузка видео...

Не удалось загрузить видео

На главную

Updated & turned my Big LLM Architecture Comparison article into a narrated video lecture. The 11 LLM architectures covered in this video: 1. DeepSeek V3/R1 2. OLMo 2 3. Gemma 3 4. Mistral Small 3.1 5. Llama 4 6. Qwen3 7. SmolLM3 8. Kimi 2 9. GPT-OSS 10. Grok...

199,947 просмотров • 1 год назад •via X (Twitter)

Комментарии: 40

Фото профиля Sebastian Raschka
Sebastian Raschka1 год назад

And here is a link to the video on YT for easier navigation with chapter marks:

Фото профиля Sebastian Raschka
Sebastian Raschka1 год назад

The just-released Qwen3 Next has a crazy large number of experts, and a shared expert. Looks like they already implemented my suggestions 😆

Фото профиля Sebastian Raschka
Sebastian Raschka1 год назад

Timely update: The Qwen3 team just released the Qwen3 Next MoE. And it does have a shared expert now!

Фото профиля mrityunjoy panday
mrityunjoy panday1 год назад

Can you write a blog on design space of llm architecture

Фото профиля Sebastian Raschka
Sebastian Raschka1 год назад

with design space you mean the different choices one can make in terms of attention variant, norm layer placement, etc? I think that's pretty much this article and video when taking the superset of all the components discussed 😊

Фото профиля mrityunjoy panday
mrityunjoy panday1 год назад

Agree, I was wondering, of you could also write about potential architecture which are not yet tested.

Фото профиля Sebastian Raschka
Sebastian Raschka1 год назад

That’d be a research paper… but perhaps one day haha

Фото профиля Dylan Lamb
Dylan Lamb1 год назад

Thanks for making this. It’s so hard to keep up with all the new models dropping every week

Фото профиля Alex Veremeyenko
Alex Veremeyenko1 год назад

this video is a solid resource, makes it easier to digest all the new models. great work turning the article into a lecture

Фото профиля ℙö𝕚𝕟𝕥∫♄⊙ρ 🚀🌎🔳
ℙö𝕚𝕟𝕥∫♄⊙ρ 🚀🌎🔳1 год назад

Great summary, thanks for sharing!

Фото профиля Ramin
Ramin1 год назад

missed our LFM2!

Фото профиля Vedant Korade
Vedant Korade1 год назад

very interesting, thanks

Фото профиля pratyush
pratyush1 год назад

The gift that keeps on giving

Фото профиля 南北西东
南北西东1 год назад

this is a fantastic overview thank you

Фото профиля Dilip Mysuru
Dilip Mysuru1 год назад

Nice! Your article on the same is a classic reference material.

Фото профиля VLT forever
VLT forever1 год назад

Another banger

Фото профиля Jishan Ahmed
Jishan Ahmed1 год назад

Thanks, Sebastian! Really appreciate you turning the detailed architecture comparison into a digestible one!

Фото профиля Tarik Hammadou
Tarik Hammadou1 год назад

Very cool mate. This is really good. 🙏 T

Фото профиля Robert Youssef
Robert Youssef1 год назад

awesome job on the video. breaking down those architectures is no easy feat. definitely helps to keep up with all the changes.

Фото профиля Ruairi ⚽🍊🇪🇺
Ruairi ⚽🍊🇪🇺1 год назад

Excellent, some entertainment for my breakfast break

Фото профиля LORD ATU
LORD ATU1 год назад

okay

Фото профиля sam ;D
sam ;D1 год назад

HG of LLM architectures 🙌🏻

Фото профиля Gerardo Salazar
Gerardo Salazar1 год назад

Goated. Can’t wait to watch this.

Фото профиля ModelDrift
ModelDrift1 год назад

Always good stuff. I had this all queued up and ready to watch yesterday, just didn't make it though the queue yet!

Фото профиля 💫ℹ️🐚▪️🌐💺🗨️🚨®️
💫ℹ️🐚▪️🌐💺🗨️🚨®️1 год назад

What about Gemma 3n and embedding-Gemma?

Фото профиля Sebastian Raschka
Sebastian Raschka1 год назад

I removed 3n from the slides due to brevity but It’s in the article

Фото профиля Wassollichhier
Wassollichhier1 год назад

why not Mistral 3.2 which is much improved compared to 3.1?

Фото профиля Sebastian Raschka
Sebastian Raschka1 год назад

Good point. It’s the same architecture though afaik. According to their model hub readme: “Mistral-Small-3.2-24B-Instruct-2506 is a minor update of Mistral-Small-3.1-24B-Instruct-2503.”

Фото профиля 𝕱𝖗𝖆𝖓𝖐𝖙𝖔𝖓𝖞 (Ø,G)
𝕱𝖗𝖆𝖓𝖐𝖙𝖔𝖓𝖞 (Ø,G)1 год назад

Crazy to see how fast all these LLM architectures are stacking up 🤯. Feels like we’re barely catching up with one before another drops. Makes me wonder if #LAM from @GetActionModel will start getting compared alongside these soon.

Фото профиля Suhrab Khan⚡️
Suhrab Khan⚡️1 год назад

Turning this comparison into a video lecture is brilliant. A must-watch for anyone serious about LLMs!

Фото профиля Siolu
Siolu1 год назад

Teacher Raschka pumping those sessions out!

Фото профиля Yuki He
Yuki He1 год назад

kinda wild how these architectures shape future tools

Фото профиля Bnaf.OG | 🟧
Bnaf.OG | 🟧6 месяцев назад

The MLA convergence across this list is the real story — Mistral 3 Large shipped with DeepSeek’s attention architecture within 3 months. When open weights enable that speed of cross-pollination, architectural moats look different. Which of these 11 has the most defensible design?

Фото профиля Vinh Nguyen
Vinh Nguyen1 год назад

Thank you for the video.

Фото профиля ✌️Stanislaw
✌️Stanislaw8 месяцев назад

Here is a video overview of these LLMs

Фото профиля Avi Bhargava
Avi Bhargava1 год назад

@rasbt This is wonderful. Would love to more of your views on diffusion models! If you have already written your thoughts on the same, would love a link to it!

Фото профиля Md Fahim
Md Fahim1 год назад

Got it! I’ll keep it casual and positive. Ready for the next comment!

Фото профиля Damien
Damien11 месяцев назад

very interesting and detailled architecture, glm-4.6 use the same like 4.5 i guess ? And u can have a -10% on Coding Plan subscription for GLM models with

Фото профиля Nguyen Ngoc Hai
Nguyen Ngoc Hai1 год назад

Will this content be a part of your updated LLM book? Thank you very much.

Фото профиля Sebastian Raschka
Sebastian Raschka1 год назад

My book is focuses on a from-scratch deep-dive coding approach of one of the architectures. Doing that for all 11 would be 11 books 😆

Похожие видео