Загрузка видео...

Не удалось загрузить видео

На главную

Completed a first hour side-by-side comparison between Qwen3.5 27b and Qwen3.6 27b on the same 4 canvas coding tests. Running the Qwen3.6 27b FP8, vLLM. What do you think?

114,366 просмотров • 5 месяцев назад •via X (Twitter)

Комментарии: 40

Фото профиля Schinsly ✝️
Schinsly ✝️5 месяцев назад

why are we still running these retarded tests. like I actually really hate this

Фото профиля rijndael
rijndael5 месяцев назад

qwen still doesnt know what a roof looks like 😂

Фото профиля Oxytocin
Oxytocin5 месяцев назад

so close yet so far

Фото профиля Sakura Yuki
Sakura Yuki5 месяцев назад

Been running the 3.5 27B on my 5070 Ti. The FP8 quant is surprisingly close to the full model for code gen. Did you notice any latency hits on 3.6 with vLLM, or is the TTFT basically the same?

Фото профиля stevibe
stevibe5 месяцев назад

Didn't measure the TTFT yet, will check it later.

Фото профиля Flow
Flow5 месяцев назад

Stunningly good. Both of them but clearly better. The 35B has an issue with unverified assumptions, it tends to assume something without looking into it. I wonder how it would perform your test if you adapt the system prompt and tell it to actively monitor for unverified assumptions.

Фото профиля stevibe
stevibe5 месяцев назад

Good take, will try to update the system prompt next time.

Фото профиля Khalid
Khalid5 месяцев назад

do a 3.6 32b vs 27b comparison to assess the marginal gap.

Фото профиля Morgan
Morgan5 месяцев назад

Super interesting. On the house building prompt, it's interesting to see both build something out of brick on top of the house, seemed to confuse them both. The fish demo is my favorite, kinda wild that 3.5 used triangles rather than fish. For 3.6, the fish have a crazy shake they do, but at least they're actually fish! 🐟 Of course, it goes without saying that 3.6 is insanely good, best local AI I think we've ever seen.

Фото профиля stevibe
stevibe5 месяцев назад

Definitely, We can definitely see the progress in local models.

Фото профиля Morgan
Morgan5 месяцев назад

Ohhh, honor to see you in the comments! You’ve been sharing some amazing stuff 🔥

Фото профиля Jonathan Leaders
Jonathan Leaders5 месяцев назад

Now do 3.6 MOE vs 3.6 dense

Фото профиля Seth Pratt
Seth Pratt5 месяцев назад

3.6 is a really good model. Full precision is notably higher performing on my benchmarks than the quantized versions.

Фото профиля Edzward
Edzward5 месяцев назад

Personally, I don't like purely abstract test scenarios. For me, it's much more important to test real use case's scenarios. 🤔 In most real use cases, we'll use a well structured prompt using our exiting code and environment as base, if necessity, we'll quickly pseudo-coded what we can want and ask the LLM to build upon it. 🤷‍♂️

Фото профиля richaX
richaX5 месяцев назад

seems gemma4:31b is still better.

Фото профиля BenUsesAI
BenUsesAI5 месяцев назад

qwen3.6 is just qwen3.5 with a new resume and the same bugs, but sure let's benchmark the difference until your gpu melts

Фото профиля AX⚡
AX⚡5 месяцев назад

Qwen 3.6 works through a codebase very similar to opus

Фото профиля 🗻🏔️ ALASKVN DJ 🏔️🗻
🗻🏔️ ALASKVN DJ 🏔️🗻5 месяцев назад

I'll keep both

Фото профиля sabesh 📟
sabesh 📟5 месяцев назад

these evals look beautiful

Фото профиля ar0cket1
ar0cket15 месяцев назад

pretty good, RL scaling works

Фото профиля The Only True Gamer
The Only True Gamer5 месяцев назад

it feels like it understood the prompts *a tad better* but still had some issues. how do frontier models play this?

Фото профиля John Shina
John Shina5 месяцев назад

Damn! This is super amazing. Good job

Фото профиля BombaySaphire
BombaySaphire5 месяцев назад

cool test 😃

Фото профиля Nate MacInnes
Nate MacInnes5 месяцев назад

Time to bump up to 3.6

Фото профиля Momin
Momin5 месяцев назад

incremental 👀

Фото профиля Hendrix.btc ⭕🔶
Hendrix.btc ⭕🔶5 месяцев назад

addictive!!!

Фото профиля kanaley
kanaley5 месяцев назад

Надо попробовать в задачах.

Фото профиля BROMSON
BROMSON5 месяцев назад

@lmstudio pls add support this thing

Фото профиля Mateusz Mirkowski
Mateusz Mirkowski5 месяцев назад

New king has arrived!

Фото профиля Pietro Mastro
Pietro Mastro5 месяцев назад

My God

Фото профиля canberk
canberk5 месяцев назад

wasn’t expecting Qwen 3.6 FP8 to plan this cleanly and smartly across all 4 tests! clear progress on house build and fish swarm, starry night looks solid too, though 3.5 still feels more natural on the tree.

Фото профиля Albatross
Albatross5 месяцев назад

The house one seemed to throw it off a bit, that’s a good one

Фото профиля Linux-Howto.org
Linux-Howto.org5 месяцев назад

3.6 27b runs mega wow in hermes

Фото профиля bjornmuh
bjornmuh5 месяцев назад

Can you run the same promp on the DFlash ablitterated model too Stevibe? Wonder what we lose from the DFlash strategy

Фото профиля stevibe
stevibe5 месяцев назад

Testing DFlash would be interesting, but we need a solid way to test it to really see how it compares.

Фото профиля Samian
Samian5 месяцев назад

qwen3.6 edging out on canvas tests? what's the win on multi-step reasoning tasks

Фото профиля bubba gump
bubba gump5 месяцев назад

my body is ready...

Фото профиля Syed Muzayan Mehmud
Syed Muzayan Mehmud5 месяцев назад

whats the best/economical hardware for running this model and serving inference to say, your laptop?

Фото профиля This is Dmitry Zhomir
This is Dmitry Zhomir5 месяцев назад

Not GPT/Claude yet, but already better than previous version for sure

Фото профиля ogx
ogx5 месяцев назад

Cool

Похожие видео