正在加载视频...

视频加载失败

Completed a first hour side-by-side comparison between Qwen3.5 27b and Qwen3.6 27b on the same 4 canvas coding tests. Running the Qwen3.6 27b FP8, vLLM. What do you think?

114,366 次观看 • 5 个月前 •via X (Twitter)

40 条评论

Schinsly ✝️ 的头像
Schinsly ✝️5 个月前

why are we still running these retarded tests. like I actually really hate this

rijndael 的头像
rijndael5 个月前

qwen still doesnt know what a roof looks like 😂

Oxytocin 的头像
Oxytocin5 个月前

so close yet so far

Sakura Yuki 的头像
Sakura Yuki5 个月前

Been running the 3.5 27B on my 5070 Ti. The FP8 quant is surprisingly close to the full model for code gen. Did you notice any latency hits on 3.6 with vLLM, or is the TTFT basically the same?

stevibe 的头像
stevibe5 个月前

Didn't measure the TTFT yet, will check it later.

Flow 的头像
Flow5 个月前

Stunningly good. Both of them but clearly better. The 35B has an issue with unverified assumptions, it tends to assume something without looking into it. I wonder how it would perform your test if you adapt the system prompt and tell it to actively monitor for unverified assumptions.

stevibe 的头像
stevibe5 个月前

Good take, will try to update the system prompt next time.

Khalid 的头像
Khalid5 个月前

do a 3.6 32b vs 27b comparison to assess the marginal gap.

Morgan 的头像
Morgan5 个月前

Super interesting. On the house building prompt, it's interesting to see both build something out of brick on top of the house, seemed to confuse them both. The fish demo is my favorite, kinda wild that 3.5 used triangles rather than fish. For 3.6, the fish have a crazy shake they do, but at least they're actually fish! 🐟 Of course, it goes without saying that 3.6 is insanely good, best local AI I think we've ever seen.

stevibe 的头像
stevibe5 个月前

Definitely, We can definitely see the progress in local models.

Morgan 的头像
Morgan5 个月前

Ohhh, honor to see you in the comments! You’ve been sharing some amazing stuff 🔥

Jonathan Leaders 的头像
Jonathan Leaders5 个月前

Now do 3.6 MOE vs 3.6 dense

Seth Pratt 的头像
Seth Pratt5 个月前

3.6 is a really good model. Full precision is notably higher performing on my benchmarks than the quantized versions.

Edzward 的头像
Edzward5 个月前

Personally, I don't like purely abstract test scenarios. For me, it's much more important to test real use case's scenarios. 🤔 In most real use cases, we'll use a well structured prompt using our exiting code and environment as base, if necessity, we'll quickly pseudo-coded what we can want and ask the LLM to build upon it. 🤷‍♂️

richaX 的头像
richaX5 个月前

seems gemma4:31b is still better.

BenUsesAI 的头像
BenUsesAI5 个月前

qwen3.6 is just qwen3.5 with a new resume and the same bugs, but sure let's benchmark the difference until your gpu melts

AX⚡ 的头像
AX⚡5 个月前

Qwen 3.6 works through a codebase very similar to opus

🗻🏔️ ALASKVN DJ 🏔️🗻 的头像
🗻🏔️ ALASKVN DJ 🏔️🗻5 个月前

I'll keep both

sabesh 📟 的头像
sabesh 📟5 个月前

these evals look beautiful

ar0cket1 的头像
ar0cket15 个月前

pretty good, RL scaling works

The Only True Gamer 的头像
The Only True Gamer5 个月前

it feels like it understood the prompts *a tad better* but still had some issues. how do frontier models play this?

John Shina 的头像
John Shina5 个月前

Damn! This is super amazing. Good job

BombaySaphire 的头像
BombaySaphire5 个月前

cool test 😃

Nate MacInnes 的头像
Nate MacInnes5 个月前

Time to bump up to 3.6

Momin 的头像
Momin5 个月前

incremental 👀

Hendrix.btc ⭕🔶 的头像
Hendrix.btc ⭕🔶5 个月前

addictive!!!

kanaley 的头像
kanaley5 个月前

Надо попробовать в задачах.

BROMSON 的头像
BROMSON5 个月前

@lmstudio pls add support this thing

Mateusz Mirkowski 的头像
Mateusz Mirkowski5 个月前

New king has arrived!

Pietro Mastro 的头像
Pietro Mastro5 个月前

My God

canberk 的头像
canberk5 个月前

wasn’t expecting Qwen 3.6 FP8 to plan this cleanly and smartly across all 4 tests! clear progress on house build and fish swarm, starry night looks solid too, though 3.5 still feels more natural on the tree.

Albatross 的头像
Albatross5 个月前

The house one seemed to throw it off a bit, that’s a good one

Linux-Howto.org 的头像
Linux-Howto.org5 个月前

3.6 27b runs mega wow in hermes

bjornmuh 的头像
bjornmuh5 个月前

Can you run the same promp on the DFlash ablitterated model too Stevibe? Wonder what we lose from the DFlash strategy

stevibe 的头像
stevibe5 个月前

Testing DFlash would be interesting, but we need a solid way to test it to really see how it compares.

Samian 的头像
Samian5 个月前

qwen3.6 edging out on canvas tests? what's the win on multi-step reasoning tasks

bubba gump 的头像
bubba gump5 个月前

my body is ready...

Syed Muzayan Mehmud 的头像
Syed Muzayan Mehmud5 个月前

whats the best/economical hardware for running this model and serving inference to say, your laptop?

This is Dmitry Zhomir 的头像
This is Dmitry Zhomir5 个月前

Not GPT/Claude yet, but already better than previous version for sure

ogx 的头像
ogx5 个月前

Cool

相关视频