正在加载视频...
视频加载失败
Completed a first hour side-by-side comparison between Qwen3.5 27b and Qwen3.6 27b on the same 4 canvas coding tests. Running the Qwen3.6 27b FP8, vLLM. What do you think?
40 条评论

why are we still running these retarded tests. like I actually really hate this

qwen still doesnt know what a roof looks like 😂

so close yet so far

Been running the 3.5 27B on my 5070 Ti. The FP8 quant is surprisingly close to the full model for code gen. Did you notice any latency hits on 3.6 with vLLM, or is the TTFT basically the same?

Didn't measure the TTFT yet, will check it later.

Stunningly good. Both of them but clearly better. The 35B has an issue with unverified assumptions, it tends to assume something without looking into it. I wonder how it would perform your test if you adapt the system prompt and tell it to actively monitor for unverified assumptions.

Good take, will try to update the system prompt next time.

do a 3.6 32b vs 27b comparison to assess the marginal gap.

Super interesting. On the house building prompt, it's interesting to see both build something out of brick on top of the house, seemed to confuse them both. The fish demo is my favorite, kinda wild that 3.5 used triangles rather than fish. For 3.6, the fish have a crazy shake they do, but at least they're actually fish! 🐟 Of course, it goes without saying that 3.6 is insanely good, best local AI I think we've ever seen.

Definitely, We can definitely see the progress in local models.

Ohhh, honor to see you in the comments! You’ve been sharing some amazing stuff 🔥

Now do 3.6 MOE vs 3.6 dense

3.6 is a really good model. Full precision is notably higher performing on my benchmarks than the quantized versions.

Personally, I don't like purely abstract test scenarios. For me, it's much more important to test real use case's scenarios. 🤔 In most real use cases, we'll use a well structured prompt using our exiting code and environment as base, if necessity, we'll quickly pseudo-coded what we can want and ask the LLM to build upon it. 🤷♂️

seems gemma4:31b is still better.

qwen3.6 is just qwen3.5 with a new resume and the same bugs, but sure let's benchmark the difference until your gpu melts

Qwen 3.6 works through a codebase very similar to opus

I'll keep both

these evals look beautiful

pretty good, RL scaling works

it feels like it understood the prompts *a tad better* but still had some issues. how do frontier models play this?

Damn! This is super amazing. Good job

cool test 😃

Time to bump up to 3.6

incremental 👀

addictive!!!

Надо попробовать в задачах.

@lmstudio pls add support this thing

New king has arrived!

My God

wasn’t expecting Qwen 3.6 FP8 to plan this cleanly and smartly across all 4 tests! clear progress on house build and fish swarm, starry night looks solid too, though 3.5 still feels more natural on the tree.

The house one seemed to throw it off a bit, that’s a good one

3.6 27b runs mega wow in hermes

Can you run the same promp on the DFlash ablitterated model too Stevibe? Wonder what we lose from the DFlash strategy

Testing DFlash would be interesting, but we need a solid way to test it to really see how it compares.

qwen3.6 edging out on canvas tests? what's the win on multi-step reasoning tasks

my body is ready...

whats the best/economical hardware for running this model and serving inference to say, your laptop?

Not GPT/Claude yet, but already better than previous version for sure

Cool
