Loading video...
Video Failed to Load
Qwen3.8 Flash I really like this model and find it great for frontend. It performs much better than Qwen3.8 27B.
61,327 views • 19 days ago •via X (Twitter)
41 Comments

I don’t know why it got kinda buried, I think what people miss about Qwen is that it has this humanistic approach to reasoning, and it’s the reason why original Qwen 3.6 27B blew everyone’s socks off. Qwen 3.8 Flash basically takes the weaker points of the 27B model and removes them, serving at decent speed and concurrency. If you read the reasoning tokens, it’s much more like a person than GLM which is more like a a terminator. It’s probably superb for Ux work because of this reason and make better non-coding agents that do work like marketing. For my agent fleet this is the default model now unless I need GLM 5.3 to go terminator on something. I think both Qwen and GLM doesn’t do hype marketing so their capabilities are greatly under advertised.

You'll soon be able to run it on a single DGX Spark on native nvfp4 👀

@yume_arasaki Oh I’d love to run it to replace Qwen 3.8 27B on my 1 spark….

@MiaAI_lab

what quant are you running, and does it still hold context on longer frontend files

nvfp4, and yes

wait this looks so nice, you should run it through the 100 html file test

This is ONE of the htmls in the 100-html test 👀

release the html files!!!

Soon!

It didn’t get buried just a lot of people don’t have 128 GB VRAM to run flash but a lot of people have enough vram to run 27B, nobody with enough vram running the 27B vs Flash, different class of model tbh

maybe cause it's 10x bigger?

Will there be a Mia-AiLab/Qwen3.8-flash-EXL3 Version ? 😀

In the photography community there was a saying: “the best camera is the one you carry”. So for many people “27B” performs better than “flash”, just because they’re able to run it.

I heard that it's about on par Sometimes worse Sometimes better

in frontend its 100% better

Can't wait for the recipe

Yea I’m trying it out on 2 sparks and liking it pretty good. It might be as good as ds4 flash

which harness is it👀

Bro who even uses these designs? I only see them in demo, these have no real world usage imo

Is it only the UI that is better than 27B? What about coding?

I'd love to see some more examples. I've yet to fully contextualize exactly what sort of real work you can get done with these models, assuming you had effectively unlimited local compute and time.

carousel is where flash usually loses me

Could you give us an example prompt for generating this?

@budgierless

There are way more people who own an RTX 5090 or 4090 than people who own a DGX Spark. And you need two Sparks just to get halfway decent performance running NVFP4 Qwen3.8-Next-Flash. Meanwhile, a single RTX 5090 can make Qwen3.8-27B absolutely fly.

27B looked smarter on paper. Flash actually ships the UI.

Yes,it’s the best model so far I have tested… Q4 on RTX 5090

How it compare with DSV4 Flash vision?

Same here. Fast enough on my Mac to be my new daily default Hermes model

Not sure it holds for complex state logic. Building multi-step form flows, I found similar flash-class models losing their edge over the full 27B. Were you testing mostly component generation?

flash over the 27b is surprising. what did you run them both on?

dgx spark

on a spark, nice. does the gap hold if you drop to something with 16gb?

@ekosproject

How you made this what's The prompt

Im getting a spamming "!!!!!!!!!" And I cant figure out why... I asked my qwen 3.8 27b to use your recipe yesterday, and with all my cli's it was happening... otherwise id use it full time prob. And advice ?

HOLY this is beautiful

Another reason is that DeepSeek-V4-Flash and GLM-5.3-Flash happen to occupy the exact same niche as Qwen3.8-Next-Flash on a dual-Spark setup. Qwen is far behind DeepSeek when it comes to creative writing, and it can’t match GLM on coding either. That leaves it in a pretty awkward position.

Flash models punching above their weight for UI is wild

best cloud provider of this model? alibaba cloud?
