正在加载视频...

视频加载失败

For the first time, i'm not even bothered about missing Fable 5.1 Because i now have Qwen Qwen3.8-Flash-Next with me🤯! PS: Single 3090 users, you might not want to skip this one, you're in for a treat ! Gap between frontier closed models and local models is getting really...

12,914 次观看 • 29 天前 •via X (Twitter)

20 条评论

Andrew I. Christianson 的头像
Andrew I. Christianson29 天前

@Alibaba_Qwen yeah and this is flash *next* --setting the groundwork for qwen 4. qwen 4 is going to be epic when it gets here

AJ 的头像
AJ29 天前

@Alibaba_Qwen Yes it is, wonder how next 27b looks like !

AJ 的头像
AJ29 天前

Should i continue with this even more, to see where i can get or should i deploy it as it is and share it so all of you can take a look? Or deploy and continue to work ;)

AJ 的头像
AJ29 天前

Oh i forgot to mention this is with reasoning budget 2048, mid level, not high or xhigh. Pretty good result if you ask me.

Linux-Howto.org 的头像
Linux-Howto.org29 天前

@Alibaba_Qwen damn it runs... what is your perfect llama.cpp call / params?

Marcel Holter 的头像
Marcel Holter29 天前

@Alibaba_Qwen @ItsmeAjayKV Join me in my attempt to normalise to always also post the exact prompts that we use for these kind of efforts!!

AJ 的头像
AJ28 天前

@Alibaba_Qwen I'm going to push it to github, with prompt as well once im done with it. Similar to this repo.

Kvn 的头像
Kvn29 天前

@Alibaba_Qwen Can only agree with you, fuck anthropic with their ultimate expensive forced downgradable hallucinated models 🤣🤣🤣

Sonic的奇思妙想 的头像
Sonic的奇思妙想29 天前

@Alibaba_Qwen Locally trained models can achieve such effects, which is great. looking forward to more sharing👍

AJ 的头像
AJ29 天前

@Alibaba_Qwen Thanks, more testing underway.

HAXXVII 的头像
HAXXVII29 天前

@Alibaba_Qwen What's the tok/s on Qwen3.8-Flash-Next on a single 3090?

AJ 的头像
AJ29 天前

@Alibaba_Qwen See this post from yesterday.

HAXXVII 的头像
HAXXVII29 天前

@Alibaba_Qwen Damn that is definitely usable. What about context rot. Is there any vibes around that with flash-next and 27b ?

Michael Waitze 的头像
Michael Waitze29 天前

@Alibaba_Qwen Local models are getting ridiculous! I'm curious whether Fable's 75% cheaper cache reads and Terminal-Bench jumps shift how you think about the frontier/local tradeoff. We actually broke this down here:

JQ 的头像
JQ29 天前

@Alibaba_Qwen WAit - Qwen3 flash can run on 3090 ? I call BS! If not -Amazing. Can it run on MLX too?

Michał Kołodziej 🇵🇱 的头像
Michał Kołodziej 🇵🇱29 天前

@Alibaba_Qwen How does such quant compare to 27b? I’d expect them to provide similar quality.

Apptor 的头像
Apptor29 天前

@Alibaba_Qwen What harness / coding environment do you use?

AJ 的头像
AJ29 天前

@Alibaba_Qwen DeepSeek Harness It's really good.

Voltage (Fella) 的头像
Voltage (Fella)28 天前

@Alibaba_Qwen I feel exactly the same way. All I want to do now is squeeze the maximum tokens per second possible on Qwen3.8-Flash-Next. It is a fierce model! 🔥👌

Yordi builds 的头像
Yordi builds29 天前

@Alibaba_Qwen Wait a min. Flash on a single 3090? What's the recipe? Will it work on 5090?

相关视频

Qwen3.8 Flash Next is starting to look ridiculous on Apple Silicon. I’m running the 4-bit MTP build locally on an M3 Ultra Studio, and the latest OMP run hit: 97.1 tok/s decode 23K context ~1,131 tok/s uncached prompt processing That first number is the one that caught my attention. Nearly 100 tokens per second from a local Qwen3.8 Flash Next setup is already fast enough that the usual “local models are slow” argument starts feeling pretty outdated. And the prompt processing speed is even crazier. Over 1,100 tok/s on an uncached prompt means the model can chew through a large amount of context before generation even starts. The setup matters here. This isn’t just Qwen3.8 Flash Next running untouched. It’s a 4-bit quantized build with MTP, and the inference stack is clearly doing a lot of work behind the scenes to make the hardware perform like this. There’s already a PR open for the implementation on oMLX, so this isn’t just a one-off local experiment either. If these optimizations make their way into the broader MLX ecosystem, running large models locally on Apple Silicon gets even more interesting. The other thing I like about numbers like this is that they put the focus back on the entire inference stack. Model size is one variable. Quantization is another. Then you have MTP, KV cache configuration, runtime optimizations and the hardware itself. Change the recipe and the same model can feel completely different. Qwen3.8 Flash Next at ~97 tok/s on an M3 Ultra is a pretty good demonstration of that.

FHILY👑

15,378 次观看 • 18 天前

hey here is the final result of octopus invaders on nvidia's flagship at full precision. nemotron super 120B on 2x H200 NVL. BF16 unquantized. 287GB of VRAM. hermes agent as the harness. 60 tok/s. first try it autonomously coded for 6 minutes straight. created 11 files. correct project structure. correct load order. started the server. i opened the browser and the result was a blank screen. i did not give up. second try i gave it a precise list of bugs and things to fix. it went back in for another 3 minutes. patched the code. served it again. still blank. so i did what any sane person would do. third try i just said the screen is blank, test it and fix it yourself. and this is where nemotron showed what it actually is. it became a debugger. you can see it in the video. realtime CSS test squares, red screen flashes, hermes agent browser tools, inspecting its own output. it built the parallax background with planets and comets. it rendered a rocket ship that tracks your mouse with fire and bullet physics. the aesthetic is real. but no enemies spawn. no collision. not playable. what surprised me is qwen 27B one shotted this exact game on a single RTX 3090 at Q4 quant. and here is nvidia's flagship at full precision on enterprise hardware needing 3 tries and still not getting there. that makes my hope high for the undisputed qwen 122B which is about to face the same test next. same hardware. same prompt and same harness. lets see if it one shots or not. full session in the video. no cuts. 5x speed.

Sudo su

11,011 次观看 • 6 个月前