Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

For the first time, i'm not even bothered about missing Fable 5.1 Because i now have Qwen Qwen3.8-Flash-Next with me🤯! PS: Single 3090 users, you might not want to skip this one, you're in for a treat ! Gap between frontier closed models and local models is getting really...

12,914 Aufrufe • vor 29 Tagen •via X (Twitter)

20 Kommentare

Profilbild von Andrew I. Christianson
Andrew I. Christiansonvor 29 Tagen

@Alibaba_Qwen yeah and this is flash *next* --setting the groundwork for qwen 4. qwen 4 is going to be epic when it gets here

Profilbild von AJ
AJvor 29 Tagen

@Alibaba_Qwen Yes it is, wonder how next 27b looks like !

Profilbild von AJ
AJvor 29 Tagen

Should i continue with this even more, to see where i can get or should i deploy it as it is and share it so all of you can take a look? Or deploy and continue to work ;)

Profilbild von AJ
AJvor 29 Tagen

Oh i forgot to mention this is with reasoning budget 2048, mid level, not high or xhigh. Pretty good result if you ask me.

Profilbild von Linux-Howto.org
Linux-Howto.orgvor 29 Tagen

@Alibaba_Qwen damn it runs... what is your perfect llama.cpp call / params?

Profilbild von Marcel Holter
Marcel Holtervor 29 Tagen

@Alibaba_Qwen @ItsmeAjayKV Join me in my attempt to normalise to always also post the exact prompts that we use for these kind of efforts!!

Profilbild von AJ
AJvor 29 Tagen

@Alibaba_Qwen I'm going to push it to github, with prompt as well once im done with it. Similar to this repo.

Profilbild von Kvn
Kvnvor 29 Tagen

@Alibaba_Qwen Can only agree with you, fuck anthropic with their ultimate expensive forced downgradable hallucinated models 🤣🤣🤣

Profilbild von Sonic的奇思妙想
Sonic的奇思妙想vor 29 Tagen

@Alibaba_Qwen Locally trained models can achieve such effects, which is great. looking forward to more sharing👍

Profilbild von AJ
AJvor 29 Tagen

@Alibaba_Qwen Thanks, more testing underway.

Profilbild von HAXXVII
HAXXVIIvor 29 Tagen

@Alibaba_Qwen What's the tok/s on Qwen3.8-Flash-Next on a single 3090?

Profilbild von AJ
AJvor 29 Tagen

@Alibaba_Qwen See this post from yesterday.

Profilbild von HAXXVII
HAXXVIIvor 29 Tagen

@Alibaba_Qwen Damn that is definitely usable. What about context rot. Is there any vibes around that with flash-next and 27b ?

Profilbild von Michael Waitze
Michael Waitzevor 29 Tagen

@Alibaba_Qwen Local models are getting ridiculous! I'm curious whether Fable's 75% cheaper cache reads and Terminal-Bench jumps shift how you think about the frontier/local tradeoff. We actually broke this down here:

Profilbild von JQ
JQvor 29 Tagen

@Alibaba_Qwen WAit - Qwen3 flash can run on 3090 ? I call BS! If not -Amazing. Can it run on MLX too?

Profilbild von Michał Kołodziej 🇵🇱
Michał Kołodziej 🇵🇱vor 29 Tagen

@Alibaba_Qwen How does such quant compare to 27b? I’d expect them to provide similar quality.

Profilbild von Apptor
Apptorvor 29 Tagen

@Alibaba_Qwen What harness / coding environment do you use?

Profilbild von AJ
AJvor 29 Tagen

@Alibaba_Qwen DeepSeek Harness It's really good.

Profilbild von Voltage (Fella)
Voltage (Fella)vor 28 Tagen

@Alibaba_Qwen I feel exactly the same way. All I want to do now is squeeze the maximum tokens per second possible on Qwen3.8-Flash-Next. It is a fierce model! 🔥👌

Profilbild von Yordi builds
Yordi buildsvor 29 Tagen

@Alibaba_Qwen Wait a min. Flash on a single 3090? What's the recipe? Will it work on 5090?

Ähnliche Videos

Qwen3.8 Flash Next is starting to look ridiculous on Apple Silicon. I’m running the 4-bit MTP build locally on an M3 Ultra Studio, and the latest OMP run hit: 97.1 tok/s decode 23K context ~1,131 tok/s uncached prompt processing That first number is the one that caught my attention. Nearly 100 tokens per second from a local Qwen3.8 Flash Next setup is already fast enough that the usual “local models are slow” argument starts feeling pretty outdated. And the prompt processing speed is even crazier. Over 1,100 tok/s on an uncached prompt means the model can chew through a large amount of context before generation even starts. The setup matters here. This isn’t just Qwen3.8 Flash Next running untouched. It’s a 4-bit quantized build with MTP, and the inference stack is clearly doing a lot of work behind the scenes to make the hardware perform like this. There’s already a PR open for the implementation on oMLX, so this isn’t just a one-off local experiment either. If these optimizations make their way into the broader MLX ecosystem, running large models locally on Apple Silicon gets even more interesting. The other thing I like about numbers like this is that they put the focus back on the entire inference stack. Model size is one variable. Quantization is another. Then you have MTP, KV cache configuration, runtime optimizations and the hardware itself. Change the recipe and the same model can feel completely different. Qwen3.8 Flash Next at ~97 tok/s on an M3 Ultra is a pretty good demonstration of that.

FHILY👑

15,378 Aufrufe • vor 18 Tagen

hey here is the final result of octopus invaders on nvidia's flagship at full precision. nemotron super 120B on 2x H200 NVL. BF16 unquantized. 287GB of VRAM. hermes agent as the harness. 60 tok/s. first try it autonomously coded for 6 minutes straight. created 11 files. correct project structure. correct load order. started the server. i opened the browser and the result was a blank screen. i did not give up. second try i gave it a precise list of bugs and things to fix. it went back in for another 3 minutes. patched the code. served it again. still blank. so i did what any sane person would do. third try i just said the screen is blank, test it and fix it yourself. and this is where nemotron showed what it actually is. it became a debugger. you can see it in the video. realtime CSS test squares, red screen flashes, hermes agent browser tools, inspecting its own output. it built the parallax background with planets and comets. it rendered a rocket ship that tracks your mouse with fire and bullet physics. the aesthetic is real. but no enemies spawn. no collision. not playable. what surprised me is qwen 27B one shotted this exact game on a single RTX 3090 at Q4 quant. and here is nvidia's flagship at full precision on enterprise hardware needing 3 tries and still not getting there. that makes my hope high for the undisputed qwen 122B which is about to face the same test next. same hardware. same prompt and same harness. lets see if it one shots or not. full session in the video. no cuts. 5x speed.

Sudo su

11,011 Aufrufe • vor 6 Monaten