Video wird geladen...
Video konnte nicht geladen werden
Loading DeepSeek v4 Flash just reminded me how good and fast this model is and how well it runs on DGX Sparks. GLM-5.2 is really awesome, but if you want fast and concurrent inference, DS4F is still a top choice IMO. Created by DS4F in single shot The coffee guy
40,794 Aufrufe • vor 1 Monat •via X (Twitter)
53 Kommentare

Prompt: Design and create a very creative, elaborate, and detailed voxel art scene of a pagoda in a beautiful garden with trees, including some cherry blossoms. Make the scene impressive and varied and use colorful voxels. Use whatever libraries to get this done but make sure I can paste it all into a single HTML file.

GLM-5.2 results

@thatcofffeeguy Specially good on running multi agents with long context due to tiny kv cache it needs.

@thatcofffeeguy 💯

Gonna run this on my GLM-5.2. Will post results.

@thatcofffeeguy Huge love for GLM-5.2 , but when only have a single DGX to work with, DeepSeek v4 Flash is the clear choice.

@thatcofffeeguy for concurrent runs, DS4F still gets my vote

@thatcofffeeguy Concurrencies and speed

@thatcofffeeguy I prefer to run DS V4 on my Sparks and use it for 95% of my tasks, then use GLM/GPT/K3 with an inference provider when it get stuck, than to run GLM with a slow speed. It's a very good model and rarely can't get a task done.

@thatcofffeeguy I've optimized the ds4 flash a bit more, and now I have a stable 400 tps for prefill and about 40 tps for decode on a 160k context on a 4x3090 I'll tell more about it soon

@thatcofffeeguy Complete Voxel Render Engine build in Go. From scratch with DS V4 Flash. :)

@TechMDAI @thatcofffeeguy What’s the best way to take advantage of the concurrency as a single user running DS4F on two sparks? Have your harness kick off subagents? I’m using pi & Hermes as harnesses mostly for local AI.

@thatcofffeeguy That’s actually cute.

@TechMDAI @thatcofffeeguy Agreed, it is working excellent in Hermes.

@thatcofffeeguy And if you want less hallucination go with Mimo 2.5. These two are the most used for ke now. DS4VF and Mimo v2.5

@thatcofffeeguy Tbh I don't have any serious hallucinations with DS4F

@thatcofffeeguy Soketimes I have. Probably the use case is different. (3dsmax sdk and other low level macro heavy stuffs) For that it's not always that reliable unfortunately. Mimo handle those cases better I think.

@thatcofffeeguy Lovely tower

@thatcofffeeguy Was this run with any harness, or just one-shot directly to the model?

@thatcofffeeguy I used Pi

@thatcofffeeguy I'd like to try to recreate the test are you using opencode?

@thatcofffeeguy DS4 is still my regular local model along with qwen.

@thatcofffeeguy DSV4 flash is goated

@thatcofffeeguy Yeah it really is

@thatcofffeeguy QWEN 3.6 35B A3B

@thatcofffeeguy looking good!

@thatcofffeeguy Even using APIs ds4f is still the cheapest and most cost efficient model for me.

@thatcofffeeguy DS4F through API is nerfed and unusable for me. My local DS4F beats it by a big margin.

@thatcofffeeguy Oh i'm sure of that, I can feel the quantized context 😭

@thatcofffeeguy Jup- Ich liebe mein DSV 4 Flash. Auf zwei Sparks für bisher das beste Arbeitstier.

@thatcofffeeguy I would have been super impressed with this in 1996

@thatcofffeeguy What do you think is better to run ds4 and mimo across 4 sparks or glm 5.2 on 3 and qwen on one

@thatcofffeeguy Personally I'd use the 4th spark for Qwen3.6 35b for agents when I care about speed.

@thatcofffeeguy That’s exactly what I’m running you know surprisingly qwen 35b has been one of the best performers from all of these new models that came out recently with all their claims , qwen never broke it just keeps doing its thing !

@thatcofffeeguy Great setup 😀

@thatcofffeeguy Yes until I get that switch then I’ll run glm 5.2 across all four and use my rtx 6000 for a super fast model which still only qwen comes to mind

@thatcofffeeguy are you using tonyd2wild's 2 spark recipe ? or one you made yourself .. or ... something else?

@thatcofffeeguy The new DS model seems to have some improvements

@thatcofffeeguy I love it. A little more speed on it and an update for the 2 Spark crowd would be awesome. Also what’s the new DeepSeek drops add to it?

@thatcofffeeguy What is the hw reqs when running on older gpus like the 30x + ram? That confuses me

@thatcofffeeguy Nice. On a single DGX?

@thatcofffeeguy I run DS4F on 2 units.

@thatcofffeeguy OK, it sounds like I need to buy one more DGX. Thy are just incredibly expensive here in Aus.

@thatcofffeeguy 저랑 똑같은 느낌을 받으셨군요 저도 이것저것 테스트 하면서 느낀게 DeepSeek v4 Flash 가 정말 좋다는 느낌을 많이 받았습니다.

@thatcofffeeguy Can you run it on 2 GB10? What TPS do you get?

@thatcofffeeguy try this new quant, check the benchmarks, it might be the first variant of qwen3.6:27b thats actually better than the original.

@thatcofffeeguy GLM 5.2 for planning and Deepseek v4 pro for implementation.

@thatcofffeeguy DS4F is so reliable and consistent for agentic work loads. Checking daily for the GA release.

@thatcofffeeguy What do you think about the Laguna s 2.1 ? have you tested it ?

@thatcofffeeguy Yes. They told me to hold on testing because they are fixing stuff, and they'll ping me back, but never got the ping...

Oh you are in touch with them. I have been using it with @mr_r0b0t config, but thinking off. It has been really great option for a single spark for agentic coding. I hope @poolsideai will achieve since if this model is this good without thinking, then it must be much better if they get the thinking fixed. That is what caused the infinite tool looping and other failures.

@MiaAI_lab @thatcofffeeguy @mr_r0b0t @poolsideai I tested Laguna S 2.1 nvfp4 mlx extensively, without thinking it broke down in the benchmarks compared to any other model. You have to force it to think, then it is fine. Great in structuring but depth of analysis is not en par with qwen3.6 35b

@MiaAI_lab @thatcofffeeguy @mr_r0b0t @poolsideai Which benchmark did you run?
