Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Loading DeepSeek v4 Flash just reminded me how good and fast this model is and how well it runs on DGX Sparks. GLM-5.2 is really awesome, but if you want fast and concurrent inference, DS4F is still a top choice IMO. Created by DS4F in single shot The coffee guy

40,794 Aufrufe • vor 1 Monat •via X (Twitter)

53 Kommentare

Profilbild von Mia
Miavor 1 Monat

Prompt: Design and create a very creative, elaborate, and detailed voxel art scene of a pagoda in a beautiful garden with trees, including some cherry blossoms. Make the scene impressive and varied and use colorful voxels. Use whatever libraries to get this done but make sure I can paste it all into a single HTML file.

Profilbild von Mia
Miavor 1 Monat

GLM-5.2 results

Profilbild von Sam // SonivaLabs
Sam // SonivaLabsvor 1 Monat

@thatcofffeeguy Specially good on running multi agents with long context due to tiny kv cache it needs.

Profilbild von Mia
Miavor 1 Monat

@thatcofffeeguy 💯

Profilbild von Mia
Miavor 1 Monat

Gonna run this on my GLM-5.2. Will post results.

Profilbild von G36_maid
G36_maidvor 1 Monat

@thatcofffeeguy Huge love for GLM-5.2 , but when only have a single DGX to work with, DeepSeek v4 Flash is the clear choice.

Profilbild von Miles S.
Miles S.vor 1 Monat

@thatcofffeeguy for concurrent runs, DS4F still gets my vote

Profilbild von Mia
Miavor 1 Monat

@thatcofffeeguy Concurrencies and speed

Profilbild von Simon Bissonnette 🇨🇦
Simon Bissonnette 🇨🇦vor 1 Monat

@thatcofffeeguy I prefer to run DS V4 on my Sparks and use it for 95% of my tasks, then use GLM/GPT/K3 with an inference provider when it get stuck, than to run GLM with a slow speed. It's a very good model and rarely can't get a task done.

Profilbild von Alexey Fateev
Alexey Fateevvor 1 Monat

@thatcofffeeguy I've optimized the ds4 flash a bit more, and now I have a stable 400 tps for prefill and about 40 tps for decode on a 160k context on a 4x3090 I'll tell more about it soon

Profilbild von Blue - Klarname (Olaf Merz)
Blue - Klarname (Olaf Merz)vor 1 Monat

@thatcofffeeguy Complete Voxel Render Engine build in Go. From scratch with DS V4 Flash. :)

Profilbild von Iman Mostafavi
Iman Mostafavivor 1 Monat

@TechMDAI @thatcofffeeguy What’s the best way to take advantage of the concurrency as a single user running DS4F on two sparks? Have your harness kick off subagents? I’m using pi & Hermes as harnesses mostly for local AI.

Profilbild von Dragos Roua
Dragos Rouavor 1 Monat

@thatcofffeeguy That’s actually cute.

Profilbild von Myrmidon Achilles
Myrmidon Achillesvor 1 Monat

@TechMDAI @thatcofffeeguy Agreed, it is working excellent in Hermes.

Profilbild von rapidTools / Mike
rapidTools / Mikevor 1 Monat

@thatcofffeeguy And if you want less hallucination go with Mimo 2.5. These two are the most used for ke now. DS4VF and Mimo v2.5

Profilbild von Mia
Miavor 1 Monat

@thatcofffeeguy Tbh I don't have any serious hallucinations with DS4F

Profilbild von rapidTools / Mike
rapidTools / Mikevor 1 Monat

@thatcofffeeguy Soketimes I have. Probably the use case is different. (3dsmax sdk and other low level macro heavy stuffs) For that it's not always that reliable unfortunately. Mimo handle those cases better I think.

Profilbild von Vincent Yiu
Vincent Yiuvor 1 Monat

@thatcofffeeguy Lovely tower

Profilbild von luke
lukevor 1 Monat

@thatcofffeeguy Was this run with any harness, or just one-shot directly to the model?

Profilbild von Mia
Miavor 1 Monat

@thatcofffeeguy I used Pi

Profilbild von Yume_X
Yume_Xvor 1 Monat

@thatcofffeeguy I'd like to try to recreate the test are you using opencode?

Profilbild von Nabeel
Nabeelvor 1 Monat

@thatcofffeeguy DS4 is still my regular local model along with qwen.

Profilbild von TechMD
TechMDvor 1 Monat

@thatcofffeeguy DSV4 flash is goated

Profilbild von Mia
Miavor 1 Monat

@thatcofffeeguy Yeah it really is

Profilbild von HiddenFella
HiddenFellavor 1 Monat

@thatcofffeeguy QWEN 3.6 35B A3B

Profilbild von Mia
Miavor 1 Monat

@thatcofffeeguy looking good!

Profilbild von Mieszkaniec Dzikolandu.
Mieszkaniec Dzikolandu.vor 1 Monat

@thatcofffeeguy Even using APIs ds4f is still the cheapest and most cost efficient model for me.

Profilbild von Mia
Miavor 1 Monat

@thatcofffeeguy DS4F through API is nerfed and unusable for me. My local DS4F beats it by a big margin.

Profilbild von Mieszkaniec Dzikolandu.
Mieszkaniec Dzikolandu.vor 1 Monat

@thatcofffeeguy Oh i'm sure of that, I can feel the quantized context 😭

Profilbild von Blue - Klarname (Olaf Merz)
Blue - Klarname (Olaf Merz)vor 1 Monat

@thatcofffeeguy Jup- Ich liebe mein DSV 4 Flash. Auf zwei Sparks für bisher das beste Arbeitstier.

Profilbild von Marshall
Marshallvor 1 Monat

@thatcofffeeguy I would have been super impressed with this in 1996

Profilbild von Pacboy22
Pacboy22vor 1 Monat

@thatcofffeeguy What do you think is better to run ds4 and mimo across 4 sparks or glm 5.2 on 3 and qwen on one

Profilbild von Mia
Miavor 1 Monat

@thatcofffeeguy Personally I'd use the 4th spark for Qwen3.6 35b for agents when I care about speed.

Profilbild von Pacboy22
Pacboy22vor 1 Monat

@thatcofffeeguy That’s exactly what I’m running you know surprisingly qwen 35b has been one of the best performers from all of these new models that came out recently with all their claims , qwen never broke it just keeps doing its thing !

Profilbild von Mia
Miavor 1 Monat

@thatcofffeeguy Great setup 😀

Profilbild von Pacboy22
Pacboy22vor 1 Monat

@thatcofffeeguy Yes until I get that switch then I’ll run glm 5.2 across all four and use my rtx 6000 for a super fast model which still only qwen comes to mind

Profilbild von Joshua Hudson
Joshua Hudsonvor 1 Monat

@thatcofffeeguy are you using tonyd2wild's 2 spark recipe ? or one you made yourself .. or ... something else?

Profilbild von Rivers
Riversvor 1 Monat

@thatcofffeeguy The new DS model seems to have some improvements

Profilbild von Brian R. O'Leary
Brian R. O'Learyvor 1 Monat

@thatcofffeeguy I love it. A little more speed on it and an update for the 2 Spark crowd would be awesome. Also what’s the new DeepSeek drops add to it?

Profilbild von ࢡď 屮♱
ࢡď 屮♱vor 1 Monat

@thatcofffeeguy What is the hw reqs when running on older gpus like the 30x + ram? That confuses me

Profilbild von Anders
Andersvor 1 Monat

@thatcofffeeguy Nice. On a single DGX?

Profilbild von Mia
Miavor 1 Monat

@thatcofffeeguy I run DS4F on 2 units.

Profilbild von Anders
Andersvor 1 Monat

@thatcofffeeguy OK, it sounds like I need to buy one more DGX. Thy are just incredibly expensive here in Aus.

Profilbild von KimSunIl
KimSunIlvor 1 Monat

@thatcofffeeguy 저랑 똑같은 느낌을 받으셨군요 저도 이것저것 테스트 하면서 느낀게 DeepSeek v4 Flash 가 정말 좋다는 느낌을 많이 받았습니다.

Profilbild von Peter Morris
Peter Morrisvor 1 Monat

@thatcofffeeguy Can you run it on 2 GB10? What TPS do you get?

Profilbild von MEs
MEsvor 1 Monat

@thatcofffeeguy try this new quant, check the benchmarks, it might be the first variant of qwen3.6:27b thats actually better than the original.

Profilbild von Clement Lumumba
Clement Lumumbavor 1 Monat

@thatcofffeeguy GLM 5.2 for planning and Deepseek v4 pro for implementation.

Profilbild von Rick deckard
Rick deckardvor 1 Monat

@thatcofffeeguy DS4F is so reliable and consistent for agentic work loads. Checking daily for the GA release.

Profilbild von Mike Hutu
Mike Hutuvor 1 Monat

@thatcofffeeguy What do you think about the Laguna s 2.1 ? have you tested it ?

Profilbild von Mia
Miavor 1 Monat

@thatcofffeeguy Yes. They told me to hold on testing because they are fixing stuff, and they'll ping me back, but never got the ping...

Profilbild von Mike Hutu
Mike Hutuvor 1 Monat

Oh you are in touch with them. I have been using it with @mr_r0b0t config, but thinking off. It has been really great option for a single spark for agentic coding. I hope @poolsideai will achieve since if this model is this good without thinking, then it must be much better if they get the thinking fixed. That is what caused the infinite tool looping and other failures.

Profilbild von Michael Master
Michael Mastervor 1 Monat

@MiaAI_lab @thatcofffeeguy @mr_r0b0t @poolsideai I tested Laguna S 2.1 nvfp4 mlx extensively, without thinking it broke down in the benchmarks compared to any other model. You have to force it to think, then it is fine. Great in structuring but depth of analysis is not en par with qwen3.6 35b

Profilbild von Mike Hutu
Mike Hutuvor 1 Monat

@MiaAI_lab @thatcofffeeguy @mr_r0b0t @poolsideai Which benchmark did you run?

Ähnliche Videos

I have been testing DeepSeek-V4-Pro with the Pi coding agent. I am mindblown by how well it works out of the box. A few notes: I spent a few hours building an LLM wiki with an agent powered entirely by DeepSeek-V4-Pro on Fireworks inference. This is the first time I feel like there is an open-weight model that can reason at the level of Claude and Codex. And it does this in a cost-effective way with support for 1M context length. To be clear, I am using DeepSeek-V4-Pro inside of Pi without any special configuration. It works out of the box. It's exciting that there is a model that can just be plugged into a basic harness like Pi, and it just works. I've never seen that before. Most models require lots of configuration and setup. DeepSeek's DeepSeek-V4-Pro is clearly good at agentic coding (probably the best from the open-weight models), but the model is also great on knowledge-intensive tasks where reasoning matters. The agent pulled agentic engineering best practices from different company docs (Anthropic, OpenAI, Google, Stripe, Meta, Modal, DeepSeek, Mistral, Cohere), searched and digested Reddit and HN threads, summarized arxiv papers, and surfaced trending GitHub repos. Then it distilled everything into actionable tips across categories. I love the Wiki it built. The quality is really good. Here is a snapshot of what the wiki looks like: DeepSeek-V4-Pro handled the task without breaking stride. Multi-step research queries, code generation for scaffolding, context-heavy reasoning across disparate sources. For coding specifically, this is the first open-weight model that genuinely feels like a Codex or Claude Code experience. It compares in capability and actual multi-turn agentic work. What made the loop feel so responsive was Fireworks' inference speed (the fastest in the market) and the fact that they actually validate models at the systems level before shipping. No corrupted reasoning traces. Just fast, reliable iteration. The hybrid CSA and HCA attention design cuts KV cache to just 10% and inference FLOPs by nearly 4x at 1M-token context. This is what makes the agent loop actually fast and cheap enough to run in practice. For devs who've been watching open-weight models close the gap but haven't found one that actually delivers in practice, this is the closest I've seen. Try it here:

elvis

60,281 Aufrufe • vor 4 Monaten

glm 5.3 vs qwen 3.8 vs gemini 3.7 vs deepseek v4 flash four models designed and built three structures each on a physics-backed site, with no dimensions anywhere in the brief the setup: our own agent loop on OpenRouter, a construction site as the tool set – footings, walls, arches, roofs, scaffold, a lamp. the site enforces physics and nothing else: unsupported brick falls, a roof needs walls under it, a worker reaches 3.2 m above whatever he stands on, an arch needs centring until the keystone is set, concrete cures before it carries. no budget ceiling – material cost is tallied and reported, never blocked. tasks: 1. house – a plot and a palette, no plan. shape, height and material are the model's call 2. lighthouse – a headland cut by a gully, with a rock stack standing 30 m offshore. the lamp must burn, it must be the highest thing built, and the keeper must be able to walk to it 3. bridge – a river with one islet and banks at different heights. cross it however you want models: Z.ai glm 5.3 flash, Qwen qwen 3.8 flash, Google DeepMind gemini 3.7 flash, DeepSeek v4 flash vision all twelve objects were finished and signed off by the models themselves. tallest lighthouse is qwen's at 38.4 m, planted on the offshore stack with a bridge run out to it – the only model that read the site that way. deepseek signed off its bridge on an empty riverbed: 0 bricks, 107 minutes, $1.16m of material tallied - total cost, three builds #1 glm 5.3 flash – $0.201 #2 gemini 3.7 flash – $0.871 #3 qwen 3.8 flash – $1.058 #4 deepseek v4 flash – $1.567 - wall clock, three builds #1 gemini 3.7 flash – 91m #2 glm 5.3 flash – 228m #3 deepseek v4 flash – 502m #4 qwen 3.8 flash – 912m - total tokens #1 gemini 3.7 flash – 3,567,052 #2 glm 5.3 flash – 4,732,748 #3 qwen 3.8 flash – 13,469,333 #4 deepseek v4 flash – 18,230,076 - defects logged by the site #1 deepseek v4 flash – 59 #2 gemini 3.7 flash – 132 #3 glm 5.3 flash – 221 #4 qwen 3.8 flash – 350 - material tallied across three builds #1 gemini 3.7 flash – $359,884 #2 glm 5.3 flash – $583,358 #3 deepseek v4 flash – $1,327,484 #4 qwen 3.8 flash – $2,188,625 observations: • glm is the cheap one and nothing here is close – $0.201 for three buildings, $0.042 per million tokens, 6x under gemini's rate • what glm spends it on is bulk, not care: 166,228 bricks in one house and 156 defect weight, the worst single object in the set • gemini is the efficiency line – 91 minutes and 3.57m tokens for all three and an eighth of qwen's clock • gemini also builds the smallest of everything. its lighthouse is 22.5 m against qwen's 38.4, its house 6.9 m against 19.3 • qwen is the maximalist: 1.18m bricks, $2.19m of material, tallest on all three tasks, and 912 minutes – 15 hours – to get there conclusion: twelve finished objects for $3.80 all in, and a 7.8x price spread between the cheapest model and the priciest! follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

26,360 Aufrufe • vor 24 Tagen