
Ivan Fioravanti ᯅ
@ivanfioravanti • 41,873 subscribers
GenAI/LLM addicted, Apple MLX, Cloud computing, Kubernetes, Technology Advisor, Investor and Co-Founder & Board Member of CoreView.
Shorts
Videos

A video showing impact of max frequency change on my 2 x DGX Spark cluster with DeepSeek V4 Flash 0731 nvfp4 running like crazy at c4 using Mia recipe! Power consumption at wall is surely higher than the one reported by nvtop, but it gives you an idea: - 2455 MHz ~ 47W ~72º C - 2300 MHz ~34W ~67º C - 2200 MHz ~32W ~66º C - 2000 MHz ~27W ~63º C - 1800 MHz ~23W ~61º C This has minimal impact on decode speed that is memory bound, but there is impact on prefill and other compute bound activities. set it with: sudo nvidia-smi -lgc 0,2200 reset it with: sudo nvidia-smi -rgc I need to thank Paul to let me discover this magical trick!
Ivan Fioravanti ᯅ76,968 görüntüleme • 18 gün önce

GLM-4.7-8bit (350GB) running at 19 toks/s on two M3 Ultra 512GB using Tensor Parallelism with EXO - MLX, versus 14 toks/s with single node. 🚀 Now context benchmarking & then OpenCode tests 🔥 Note: this is from sources, I had to change things to run it.
Ivan Fioravanti ᯅ327,687 görüntüleme • 8 ay önce

So what's your feeling on DeepSeek V4 Flash 0731? Surely great model for Local AI and coding but clearly not at GLM 5.2 or Kimi K3 level. Obviously, because it's a different category. For Frogger I tried 3 times using omp, but I always get a game too complex to play or with small issues here and there. I've used the mxfp4 version. But it worked great with web search, html and CSS issues with custom plugins on TMRNL X with Hermes Agent.
Ivan Fioravanti ᯅ46,353 görüntüleme • 1 ay önce

MLX GLM 5.2 Distributed on two M3 Ultra 512GB 🔥 One M3 Ultra: 18.8 tokens/sec Two M3 Ultra: 23.4 tokens/sec Context: - PR by Pedro Cuenca is still open and probably there is room for improvement: - basic generation test to measure decoding performance here, I will do a full context benchmarking once PR is more mature - nvfp4 quantization used - Video alternates standard speed and x20, with one Mac first and distributed later. Enjoy! 🙌🏻
Ivan Fioravanti ᯅ87,805 görüntüleme • 2 ay önce

I respectfully disagree for several reasons. Calling a customer, whether free or paying, an idiot is simply wrong. OpenCode, like any other coding agent, clearly tries to preserve the prompt cache as much as possible. Otherwise, it would be painfully slow. The “Stop Using OpenCode” article, which I believe Dax is referring to, focused on specific cases that can invalidate the cached prefix. These issues become much more visible and painful when using local models. Modifying AGENTS.md is one example, as shown in this video. During this coding session, cache efficiency is extremely high. But the moment I modify AGENTS.md, boom: 62K tokens need to be processed again as new input. The date change is another valid point raised in the article. I understand that both issues may sound trivial, and personally I've never been impacted by them, but their impact on the OpenCode end-user experience can be significant when using local models. I agree that both the feedback and the article were too direct and were probably written by the author out of frustration. However, they also contained constructive points that could help improve the product. I use OpenCode alongside several other tools, and I like it overall. But I really dislike this kind of reaction to user feedback, regardless of how that feedback is expressed. I would have preferred a response focused more on listening, understanding the problem, and improving the product.
Ivan Fioravanti ᯅ37,594 görüntüleme • 1 ay önce

Where there's a will, there's a way 😎 DwarfStar DeepSeek V4 Flash 0731 mxfp4 on M3 Ultra single request no MPT/DSpark ~41 toks/s decode!!! 🚀 40 tps barrier broken! Same logits, 100% same result: this is The Mandatory Rule for any agents improving kernels that I think everyone should follow. Speed without mathematical precision is useless. Thanks 20% Opus 5 and 80% Kimi.ai K3 (what a model!) And just looking at the thinking trace of K3 you can learn tons of stuff! Here usually I have a separate chat with K3 and when I see something I want to dig deeper or ask about from the main thinking trace I do and in some cases I stop the experiment and pivot immediately to something else. Now let's see if I can apply learned lessons to the M5 kernel! 🚀 Branch for any braves soul willing to test it:
Ivan Fioravanti ᯅ18,381 görüntüleme • 26 gün önce

DwarfStar DeepSeek v4 Flash 0731 pre-M5 kernel keeps improving 💪🏻 - prefill +11% - decode +10% - single request no DSpark/MTP. - 100% same logics between baseline and now. Today is QA like crazy. I hope to file PR by EOD antirez Here are results on the M3 Ultra and a video now vs baseline using ds4-eval.
Ivan Fioravanti ᯅ17,121 görüntüleme • 26 gün önce

Repo Prompt + o3 = Mind Blowing results! Plain video, no edit!
Ivan Fioravanti ᯅ180,129 görüntüleme • 1 yıl önce

Here is Kimi K3 frogger version to show the crazy jump between the two models! Same prompt. 🤩
Ivan Fioravanti ᯅ25,310 görüntüleme • 1 ay önce