Ivan Fioravanti ᯅ's banner
Ivan Fioravanti ᯅ's profile picture

Ivan Fioravanti ᯅ

@ivanfioravanti41,873 subscribers

GenAI/LLM addicted, Apple MLX, Cloud computing, Kubernetes, Technology Advisor, Investor and Co-Founder & Board Member of CoreView.

Shorts

Ox Alpha will be unveiled soon, are you ready? Ox Alpha 很快就要发布了,你准备好了吗?

Ox Alpha will be unveiled soon, are you ready? Ox Alpha 很快就要发布了,你准备好了吗?

31,803 görüntüleme

MLX MiniMax 2.5 running LOCALLY on a single M3 Ultra 512GB! Writing a poem on LLMs at 6bit quantization! 🔥 Let's start some coding, context and distributed tests! Generation: 40.2 tokens-per-sec Peak memory: 186 GB

MLX MiniMax 2.5 running LOCALLY on a single M3 Ultra 512GB! Writing a poem on LLMs at 6bit quantization! 🔥 Let's start some coding, context and distributed tests! Generation: 40.2 tokens-per-sec Peak memory: 186 GB

226,466 görüntüleme

Another local video created with MiniMax H3 on DGX Spark, it took forever 1 hours and 20 minutes 😱 For: - 1280x736 (16:9) - 15 seconds - DGX Spark - ComfyUI Template Text to Video Thanks MiniMax (official) for this model!

Another local video created with MiniMax H3 on DGX Spark, it took forever 1 hours and 20 minutes 😱 For: - 1280x736 (16:9) - 15 seconds - DGX Spark - ComfyUI Template Text to Video Thanks MiniMax (official) for this model!

40,613 görüntüleme

It worked! Hermes Agent + Exo + Qwen3 Coder Next 8bit to create an incredible snake game, with model following 100% specifications passed in prompt! Let's load something bigger now!

It worked! Hermes Agent + Exo + Qwen3 Coder Next 8bit to create an incredible snake game, with model following 100% specifications passed in prompt! Let's load something bigger now!

160,169 görüntüleme

I'm going crazy on this one, everyone posting incredible benchmarks of Laguna S 2.1 on Apple Silicon. But on M5 Max 128GB using mlx-community/Laguna-S-2.1-oQ4e I'm getting 20 tps when I'm lucky. Using macOS 27 beta 4 here, but I don't think this is the culprit.

I'm going crazy on this one, everyone posting incredible benchmarks of Laguna S 2.1 on Apple Silicon. But on M5 Max 128GB using mlx-community/Laguna-S-2.1-oQ4e I'm getting 20 tps when I'm lucky. Using macOS 27 beta 4 here, but I don't think this is the culprit.

25,524 görüntüleme

MiniMax H3 960 × 544 - 10 seconds - generated in 16:57 mins on DGX Spark. Time to test on Apple Silicon. Love this Nolan/Inception style videos!

MiniMax H3 960 × 544 - 10 seconds - generated in 16:57 mins on DGX Spark. Time to test on Apple Silicon. Love this Nolan/Inception style videos!

18,343 görüntüleme

First real test on M4 Max 40GPU to transcribe 3 hours of video with mlx_whisper 0.4.1 - M4 Max 2:19 mins (Max Fan no throttling) - M4 Max 2:29 mins (System Fan throttling towards end) GPU freq goes 1.578 --> 952 Here the video (8x) showing throttling kicking in at the end.

First real test on M4 Max 40GPU to transcribe 3 hours of video with mlx_whisper 0.4.1 - M4 Max 2:19 mins (Max Fan no throttling) - M4 Max 2:29 mins (System Fan throttling towards end) GPU freq goes 1.578 --> 952 Here the video (8x) showing throttling kicking in at the end.

216,328 görüntüleme

Wait, what? 😱 Apple Foundation Model on Private Cloud Compute is working on fm CLI! Here's a video!

Wait, what? 😱 Apple Foundation Model on Private Cloud Compute is working on fm CLI! Here's a video!

38,712 görüntüleme

Qwen 3 0.6B is Ultra powerful if fine-tuned! Here I got 83.5% on classification tasks! Only Gemini 2.5 Pro does better with 85%! Video is in real time on M3 Ultra!

Qwen 3 0.6B is Ultra powerful if fine-tuned! Here I got 83.5% on classification tasks! Only Gemini 2.5 Pro does better with 85%! Video is in real time on M3 Ultra!

143,143 görüntüleme

This is on M3 Ultra; M5 Ultra will be crazy!

This is on M3 Ultra; M5 Ultra will be crazy!

98,247 görüntüleme

GLM-4.7 Flash by Z.ai running on M3 Ultra using MLX! 4bit at 81 toks/s 🔥 8bit at 64 toks/s 🔥 (video below) Model converted from transformers to MLX on the fly with Codex + GPT-5.2 HIgh!

GLM-4.7 Flash by Z.ai running on M3 Ultra using MLX! 4bit at 81 toks/s 🔥 8bit at 64 toks/s 🔥 (video below) Model converted from transformers to MLX on the fly with Codex + GPT-5.2 HIgh!

75,425 görüntüleme

Another incredible p5js animation created by the new King: Sonnet 3.7 Thinking! It's a kind of magic!

Another incredible p5js animation created by the new King: Sonnet 3.7 Thinking! It's a kind of magic!

137,251 görüntüleme

The last thing you see in your first and last Space Travel MiniMax H3 prompt: amateur handheld pov footage of a tourist in their plush and comfortable room of a space cruise, carpeted floor and comfy bed, large window shows a view of Gargantua blackhole, you can see their reflection in the window, they turn off the light half way through so we can see outside better and say: "Wow look at that!", then they show back the room

The last thing you see in your first and last Space Travel MiniMax H3 prompt: amateur handheld pov footage of a tourist in their plush and comfortable room of a space cruise, carpeted floor and comfy bed, large window shows a view of Gargantua blackhole, you can see their reflection in the window, they turn off the light half way through so we can see outside better and say: "Wow look at that!", then they show back the room

10,416 görüntüleme

MLX news: MOSS-TTS-Local Transformer 1.5 available on mlx-audio now! Thanks to Prince Canuma and Lucas Newman Audio on! 🔉

MLX news: MOSS-TTS-Local Transformer 1.5 available on mlx-audio now! Thanks to Prince Canuma and Lucas Newman Audio on! 🔉

20,292 görüntüleme

I tested MTPLX v2 with QWEN 3.6 27B and compared it with oMLX without cache on M5 Max and DGX Spark on vllm using nvfp4 model version. More details in 🧵 I've reached 82.8 tps of max decoding speed! 🔥 Custom Metal Kernel design specifically for this model and for Apple Silicon is just perfect! This is the way forward! Great job Youssof Al Toukhi Look at the website here! 👇 Here a website with recap, built with GLM 5.2 running locally 💪 First chart and preview from the website.

I tested MTPLX v2 with QWEN 3.6 27B and compared it with oMLX without cache on M5 Max and DGX Spark on vllm using nvfp4 model version. More details in 🧵 I've reached 82.8 tps of max decoding speed! 🔥 Custom Metal Kernel design specifically for this model and for Apple Silicon is just perfect! This is the way forward! Great job Youssof Al Toukhi Look at the website here! 👇 Here a website with recap, built with GLM 5.2 running locally 💪 First chart and preview from the website.

15,885 görüntüleme

Eulerian Fluid simulation test! Zero-shot! Opus 4.6 vs GPT-5.3 vs Gemini 3 Deep Think! My personal preference: 🥇 Gemini 3 Deep Think (really strong!) 🥈 Opus 4.6 🥉 GPT 5.3 High

Eulerian Fluid simulation test! Zero-shot! Opus 4.6 vs GPT-5.3 vs Gemini 3 Deep Think! My personal preference: 🥇 Gemini 3 Deep Think (really strong!) 🥈 Opus 4.6 🥉 GPT 5.3 High

38,277 görüntüleme

GLM-4.7 and MiniMax M2.1 side by side, both powered by Claude Code! Prompt below. Trick: one setting file per provider: - claude --settings ~/.claude/minimax_settings.json - claude --settings ~/.claude/zai-settings.json

GLM-4.7 and MiniMax M2.1 side by side, both powered by Claude Code! Prompt below. Trick: one setting file per provider: - claude --settings ~/.claude/minimax_settings.json - claude --settings ~/.claude/zai-settings.json

44,866 görüntüleme

Running Gemma 4 12B on your iPhone? Yes! 🧵 LM Studio + Locally AI latest version with LM Link is really cool! This opens up additional scenarios! My brain is on fire 🔥

Running Gemma 4 12B on your iPhone? Yes! 🧵 LM Studio + Locally AI latest version with LM Link is really cool! This opens up additional scenarios! My brain is on fire 🔥

19,292 görüntüleme

Playing with Qwen3-TTS is and MLX-Audio locally on Mac Studio M3 Ultra 🔥 Amazing model by Qwen and great work by Prince Canuma bringing this magic to MLX! Tuning the voice with instructions feels like magic! Command and prompts to run it are in the video.

Playing with Qwen3-TTS is and MLX-Audio locally on Mac Studio M3 Ultra 🔥 Amazing model by Qwen and great work by Prince Canuma bringing this magic to MLX! Tuning the voice with instructions feels like magic! Command and prompts to run it are in the video.

35,720 görüntüleme

Videos

ivanfioravanti's profile picture

I respectfully disagree for several reasons. Calling a customer, whether free or paying, an idiot is simply wrong. OpenCode, like any other coding agent, clearly tries to preserve the prompt cache as much as possible. Otherwise, it would be painfully slow. The “Stop Using OpenCode” article, which I believe Dax is referring to, focused on specific cases that can invalidate the cached prefix. These issues become much more visible and painful when using local models. Modifying AGENTS.md is one example, as shown in this video. During this coding session, cache efficiency is extremely high. But the moment I modify AGENTS.md, boom: 62K tokens need to be processed again as new input. The date change is another valid point raised in the article. I understand that both issues may sound trivial, and personally I've never been impacted by them, but their impact on the OpenCode end-user experience can be significant when using local models. I agree that both the feedback and the article were too direct and were probably written by the author out of frustration. However, they also contained constructive points that could help improve the product. I use OpenCode alongside several other tools, and I like it overall. But I really dislike this kind of reaction to user feedback, regardless of how that feedback is expressed. I would have preferred a response focused more on listening, understanding the problem, and improving the product.

Ivan Fioravanti ᯅ

37,594 görüntüleme • 1 ay önce