
FHILY👑
@Oluwaphilemon1 • 8,945 subscribers
Defi + Content Creator + Entertainment + Claude News + AI
Shorts
Videos

Qwen3.8-Flash-Next is still going strong at 364.7K tokens of context on an M5 Max. And this isn’t just a static long-context test. The model was reasoning about how to speed up its own workflow while using tools, and the tool calls kept working without misses. Setup: • Qwen3.8-Flash-Next • M5 Max • 128GB unified memory • MLX-Serve PR #363 • OpenCode 2 • 364.7K context The interesting part isn’t simply getting hundreds of thousands of tokens into memory. It’s what happens once the context gets this large. Long-context inference usually comes with a painful tradeoff. As the KV cache grows, memory pressure increases and generation can slow down. But this setup is still pushing through 364K tokens while maintaining a usable agent workflow. The model can reason, call tools, inspect results, continue working, and keep the session moving. And the tool calls reportedly haven’t missed so far. That’s important for agentic coding. A huge context window is only useful if the model can actually operate reliably inside it. A 400K-token context that constantly breaks tool calls isn’t very useful. A 364K session that can keep reasoning and executing tools is a different story. And the test isn’t finished yet. The current run is approaching 400K tokens, with the expectation that it can keep going. This is also another interesting example of why Apple Silicon keeps showing up in local LLM experiments. The M5 Max’s unified memory gives a large model and its growing KV cache access to one shared memory pool. With MLX-Serve continuing to improve, these machines are becoming surprisingly capable long-context inference boxes. The bigger takeaway: Context length is becoming a workload, not just a model specification. Running a model at 256K is one thing. Keeping an agent alive at 300K+ while it reasons and uses tools is much more interesting. And Qwen3.8-Flash-Next is showing that this can be pushed surprisingly far on a single 128GB Mac. 364.7K and counting. Next stop: 400K.
FHILY👑39,982 Aufrufe • vor 4 Tagen

125B on a single RTX 4090 at 25 tok/s is actually insane Qwen3.8-Flash-Next + MTP is hitting: 80K context 25.35 tok/s decode 471 tok/s prefill On one consumer GPU. Yes, it’s an MoE, so you’re not running all 125B parameters for every token. But still, getting a model this large to run this fast on a 4090 is wild. MTP is doing a lot of the heavy lifting here too. Instead of generating one token at a time, it can speculate on multiple tokens and verify them together. And 25 tok/s is the important number.
FHILY👑38,386 Aufrufe • vor 8 Tagen

Сlaude Design is so cracked at marketing. 🤯 and anyone can achieve the same results. Here's the prompt ↓
FHILY👑219,605 Aufrufe • vor 4 Monaten

Qwen3.8-27B vs Ornith-1.5-35B on the Bioluminescent Abyssal Temple challenge. Same prompt. Same Grok Build harness. 4-bit MLX. → Qwen: ~24 tok/s · 33 min → Ornith: ~55 tok/s · 16 min Ornith was over 2× faster and finished in roughly half the time. This is a great example of why raw parameter count doesn’t tell the whole story. Qwen vs Ornith is getting seriously interesting.
FHILY👑22,929 Aufrufe • vor 21 Tagen

Built this with just 3 tools; Materials¹ → Create 3D visuals Seedance 2.0 → Animation Claude Code → Build
FHILY👑50,427 Aufrufe • vor 2 Monaten

changed how I design with AI. I show how to build faster by adding a design system directly into the prompt, importing any website URL to get a strong starting point, then polishing everything in my own style. I also mix it with components and skills to keep the design consistent and more unique. GPT-5.5 cloning websites is honestly crazy good. The result can get very close to the original, then you can refine it and make it your own. Full breakdown in the comments.
FHILY👑57,007 Aufrufe • vor 3 Monaten

This is how I create a motion design workflow fully AI-powered with Claude and Higgsfield MCP. 1. Start by pulling references via Pinterest API in Claude Code 2. Generate a 6-scene storyboard through GPT Images 2.0 3. In Higgsfield MCP, then render it straight to video in Seedance 2.0 All connected in one seamless pipeline.
FHILY👑45,001 Aufrufe • vor 3 Monaten

Everyone is talking about Claude Design… but I keep going back to Gemini. Gemini 3.1 Pro is seriously impressive when it comes to motion design. I actually want to try Claude Design, but after only two prompts, I was already running out of credits. If you care about anti AI slop, Gemini is still winning. With a single prompt, it can transform a static HTML layout into dynamic, polished motion graphics.
FHILY👑25,019 Aufrufe • vor 2 Monaten

Two frontier models. Two very different demos. (GPT-5.6 Sol and Claude Fable 5) GPT-5.6 Sol Ultra was pushed in a mathematical direction and generated a Minecraft-style clone in Lean. Yes, that Lean. Meanwhile, Claude Fable 5 took a game development route and built a Minecraft mod featuring a Dark Souls-inspired enemy with multiple attack patterns, combos, and combat logic. They're solving completely different problems, but both are great examples of how capable today's models have become. The pace of progress is remarkable.
FHILY👑19,336 Aufrufe • vor 2 Monaten

Motion design with Claude + Higgsfield MCP 🧩 Turn a single text prompt into pro-level motion videos. It’s actually insanely simple: 1. Describe your idea to Claude 2. Ask it to generate a storyboard 3. Pick your favorite concept 4. Let Higgsfield MCP render the final video
FHILY👑25,051 Aufrufe • vor 3 Monaten

Hey Claude, redesign my landing page and add some 3D depth to it 🔥 — Claude Opus 4.8 ultracode
FHILY👑15,613 Aufrufe • vor 2 Monaten

How I create 3D-feel websites with expanding images using Claude Code👇 1. Use Claude to plan scroll-based animations + scene flow 2. Start with full-width images that expand on scroll 3. Add zoom-in + scale effects to create depth illusion 4. Layer images to simulate foreground/background (fake 3D) 5. Use parallax so elements move at different speeds 6. Add smooth transitions + ambient motion for immersion 7. Fine-tune timing (0.6s–1.2s) for cinematic feel Save this if you want next-level websites 🚀
FHILY👑10,268 Aufrufe • vor 3 Monaten
Keine weiteren Inhalte verfügbar