正在加载视频...

视频加载失败

MLX Swift LLM example works with: - Mistral / Llama - Phi-2 - Qwen 1.5 - Starcoder 2 Quick-start: Qwen 1.5 0.5B runs pretty fast in 16-bit on my iPhone 14, no quantization needed:

30,441 次观看 • 2 年前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

Holy shit... Microsoft open sourced an inference framework that runs a 100B parameter LLM on a single CPU. It's called BitNet. And it does what was supposed to be impossible. No GPU. No cloud. No $10K hardware setup. Just your laptop running a 100-billion parameter model at human reading speed. Here's how it works: Every other LLM stores weights in 32-bit or 16-bit floats. BitNet uses 1.58 bits. Weights are ternary just -1, 0, or +1. That's it. No floats. No expensive matrix math. Pure integer operations your CPU was already built for. The result: - 100B model runs on a single CPU at 5-7 tokens/second - 2.37x to 6.17x faster than llama.cpp on x86 - 82% lower energy consumption on x86 CPUs - 1.37x to 5.07x speedup on ARM (your MacBook) - Memory drops by 16-32x vs full-precision models The wildest part: Accuracy barely moves. BitNet b1.58 2B4T their flagship model was trained on 4 trillion tokens and benchmarks competitively against full-precision models of the same size. The quantization isn't destroying quality. It's just removing the bloat. What this actually means: - Run AI completely offline. Your data never leaves your machine - Deploy LLMs on phones, IoT devices, edge hardware - No more cloud API bills for inference - AI in regions with no reliable internet The model supports ARM and x86. Works on your MacBook, your Linux box, your Windows machine. 27.4K GitHub stars. 2.2K forks. Built by Microsoft Research. 100% Open Source. MIT License.

Guri Singh

2,180,357 次观看 • 5 个月前

Qwen 3.8 27B Q4_K_M - 90 tokens/sec on a single NVIDIA RTX 4090 (24 GB VRAM) with Dflash2! (MTP 60 tps -> 90 tps Dflash2!!!!) Local AI moves so fast (literally!) it’s terrifying. Z lab just dropped DFlash 2 for Qwen 3.8 27b and Muse Glimmer. I patched llama.cpp (PR #27342) and paired it with Unsloth’s Qwen 3.8 27B UD-Q4_K_XL quant. The result? Lossless 90 tokens/s decode. My last post highlighted native MTP hitting 60 t/s at 130,000 context. But DFlash 2 just completely shattered that ceiling. By using parallel block diffusion drafting (predicting whole blocks of tokens in a single pass using dynamic convolutions), DFlash achieves a massive 5.39 token acceptance rate. THE ALPHA TWEAK: `n-max 7` eats too much VRAM for draft states. But if you drop the draft limit to `--spec-draft-n-max 4`, you slash the VRAM overhead and actually increase the throughput. Here is the new 24GB VRAM Physics Matrix (DFlash 2 @ n-max 4): - 30k Context: 1,725 t/s prefill | 87.05 t/s decode | 22.2 GB VRAM - 80k Context: 1,789 t/s prefill | 84.20 t/s decode | 23.3 GB VRAM - 110k Context: 1,767 t/s prefill | 83.35 t/s decode | 23.96 GB VRAM (110k context at 83+ tokens a second sitting exactly on the 24GB hardware limit is absolute wizardry). How to compile the PR today: git clone cd llama.cpp git fetch origin pull/27342/head:pr-27342 git switch pr-27342 cmake -B build -DGGML_CUDA=ON && cmake --build build -j Llama.cpp flags for Dflash (110k Context Ceiling): ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf --spec-type draft-dflash --spec-draft-n-max 4 -c 110000 -ngl 99 --port 8080 -ctv q4_0 -ctk q4_0 The fact that the open source community is shipping block diffusion drafters so quickly that run entirely locally on a gaming GPU is unbelievable. If you own a single RTX 3090 or 4090, it is officially time to upgrade to qwen 3.8 27b with dflash 2 and cancel your API subscriptions and let local silicon eat the cloud. This model beats GPT 5.6 Terra, GLM 5.2 DeepSeek V4 Pro, Muse Spark 1.2 and Claude Opus 4.8 on the artificial analysis agentic index (details in the replies) Hugging Face GGUF links (Base + DFlash2) and the full visual VRAM scaling and Dflash2 vs MTP graphs are also in the replies below. are you sticking to native MTP for the 130k context, or sacrificing 20k context to redline your decode speed? How many tokens/sec are you pushing on your current local rig?

Alok

102,967 次观看 • 8 天前

Made with Seednace 2 prompt: Idol dressing room at an MV filming studio, midday. In a brief lull while preparing for a fan meeting, a female idol @ image is left alone, filming a behind-the-scenes vlog on her own phone. Bright, bubbly idol energy — genuine smiles, playful expressions, eye contact with the lens, small laughs. Handheld from start to finish, quick and lively, alternating between selfie and mirrorless shots, capturing small moments of the prep process. Eight short cuts flow together, bouncy, within one continuous space and time. Each cut follows the rhythm of a vlog: greeting → introducing the gift bags → putting on a hairpin → detail insert → practicing fan-service poses → a member's voice on a phone call → laughing → excitedly exiting at a staff call. A light ending: excitement as she heads off to meet her fans. Characters CHASE — a Korean idol in her 20s. Long straight black hair (past her chest), elegant yet lovely Korean features, dewy glass skin, coral pink lips, big eyes. Slim yet curvy proportions at 33C-26-34. Wearing a light pink slip dress (thin shoulder straps, V-neckline, shirred waist). Pearl-silver drop earrings. The bright, bubbly star of the vlog, filming herself facing the lens. Storyboard Fan Meeting Prep Vlog (midday, indoors, handheld) The idol dressing room at an MV filming studio. On camera left, a small table stacked with gift bags sent by fans, behind it a rolling garment rack with covered costumes, on the right her phone and hair accessories on a vanity. Above, a door leading to a bright studio hallway. She's alone, hair and makeup done, preparing for the fan meeting while filming a vlog on her phone. Quick, cheerful handheld. (Cut 1 · ~2 sec · front-facing selfie, arm's length) She leans in close to the lens with a bright smile and a quick hand-heart. CHASE: "Morning, everyone — getting ready for the fan meeting!" (Cut 2 · ~2 sec · quick whip pan, handheld POV) Phone swings from the table of gift bags → to the garment rack → to the vanity, then snaps back to her face. CHASE (off-screen, playful): "These are all gifts from you guys!" (Cut 3 · ~2 sec · medium handheld, in front of the vanity) She picks up a ribbon-shaped hairpin from the vanity and clips it into her fringe, tilting her head as she checks the mirror. CHASE: "What do you think, cute?" (Cut 4 · ~1.5 sec · macro insert, tight close-up, shallow depth of field) Detail shot: fingers adjusting the hairpin's position, the pearl-silver earrings swaying in the light. No dialogue — just the soft rustle of fabric. (Cut 5 · ~2 sec · medium handheld, phone tracking her) She practices fan-service poses, making finger hearts at a few different angles, laughing at herself. CHASE: "I think this angle works!" (Cut 6 · ~2 sec · snappy cut, close handheld on the sofa) She sits on the sofa, her phone rings, she answers on speaker, and bursts out laughing at a member's voice. CHASE (laughing): "Hey, I was literally filming right now!" (Cut 7 · ~1.5 sec · quick punch-in, tight selfie) Still on the call, she shoots a playful side-eye at the lens and pouts. CHASE: "Why is she like this—" (Cut 8 · ~2 sec · arm's-length selfie finish) An off-screen staff knock — her eyes widen with excitement. A quick wave, a big wink with both hands making heart shapes, then she jumps out of frame — the camera lingers on the vanity for half a beat. CHASE (shouting as she jumps up): "Off to meet my fans — bye~!"

WasifAI

47,804 次观看 • 1 个月前

50% more context unlocked for Qwen 3.8 27b Q4_K_XL dflash 2 on a single RTX 4090 (24 GB VRAM) I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B is now running a massive 250,000 context at 75 tokens/s on a single RTX 4090. Here is the secret: By default, `llama-server` reserves massive chunks of your VRAM to handle multiple concurrent users (batching). If you are running a single user session, you are bleeding memory for features you aren't using. By passing the `--parallel 1` flag, you force the engine to dedicate 100% of your 24GB VRAM buffer to a single user. When we combine the VRAM saved by our Q2_K 2-bit drafter with the VRAM saved by `--parallel 1`, the context ceilings absolutely explode: Note: all benchmarks carried out with a massive 28k prompt. Ubuntu 22. ### THE NEW 24GB PHYSICAL LIMITS (Single RTX 4090): # 1. The "Repo Swallower" (Q4 KV Cache): - Context: 250,000 tokens (Up from 170k!) - Speed: 73.66 t/s decode | 1,608 t/s prefill - Peak VRAM: 23.8 GB # 2. The "High-Precision SWE" (Q8 KV Cache): - Context: 150,000 tokens (Up from 100k!) - Speed: 75.01 t/s decode | 1,667 t/s prefill - Peak VRAM: 23.9 GB # 3. The "Pristine Attention" (Unquantized FP16 KV): - Context: 90,000 tokens - Speed: 80.58 t/s decode | 1,699 t/s prefill - Peak VRAM: 23.92 GB ### HOW TO RUN THE 250K GOD STACK TODAY: (Requires PR #27342 + my Q2_K Hugging Face drafter) llama.cpp flags: ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q2_K.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 250000 -ngl 99 --parallel 1 --port 8080 -ctv q4_0 -ctk q4_0 We are pushing a quarter million tokens of context with speculative DFlash 2 decoding at 73 tokens/second on a single consumer gaming GPU. I dropped my custom 2 bit Hugging Face GGUF links, visual performance graphs, and the PR #27342 build instructions in the replies below. If you own a single RTX 3090 or 4090, it is officially time to cancel your API subscriptions and let local silicon eat the cloud. how much monthly API spend does an optimized 4090 rig like this actually replace for you?

Alok

39,189 次观看 • 6 天前

Luxury perfume commercials do not need a full production crew anymore. I created this premium fragrance ad using Nano Banana 2 and Seedance on Creatify AI , combining cinematic product shots, realistic motion, and commercial-quality visuals in minutes. Prompt: Shot 1 (0:00–0:01.5) — Quick Establishing Bright, sunlit vanity scene, bottle standing on a blush-gold pedestal surrounded by lavender sprigs, a white orchid, and honeycomb. Quick push-in — energetic, not lingering. BGM: Upbeat, light acoustic-pop instrumental kicks in immediately, bright and warm, mid-tempo with a confident rhythm — loud enough to carry the whole spot. Shot 2 (0:01.5–0:03) — Jessica Arrival Jessica walks into frame toward the vanity with natural, real-time movement — no slow-mo — reaching for the bottle with a light smile. Camera tracks briskly alongside her. Shot 3 (0:03–0:04.5) — Macro Bottle Turn Quick macro shot: bottle rotates in Jessica's hand at natural speed, gold cap and floral artwork catching light, honeycomb and lavender blurred in the foreground. Shot 4 (0:04.5–0:06) — Lifting to Apply Jessica lifts the bottle toward her neck/collarbone at natural speed, wrist turning, confident everyday gesture — not slow, not hesitant. Shot 5 (0:06–0:08) — The Spray (brief slow-mo accent only) Jessica presses the gold pump at her neck — a fine, visible mist releases and lands on her collarbone/upper chest area. Only this exact release moment gets a brief 0.3–0.5 sec slow-mo accent for visual impact, then instantly resumes normal speed as she lowers the bottle with a satisfied, natural exhale and slight smile — no eyes-closed lingering pause. BGM: Beat swells slightly on the spray release for emphasis, then continues driving forward. Shot 6 (0:08–0:09.5) — Wrist Application Quick cut: Jessica dabs/sprays lightly at her inner wrist, natural real-time motion, then rubs wrists together briefly — fast, confident, commercial-standard gesture. Shot 7 (0:09.5–0:11) — Notes Visualized Fast macro cutaway: lavender petals and a drop of golden honey caught mid-air in crisp, quick motion (not slow-mo) — visualizing lavender, honey, and orchid notes in under 1.5 seconds. Shot 8 (0:11–0:12.5) — Product Detail Quick tilt down the bottle from cap to base, floral artwork sharp and legible, brisk camera movement matching commercial pace. Shot 9 (0:12.5–0:15) — Final Hero Reveal Jessica turns toward camera with the bottle in hand, natural confident stance, quick dolly-out revealing the full vanity scene. Clean, crisp final frame — no slow lingering fade. BGM: Track builds to its peak in the final second, ending on a bright, resolved note — loud and full, not trailing off quietly. Voiceover (minimal, energetic, matched to BGM — not whispered ASMR): (0–1.5 sec) "Meet Lollia Relax." (6–7 sec, on the spray) "Lavender. Honey. Pure calm." (13–15 sec) "Relax into it." Delivered in a warm, clear, confident Western English accent — natural speaking volume and pace, energetic and inviting, not breathy or drawn out. #Creatify #aimediabuyer #aitools #MarketingTips #D2CMarketing

Jessica Collins

51,987 次观看 • 1 个月前