MLX Swift LLM example works with: - Mistral /... Llama - Phi-2 - Qwen 1.5 - Starcoder 2 Quick-start: Qwen 1.5 0.5B runs pretty fast in 16-bit on my iPhone 14, no quantization needed:show more

Awni Hannun
30,441 次观看 • 2 年前
DeepSeek R1 distilled to Qwen 1.5B easily runs on... my iPhone 16 with MLX swift. Here's the 4-bit model reasoning entirely on device at almost 60 toks/sec:show more

Awni Hannun
1,139,656 次观看 • 1 年前
I wanted to try Maple-Preview by DeepGrove on iPhone... so I simply ported the model to MLX Swift using Bionic Maple runs at ~80tk/s on iPhone 17 Pro with my quick port, it’s a 20B-A1B model at 2-bit The model is close to Qwen3.5 35-A2B in many benchs and runs easily on iPhoneshow more

Adrien Grondin
13,376 次观看 • 25 天前
Omg.. this is wild... This Github repo removes LLM... censorship permanently in 45 minutes. It's called Heretic - 100% Open Source. No jailbreaks, prompts, and just One command. ↳ Zero configuration ↳ Keeps model intelligence intact ↳ Works with Llama, Qwen, Gemma ↳ Runs locally But it won't stay under the radar forever.show more

Kanika
21,170 次观看 • 1 个月前
BOOM 🚨: Tencent’s Hunyuan Turbo S just landed in... the Top 8 globally on Chatbot Arena and is now #2 in China, just behind DeepSeek. Even Google ex-CEO Eric Schmidt says China had no foundation models two years ago — now it has three: DeepSeek, Qwen, Hunyuan — on par with OpenAI’s O.1. At #GoogleIO, all three scored high on the global leaderboard like #DeepSeek, #Qwen, #Hunyuan Tencent Hy isn’t just racing — it’s rewriting the leaderboard. #HunyuanTurboS #TencentAI #LLM #AIRevolutionshow more

Emily Watson | AI Tools & Tech News
24,853 次观看 • 1 年前
Apple iPhone 16 vs Pixel 9: Which one should... you buy? The iPhone 16 is more of an iterative upgrade over its predecessor, but the Pixel 9 brings several notable upgrades, especially in the camera, design and chipset. While the Tensor G4 may not be as fast as the iPhone 16’s A18 chipset, it no longer heats up and is pretty fast when it comes to multitasking and completing AI-powered tasks. If you are already a part of the Apple ecosystem and have the iPhone 14 or an older device, the iPhone 16 might make sense. On the other hand, the Pixel 9 with its smart AI features, solid design and capable cameras might appeal to to those who want a phone that will get software updates for years to come and want to enjoy the clean user interface stock Android has to offer.show more

cutie 🥰
23,995 次观看 • 1 年前
No makeup artist. No complicated routine. Just a quick... everyday glow-up in seconds Created this with Seedance 2.0 on BudgetPixel AI Prompt Storyboard: “15-Second UGC Makeup Tutorial Script (Natural, Fast-Paced, Viral Style) 0.0 – 1.5s | Hook (Front Camera, natural light) “Stop scrolling— a girl without makeup " 1.5 – 4.0s | Base (Close-up cuts) “Start with a lightweight moisturizer… then a skin tint, just dab and blend with fingers.” 4.0 – 7.0s | Face (Fast transitions) “Cream blush on cheeks and nose—this gives that soft, fresh girl look instantly.” 7.0 – 10.0s | Eyes + Brows (Quick mirror shots) “Brush brows up, add a tiny bit of mascara… no heavy eyes, just clean and lifted.” 10.0 – 12.5s | Lips (Close-up) “Lip tint or gloss—blur it slightly with fingers for that natural stain effect.” 12.5 – 15s | Final Reveal (Smile, camera pull-back) “And that’s it—everyday glow in seconds. Save this if you love easy makeup, party look"show more

Ai Girllie
34,493 次观看 • 2 个月前
Holy shit... Microsoft open sourced an inference framework that... runs a 100B parameter LLM on a single CPU. It's called BitNet. And it does what was supposed to be impossible. No GPU. No cloud. No $10K hardware setup. Just your laptop running a 100-billion parameter model at human reading speed. Here's how it works: Every other LLM stores weights in 32-bit or 16-bit floats. BitNet uses 1.58 bits. Weights are ternary just -1, 0, or +1. That's it. No floats. No expensive matrix math. Pure integer operations your CPU was already built for. The result: - 100B model runs on a single CPU at 5-7 tokens/second - 2.37x to 6.17x faster than llama.cpp on x86 - 82% lower energy consumption on x86 CPUs - 1.37x to 5.07x speedup on ARM (your MacBook) - Memory drops by 16-32x vs full-precision models The wildest part: Accuracy barely moves. BitNet b1.58 2B4T their flagship model was trained on 4 trillion tokens and benchmarks competitively against full-precision models of the same size. The quantization isn't destroying quality. It's just removing the bloat. What this actually means: - Run AI completely offline. Your data never leaves your machine - Deploy LLMs on phones, IoT devices, edge hardware - No more cloud API bills for inference - AI in regions with no reliable internet The model supports ARM and x86. Works on your MacBook, your Linux box, your Windows machine. 27.4K GitHub stars. 2.2K forks. Built by Microsoft Research. 100% Open Source. MIT License.show more

Guri Singh
2,180,357 次观看 • 5 个月前
> 8 GPUs in one server rig > dude... went homeless to build it > electrical bill costs more than rent now > while everyone else pays $400/month to openai > a 2 GPU desktop kills the api bill forever > rtx 4080 super + rtx 5060 ti = 32gb vram > runs qwen 3.6 with 100k context locally > no rate limits, no api keys, no data leaving the room > agents loop 400 times for free > claude opus still wins on hard reasoning > but local handles 90% of daily work > $1,200 setup pays itself off in 4 months > bookmark this and read the article belowshow more

starmex
167,058 次观看 • 3 个月前
Yup, a football video. The World Cup made us... do it Luma rebuilt image generation from scratch — reasoning first, pixels second. And it beats Google's Nano Banana 2 and GPT Image 1.5 on reasoning benchmarks All 3 new models are now live on AI/ML API luma/uni-1 plans before it draws. The model generates autoregressively: it works out layout, composition and text placement first, then renders the pixels. $0.052/image luma/uni-1-max — same prompts, same params, max fidelity. 2K output + editing with up to 9 reference images. Built for hero shots and ad creative. $0.13/image luma/ray-3-2 — up to 16 keyframes per clip, 20s, 1080p, native HDR + 16-bit EXR export. The video in this post came straight out of it model ids "luma/uni-1" "luma/uni-1-max" "luma/ray-3-2" Luma cooked. We serveshow more

AI/ML API
19,375 次观看 • 1 个月前
🚨 Alibaba just open sourced a GUI agent that... lives inside your webpage and controls it with natural language. It's called Page Agent and it's not a browser extension. It's pure JavaScript no Python, no Puppeteer, no headless browser, no screenshots. Just one script tag and your web app understands natural language. Here's what it actually does: → Embed it with a single tag or npm install → Control any web interface with plain English commands → Text-based DOM manipulation no OCR, no vision models needed → Bring your own LLM (GPT, Claude, Qwen, anything) → Ships a built-in UI with human-in-the-loop support → Turn 20-click ERP/CRM workflows into one sentence → Optional Chrome extension for multi-tab agent tasks → Works on any web app SaaS, admin panels, internal tools Companies are charging $30/month for AI copilots built on this exact idea. This is 3 lines of code. Your users. Your interface. The AI copilot layer for every web app just got open sourced. 1.6K stars. 100% Open Source. (Link in the comments)show more

Ihtesham Ali
135,634 次观看 • 5 个月前
Most robots still need markers, checkerboards, or long calibration... rituals just to know where their arms are. Now it works from raw images in seconds. roboreg is a markerless multi arm localization toolkit that plugs into ROS 2 and RViz. No special hardware. No custom setup. You toggle between robot descriptions and the system figures out the rest. The idea is simple: ✅ Hand eye calibration from plain RGB or RGB D images ✅ Only three robot poses needed for millimeter accuracy ✅ Works with any ROS 2 compatible robot and camera ✅ Fully open source under Apache 2.0 It is powered by Hydra, a new marker free ICP variant that converges far more reliably than classical baselines and runs in under a second. If you want to try it: roboreg: ROS 2 roboreg: Hydra paper: pip install roboreg More details and discussion on Open Robotics Discourse:show more

Ilir Aliu
18,406 次观看 • 9 个月前
Qwen 3.8 27B Q4_K_M - 90 tokens/sec on a... single NVIDIA RTX 4090 (24 GB VRAM) with Dflash2! (MTP 60 tps -> 90 tps Dflash2!!!!) Local AI moves so fast (literally!) it’s terrifying. Z lab just dropped DFlash 2 for Qwen 3.8 27b and Muse Glimmer. I patched llama.cpp (PR #27342) and paired it with Unsloth’s Qwen 3.8 27B UD-Q4_K_XL quant. The result? Lossless 90 tokens/s decode. My last post highlighted native MTP hitting 60 t/s at 130,000 context. But DFlash 2 just completely shattered that ceiling. By using parallel block diffusion drafting (predicting whole blocks of tokens in a single pass using dynamic convolutions), DFlash achieves a massive 5.39 token acceptance rate. THE ALPHA TWEAK: `n-max 7` eats too much VRAM for draft states. But if you drop the draft limit to `--spec-draft-n-max 4`, you slash the VRAM overhead and actually increase the throughput. Here is the new 24GB VRAM Physics Matrix (DFlash 2 @ n-max 4): - 30k Context: 1,725 t/s prefill | 87.05 t/s decode | 22.2 GB VRAM - 80k Context: 1,789 t/s prefill | 84.20 t/s decode | 23.3 GB VRAM - 110k Context: 1,767 t/s prefill | 83.35 t/s decode | 23.96 GB VRAM (110k context at 83+ tokens a second sitting exactly on the 24GB hardware limit is absolute wizardry). How to compile the PR today: git clone cd llama.cpp git fetch origin pull/27342/head:pr-27342 git switch pr-27342 cmake -B build -DGGML_CUDA=ON && cmake --build build -j Llama.cpp flags for Dflash (110k Context Ceiling): ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q4_K_M.gguf --spec-type draft-dflash --spec-draft-n-max 4 -c 110000 -ngl 99 --port 8080 -ctv q4_0 -ctk q4_0 The fact that the open source community is shipping block diffusion drafters so quickly that run entirely locally on a gaming GPU is unbelievable. If you own a single RTX 3090 or 4090, it is officially time to upgrade to qwen 3.8 27b with dflash 2 and cancel your API subscriptions and let local silicon eat the cloud. This model beats GPT 5.6 Terra, GLM 5.2 DeepSeek V4 Pro, Muse Spark 1.2 and Claude Opus 4.8 on the artificial analysis agentic index (details in the replies) Hugging Face GGUF links (Base + DFlash2) and the full visual VRAM scaling and Dflash2 vs MTP graphs are also in the replies below. are you sticking to native MTP for the 130k context, or sacrificing 20k context to redline your decode speed? How many tokens/sec are you pushing on your current local rig?show more

Alok
103,895 次观看 • 15 天前
JENSEN HUANG UNVEILED A BOARD THAT RUNS 1 TRILLION... PARAMETER AI MODELS. THE $249 NVIDIA BOX UNDER YOUR DESK KILLS A $200/MONTH AI BILL FOR $5 IN ELECTRICITY jensen held it up on stage with one hand and called it the architecture that runs the future of ai. that same technology now ships in a $249 box smaller than your wallet the jetson orin nano super pulls 7-25 watts and does 67 trillion ai operations per second. llama 3, mistral and deepseek run locally with no api fees and no data leaving your machine most developers pay $2,400 a year across chatgpt, openai api, claude pro and cursor. the jetson costs $314 in year one and $60 a year after. 2 year savings hit $4,431 install ollama with one command, change one line of code to point at localhost, and every tool built for openai works identically. zero rewrites, zero rate limits cloud subscriptions keep getting more expensive and rate limits keep getting tighter. the people who own the box in 2026 are going to look very far ahead in 2028 bookmark this and read the article belowshow more

starmex
54,448 次观看 • 3 个月前
Made with Seednace 2 prompt: Idol dressing room at... an MV filming studio, midday. In a brief lull while preparing for a fan meeting, a female idol @ image is left alone, filming a behind-the-scenes vlog on her own phone. Bright, bubbly idol energy — genuine smiles, playful expressions, eye contact with the lens, small laughs. Handheld from start to finish, quick and lively, alternating between selfie and mirrorless shots, capturing small moments of the prep process. Eight short cuts flow together, bouncy, within one continuous space and time. Each cut follows the rhythm of a vlog: greeting → introducing the gift bags → putting on a hairpin → detail insert → practicing fan-service poses → a member's voice on a phone call → laughing → excitedly exiting at a staff call. A light ending: excitement as she heads off to meet her fans. Characters CHASE — a Korean idol in her 20s. Long straight black hair (past her chest), elegant yet lovely Korean features, dewy glass skin, coral pink lips, big eyes. Slim yet curvy proportions at 33C-26-34. Wearing a light pink slip dress (thin shoulder straps, V-neckline, shirred waist). Pearl-silver drop earrings. The bright, bubbly star of the vlog, filming herself facing the lens. Storyboard Fan Meeting Prep Vlog (midday, indoors, handheld) The idol dressing room at an MV filming studio. On camera left, a small table stacked with gift bags sent by fans, behind it a rolling garment rack with covered costumes, on the right her phone and hair accessories on a vanity. Above, a door leading to a bright studio hallway. She's alone, hair and makeup done, preparing for the fan meeting while filming a vlog on her phone. Quick, cheerful handheld. (Cut 1 · ~2 sec · front-facing selfie, arm's length) She leans in close to the lens with a bright smile and a quick hand-heart. CHASE: "Morning, everyone — getting ready for the fan meeting!" (Cut 2 · ~2 sec · quick whip pan, handheld POV) Phone swings from the table of gift bags → to the garment rack → to the vanity, then snaps back to her face. CHASE (off-screen, playful): "These are all gifts from you guys!" (Cut 3 · ~2 sec · medium handheld, in front of the vanity) She picks up a ribbon-shaped hairpin from the vanity and clips it into her fringe, tilting her head as she checks the mirror. CHASE: "What do you think, cute?" (Cut 4 · ~1.5 sec · macro insert, tight close-up, shallow depth of field) Detail shot: fingers adjusting the hairpin's position, the pearl-silver earrings swaying in the light. No dialogue — just the soft rustle of fabric. (Cut 5 · ~2 sec · medium handheld, phone tracking her) She practices fan-service poses, making finger hearts at a few different angles, laughing at herself. CHASE: "I think this angle works!" (Cut 6 · ~2 sec · snappy cut, close handheld on the sofa) She sits on the sofa, her phone rings, she answers on speaker, and bursts out laughing at a member's voice. CHASE (laughing): "Hey, I was literally filming right now!" (Cut 7 · ~1.5 sec · quick punch-in, tight selfie) Still on the call, she shoots a playful side-eye at the lens and pouts. CHASE: "Why is she like this—" (Cut 8 · ~2 sec · arm's-length selfie finish) An off-screen staff knock — her eyes widen with excitement. A quick wave, a big wink with both hands making heart shapes, then she jumps out of frame — the camera lingers on the vanity for half a beat. CHASE (shouting as she jumps up): "Off to meet my fans — bye~!"show more

WasifAI
47,804 次观看 • 1 个月前
50% more context unlocked for Qwen 3.8 27b Q4_K_XL... dflash 2 on a single RTX 4090 (24 GB VRAM) I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B is now running a massive 250,000 context at 75 tokens/s on a single RTX 4090. Here is the secret: By default, `llama-server` reserves massive chunks of your VRAM to handle multiple concurrent users (batching). If you are running a single user session, you are bleeding memory for features you aren't using. By passing the `--parallel 1` flag, you force the engine to dedicate 100% of your 24GB VRAM buffer to a single user. When we combine the VRAM saved by our Q2_K 2-bit drafter with the VRAM saved by `--parallel 1`, the context ceilings absolutely explode: Note: all benchmarks carried out with a massive 28k prompt. Ubuntu 22. ### THE NEW 24GB PHYSICAL LIMITS (Single RTX 4090): # 1. The "Repo Swallower" (Q4 KV Cache): - Context: 250,000 tokens (Up from 170k!) - Speed: 73.66 t/s decode | 1,608 t/s prefill - Peak VRAM: 23.8 GB # 2. The "High-Precision SWE" (Q8 KV Cache): - Context: 150,000 tokens (Up from 100k!) - Speed: 75.01 t/s decode | 1,667 t/s prefill - Peak VRAM: 23.9 GB # 3. The "Pristine Attention" (Unquantized FP16 KV): - Context: 90,000 tokens - Speed: 80.58 t/s decode | 1,699 t/s prefill - Peak VRAM: 23.92 GB ### HOW TO RUN THE 250K GOD STACK TODAY: (Requires PR #27342 + my Q2_K Hugging Face drafter) llama.cpp flags: ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q2_K.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 250000 -ngl 99 --parallel 1 --port 8080 -ctv q4_0 -ctk q4_0 We are pushing a quarter million tokens of context with speculative DFlash 2 decoding at 73 tokens/second on a single consumer gaming GPU. I dropped my custom 2 bit Hugging Face GGUF links, visual performance graphs, and the PR #27342 build instructions in the replies below. If you own a single RTX 3090 or 4090, it is officially time to cancel your API subscriptions and let local silicon eat the cloud. how much monthly API spend does an optimized 4090 rig like this actually replace for you?show more

Alok
39,189 次观看 • 12 天前
How to Make Homemade Mayo. It’s creamy, delicious, and... avoids the usual risks of raw eggs. Ingredients (makes about 1–1.5 cups) 🥚 4–5 boiled eggs (peeled) 💧 Water (a splash, to help blending — roughly 2–4 tbsp) 🍋 Lemon juice (fresh, about 1–2 tbsp or to taste) 🫒 Olive oil (several tablespoons — added gradually, maybe ¼–⅓ cup total) 🧂 Pinch of salt (to taste) Optional add-ins (common for flavor): mustard, garlic powder, black pepper, or a touch of vinegar. Step-by-Step Instructions 1. Place the peeled boiled eggs in a wide-mouth jar or blending container. 2. Add a splash of water to loosen things up. 3. Pour in lemon juice and a pinch of salt. 4. Add olive oil (start with less and add more as needed for richness and emulsification). 5. Insert an immersion blender (stick blender) and blend everything until completely smooth and creamy. It transforms quickly into thick, classic-looking mayonnaise. 6. Taste and adjust seasoning (more salt, lemon, or oil if needed). Done! Store in the fridge for up to a week. Pro Tip: Use a narrow jar for the immersion blender so everything blends evenly without splatter. If it’s too thick, add a tiny bit more water or lemon juice.show more

🦅 Eagle Wings 🦅
32,448 次观看 • 3 个月前
Play has been called at DLF and the driving... range is about to become a concert area, but one man is out there still working on his game under the lights, Bryson DeChambeau. He’ll return tomorrow morning at 7:35 (visibility depending) with a 16 foot putt for par on 17, before completing his first round. Then there’s a quick turnaround before the start of his second. He’s currently 2 under par but says he’s playing great, just “rusty”, and they made some bad club selections on an unfamiliar course which led to mistakes. But make no mistake, he is well in this golf tournament.show more

Flushing It
260,485 次观看 • 1 年前
Luxury perfume commercials do not need a full production... crew anymore. I created this premium fragrance ad using Nano Banana 2 and Seedance on Creatify AI , combining cinematic product shots, realistic motion, and commercial-quality visuals in minutes. Prompt: Shot 1 (0:00–0:01.5) — Quick Establishing Bright, sunlit vanity scene, bottle standing on a blush-gold pedestal surrounded by lavender sprigs, a white orchid, and honeycomb. Quick push-in — energetic, not lingering. BGM: Upbeat, light acoustic-pop instrumental kicks in immediately, bright and warm, mid-tempo with a confident rhythm — loud enough to carry the whole spot. Shot 2 (0:01.5–0:03) — Jessica Arrival Jessica walks into frame toward the vanity with natural, real-time movement — no slow-mo — reaching for the bottle with a light smile. Camera tracks briskly alongside her. Shot 3 (0:03–0:04.5) — Macro Bottle Turn Quick macro shot: bottle rotates in Jessica's hand at natural speed, gold cap and floral artwork catching light, honeycomb and lavender blurred in the foreground. Shot 4 (0:04.5–0:06) — Lifting to Apply Jessica lifts the bottle toward her neck/collarbone at natural speed, wrist turning, confident everyday gesture — not slow, not hesitant. Shot 5 (0:06–0:08) — The Spray (brief slow-mo accent only) Jessica presses the gold pump at her neck — a fine, visible mist releases and lands on her collarbone/upper chest area. Only this exact release moment gets a brief 0.3–0.5 sec slow-mo accent for visual impact, then instantly resumes normal speed as she lowers the bottle with a satisfied, natural exhale and slight smile — no eyes-closed lingering pause. BGM: Beat swells slightly on the spray release for emphasis, then continues driving forward. Shot 6 (0:08–0:09.5) — Wrist Application Quick cut: Jessica dabs/sprays lightly at her inner wrist, natural real-time motion, then rubs wrists together briefly — fast, confident, commercial-standard gesture. Shot 7 (0:09.5–0:11) — Notes Visualized Fast macro cutaway: lavender petals and a drop of golden honey caught mid-air in crisp, quick motion (not slow-mo) — visualizing lavender, honey, and orchid notes in under 1.5 seconds. Shot 8 (0:11–0:12.5) — Product Detail Quick tilt down the bottle from cap to base, floral artwork sharp and legible, brisk camera movement matching commercial pace. Shot 9 (0:12.5–0:15) — Final Hero Reveal Jessica turns toward camera with the bottle in hand, natural confident stance, quick dolly-out revealing the full vanity scene. Clean, crisp final frame — no slow lingering fade. BGM: Track builds to its peak in the final second, ending on a bright, resolved note — loud and full, not trailing off quietly. Voiceover (minimal, energetic, matched to BGM — not whispered ASMR): (0–1.5 sec) "Meet Lollia Relax." (6–7 sec, on the spray) "Lavender. Honey. Pure calm." (13–15 sec) "Relax into it." Delivered in a warm, clear, confident Western English accent — natural speaking volume and pace, energetic and inviting, not breathy or drawn out. #Creatify #aimediabuyer #aitools #MarketingTips #D2CMarketingshow more

Jessica Collins
51,987 次观看 • 1 个月前
"I am a free bird in water. I can... move my entire body inside water, thanks to zero gravity. But my wheelchair keeps me caged on land." Geetha Kannan (42) is a wheelchaired person who LOVES swimming independently. Look at her swimming in the Nehru Stadium pool (video 1). She describes how she began doing exercises in water (hydrotherapy), & slowly began swimming 6 months ago. Now, she has bagged three gold medals in a state-level swimming competition in March. (Video 2) But she travels for 1.5 hours from Tambaram to Nehru Stadium every day for her two-hour swim, because #Chennai has barely two accessible pools out of seven public pools. There are around 1.5 lakh people with disabilities in the city and only 150 are trained in swimming lately. There are three Paralympic swimming gold medallists from the city, including Geetha. Not much money or efforts are needed to make a pool accessible. Ramps, tactile flooring, seating lifts, spacious toilets. Despite courts insisting on making every building accessible to all people, many spaces are not yet accessible, including pools, where disabled can feel 'abled'. "If there are more accessible swimming pools, more golds, silvers and bronzes will come to our city from many PwDs," said Geetha.show more

Padmaja J
32,303 次观看 • 2 年前