Mac owners don't miss this: MLX LM is now... integrated directly within Hugging Face 🤯 ⬇️ Run 4,400+ LLMs locally on Apple Silicon at max speed, no cloud, no wait.show more

Victor M
204,678 次观看 • 1 年前
LM Studio 0.3.4 ships with Apple MLX 🚢🍎 Run... on-device LLMs super fast, 100% locally and offline on your Apple Silicon Mac! Includes: > run Llama 3.2 1B at ~250 tok/sec (!) on M3 > enforce structured JSON responses > use via chat UI, or from your own code > run multiple models simultaneously > download any model from Hugging Face Video at 1x speed.show more

LM Studio
171,813 次观看 • 1 年前
🚨 BREAKING: Apple just acquired Meta's SAM team! On-device... segmentation now ships natively in macOS 26.x This is what it looks like when you run SAM3 on MLX locally. Every car. Every fire. Real-time. M2 laptop. No cloud.show more

Maziyar PANAHI
329,986 次观看 • 5 个月前
Today, we're shipping MLX support for TADA, our open-source... text-to-speech model, which means the entire pipeline (LLM, flow-matching, and decoder) can now run locally on any Apple Silicon device. We're seeing a 45% reduction in memory usage and a 10x speed-up when using it quantized. With these improvements, you can use TADA on-device for OpenClaw or any personal chatbot. If you own a MacBook, Mac Mini, or Mac Studio, record a 10-second clip of any voice, type any text, and get high-quality, natural and expressive speech in real-time. Completely offline, completely free.show more

Hume AI
24,684 次观看 • 5 个月前
I tested MTPLX v2 with QWEN 3.6 27B and... compared it with oMLX without cache on M5 Max and DGX Spark on vllm using nvfp4 model version. More details in 🧵 I've reached 82.8 tps of max decoding speed! 🔥 Custom Metal Kernel design specifically for this model and for Apple Silicon is just perfect! This is the way forward! Great job Youssof Al Toukhi Look at the website here! 👇 Here a website with recap, built with GLM 5.2 running locally 💪 First chart and preview from the website.show more

Ivan Fioravanti ᯅ
15,885 次观看 • 1 个月前
So right now basically anyone with a 16GB VRAM... card can go on Hugging Face and download a model which BEATS Claude Sonnet 4.5, all running locally 🤯 LOOK AT THE WATER PARTICLES?? What are they doing in these labs man What is this Unsloth AI UD quantization magic?? I'm running the lowest Q3 here, it oneshots everything 😭 Bros what is happening?show more

left curve dev
90,204 次观看 • 4 个月前
This is huge. (And we don't just throw that... around.) Starlight Mini—with local rendering—just launched in Video AI 7. That means no pay-per-render. No cloud. Just pure and powerful Starlight video upscaling and enhancing—right on your machine. It's available locally for high-powered Windows + NVIDIA systems for now, with more coming soon. Or, you can access it in the cloud for 3x cheaper than Starlight. Comment an old, low-res video of your own below and if selected, we'll run it in Starlight Mini.show more

Topaz Labs
40,404 次观看 • 1 年前
Cancelled ChatGPT -> Built JARVIS -> Pays $0 ->... it works offline + it's smarter than the $20/month version. No WiFi needed, no cloud, no API keys, no rate limits, no queues, no $20/month just to ask a server in Virginia for the weather. Just a local model running directly on the laptop hardware, voice activated, system integrated, controlling apps, answering questions, doing the work. Iron Man had JARVIS embedded in his suit, this guy has it embedded in his MacBook and it works on a plane, in a basement, on a remote cabin with zero signal. OpenAI is burning $700,000 a day on infrastructure to deliver something this guy runs for free. Anthropic charges $200/month for unlimited Claude access, microsoft built Copilot into every product they sell. This guy skipped all of it, downloaded a model and made his laptop the smartest device in the room. No subscription. No login. No internet. No data sent anywhere ever. The most powerful AI assistant on earth is now the one running locally on hardware you already own. ChatGPT charges you to think slower, he pays nothing and thinks alone, he made it himself.show more

Defileo🔮
154,607 次观看 • 4 个月前
Free NVIDIA GPU with 16 GB VRAM GPU for... Running Local LLMs! If you want to master local LLMs but you're waiting until you can afford a $1,500 GPU, you're honestly not going to make it. The open source AI ecosystem is moving way too fast for you to wait on your budget to catch up. Especially when you can build a bleeding edge inference engine from scratch right now, completely for free. You don't need a heavy local rig to start. Google is literally letting you use an enterprise grade NVIDIA Tesla T4 GPU for $0/hour. At standard cloud computing rates (~$0.20/hr), Google Colab’s 4 hour daily free tier hands you roughly $24 worth of data center tier GPU compute every single month. And most people just waste it. Let’s talk about the hardware you get access to for free. The NVIDIA Tesla T4 is an absolute workhorse: - Architecture: NVIDIA Turing (TU104) - VRAM: 16GB GDDR6 (320 GB/s bandwidth) - Compute: 320 Tensor Cores | 2560 CUDA Cores - Performance: 130 TOPS INT8 | 8.1 TFLOPS FP32 - Power: Sipping energy at a max 70W TDP This is the exact same hardware I used to run DeepMind's Gemma 4 26B A4B QAT MoE at a 250,000 context window without a single Out Of Memory (OOM) crash. If you have a web browser and 10 minutes, you have everything you need. I’ve put together a fully documented, cell by cell Google Colab notebook that teaches you exactly how to do this. Here is what the notebook actually teaches you: - How to provision an Ubuntu Linux environment with CUDA 13.0 and verify your driver stack. - How to pull the source code and compile the latest llama.cpp C++ binaries from scratch, specifically optimizing the build for your exact GPU using the -DCMAKE_CUDA_ARCHITECTURES=native flag. - How to directly download quantized local LLMs (GGUF format) straight from HuggingFace using the CLI. - How to manage 16GB VRAM limits, offload neural network layers to the GPU, and push massive context windows. Compile raw llama.cpp, ollama run a model, or spin up the LM Studio CLI. Pick whatever stack you are comfortable with. just start building. No hardware. No credit card. No excuses. Bookmark this post right now so you don't lose the tutorial. Even if you don't have time to run it today, you are going to want this workflow in your engineering toolkit. The link to the free Colab Notebook is in the comments below. Lemme know if you need more tutorials like this.show more

Alok
178,744 次观看 • 2 个月前
This is the fastest way I’ve found to send... huge files without compression across the world. It’s called Blip. A file transfer app that lets you send massive files directly from your device to someone else’s. No cloud upload first. No waiting for a download link. No file size panic. No video quality getting destroyed. No “upgrade your storage” nonsense. Just pick the file and send it. What it can do: • Send files of any size • Send entire folders • Resume interrupted transfers • Keep original quality • Work across long distances • Transfer between desktop and mobile • Use end-to-end encrypted transfers • Run on Mac, Windows, Linux, iPhone, iPad, and Android The wild part: Blip says it can handle files up to 99TB. Because it is not trying to be Dropbox. It does not make you upload the file to a cloud drive first, then force the other person to download it later. The receiver starts getting the file while you are still sending it. WeTransfer makes you upload. AirDrop only works nearby. Cloud drives turn sharing into storage management. Blip just sends the file. This is the file transfer app creators should have had years ago. Website:show more

Hasan Toor
13,941 次观看 • 3 个月前
Nvidia just put a $250,000 cloud workload on your... desk for $2,999 - and killed your $1,900/month AWS bill in the process You don't rent it, you don't manage it, you don't pay a single cloud bill - you just plug it in and let it eat the workloads you used to wire to AWS every month It looks like a small Mac mini, it's actually a full GB10 Grace Blackwell stack with 128GB of unified memory running models up to 200B parameters It's called DGX Spark, the consumer version of the rack Nvidia ships to OpenAI The reason Nvidia did this is simple Cloud GPU pricing is a tax on every developer building AI right now $1,900/month per seat, billions in margin flowing to AWS, Lambda, and CoreWeave Nvidia just cut themselves in by removing the cloud entirely Their solution is to skip the middleman, ship the rack to your desk, and let you keep every dollar of margin you used to wire to a hyperscaler This is much cheaper, faster, and you own the asset at the end But there is still a question nobody is answering yet, what happens to AWS, GCP, and Lambda when 500,000 developers move their inference back to a $2,999 box on their desk Also, technically you can stack four of these and run a 1.6 trillion parameter model locally for under $12,000 Even a single Spark out-performs the cloud subscription Anthropic engineers were running two years ago bookmark this, it pays back in 60 days 👇show more

ZEUS⚡️
85,803 次观看 • 3 个月前
China open-sourced a peanut-sized OCR that parses entire 100-page... PDFs in one shot.. It's called Unlimited-OCR. Only 3B params. Runs locally. Every other OCR tool chops your doc into pages and loses the thread. this one reads the whole thing in a single pass. → One-shot "long-horizon" parsing (32K context window) → Multilingual, out of the box → 93% on the standard parsing benchmark (+6 over baseline) → <0.11 error rate past 40 pages → Runs 100% locally on your own hardware → Works with Transformers, vLLM, SGLang, Docker, Ollama, llama.cpp Traditional cloud OCR (Textract, Google Vision, Azure Doc Intelligence) costs $1.50–$15 per 1,000 pages. This runs on your machine. For free. Forever. Baidu built it explicitly to push DeepSeek-OCR one step further. Already at 1.9M downloads on Hugging Face and most people have no idea it exists yet. 100% open source.show more

Superman
1,102,443 次观看 • 1 个月前
BREAKING: LLMs just learned to COMPUTE for real, it's... mean NO MORE GUESSING math. Chinese college kid Guo Hanjiang vibe-coded MiroFish in 10 days (23k+ GitHub stars, $4.1M from Shanda in 24h) - the AI swarm simulator that’s already printing. ByteDance (VolcEngine) dropped the nuclear upgrade: OpenViking - structured viking:// filesystem memory (L0 ultra-summary -> L2 full details) - agents now run 100+ steps with zero amnesia or hallucinations, 11.6k stars and climbing. Now this just dropped and the entire AI timeline is shaking. Startup Percepta embedded a full WASM virtual machine directly into Transformer weights. No more external Python sandboxes. No more hallucinations in exact tasks. The model streams raw machine code at 30,000+ tokens/sec on CPU, executes millions of steps, and solves the world’s hardest Sudoku via real backtracking + constraint propagation - 100% accurate, zero bullshit. They killed the Attention Bottleneck with Exponentially Fast Attention (HullKVCache + 2D heads + convex hull queries in log time). What used to die at 1k steps now flies. This is the bridge: System 1 intuition (normal LLMs) + System 2 deterministic logic (native code execution) in ONE brain. Agents won’t need tools anymore. Heavy simulations will run inside the weights. Check out: Now put it all together: MiroFish swarms + OpenViking infinite memory + Percepta native flawless compute = agents that can hardcore simulate millions of future scenarios, run perfect logic loops for days, and predict events/markets/reality with god-tier accuracy. No drift. No bullshit. Just pure foresight. This combo will change everything, imo. The era of predictive super-agents that actually print the future is here. We’re watching this one closely. Save this combo.show more

slash1s
156,699 次观看 • 5 个月前
i just ran Google's brand new Unsloth Gemma4 12B... dense GGUF on my RTX 4060 using llama.cpp + CUDA 13.2 21 tokens per second. on a budget consumer GPU. locally. no API. no cloud. no subscription. and the benchmarks are absolutely cooked # first let's talk architecture because this is genuinely different every multimodal model you've used has a frozen vision encoder + frozen audio encoder + LLM backbone glued together Gemma 4 12B is different it's a single decoder only transformer. that's it. vision? raw 48×48 pixel patches → one matmul → projected directly into the LLM audio? raw 16kHz signal sliced into 40ms frames → linear projection → same LLM input space no encoder tax. no latency penalty. no fragmented memory to put the encoder savings in perspective: old Gemma 4 26B approach: - 550M param vision encoder (frozen) - 300M param audio encoder (frozen) - LLM backbone Gemma 4 12B: - 35M param vision embedder (a single matmul) - no audio encoder at all - LLM backbone handles EVERYTHING 550M → 35M for vision alone. that's a 15x reduction this is why the gemma-4-12b-it-Q4_K_M.gguf is just 6.6 GBs!!! and it has 256K native context context # Benchmarks: AIME 2026 (math olympiad): 77.5% GPQA Diamond (expert science): 78.8% LiveCodeBench v6 (real code): 72% Codeforces ELO: 1659 MMLU Pro: 77.2% MATH-Vision: 79.7% BigBench Extra Hard: 53% inference → llama.cpp, LM Studio, vLLM, SGLang llamacpp flags: -m "gemma-4-12b-it-Q4_K_M.gguf" -ngl 99 -c 8000 -v --port 8080 Available on huggingface now! Link belowshow more

Alok
281,007 次观看 • 3 个月前
🚨 One photo of your face. That's all someone... needs to become you on a live video call. In real time. Right now. The tool is free and open source. It's called Deep-Live-Cam. One image. One click. You become anyone on a live webcam feed. No training. No datasets. No waiting. Instant. Your face. Your expressions. Your mouth movements. All stolen from a single photo. Here's what this thing does: → Upload one photo of any face → Turn on your webcam → You are now that person. Live. In real time. → It matches your pose, your expressions, even your lighting → Mouth masking so the swapped face moves its lips when you talk → Multi-face mapping. Swap different faces on different people in the same call. → Virtual camera output. Plug it into Zoom, Google Meet, Teams. Nobody knows. → Works on NVIDIA, AMD, Intel, and Apple Silicon Here's the part that should terrify you: Your boss could be on a Zoom call with someone wearing your face right now. A scammer could call your parents looking exactly like you. A stranger could take your LinkedIn photo and become you in a video meeting. IShowSpeed's reaction when he saw it: "What the F**! This shit is crazy!" SomeOrdinaryGamers: "That's fucking freaky dude... that's so wild." This was the #1 trending repo on GitHub the day it launched. 1,600 stars in 24 hours. 80K+ stars today. No one is ready for what this means. And it's already out there. 100% Open Source.show more

Nav Toor
305,709 次观看 • 5 个月前
Microsoft made 100B parameter models run on a single... CPU. bitnet.cpp: The official inference framework for 1-bit LLMs. The math behind 1-bit LLMs is what makes them revolutionary. Traditional LLMs use 16-bit floating point weights. Every parameter is a number like 0.0023847 or -1.4729. When you run inference, you multiply these floats together. Billions of times. That's why you need GPUs, they're optimized for floating point matrix multiplication. BitNet b1.58 uses ternary weights: {-1, 0, 1}. That's not a simplification. That's a fundamental change in the math. When your weights are only -1, 0, or 1: → Multiply by 1 = keep the value → Multiply by -1 = flip the sign → Multiply by 0 = skip entirely Matrix multiplication becomes addition and subtraction. No floating point operations. No GPU required. This is why bitnet.cpp achieves: → 2.37x to 6.17x speedup on x86 CPUs → 1.37x to 5.07x speedup on ARM CPUs → 71.9% to 82.2% energy reduction on x86 → 55.4% to 70.0% energy reduction on ARM The speedups scale with model size. Larger models see bigger gains because there are more operations to simplify. A 100B parameter model running at human reading speed (5-7 tokens/second) on a single CPU. That's not optimization. That's a different paradigm. Why 1.58 bits? Because log₂(3) ≈ 1.58. Three possible values = 1.58 bits of information per weight. The key insight: These models aren't quantized after training. They're trained from scratch with ternary weights. The model learns to work within the constraint. No precision loss. No quality tradeoff.show more

Tech with Mak
23,036 次观看 • 4 个月前
Apple is now sending urgent alerts directly to iPhones... targeted by government-grade spyware attacks. These "Apple threat notifications" warn specific individuals being singled out by sophisticated mercenary spyware like Pegasus. These alerts arrive via email, iMessage, and appear as a banner on the user’s Apple ID webpage. This reveals that state-funded digital surveillance remains a serious global reality in 2026. While average users are not at risk from these very expensive attacks, it is a critical safety feature for high-profile journalists, activists, and politicians. Upon receiving an alert, high-risk users must immediately enable Apple's Lockdown Mode. Relying on standard security settings is no longer enough when state-level resources are deployed to hack a mobile device.show more

Anonymous
59,012 次观看 • 20 天前
Right now, you may not have access to models... like GPT‑5.6 Sol, GPT‑4.6 Terra, GPT‑5.6 Luna, Claude Mythos 5, or Claude Fable 5. But you can run something surprisingly powerful today, locally, and completely free. in the next 10 mins on your 8 GB VRAM gaming laptop. Gemma 4 26B A4B QAT (MoE) delivers strong performance on a standard 8 GB VRAM GPU using Ollama, with no API, no usage limits, and no external dependencies. Out of the box, it reaches around 20 tokens per second without any optimizations. Only one command in your terminal: Ollama run gemma4:26b This means: Full offline capability (privacy by default) Zero recurring cost Competitive performance for many real world tasks Fast enough for interactive use on cheap consumer hardware If you're waiting for cutting edge cloud models, you're missing what is already practical today: a capable, local LLM that runs entirely on your own machine.show more

Alok
65,387 次观看 • 2 个月前
🇺🇸 HERE'S WHAT NO ONE IS TELLING YOU ABOUT... THE 2026 MIDTERMS! The GOP is doing something no party has ever done before. A full presidential-style convention. For the midterms. Let that sink in. While Republicans are building a stage, firing up a crowd, and showing America exactly what they stand for, Democrats are sitting in the locker room with no game plan and no coach. What are they going to run on? Grocery prices that crushed working families? The open border they cheered for? More government control dressed up as "progress"? Go ahead, step to the podium and make the case for socialism. We'll wait. This convention isn't just a rally. It's a statement. The Republican Party is playing offense while the left is still trying to figure out what team they're on. Energy wins elections. Momentum wins elections. And right now the GOP has both. The machine is warming up. The base is ready. And come November 2026, this could be the move that changes everything. History is being made in real time. Don't miss it!show more

Bill Mitchell
99,951 次观看 • 2 个月前
The of #Ethereum is finally here! Craze Token 🔥... FAIRLAUNCH LIVE NOW!🚀 BUY $CRAZE NOW! SC already filled under 1 hour, 25 $ETH filled on PinkSale NOW, still 24 hours to go! So don't miss this PreSale, LIVE NOW until December 10th, 18:00h UTC! DYOR 👇 🌐: 💬: ❓WHY JOIN CRAZE 🌟 Gasless Launches: No more crazy fees! 🌟 Revenue Sharing: 50% of platform revenue. 🌟 Future-Proof: Built for the Ethereum bull run. 🌟 Transparent & Secure: Fully audited and KYC verified. 🔒 SAFE & SECURE ✅ Audit: View Audit ✅ KYC: Verified by Assure DeFi 🔥 Be part of the gasless revolution – don’t miss this chance to change the game and start the #bullrun on #ETH! 🔥show more

Big Bunny Crypto
11,426 次观看 • 1 年前
A preventable tragedy on Veppur Coot Road. A biker... died yesterday evening in a collision with a bus at the Kattumayilur junction - a spot locals say has been a known risk for a while now. Why? Vehicles entering from Salem reportedly speed through this junction because there's no barricade to slow them down - unlike the Kandapankurichi–Nallur junction nearby, which does have one. 🚧 Residents are now asking: why hasn't the same safety measure been extended here? 📍 Kattumayilur Junction, Veppur Coot Road If you're near this stretch, please slow down at unmarked junctions - and if you know the local authorities, this is worth flagging directly. 🙏 RoadSafetyIndia TamilNaduRoads BarricadeNowshow more

Rajneeti Tadka 🌶️ 🔥
38,071 次观看 • 1 个月前