Microsoft just a 1-bit LLM with 2B parameters that... can run on CPUs like Apple M2. BitNet b1.58 2B4T outperforms fp LLaMA 3.2 1B while using only 0.4GB memory versus 2GB and processes tokens 40% faster. 100% opensource.show more

Shubham Saboo
260,049 次观看 • 1 年前
Microsoft is testing a new feature in Windows 11... called Low Latency Profile that gives a short boost to your CPU for 1 to 3 seconds when you open apps or menus. Early tests show Microsoft apps like Edge and Outlook open up to 40% faster, while the Start menu and other menus can load up to 70% faster, with CPU use jumping high for a few seconds then dropping right back down. Microsoft says this is normal because Apple does it in macOS and Linux, plus phones use the same type of quick CPU boosts for better speed. The feature is only in early Insider testing now, so it is not out for everyone yet.show more

Pirat_Nation 🔴
145,459 次观看 • 3 个月前
Microsoft's new Outlook takes more than 10 seconds to... open an email via notifications, while Outlook Classic opens instantly on Windows 11! The funniest part? If I ignore the notification, open Outlook manually, find the email myself, and click it, I can finish faster than the notification flow. It's all because Outlook for Windows is literally Microsoft Edge running Outlook .com in a browser window. New Outlook runs through WebView2 with around 10 separate processes: GPU process, service worker, utility processes, manager, and more. It uses roughly 490MB to 636MB RAM while idle, compared to Outlook Classic sitting around 117MB to 148MB. CPU usage is also worse: around 4% idle on new Outlook versus under 1% on Classic in my testing. Microsoft shut down Mail and Calendar, keeps pushing enterprises toward this web wrapper, and still wants people to believe it is the future of email on Windows. SHAME!show more

Windows Latest
101,988 次观看 • 2 个月前
THAT $70 "RUN YOUR OWN LLMS" PI KIT CAN'T... RUN A SINGLE LLM. IT'S A VISION CHIP WITH NO RAM. that clip sells a raspberry pi 5 in a slick case with an ai accelerator and the caption "your own llms." clean build, fun kit. the claim is where it breaks. the fine print: the popular $70 pi ai kit uses a hailo-8l, 13 tops. it's built for vision, object detection and image processing, and it has no memory of its own. so it cannot run large language models. full stop the board that actually can is a different one: the newer ai hat+ 2, hailo-10h, 40 tops, with 8gb of dedicated ram. that's $130, not $70 and even that runs only tiny models. llama 3.2 at 1b, qwen 2.5 at 1.5b, deepseek r1 at 1.5b. edge llms live in the 1-7b range, against cloud models at 500b to 2 trillion so the honest pitch: for $130 you can run a very small language model on a pi, slowly, as a fun learning project. that's real and it's cool. "your own llms" on a $70 vision kit is not. why this keeps happening: "ai kit" and a big "tops" number sell. tops sounds like intelligence. but tops measures vision-style math, not whether the chip has the memory to hold a language model. the spec that matters for llms is ram, and the cheap kit has none. the honest caveats, both ways: the $70 kit is genuinely great, just at vision. cameras, object detection, that's its job the $130 hat really does run small llms locally, which a pi couldn't do at all two years ago. that's progress "small" is the load-bearing word. don't expect gpt at home on a pi the takeaway: before you buy a kit because the caption says llm, check two numbers. not the tops. the ram, and the size of the model it can actually load. no 70-dollar miracle, no gpt in a pi case, no tops number that means what you think. save this before you buy the wrong kit for the word on the box.show more

RetroChainer
11,100 次观看 • 1 个月前
1/5 𝗧𝗵𝗲 𝗪𝗼𝗿𝗹𝗱'𝘀 𝗙𝗶𝗿𝘀𝘁 𝗢𝗻-𝗰𝗵𝗮𝗶𝗻 𝗧𝗼𝗸𝗲𝗻𝗶𝘇𝗲𝗱 𝗣𝗼𝗿𝘁𝗳𝗼𝗹𝗶𝗼𝘀 (𝗢𝗧𝗣𝘀) 𝗯𝘆... 𝗜𝗫 𝗦𝘄𝗮𝗽! 🚀 #OTPs provide access to a diversified portfolio of assets managed by an investment professional. These assets range from digital assets like crypto and tokens to real estate, private equity, and more. Essentially, its a tokenized fund that can be quickly set up at minimal cost. Investors can begin with just 1 USDT, and all transactions are conducted on-chain using stablecoins. After funding closes, investors receive real-world asset tokens (fund shares) directly in their wallet and can immediately trade on the secondary market (IXS DEX). It’s a cost-effective, fast, and flexible alternative to traditional funds, with no long lock-ups.show more

IXS
11,094 次观看 • 1 年前
1/3 🤖 Meet AgentOS: A Token-efficient, Microkernel AI agent... with on-device model routing across CLI, Web UI, and chat. A local router reads every message on your device and sends it to the cheapest model that can still do the job well. You stop overpaying for AI. ⚡ What makes it stand out: ■ Smart on-device model routing across 20+ providers ( bankrbot LLM Gateway, OpenRouter, OpenAI, Anthropic, Ollama, and more). ■ Persistent local memory that survives restarts. ■ A layered security sandbox (Standard, Strict, Locked). ■ 37 built-in skills plus MCP, loaded only when a task needs them, and more ■ One unified gateway for CLI, Web UI, Slack, Telegram, Discord, and more Remember this: AgentOS has been integrated with the Bankr LLM Gateway since day one. That means any Bankr user with Bankr API key can start using AgentOS in minutes. Let the router cook.show more

AgentOS
203,414 次观看 • 1 个月前
YOU CAN NOW CONTROL ALL YOUR HERMES AGENTS FROM... ONE INTERFACE. 🚨 Someone just built Hermes Workspace - the most complete GUI for Hermes Agent. Instead of jumping between tools, everything lives in one place. What's inside: > Chat, terminal, memory browser, and skills manager > Run unlimited agents with 1 orchestrator > Kanban board to track everything your agents are doing > 2,000+ skills to install without leaving the workspace > Works with any LLM provider Stop managing agents manually and start running them like a pro 👀 100% Free. Open Source.show more

Oliver Prompts
19,409 次观看 • 1 个月前
1-bit Kimi K3 performs at Opus 5 level on... 3D physics! We ran our Atomic Chat quant of Kimi K3 locally on 4x B200 against three cloud models and gave them all the same task, to build a giant anvil drop test as a single HTML file with real physics Outputs: K3 1bit (local): 15.8K tokens, $0 API cost Kimi K3 (API): 15.3K tokens, $0.30 API cost Opus 5: 22.8K tokens, $0.77 API cost GPT 5.6: 14.5K tokens, $0.72 API cost All four got the physics right. But only Kimi made a working winch. The drum turns and the chain drags the flat car off the pad. Opus 5 drew the most detail, road markings and sparks on the hit. And you can run a model at this level on your own box now. That still feels insane to usshow more

atomic.chat
53,845 次观看 • 29 天前
A peanut-sized Chinese model just dethroned Gemini at reading... documents. GLM-OCR is a 0.9B parameter vision-language model. It scores 94.62 on OmniDocBench V1.5, ranking #1 overall. For context, it outperforms models 100x its size. 100% open-source. It works in two stages. 1. A layout engine detects every region in a document. 2. Each region gets read in parallel. The model predicts multiple tokens per step instead of one. That's what makes it so fast at small size. It handles things most OCR tools struggle with: > Complex tables and nested layouts > Handwritten text and stamps > Math formulas and code blocks > Mixed image-and-text documents You can run it locally through Ollama. It fits on edge devices with limited compute. Every expensive OCR API just got a free competitor.show more

AlphaSignal
92,071 次观看 • 4 个月前
A peanut-sized Chinese model just dethroned Gemini at reading... documents. GLM-OCR is a 0.9B parameter vision-language model. It scores 94.62 on OmniDocBench V1.5, ranking #1 overall. For context, it outperforms models 100x its size. 100% open-source. It works in two stages. 1. A layout engine detects every region in a document. 2. Each region gets read in parallel. The model predicts multiple tokens per step instead of one. That's what makes it so fast at small size. It handles things most OCR tools struggle with: > Complex tables and nested layouts > Handwritten text and stamps > Math formulas and code blocks > Mixed image-and-text documents You can run it locally through Ollama. It fits on edge devices with limited compute. Every expensive OCR API just got a free competitor.show more

Jafar Najafov
13,630 次观看 • 4 个月前
🚨 Do you understand what Claude just quietly dropped... while everyone was distracted? 1 million tokens. Let me explain what that actually means because the number alone doesn't hit right. > A senior engineer joins a company and spends 3 to 6 months just reading code.. Understanding how things connect. Learning where the bugs hide. Why that one file nobody touches exists. It takes months because a codebase is massive and human memory is small. > Claude just loaded the entire thing in one prompt. 30 seconds. Every file, Every function, Every line. All of it. Sitting in memory like it's been working there for years. And it scored highest among every single frontier model. Not GPT.. Not Gemini, Nobody. > Yesterday Amazon's AI nuked production because it couldn't see the full picture - it made a decision with partial context and deleted everything. Today an AI can hold 1 million tokens of context at once. That's the fix. That's the "before and after" moment for AI coding. > 600 images in one request. Entire PDFs. Full repos. And they dropped it on a Friday on all plans like it was a patch note. The scariest AI updates aren't the ones with press conferences. They're the ones that drop in a tweet at 6pm and change everything by Monday morning.show more

Tuki
206,309 次观看 • 5 个月前
50% more context unlocked for Qwen 3.8 27b Q4_K_XL... dflash 2 on a single RTX 4090 (24 GB VRAM) I found a hidden VRAM tax in llama.cpp. By combining my custom 2 bit DFlash 2 drafter with one overlooked server flag, I just unlocked another +80,000 tokens of context. Qwen3.8-27B is now running a massive 250,000 context at 75 tokens/s on a single RTX 4090. Here is the secret: By default, `llama-server` reserves massive chunks of your VRAM to handle multiple concurrent users (batching). If you are running a single user session, you are bleeding memory for features you aren't using. By passing the `--parallel 1` flag, you force the engine to dedicate 100% of your 24GB VRAM buffer to a single user. When we combine the VRAM saved by our Q2_K 2-bit drafter with the VRAM saved by `--parallel 1`, the context ceilings absolutely explode: Note: all benchmarks carried out with a massive 28k prompt. Ubuntu 22. ### THE NEW 24GB PHYSICAL LIMITS (Single RTX 4090): # 1. The "Repo Swallower" (Q4 KV Cache): - Context: 250,000 tokens (Up from 170k!) - Speed: 73.66 t/s decode | 1,608 t/s prefill - Peak VRAM: 23.8 GB # 2. The "High-Precision SWE" (Q8 KV Cache): - Context: 150,000 tokens (Up from 100k!) - Speed: 75.01 t/s decode | 1,667 t/s prefill - Peak VRAM: 23.9 GB # 3. The "Pristine Attention" (Unquantized FP16 KV): - Context: 90,000 tokens - Speed: 80.58 t/s decode | 1,699 t/s prefill - Peak VRAM: 23.92 GB ### HOW TO RUN THE 250K GOD STACK TODAY: (Requires PR #27342 + my Q2_K Hugging Face drafter) llama.cpp flags: ./build/bin/llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf -md Qwen3.8-27B-DFlash2-Q2_K.gguf --spec-type draft-dflash --spec-draft-n-max 3 -c 250000 -ngl 99 --parallel 1 --port 8080 -ctv q4_0 -ctk q4_0 We are pushing a quarter million tokens of context with speculative DFlash 2 decoding at 73 tokens/second on a single consumer gaming GPU. I dropped my custom 2 bit Hugging Face GGUF links, visual performance graphs, and the PR #27342 build instructions in the replies below. If you own a single RTX 3090 or 4090, it is officially time to cancel your API subscriptions and let local silicon eat the cloud. how much monthly API spend does an optimized 4090 rig like this actually replace for you?show more

Alok
39,189 次观看 • 8 天前
Right now, you may not have access to models... like GPT‑5.6 Sol, GPT‑4.6 Terra, GPT‑5.6 Luna, Claude Mythos 5, or Claude Fable 5. But you can run something surprisingly powerful today, locally, and completely free. in the next 10 mins on your 8 GB VRAM gaming laptop. Gemma 4 26B A4B QAT (MoE) delivers strong performance on a standard 8 GB VRAM GPU using Ollama, with no API, no usage limits, and no external dependencies. Out of the box, it reaches around 20 tokens per second without any optimizations. Only one command in your terminal: Ollama run gemma4:26b This means: Full offline capability (privacy by default) Zero recurring cost Competitive performance for many real world tasks Fast enough for interactive use on cheap consumer hardware If you're waiting for cutting edge cloud models, you're missing what is already practical today: a capable, local LLM that runs entirely on your own machine.show more

Alok
65,387 次观看 • 2 个月前
This Chinese developer launched Llama 70B locally on a... MacBook on a plane and for a full 11 hours without internet ran client projects. He was sitting by the window on a transatlantic flight with a MacBook Pro M4 with 64 GB of memory. WiFi on board cost $25 for the flight. He declined. No cloud API, no connection to Anthropic or OpenAI servers, no internet at all. Just a local Llama 3.3 70B on bf16 and his own orchestrator script. The model runs through llama.cpp. Generation speed, 71 tokens per second. Context around 60,000 tokens. Memory usage, 48.6 GiB out of 64. Battery at takeoff, 3 hours 21 minutes. And he gave the orchestrator this system prompt before takeoff: "You are an offline orchestrator running on a single MacBook. There is no network. The only resources you have are local files in /Users/dev/work, the Llama 70B inference server at localhost:8080, and a battery budget of 3 hours 21 minutes. Process the queue at /Users/dev/work/queue.jsonl (one client task per line). For each task: draft → run local evals → save artefact to /Users/dev/work/done/. Save context checkpoints every 12 tasks so you can resume after a battery swap. Stop only on empty queue or when battery drops below 5%." So the system knows exactly what resources it is running on. It knows it has no connection to the outside world for the next 11 hours. It knows it has finite memory and a finite battery. It knows the human will not intervene until the plane lands. The system runs in 1 loop. Takes a task from the queue, runs it through inference, saves the artifact, writes a checkpoint. Task after task, just like that. And only when the battery drops below 5% does the orchestrator automatically pause, waits for the laptop to switch to the backup power bank, and continues from the last checkpoint. Here is what the system actually writes in his log during the flight: "saved context checkpoint 8 of 12 (pos_min = 488, pos_max = 50118, size = 62.813 MiB)" "restored context checkpoint (pos_min = 488, pos_max = 50118)" "prompt processing progress: n_tokens = 50 / 60 818" "task 37016 done | tps = 71 s tokens text → /Users/dev/work/done/proposal_westside.md" Outside the window, clouds, blue sky, and no WiFi. On the tray, 1 MacBook, an open terminal on 2 screens, and an inference server on localhost. From what I have observed, this is the cleanest offline AI workflow I have seen in the past year: 11 hours of flight, $0 for WiFi, and the entire client queue closed before landing.show more

Blaze
1,841,161 次观看 • 4 个月前
GPT-5.6 Sol is unbelievably good at creating and editing... videos. It can do motion design, product demos, and animations like this one I made by simply giving it a screen recording. GPT 5.6 has the best design taste and significantly outperforms Fable, which relies heavily on repetitive design patterns. To help you experiment with video editing on it, we just launched a collection of 100 ready-to-use skills that show what’s possible and help you get started with video editing using GPT-5.6. These skills can create anything from motion graphics launch videos for your product to a 3B1B-style science explainer video. You can also use them to edit existing videos: add captions, generate motion graphics, create voiceovers, redesign visual styles, translate into new languages, and much more. If you want access to the full library, comment “VIDEO SKILLS” and I’ll share it with you. (You'll have to follow me so I can DM you.)show more

Akash Anand
515,487 次观看 • 1 个月前
Create a short film like this in just 1... minute with GPT Image 2.0 + Seedance 2.0. GPT Image 2.0 can naturally combine multiple photos into one single image, while Seedance 2.0 can use that image as a reference to automatically separate the scenes, generate a coherent video sequence, and add suitable background music. This workflow greatly improves the overall creative efficiency. When using this method, simply provide the merged image as a reference for Seedance 2.0 and briefly describe each scene with a simple prompt. This can significantly increase the success rate of the final video. All of the above was created on GPT Image Prompt: Seedance Prompt:show more

Midjourney Sref and prompt Library
40,572 次观看 • 4 个月前
Researchers made KMeans 200x faster. And the new technique... also beats approaches like cuML and FAISS. Flash-KMeans is an IO-aware implementation of exact KMeans that redesigns the algorithm around modern GPU bottlenecks. By attacking the memory bottlenecks directly, Flash-KMeans achieves: - 33x speedup over cuML - 200x speedup over FAISS This speedup comes from how it moves through GPU memory. Standard KMeans runs in two steps, and both are bottlenecked by reads and writes to GPU memory: 1) The first step matches every point to its nearest centroid. Standard KMeans computes the full point-to-centroid distance matrix, writes it out to GPU memory, then reads it back to find each nearest centroid. That write-then-read round trip is the bottleneck. Flash-KMeans combines the distance calculation with the nearest-centroid step, so the result is computed on-chip and the full matrix is never written out. 2) The second step recomputes each centroid by averaging the points assigned to it. Standard KMeans has thousands of threads writing into the same centroid slots at once, so they stall waiting for their turn. Flash-KMeans sorts points by cluster first, turning scattered writes into sequential reductions that read and write memory in one efficient pass. Using these two optimizations at the million-scale, Flash-KMeans completes a standard KMeans iteration in a few milliseconds. The video below depicts this in action. Several reasons why this is important: KMeans has always been an offline primitive. Something you run once to preprocess data and move on. These speedups make the approach viable in several runtime-critical systems. ↳ Vector indices like FAISS use KMeans to build search indices. Faster KMeans means you can re-index dynamically as data changes. ↳ LLM quantization methods need KMeans to find optimal weight codebooks, per layer, repeatedly. What takes hours could now take minutes. ↳ MoE models need fast token routing at inference time. Flash-KMeans makes it viable to run this inside the inference loop, not just in preprocessing. I have shared the paper in the replies. That said, memory is the real constraint Flash-KMeans solves, and the problem is not just limited to clustering. The vectors a RAG system stores after indexing create similar bottlenecks. I wrote a detailed walkthrough recently on cutting this vector memory by 32x with binary quantization, querying 36M+ vectors in a few milliseconds. Read it below.show more

Avi Chawla
89,234 次观看 • 2 个月前
🚨 Insider wallet alert: The same wallet that nailed... the exact timing of the US–Iran strike with a $500K bet just dropped $700K on a ground invasion of Iran. I built a step‑by‑step guide to track and copy‑trade insider wallets like this on Polymarket. Free for 24 hours: - Comment: '' Market '' - Like + Retweet - Follow Theo (so I can DM you) All you need: Claude + laptop + 1 hour/day. While you’re still “analysing charts,” this wallet is using AI and betting millions before headlines hit CNN. They don’t guess. They don’t hesitate. They move first—and win. You can keep watching from the side-lines… or learn how to move with them. Don't forget to bookmark and follow me Theoshow more

Theo
27,827 次观看 • 4 个月前