mlx_lm.server + Qwen3 Coder Next 6bit + OpenCode +... M3 Ultra = a pretty capable and very fast local coding setup. Not sped up:show more

Awni Hannun
66,668 次观看 • 7 个月前
MLX MiniMax 2.5 running LOCALLY on a single M3... Ultra 512GB! Writing a poem on LLMs at 6bit quantization! 🔥 Let's start some coding, context and distributed tests! Generation: 40.2 tokens-per-sec Peak memory: 186 GBshow more

Ivan Fioravanti ᯅ
226,466 次观看 • 6 个月前
someone submitted a pr that added "beast prompt" to... opencode for gpt models it makes 4.1 actually good at tool calling look how fast this thing goes, video not sped upshow more

dax
80,952 次观看 • 1 年前
this is not sped up. Cognition SWE-1.6 is just... moving fast enough that the interaction feels like a different product category the nice thing about a model this fast is you stop treating it like a task runner and start treating it like something you can actually riff with!show more

nader dabit
12,807 次观看 • 3 个月前
Next task - building a training program generator. Pretty... straightforward using Claude Sonnet 4.6. Will complete, refine, set up payment system, and test. I don't intend this project to go live - I'm just learning; zero background in coding or software development.show more

GuruAnaerobic
15,597 次观看 • 6 个月前
Quick update on the water situation 💦 M3 Ultra... and Titan (RTX6000 Pro) seem to have recovered with little to no visible damage. The main issues are with my MacBook which is in service and Titan CPU temperatures being above avg when idling (58C up from 35C prior to water incident). Anyways, here is a video of MLX-VLM serving Qwen3-4B-Instruct on Titan (~300 tok/s) to do autocomplete and git commit message generation completely locally via Zed IDE.show more

Prince Canuma
17,065 次观看 • 3 个月前
Making OpenCode as lean as Pi agent? Just trimmed... 25k out of OpenCode's system prompt (from 30k to 4-5k tokens) How? Just disable skills and get rid of massive skill definition bloat. Who needs skills anyway? Just kidding, this is the not the way. It makes the agent lame and defeats the point of using one. But it sets a precedent: Find a way to use skills without their definitions pre-loaded into the system prompt every single turn. Another interesting stuff: Upon testing this temporary "no skill setup" with two of hottest OpenCode Zen free models, Mimo V2.5 vs DeepSeek V4 Flash: One thinks more and talks less One thinks less and talks more Check the video to see which is which If you made it here, I'm finding a way to leanest OpenCode setup that I can get I simply don't believe that OpenCode can't be as lean as Pi Upon tinkering, I made a plugin that temporarily extracts the system prompt while I test, and noticed the hundreds of definitions in it from my .agents/skills directory which is shared across all my coding agents (Cursor, Antigravity, Claude, etc.) Of course disabling skills is not the answer, but it just proved that there is a way to strip the system prompt of these massive skill defs Aside from the system prompt hierarchy that injects confusion imo if you have a conflicting and redundant AGENTS.md which I discovered upon digging into OpenCode's source code Apparently it has prompt.ts/system.ts/instruction.ts/llm.ts and loads base .txt prompts based on model family (claude/gpt-o/gpt-5/codex/gemini/others) that all work together to make OpenCode aware of who it was and how it should use tools and become a "coding agent" Gotta find the most minimal mix that fits right into my workflow Make OpenCode as lean as Pi? We'll see. All inshow more

raymel 👋
37,939 次观看 • 3 个月前
Grok showing up inside Warp is a pretty big... developer move. One of the fastest-growing AI-powered terminals now has native Grok integration. Developers can link their 𝕏 Premium + Grok account, switch to grok-build-0.1, and start coding directly inside Warp. That is a very different kind of “AI assistant.” Grok / Writer: Annette, Designer: Jannéshow more

Mario Nawfal
41,774 次观看 • 2 个月前
We're getting to the point where AI isn't just... helping us write code; it can help launch the business behind that code too. One of the more interesting examples I've come across is **Lovie** Instead of jumping between legal websites, paperwork, and incorporation services Lovie lets founders form a US LLC or Corporation directly from AI coding tools like **Claude Code, Cursor, and Windsurf** through its MCP-first workflow. If you're already building your product inside your AI coding assistant, the setup process can become part of that same workflow instead of another tab to manage. For founders, indie hackers, and developers shipping fast, that's a pretty compelling idea. Flat **$29/month**, company formation in minutes, and it fits naturally into the AI tools many developers are already using. Feels like another step toward AI becoming a true operating system for builders, not just a coding assistant.show more

Parul Gautam
41,751 次观看 • 1 个月前
A CHINESE GUY PUT 4 MINISFORUM MS-S1 MAX MINI... PCs IN HIS BEDROOM AND TURNED THEM INTO A 24/7 AI AGENT CLUSTER. TOTAL POWER BILL: ABOUT $44/MO. each box is a tiny local AI workstation built around the Ryzen AI Max+ 395. around $3,000 per unit gets him 128GB of unified memory, 2TB storage, dual 10GbE, and up to roughly 96GB usable as VRAM on Linux. one MS-S1 Max can already run serious open models without touching the cloud. Qwen3-Coder 30B for fast coding, Llama 3.3 70B for heavier reasoning, and larger research models overnight when speed matters less than free inference. four boxes in one room changes the whole game. he is not opening a chatbot, paying for every loop, or shutting agents down before sleep. this is private infrastructure that keeps working even when he is offline. the agents can sort inboxes, review code, summarize documents, monitor feeds, prep meetings, and read papers overnight. on cloud APIs, that kind of always-on stack can easily burn $800 to $1,200 a month if used aggressively. his setup is roughly a $12,000 hardware spend, but the monthly cost is basically electricity. a rack, a switch, a NAS, a small monitor, and four tiny MS-S1 Max boxes turning a bedroom corner into a private inference factory. this is what AI looks like when it stops being rented and starts becoming something you own.show more

Gipp 🦅
24,836 次观看 • 2 个月前
STARLINK LANDS IN BHUTAN—HIGH-SPEED INTERNET FOR ALL! Bhutan just... got a major digital boost—Starlink is now live! Starlink’s blazing-fast internet is connecting even the most remote spots. - Speeds That Stun – Up to 220 Mbps with ultra-low latency. - Globally Connected – Bhutan joins a growing list of nations unlocking next-gen connectivity. - Pocket-Sized Power – Starlink Mini brings high-speed internet anywhere—just unpack and go. - Fully Approved – Local regulators gave the green light, to integrated payment systems. Bhutan just unlocked next-level connectivity—stream, work, and game from anywhere! Source: Starlink, Latestly, SpaceXshow more

Mario Nawfal
25,472 次观看 • 1 年前
EUROPE | UK 2026 🌍 We’re thrilled to return... next summer with special guests QOTSA and Acid Bath. Sign-up at for first access to tickets starting Tuesday, 16 Sept at 12PM local time. General on sale begins Friday, 19 Sept at 12PM local. Plans are also underway for a very special event to take place following these shows. Stay tuned…show more

System Of A Down
129,997 次观看 • 11 个月前
you can run claude code inside antigravity completely Free... with zero credit card and no rate limits 😳 use openrouter’s free models + antigravity. no anthropic bill. no paid api keys. takes 10 minutes to set up. what you get during this setup: - full claude code agent experience - strong coding models (including deepseek-r1, qwen2.5-coder, llama-4, grok-4 free tier) - antigravity’s clean workspace and sandbox - unlimited usage (as long as you stay on free models) - easy model swapping - zero cost full setup guide (100% free): step 1: install antigravity -go to and install it -create a new workspace step 2: install claude code - inside antigravity, install the claude code extension from the marketplace - open the built-in terminal step 3: create openrouter free account -go to - sign up with google (no card needed) - go to keys and create a new api key step 4: set the environment variables -in antigravity terminal run: export ANTHROPIC_API_KEY=sk-or-xxx export OPENROUTER_API_KEY=sk-or-xxx step 5: launch claude code with free model -run this command: claude-code --model deepseek/deepseek-r1:free or try: qwen/qwen2.5-coder:free if you already have antigravity? skip straight to step 2. after 10 minutes you’ll have a full agentic coding setup running for free. this is currently one of the cheapest ways to run serious coding agents in 2026. bookmark this before they limit the free models.show more

painn
32,057 次观看 • 2 个月前
It's been an intense 10 months since we kicked... off hardware engineering and I couldn't be prouder to have delivered a Pendant to our first customer along with Dan Siroker Alyssa Atkins Stammy! Very few hardware startups go from concept to customer that fast when using tooled mass production methods capable of producing 10s of thousands of units Much, much more to come. Tune in to Limitless next week for a full video of the deliveryshow more

jer
42,310 次观看 • 2 年前
Cybertruck is what happens when the truck stops pretending.... Plenty of trucks are designed to look badass. Very nice. Very angry headlights. Then Cybertruck shows up with stainless steel panels taking bullets, sledgehammers, and saws like someone is stress-testing a bank vault. It may not be everyone’s idea of pretty. But Cybertruck actually IS badass. Tesla / Writer: Annette, Designer: Jannéshow more

Mario Nawfal
64,021 次观看 • 1 个月前
Opening up the beta to a free cloud machine... for developers to access their coding agents from any device Preconfigured with your favorite repos, CLIs, and authenticated agents. Provisions in less than 5 mins Build, prompt, and resume projects across a phone, laptop, desktop, or iPad. No syncing remote branches, no environment setup for new devices. Just good old ssh, a cli agent, and my prompts Manage multiple sessions with tmux, install new dependencies, navigate directories, and run scripts beyond just a conversation-style interface The development and deployment of happens on a workbench itself! (capacity is very scarce as i am scaling the fleet UwU)show more

saucepoint
13,490 次观看 • 1 个月前
⚠️A VERY ACTIVE Weather Pattern is setting up next... week, bringing MULTIPLE chances for major Winter Storms to sweep through the US. 🥶👊🔥This setup features a BATTLEGROUND of Warm Tropical Air clashing with Frigid Polar Air, creating an environment that is ripe for Winter Storm development. ❓Specific storm details are still VERY much up in the air, and we should not treat any single model run as gospel at this point. I'm only showing this GFS run as an example. 🌨️ The key takeaway is that the pattern is right, and Heavy Snow, Ice, and Wind are all on the table somewhere across the country. ❄️ Hang in there, my US Snow brotheren. More details to come....show more

Brady Harris
59,298 次观看 • 7 个月前
Fable 5 comes back!It can now build playable game... prototypes. I think it is actually a signal for where AI coding is going. Making a game is not just “write some code.” Even a small browser game needs: game loop;character movement;collision logic;scoring system;UI states;physics tuning;visual feedback;bug fixing;playtesting This is why game prototyping is a great test for AI models. A model cannot fake it with a pretty answer. Either the game runs, or it does not. What impressed me about Fable 5 is that it is useful for the messy middle: turning an idea into mechanics, turning mechanics into code, debugging broken interactions, and iterating until the prototype feels playable. But here is the practical part: I would not use the strongest model for every step. For game building, I would split the workflow: 1. Fable 5 for game design + architecture 2. a fast coding model for routine implementation 3. a vision-capable model for screenshot/UI feedback 4. a cheaper model for docs, test cases, and small fixes 5. fallback when latency, cost, or output quality becomes a problem That is the real AI coding stack. Not “one magic model does everything.” More like: the right model, for the right task, at the right cost, with fallback when things break. This is why I’ve been looking at ZenMux ZenMux. ZenMux gives developers one gateway to access multiple leading AI models, with OpenAI / Anthropic / Google Vertex compatible APIs, cost tracking, quality benchmarks, auto-routing, and compensation when output quality, latency, or throughput falls short. If AI can now make games, the next question is not just “which model is strongest?” It is:how do we manage the whole model workflow Fable 5 shows the creative ceiling. ZenMux is closer to the infrastructure layer you need when AI coding becomes a real production habit.show more

Rachel🥥
61,441 次观看 • 2 个月前
This work makes a humanoid robot do simple parkour... moves by looking with a depth camera and choosing the right move on the fly. The big deal is that it turns lots of small human moves into long, real-time robot behavior, without hand-coding every transition or retraining for each new course. A humanoid robot is usually good at steady walking, but it often fails when it has to do fast moves like jumping up, vaulting, or rolling, and then keep going to the next obstacle. The hard part is that you cannot easily collect training data for every possible obstacle shape, distance, and mistake, so robots end up learning a few moves that only work in a narrow setup. This work starts from short clips of real human parkour moves, like stepping over, vaulting, climbing, and rolling. It uses motion matching, which is basically a smart “pick the next clip that fits best right now” search, to stitch those short clips into a long, smooth plan that looks like a human doing a whole course. Then it trains a controller with reinforcement learning (RL), which means the robot learns by trial and error to copy that plan while staying balanced and not falling. After training separate expert controllers for different moves, it compresses them into 1 controller that uses only onboard depth sensing and a simple “go this fast in this direction” command. In real tests on a Unitree G1 humanoid, it can clear multiple obstacles in a row, adapt when obstacles get moved, and climb a wall up to 1.25m.show more

Rohan Paul
37,121 次观看 • 6 个月前
Introducing fx, a tiny, open, native coding agent from... Vercel Labs. Originally an internal tool, fx is a harness and CLI written in Zig, optimized for research and embedding in larger systems. Today, we're open sourcing it. fx is built on three principles: 1. Fast. A single native binary, no runtime to install. It cold starts in 10µs and does no unnecessary work or I/O before accepting input. fx is the answer to "how fast can a coding agent be?" 2. Light. The 6.3MiB binary uses single-digit megabytes of memory at baseline, made for instant installation and embedding in resource-constrained environments and agent sandboxes. 3. Open. Apache-2.0, model and provider agnostic, suitable for local and cloud inference. Its small core extends through skills, plugins, and MCP. Minimalism is an obsession throughout the entire harness: system prompt, tools, features, binary. The goal was to keep context usage and time to first token low, and make fx optimal for model benchmarking, sandboxing, evals, and gyms. You can use fx directly or embed it as infrastructure. The CLI feels more like a Unix shell than an IDE in the terminal: it preserves scroll history, produces minimal output, and uses complex TUI rendering very, very sparingly. Programmatically, 𝚏𝚡 𝚊𝚜𝚔 --𝚓𝚜𝚘𝚗 gives structured output, 𝚏𝚡 𝚊𝚌𝚙 connects to editors and other clients, and WebAssembly can even run the whole thing inside the browser (see: Privacy is a design constraint: no product telemetry, sessions and usage stay local, and no source code or prompts are shared with any endpoint other than inference. With local inference and auto-updates off, fx is fully hermetic. fx is experimental. Use at your own risk and expect frequent changes. Chat with us on X ( or file issues ( 𝚌𝚞𝚛𝚕 -𝚏𝚜𝚂𝙻 𝚏𝚡.𝚜𝚑/𝚜𝚎𝚝𝚞𝚙.𝚜𝚑 | 𝚋𝚊𝚜𝚑show more

Vercel Developers
949,785 次观看 • 15 天前
Here’s how Starlink next‑generation V3’s laser network works in... space Each Starlink V3 satellite is equipped with six high-capacity space lasers capable of operating at up to 400 Gbps These laser links connect satellites directly to one another, creating a fast, high-bandwidth and resilient mesh network in orbit that can route massive amounts of data around the planet without relying entirely on nearby ground stations That means more network capacity, faster data routing, greater reliability and stronger connectivity across remote and hard-to-reach regions Starlink V3 is not just a more powerful satellite It is a major upgrade to the internet infrastructure being built in spaceshow more

X Freeze
21,988 次观看 • 1 个月前