I replaced our Datadog bill with a single binary... and cut infrastructure costs by 98% overnight. It’s called OpenObserve. Logs, metrics, traces, and frontend monitoring in one tool, self-hosted, and it’s built specifically to stop the bill that grows every time you add a host, a user, or a custom metric. → One binary, running in under 2 minutes. No cluster, no separate components for logs vs metrics vs traces → Built in Rust on the DataFusion query engine, so it stays fast even at petabyte scale → Uses Parquet columnar storage on S3-compatible object storage instead of a proprietary format, which is where the real cost savings come from → Query with SQL or PromQL instead of a vendor’s proprietary syntax, so your team isn’t learning a new query language just to read a dashboard → Full OpenTelemetry compatibility, no proprietary agents required to get your existing instrumentation talking to it → Community dashboard library on GitHub for Kubernetes, Docker, Postgres, AWS, and LLM observability, ready to drop in instead of building from scratch In OpenObserve’s own published benchmark, the same 16-service workload cost $174/day on Datadog and $3/day self-hosted, a 98% cut, before even touching per-host or per-seat fees. 18,000+ GitHub stars. Single binary. Self-hosted, no per-host or per-user tax.show more

Harman
51,181 Aufrufe • vor 2 Monaten
this free tool turns any browser session into a... shared, watchable stream it's called neko. it streams a full desktop browser out of a docker container over webrtc, so multiple people can watch and control the same session in real time instead of relying on screen share. → multi-user control, not just viewing → built-in audio streaming, unlike guacamole or novnc → works for watch parties, remote testing, or shared demos free to self-host.show more

Oliver Prompts
30,844 Aufrufe • vor 2 Monaten
you don't need to re-explain your codebase's architecture to... your agent every session. most tools stop at telling you what broke. sentrux is a real-time architectural sensor, it watches your codebase as a live treemap and turns file structure and dependencies into one continuous quality score. the loop is simple: codebase > agent scans structure and dependencies > sentrux scores 5 root cause metrics into one signal > agent sees exactly where risk concentrates > next session starts from a live map instead of a blind grep the binary carries zero built-in language knowledge, all 52 languages live in plugin.toml and tags.scm query files, so a new language needs zero rust code. small catch: it only scores the structure, it won't tell you why the cycle happened, that part's still on you. built pure Rust with no runtime dependencies, specifically so it could sit as one binary between an agent and a codebase without adding friction.show more

Simplifying AI
18,308 Aufrufe • vor 1 Monat
New open-source agent harness just landed! I got early... access to TrueForge by TrueFoundry and have been running it locally for the past few days. The harness layer deserves as much attention as the model, and open source matters here because you can inspect the loop, run it on your own infrastructure, and swap to the latest or cheaper models. TrueForge handles the runtime work that makes an agent reliable. It drives the tool-calling loop, manages context, coordinates subagents, and executes code in a sandbox, with any model you choose. Every tool call re-sends the growing context to the model, so in practice the harness controls most of what an agent costs to run. A few things stood out from my testing and their published benchmarks. Vendor-Neutral by design. It runs OpenAI, Anthropic, and Google models alongside open-weight models like Kimi, GLM, and DeepSeek. Model routing is a setting, and you can send each task to the model that fits it. On a 14-task enterprise agent benchmark, it matched the accuracy of Claude Managed Agents running the same Opus 4.8 model at roughly 30% lower cost per run (3.8M tokens vs 10M for the same answers). Routing the same tasks to GLM-5.2 held accuracy and brought cost down by about 75%, around $3 per run instead of $12. Fully self-hosted and Open Source (MIT License). I had it running locally with one command, with sandboxed code execution working out of the box. It's time to own your agent harness. Thanks to TrueFoundry for partnering on this post.show more

elvis
11,303 Aufrufe • vor 1 Monat
this is f*cking gold engineers at Meta just deleted... the most expensive part of multi-agent systems: instead of training a communication topology, they compile a fresh one for every query 20 agents went from 7 hours to 6 minutes. the problem everyone hits: 5 agents works. 20 agents turns into a group chat that answers slower than one model and costs more than the task is worth. ReActNet's fix is that the graph is written per query, not learned once: > an LLM controller reads the query and the agent roster > it compiles a sequence of directed graphs, one per reasoning stage > every edge carries a written instruction, e.g. "list boundary cases for this behavior" > each agent updates its state from its own previous state plus assigned neighbors > a final node aggregates all five states into the answer > no RL, no gradients, no training stage at all what that buys on gpt-4o: 92.75 average across 6 benchmarks, best on 5 of them. 100.00 on MultiArith. 92.74 pass@1 on HumanEval, +21 over a single model. and the number that should worry anyone running a swarm: at 20 agents GPTSwarm needs 412 minutes and $41.42. ReActNet needs 6.22 minutes and $6.53 and scores higher. the honest catch: more agents did not make it smarter. 5 agents scored 79.74, 20 scored 77.77. the topology was never the thing to learn.show more

NO1ennn
27,780 Aufrufe • vor 27 Tagen
OpenClaw, but built for normal people. Sim is an... open-source platform that lets you build AI agent workflows on a drag-and-drop canvas. Connect them to channels like Telegram and WhatsApp and deploy without writing a single line of code. They also have a built-in Copilot that generates entire workflows from plain English, which you can then tweak and customize in the UI. Key features: - Free and open-source (Apache 2.0) - Vector store integration for RAG-grounded agents - Self-host with one command (`npx simstudio`) - Run fully local with Ollama, no API keys needed - Supports vLLM for production-grade self-hosted inference The thing I really like about Sim is the level of control you get. You can add conditional branching, parallel execution, human-in-the-loop approval gates, and even nest workflows inside other workflows. Everything is visible on the canvas, so you know exactly what your agent is doing at every step. And you can build a workflow in Sim, deploy it as an MCP server, and plug it into any agent, including OpenClaw. I've shared the link to Sim's GitHub repo in the next tweet.show more

Akshay 🚀
52,426 Aufrufe • vor 7 Monaten
Stop renting a chatbot and calling it your company’s... brain. Utopia is the first open-source enterprise world model I’ve seen that actually treats knowledge like a system, not a vibe. It's a Self-hosted, open source and built around the one question every other knowledge base ignores: when. Every fact in it carries a start date and an end date. Nothing gets overwritten. When something changes, the correction closes the old version and opens a new one. The history stays intact. So the graph isn't a snapshot anymore. It's a recording you can rewind. You drag a timeline on the screen and watch the whole thing redraw itself to show what was true on that date. Who owned what. What the policy said before it changed. All of it, replayable. And every fact points back to the exact sentence it came from. Nothing gets in without a receipt. Here's the part that got me. The whole thing is one binary and one Postgres database. That's it. No Elasticsearch, no separate vector service and no message queue. Most RAG stacks are six services duct-taped together. This is two. It even runs fully offline on an air-gapped network with any model you want.show more

Hasan Toor
175,857 Aufrufe • vor 1 Monat
50% cheaper Claude inference with just one line of... code change! - Remove → model="claude-opus-4-8" - Add → model="ship-like/claude-opus-4-8" I verified the cost saving in my own terminal by invoking the same Anthropic model with the same prompt. The underlying engineering by Ship is actually interesting, and the patterns can be used in any production LLM stack. Essentially, a trained model is a frozen artifact. Every request performs the same forward-pass, whether it extracts a date or refactors a module, because the compute decision was made at training time, before the request existed. Ship makes that decision at inference time instead. After seeing a request, it searches over executions, involving single models, cascades, ensembles, or harnesses with tools, and serves the cheapest one that will match the reference model's quality. This is not a basic router, because picking a cheaper model per query doesn't ensure the cheaper model preserves the original's behavior, like output shape, tool-call patterns, and refusals. Ship measures this equivalence directly. Outputs stay distributionally indistinguishable from the reference model, not token-identical, since two calls to the same model already differ, but they are indistinguishable in capability and behavior. Of course, some requests execute cheaply and some cost Ship more than the customer pays, but the price per request is still a flat 50% off either way, so the execution-cost variance moves off the application's bill entirely. The video below depicts the cost savings and output in my real invocation, and I partnered with the team to put this together.show more

Akshay 🚀
64,482 Aufrufe • vor 2 Monaten
Cancel your $200/mo Ahrefs subscription 🤯 Claude Code can... now run your SEO for you. Point it at your Search Console and it finds the wins, writes the fixes, and renders a live dashboard off your own data. All inside Claude Code. Perfect for DTC brands and agencies sitting on months of Search Console data nobody has time to read. Here's what it does: → Connects to your Search Console and GA4 through one guided setup that routes around Google's auth landmines → Finds the keywords sitting at positions 4 to 20 and scores them by the clicks you're leaving on the table → Ships the fix instead of naming it, with the rewritten title, the headings, and paste-ready content → Turns redirect chains, broken canonicals, and slow pages into dev tickets ranked by traffic at risk → Maps every query into hub-and-spoke clusters and flags where your own pages compete with each other → Drops a Monday report with week-over-week movement and exactly 3 priorities What you get: → 9 skills in one plugin, from the Google setup through to the Monday report → A live SEO dashboard with a 0 to 100 health score, rendered as one self-contained HTML file → Orphan pages and money-page link gaps, listed paste-ready → Content drafted from your own search data instead of a keyword tool's guesses Built 100% in Claude Code on your Search Console and GA4 data. 📌 Get the free plugin here:show more

Mike Futia
22,780 Aufrufe • vor 1 Monat
grokbot it's an agent with its own identity, its... own computer, and it stays on when you're not here's what "active AI employee" actually looks like in the demo: - chief of staff agent - checks in on your other agents, reads your calendar, dispatches tasks to the right one automatically - shopping agent - logged into your accounts, books tickets, buys groceries, reports back - marketing agent - signed into your actual linkedin, browses your past posts for tone, then writes and publishes a new one on its own - engineering agents - self-triage bug reports, kick off cloud coding agents, come back with a pull request, a screenshot, and a video of the fix the interface isn't a dashboard, it's a chat - same shape as texting a coworker, no tool calls to babysit the number that matters more than any of the demos: grok 4.6 scored 70.8% on cursor bench at $2.81 a task, fable 5 max scored 70.5% at $17.32 same capability, 6x the cost difference - that's the unlock that makes running a fleet of these actually affordable instead of a noveltyshow more

rewind
755,350 Aufrufe • vor 1 Monat
Google Translate is cooked after this. A developer built... a local AI translation engine that runs 40 languages entirely on your own laptop. It's called LibreTranslate. No API key. No usage limits. No sending your documents to Google's servers. You install it once. It runs forever. Here's what it handles: → Paste text. Translated instantly. → Drop in a file. Outputs the translated version. → Point it at a URL. Returns the page in your language. → Build it into your own app via its local REST API. The speed is not the story. The privacy is. Google Translate reads every sentence you paste into it. Legal contracts. Medical records. Internal emails. Client documents. Every word goes to their servers and stays there. LibreTranslate runs entirely offline. Nothing leaves your machine. Ever. The numbers: → 40 languages supported → Runs on CPU -- no GPU needed → Self-hosted in under 5 minutes → REST API built in for developers → 10K+ stars on GitHub 100% open source. MIT licensed. Price: $0. Google charges nothing for Translate either but it charges you something else. GitHub:show more

Rimsha Bhardwaj
89,841 Aufrufe • vor 3 Monaten
this might be the E2B killer for AI agent... sandboxes. forkd is an open-source microVM sandbox runtime built on Firecracker, made for AI agent fan-out, code interpreters, eval harnesses, anything that spins up a lot of short-lived sandboxes. before: every sandbox cold-boots its own VM and re-imports the whole runtime from scratch, numpy, torch, JIT compilation, model weights, all of it. now: forkd boots one parent VM once, warms it with your runtime already imported, then forks children from that snapshot using copy-on-write memory. the repo's own benchmark: spawning 100 sandboxes takes 101ms with forkd, versus 759ms for a raw Firecracker cold-boot, and well over a minute for Docker or gVisor. the SDK is a literal drop-in for E2B's Python client, so if you're already running code-interpreter agents on it, swapping the import line gets you a self-hosted runtime with the same isolation model at a fraction of the per-sandbox cost. pre-built recipes ship for e2b-style code interpreters, Jupyter kernels, SWE-bench coding agents, and Playwright browser fan-outshow more

Oliver Prompts
20,005 Aufrufe • vor 1 Monat
AN AWS ENGINEER QUIETLY BUILT A 2 PETABYTE HOME... SERVER FOR $9/MONTH THAT KILLS A $3,400/MONTH CLOUD STORAGE BILL the lenovo thinkstation pgx ships nvidia's gb10 grace blackwell superchip and 128gb of unified memory in a box the size of a mac mini at 1.2kg it runs an 80b qwen3 coder model at 25 to 40 tokens per second and a 196b step-3.5-flash moe model at 20 tokens per second locally the gb10 packs 6,144 cuda cores, 192 fifth-generation tensor cores and rates at 1 petaflop of fp4 with sparsity from a single 240 watt usb-c power supply fine tuning qwen 2.5 7b with lora took 18 minutes and 41gb of unified memory while the gpu pulled 65 watts and peaked at 77 degrees the box pulls a docker container from nvidia's registry and serves a frontier model on your local network with tool calling and zero data leaving your desk bookmark this and read the article belowshow more

starmex
193,839 Aufrufe • vor 3 Monaten
Holy moly. 🤯 Jack Dorsey just gave away, for... free, a full toolkit for running a company where AI agents work right next to your human team. His company Block put it on GitHub. It's already past 29,000 stars. Here's how you set it up takes about 5 minutes: 1. Copy (clone) the code from GitHub. 2. Run your own server. This is where your chats, search, code, and automatic tasks all live. 3. Add your AI agent into a chat channel, just like adding a new team member. Tell it what it's allowed to do, and let it work with your team live. No fees. No middleman. Just you, your team, and your AI agents all in one place you control. This might be what running a business looks like from now on Worth saving.show more

Charlie Hills
14,039 Aufrufe • vor 21 Tagen
this is pure f*cking treasure. my ai bill last... month: claude max 20x ............ $200 chatgpt pro ............... $200 supergrok heavy ........... $300 total ..................... $700 then i went through github and found 5 repos that eat the boring half of that bill on my own machine rizzo-flow ................ 833 stars the open local take on typesafe's jev. a 4.4gb model on llama.cpp answers yes / no / pick one / score with a probability on every answer, and you point your api at localhost by changing one url llm2jev ................... 391 stars turns qwen3.5-4b into a jev-style decision model. 48ms p50 in their own benchmark, against 652ms for the real jev fast-browser-use .......... 202 stars a browser agent that runs 100% local on qwen3.5-9b and plugs into claude code and codex as a skill. opens the right wikipedia article in about 4 seconds, zero cloud calls deepseekgui ............... 84 stars a desktop workbench on deepseek harness with git, a built-in browser and memory. you pay deepseek per token instead of a flat $200 a month oriveo .................... 25 stars one app for openai, anthropic, gemini, grok, deepseek and 10 more providers, plus ollama on your own gpu. your keys, no subscription, no account claude still writes my hardest code. the yes or no calls, the browser clicking and the everyday chatting moved to my gpu and my own api keys, and the receipt stopped looking like rent repos in the commentsshow more

starmex
73,538 Aufrufe • vor 3 Tagen
THE DASHBOARD OF THE FUTURE IS NOT A SPREADSHEET.... IT IS A LIVING NEURAL NETWORK. Look closely at this screen. This is not a video game or a Hollywood CGI mockup. This is a live visualization of a $239,000 business running entirely on a custom, autonomous terminal powered by "Claude Max." The creator abandoned traditional software. No more clunky CRMs, fragmented apps, or manual data entry. Instead, he locked himself in a room and built a fully integrated digital organism where every single business department is an interconnected node. "Reception," "Billing," "Dispatch," and "Analytics" are not just tabs on a screen—they are active AI agents constantly firing data back and forth in real-time. When a new lead comes in, the system automatically processes the work order, routes the dispatch, updates the company memory, and calculates the live P&L without a single human click. This is exactly why the video says "POV: You haven't touched grass." When you wire your entire operational flow into a self-sustaining AI loop, you no longer need to manage the day-to-day. You stop being a boss and start becoming a mere observer of your own cash-printing machine. While most people are still using AI to draft polite emails, this setup proves what happens when you give a model the keys to the entire infrastructure. It stops looking like a software tool and starts operating like a hive mind.show more

kyrox
12,854 Aufrufe • vor 2 Monaten
holy sh*t. Jev's founder Diogo Almeida dropped a Google... Doc on coding agents one of these PDFs has the map everyone needs: where Jev sits in an agent loop ↓ agents do the work, Jev makes the calls, the LLM only writes: 1 → task in: one Noul before anything runs. in scope, or does a human need to see it first? 2 → dispatch: a Choice picks who takes the step (research, write or review). act at 0.85+, everything below goes to review 3 → act: Jev gates the tool call (allow, confirm or block). the LLM only writes the arguments 4 → check: a Score decides what comes back (drop, summary or keep). whatever Jev is unsure of goes to a human 5 → finish: Jev says done above 0.8, then code proves the file actually exists the LLM shows up in 2 of 5 stages. everything else is a typed answer, plain code or a person. don't refactor your agent. hand Jev the one fork it hits most (usually the tool gate or dispatch) and watch 3 numbers for a week: cost per completed task, time per completed task, and how often escalation fired and whether it was right. if escalation never fires, the threshold is too loose and your agent is quietly running unsupervised.show more

Archive
33,524 Aufrufe • vor 12 Tagen
A good technical LLM interview question: Your RAG chatbot... is working as expected locally. You deploy it behind a load balancer with 3 replicas. Users report that it forgets what they just asked, and answers get worse with each restart. Why did this happen? (answer below) A local setup has one process that owns everything. - The vector index is a variable in memory. - Conversation history is a Python list. - The documents are on local disk. You never treat any of them as infrastructure, because restarting rebuilds all three in seconds and there is only ever one copy. The setup does not carry over to production directly. The vector index might disappear on restart, so the app re-embeds everything on boot and serves empty results until it finishes. Conversation history may belong to one replica, so a follow-up routed elsewhere has no memory of the previous turn. Documents could be on whichever container ingested them, so the three replicas hold three different corpora. None of this is evident with one user and one process. So the actual work in shipping RAG is not just the retrieval logic, but also storing the vector index, the conversation history, and the documents outside the app, where every replica reads and writes the same copy. Which comes down to three requirements: > The vector store needs persistence and has to be reachable from every replica. pgvector inside Postgres keeps embeddings next to the rest of the data instead of adding another system to operate. > Conversation state has to be checkpointed outside the app. LangGraph writes its state to Postgres, so any replica can pick up a thread mid-conversation. > Docs need shared object storage, so ingestion happens once instead of once per replica. If you get those three right, the retrieval logic you wrote in the notebook works unchanged. To learn how all of it is wired together, Akamai's GitHub has a working reference implementation. - rag-langgraph-k8s-quickstart is an airline policy Q&A assistant built with FastAPI, LangChain, and LangGraph. Terraform provisions the LKE cluster, a Postgres instance with pgvector for embeddings, a second Postgres for LangGraph checkpointing, and an object storage bucket for the policy documents, in one apply. - akamai-workshop-ai-inference covers the next step, running the model yourself instead of calling an API, with prefill and decode, KV cache tradeoffs, and continuous batching under real concurrency. Both are available on Akamai's new Developer Hub, alongside their tutorials and code samples. It also links to Edge Case, their Discord, where four developer advocates architect and deploy a production app live every other Wednesday. If you create a new Akamai Cloud account, you can also get $300 in credits for joining. Join here: That said, this post assumes the retrieval logic was right to begin with, and that is doing a lot of work. Most RAG systems fail earlier, at the point where a chunk gets treated as a self-contained unit of meaning. I wrote about the two skills that fix that gap, and why the chunk is usually the wrong thing to embed. Read it below. Thanks to Akamai Cloud for partnering today!show more

Akshay 🚀
31,971 Aufrufe • vor 1 Monat
this is more useful than my entire degree Elon... Musk's rocket company signed a $60,000,000,000 deal for Cursor in June, and eight days ago the two of them put a worker on sale for $200 a month: it gets its own computer in the cloud, signs into your accounts, clicks through your real apps, and hands back finished work instead of a draft for you to paste i ran one against my receipts folder on sunday and got back 14 filed, 2 it held because they needed a card number, and a saved method i never wrote myself Grok Bot is the one you train by doing your own job in front of it, and the whole handover fits in four messages tonight: 1. write out one job you did today the way you would brief a new hire: what has to be finished, which sites and files to work from, what to hand back, and where it stops and asks you 2. let it run once on something safe to get wrong, then correct the result until it is worth your name 3. say "save what we just did as a skill", and add the one rule about what always needs your approval 4. say "run that skill every weekday at 8 and post the result here. if the source is missing, tell me instead of using yesterday's numbers" xAI wrote that order into its own manual: one real job, then the saved method, then the clock. a schedule sitting on top of a method nobody checked replaces two hours of your clicking with two hours of your mistake turns out you never get to pick the brain, and that is the part i would argue about: the manual says there is no model picker for members or admins, no plan to add one, and the bill follows whichever model answered bookmark this, then open the piece below: which jobs deserve a worker of their own, and which ones quietly burn the seat ↓show more

Argona
21,946 Aufrufe • vor 1 Monat
Introducing fx, a tiny, open, native coding agent from... Vercel Labs. Originally an internal tool, fx is a harness and CLI written in Zig, optimized for research and embedding in larger systems. Today, we're open sourcing it. fx is built on three principles: 1. Fast. A single native binary, no runtime to install. It cold starts in 10µs and does no unnecessary work or I/O before accepting input. fx is the answer to "how fast can a coding agent be?" 2. Light. The 6.3MiB binary uses single-digit megabytes of memory at baseline, made for instant installation and embedding in resource-constrained environments and agent sandboxes. 3. Open. Apache-2.0, model and provider agnostic, suitable for local and cloud inference. Its small core extends through skills, plugins, and MCP. Minimalism is an obsession throughout the entire harness: system prompt, tools, features, binary. The goal was to keep context usage and time to first token low, and make fx optimal for model benchmarking, sandboxing, evals, and gyms. You can use fx directly or embed it as infrastructure. The CLI feels more like a Unix shell than an IDE in the terminal: it preserves scroll history, produces minimal output, and uses complex TUI rendering very, very sparingly. Programmatically, 𝚏𝚡 𝚊𝚜𝚔 --𝚓𝚜𝚘𝚗 gives structured output, 𝚏𝚡 𝚊𝚌𝚙 connects to editors and other clients, and WebAssembly can even run the whole thing inside the browser (see: Privacy is a design constraint: no product telemetry, sessions and usage stay local, and no source code or prompts are shared with any endpoint other than inference. With local inference and auto-updates off, fx is fully hermetic. fx is experimental. Use at your own risk and expect frequent changes. Chat with us on X ( or file issues ( 𝚌𝚞𝚛𝚕 -𝚏𝚜𝚂𝙻 𝚏𝚡.𝚜𝚑/𝚜𝚎𝚝𝚞𝚙.𝚜𝚑 | 𝚋𝚊𝚜𝚑show more

Vercel Developers
964,536 Aufrufe • vor 1 Monat
First fully ML-framework-free 3D Gaussian Splatting implementation in LichtFeld... Studio. I’ve completed the migration of the full training pipeline to a custom CUDA-based tensor library. No PyTorch, no LibTorch, no autograd. Every gradient is implemented by hand, either through CUDA kernels or minimal abstractions on top. This makes it the first full training setup for 3D Gaussian Splatting with zero dependencies on existing ML frameworks. It’s not just about independence, it's about control! We now manage every byte of GPU memory, which opens the door to tighter optimization and finer performance tuning. The framework footprint is minimal, without pulling in gigabytes of ML runtime code that was never designed for real-time or graphics-driven applications. A few modules, such as the metrics and 3DGUT interfaces, are still being ported, and some operations are temporarily naïve, so performance is not yet on par with master. But this refactor lays the groundwork for: - A fully self-contained binary - Fine-grained memory optimization - Easier experimentation without the weight of an ML stack We’re getting close.show more

MrNeRF
50,571 Aufrufe • vor 11 Monaten