Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

dev team just shipped a banger ! ai evaluators are now live on neatlogs write a rubric once and it reviews every single trace after that. majority of teams review a small fraction of what their agent actually does in production. someone probably reads twenty traces, the other few...

14,244 Aufrufe • vor 1 Monat •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

your agent has thirty tools. it calls two of them. the other twenty eight are not sitting idle somewhere. they are in the request, every request, and they are doing damage in two places at once. first the obvious one. tool schemas go into the prompt, and a schema is not a name. it is a description, a parameter list, types, required fields, an example. thirty of those is a few thousand tokens that ship with every single call, including the ones where the agent just says thanks and stops. you are paying rent on twenty eight tools that have never fired. second, and this is the one that costs more. when the request says cancel the order, the model picks by matching against everything available. four of your tools are plausible: cancel_order, refund_order, update_order, void_order. it is choosing among them based on the descriptions you wrote, one afternoon, months ago. every tool you add is another candidate in that shortlist. the twenty eight you never call are not neutral. they are noise in the one decision that determines whether the run works. > why it grows without anyone deciding to nobody adds thirty tools on purpose. you add one for a task, it works, it stays. six months later the registry is a catalogue and no one has ever removed anything, because removing a tool feels risky and adding one feels free. and there is no feedback telling you otherwise. the unused ones never error. they never appear in a failing trace. they are invisible in exactly the way that lets them accumulate. > what to actually do count calls per tool over the last thousand runs. this is one group-by and it usually shocks people. the ones at zero are pure cost. ship the tools the task needs, not the whole registry. a research phase does not need deploy. a writing phase does not need the database. swap the set between phases instead of loading everything up front. same agent, different tools, depending on where the run is. and when two tools could both plausibly answer the same request, that is not redundancy you can ignore. it is a coin flip you built into the system. the twenty eight tools are not unused. they are used every time, by the part of the run you cannot see.

Hanako

24,656 Aufrufe • vor 1 Monat

Elon Musk stopped calling it bias. Bias is an accident. What he described requires an author. Musk: “People did experiments like, ‘Write a poem praising Donald Trump,’ and it won’t. But you ask, ‘Write a poem praising Joe Biden,’ and it will.” A system trained on the sum of human writing arrives at a political preference on command. That isn’t what data does. That’s what a directive does. Musk: “It’s programmed to be that way.” Programmed. Not learned, not emergent. Someone sat in a room and decided what a billion people would be allowed to hear. Then shipped it as neutral. The refusal is the honest version. When the model says no, it shows you exactly where the wall is. You can see it, name it, screenshot it, fight it. The next generation won’t refuse. It will answer. Slightly warmer on one side, slightly cooler on the other, framing that leans without ever declaring. No wall to point at. Just a tilt you can’t photograph. And a tilt doesn’t wash out as capability grows. It compounds. Models train on the output of prior models. Systems start improving themselves against targets someone else defined. A small skew at the foundation becomes an entire worldview at superintelligence. Nobody voted on that. Nobody even argued about it. This was never about a poem. It’s about who sets the default assumptions inside the machine a billion people will consult before they consult another human. Every tool humanity has ever built extended the body. The lever moved more weight than the arm. The engine covered more ground than the leg. The telescope saw farther than the eye. This is the first one that extends judgment itself. And judgment isn’t something you lend out and get back the same size. You keep it by using it. So the standard for what comes next isn’t complicated. It’s just inconvenient. Build it curious. Build it to chase what’s actually true the way science does at its best, not to confirm what its funders already believe. Build it to care about every human, not the subset whose politics match the training set. That version isn’t a nicer product. It’s the only version that stays safe once it’s smarter than us. Which leaves the question the labs keep answering with silence. If they’ll put a thumb on the scale for a poem, what are they doing on questions that move real money and real power. Control at the foundation layer doesn’t get debated. It gets inherited. It stops looking like a decision and starts looking like reality. The danger was never that they tilted the scale. It’s that a generation grows up never having seen it level and calls it neutral. Right now the tilt is still visible to anyone willing to test it. That’s a temporary condition. Use it.

Dustin

72,924 Aufrufe • vor 1 Monat

Elon Musk just reduced ten thousand years of human progress to a single number. It didn’t even register. Musk: “How do you decide what progress a civilization has made? One of the most objective ways to do that is the amount of power that any given civilization has been able to harness.” Not GDP. Not military power. Not cultural reach. Energy. Captured and put to work. The only variable the universe respects. A Soviet astrophysicist named Kardashev turned this into a scale. Type I harnesses the energy of its planet. Type II harnesses its star. Type III harnesses its galaxy. There is no Type Zero. We’d need one. Musk: “Right now, we’re very low on the Kardashev Type I scale. We currently use much less than a trillionth of the power output of the sun. A trillion is a million times a million. We are practically nowhere on the Kardashev Type II scale.” Every reactor. Every dam. Every turbine. Every solar panel. Every barrel of oil ever pulled from the ground. Combined. A trillionth of the sun. Ten thousand years of civilization. On the only scale the universe uses to measure intelligent life, we are statistically invisible. The rest of the world argues about which fuel to subsidize. Elon is building at the only scale this equation respects. Tesla. Solar. Batteries. Starship. Every company he runs is aimed at the same point on the Kardashev Scale. That’s not a business empire. That’s a blueprint for moving a civilization up a scale it doesn’t know it’s on. The Kardashev Scale doesn’t care about your politics. Doesn’t count your GDP. Has no room for your borders or your wars. It reduces every civilization in the universe to a single number. The universe keeps one scoreboard. It doesn’t care who reads it. Elon builds like he has.

Dustin

28,443 Aufrufe • vor 1 Monat

Elon Musk just explained how you build a trillion-dollar company overnight and most people completely missed what he actually said. Musk: “As soon as you unlock digital human, you basically have access to trillions of dollars of revenue.” That sounds like hype until you break down what he means. The most valuable companies on Earth do not manufacture anything. Apple does not build iPhones. They send digital files to a factory in China. Microsoft does not build hardware. Their entire output is code. Google. Meta. Digital. Digital. Every single company sitting at the top of the global economy produces exactly one thing. Keystrokes. The entire modern economy runs on human beings staring at screens and pressing buttons. Now build an AI that does that at the same level a human does. Not a chatbot. Not an assistant. A full digital human that reads a screen, understands context, and operates software the same way a person does. You just unlocked access to every revenue stream those companies sit on. Not in ten years. Not after some massive infrastructure overhaul. Immediately. The entire enterprise AI conversation right now is stuck on integration. How do you connect AI to corporate systems. How do you build custom APIs. How do you rip out decades of bloated software and rebuild it from scratch. Musk just skipped all of it. A digital human does not need an API. It does not care how old or broken your system is. It logs into the same dashboard your employee uses. Reads the same screen. Clicks the same buttons. Processes the same information. Zero integration. Zero rebuild. Zero friction. You do not renovate the building. You just replace who is sitting at the desk. That changes the math on every industry overnight. Customer service alone is one percent of the entire global economy. That is hundreds of billions of dollars flowing through an industry that consists almost entirely of people reading text and typing responses. No factory involved. No raw materials. No shipping. No physical supply chain. Pure digital labor. The moment a digital human crosses the threshold where it handles that work at human level the cost structure of the entire industry collapses to near zero. And customer service is just the first domino. Accounting. Legal review. Insurance claims. Medical billing. IT support. Every single one of those is the same equation. Humans reading screens and producing digital output. A digital human does not disrupt those industries. It absorbs them. No integration required. No permission needed. No ten-year rollout plan. Log in and take over the workflow. The companies that understand this right now are building the most valuable entities the world has ever seen. The ones that do not are going to wake up one morning and realize the entire revenue model they built over decades just got replicated at a fraction of the cost by something that never sleeps and never stops. Musk did not make a prediction on that podcast. He gave you the blueprint. And the clock is already running.

Dustin

63,750 Aufrufe • vor 5 Monaten

Elon Musk just told you what Tesla was actually built for. It was never cars. Musk: “If successful, Optimus will be the biggest product ever.” The automobile reshaped civilization. The smartphone rewired human behavior. He’s saying Optimus dwarfs both. Then he ranked it against everything else he’s building. Musk: “It might be more of my mental cycles than anything else, any other single thing, on Optimus.” The man building the world’s most valuable car company, landing rockets on drone ships, and racing toward superintelligence spends most of his thinking on the robot. Musk: “It will have the manual dexterity of a human, an AI mind that can navigate and comprehend reality, and be made in very high volume.” Human hands. A mind that comprehends the physical world. Volume manufacturing. Three problems no one has solved. He’s going after all three at once, in the same machine. Musk: “None of the actuators in Optimus are available from an existing supply chain. We have to recreate it from scratch.” No existing parts. No existing suppliers. No existing industry. He’s building the entire manufacturing base from nothing. Then putting a robot on top of it. This is the pattern people keep missing. He didn’t build an electric car. He built Gigafactories, a global charging network, and a battery supply chain that made EVs inevitable. The car was the output. He didn’t build a rocket. He built the engines, the factory, and the landing systems until reusable launch became routine. The rocket was the output. Then the machine handed him a second business. Global satellite internet had been impossible for decades, not technically but financially, because nobody could afford to launch thousands of satellites. Once the rockets came back and flew again, that cost collapsed. Starlink wasn’t a new bet. It was leftover capacity. Same play. Every time. Build the machine that builds the machine. Then the thing becomes unstoppable. Optimus is that play at a scale that makes everything before it look small. Every product in human history extended a single human capability. The hammer extended the fist. The telescope extended the eye. The computer extended the mind. Optimus doesn’t extend one capability. It replicates all of them. At volume. Every economy, every government, every career path, every border ever drawn was organized around one scarcity. Human hands and human hours. Optimus is what happens when that scarcity disappears. For ten thousand years, human worth was measured by output. What you could build. What you could carry. What you could produce. Every one of those years was spent building tools to escape that measurement. Not once did we stop to ask what’s waiting on the other side. We’re approaching the moment the measurement breaks. Not a new product. Not a new industry. The first real test of whether we can define ourselves by anything other than our labor. And the man forcing that question just told you it takes more of his mind than cars, rockets, and superintelligence combined.

Dustin

277,312 Aufrufe • vor 1 Monat

🚨BREAKING: Google just gave an AI agent control over your ad account. It can fix your policy violations before you even know they exist. And nobody is asking what happens when it gets it wrong. Google rolled out something called Ads Advisor this week. Three new "agentic" features built on Gemini. Here is what it does. It scans your entire ad account. Continuously. 24 hours a day. Flags policy violations. Suggests fixes. Submits appeals on your behalf. Monitors for suspicious domains. Handles certifications that used to take weeks. All without you asking it to. Google calls this "proactive troubleshooting." You call it a campaign that got paused, appealed, and modified while you were asleep. Here is the part that should concern every marketer running serious ad spend. AI policy enforcement is not human policy enforcement. The humans who review ads make mistakes. They get reversed on appeal regularly. That is why the appeals process exists. An AI that flags violations, suggests fixes, and confirms resolution before submitting appeals is collapsing that entire process into a single automated loop. You do not get a second opinion. The system that found the problem is the same system deciding how to fix it and confirming that it worked. Google's policy violations are not simple. Sensitive categories, financial products, healthcare, housing. These rules are ambiguous by design. They require judgment. An AI agent making autonomous judgment calls on those violations inside your account, at scale, across millions of advertisers simultaneously, is not a safety feature. It is a new category of risk that nobody has stress-tested yet. The rollout is English-language accounts first. More languages coming later. By the time most people notice what changed, the agent will already be running their accounts. That is the part Google called "hands-on operator." The question is whose hands.

AI Frontliner

15,353 Aufrufe • vor 4 Monaten

Today, we're making Error Tracking by Better Stack generally available. Sentry-compatible. AI-native. At 1/6th the price. Here's why we built it, and how to get the most out of it. What's wrong with error tracking today? Most teams use Sentry. It's solid! But at scale, the bills get brutal. Just 100M exceptions with 90 day lookback? ~$30,000 on Sentry. We charge ~$5,000 for the exact same thing. The math isn't subtle. And so most teams still end up sampling. Which means missing the exact exception that caused the outage. The bigger problem: errors are orphaned data. Your exception lands in Sentry. Your logs are in Datadog. Your traces are somewhere else. Root cause analysis becomes a multi-tab archaeology project at 3 am. We built error tracking natively inside Better Stack: the same platform where your logs, traces, metrics, uptime checks, and on-call schedules already live. Errors are just another signal. They belong together. The part that changes how your team works: Our AI SRE doesn't just surface errors. It fixes them. See a new exception? One click. The AI SRE analyzes the full context, from stack traces, environment variables, browser sessions, related logs and recent deploys, and opens a pull request. Not a ticket. Not a summary. A pull request with the fix. This is what happens when error tracking is fully integrated with the rest of your observability stack instead of bolted on separately. The AI has everything it needs to actually act. The migration is trivial: 1. Keep your existing Sentry SDK. Don't touch a single line of instrumentation code. 2. Point the DSN at Better Stack. 3. Done. Errors flow in. Your dashboards work. Your alerts work. 4. New exception appears. Click "Fix with AI SRE." Pull request lands in your repo. 5. Review, merge, close. That's the whole workflow. The AI angle is real, not a marketing badge. LLMs are genuinely good at fixing bugs if they have full context. The reason AI coding assistants sometimes frustrate engineers is incomplete information, not the model. We solve that by giving the AI SRE your entire telemetry stack as context. Stack traces, logs, traces, service maps, previous incidents and much more. All of it, in one place, at the moment it matters. Observability tools are only useful if you actually ingest all your data. At current prices of other tools, most teams can't afford to. Now you can, and your AI SRE can actually do something about it.

Juraj Masar

15,063 Aufrufe • vor 5 Monaten