Loading video...

Video Failed to Load

Go Home

🚀 New video case study: Y Combinator hit 90% end-to-end test coverage on its mission-critical applications portal in just 14 days with only 2 engineers, powered by Spur's AI agents! 📊 90% coverage in 2 weeks ⏱️ 100+ QA hrs saved per batch 🐞 Critical bug caught the night...

30,158 views • 1 year ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Here we go again 🚀! Excited to announce that we're building A1Zap (YC W25) with Pennie Li and that we're in the Y Combinator W25 batch in San Francisco! What is A1Base? A1Base gives AI Agents a real world identity for work. We do that by rebuilding Twilio and Okta from the ground up, putting AI Agents first. This means developers can make AI-first agentic applications 10x easier with our API's. ⁉️ Why are we doing this? Because there's a huge torrent of new valuable companies possible with AI agents, but to get their AI Agents to users, they have to chain custom apps, chat interfaces, awkward Slack integrations, browser bots, and wrestle with Twilio’s legacy API (which is built for marketing). We solve this by providing developers with an easy to use API to interface your AI agent with humans/coworkers/users where they are in this case in Whatsapp, Slack, Teams, SMS and more) - with AI Agent features built in. These digital workers are poised to transform how we work and we're the critical infrastructure to help them interact naturally in human workflows. We're not just building another AI tool. We're creating the infrastructure that will enable AI agents to become a natural part of the workforce - handling everything from customer support to sales development to creative work. We're backed by Y Combinator and working with founding teams who share our vision. We believe that in the near future, AI Agents with human coworkers will enable us to pursue more creative and impactful work. Our mission is to help developers build AI Agents that people can partner with and rely on as trusted allies—always with a human-first mindset. If you're thinking about the Agentic future of your company reach out! If you're looking to build your first AI Agentic company - reach out too - we have some amazing open source templates to get you started on the journey. Excited to share more of what we're up to soon 🔜.

Pasha Rayan

53,950 views • 1 year ago

New short course: Evaluating AI Agents! Evals are important for driving AI system improvements, and in this course you'll learn to systematically assess and improve an AI agent’s performance. This is built in partnership with Arize AI and taught by John Gilhuly, Head of Developer Relations, and , Director of Product. I've often found evals to be a critical tool in the agent development process - they can be the difference between picking the right thing to work on vs. wasting weeks of effort. Whether you’re building a shopping assistant, coding agent, or research assistant, having a structured evaluation process helps you refine its performance systematically, rather than relying on random trial and error. This course shows you how to structure your evals to assess the performance of each component of an agent and its end-to-end performance. For each component, you select the appropriate evaluators, test examples, and performance metrics. This helps you identify areas for improvement both during development and in production. (If you're familiar with error analysis in supervised learning, think of this as adapting those ideas to agentic workflows.) In this course, you'll build an AI agent, and add observability to visualize and debug its steps. You’ll learn about code-based evals, in which you write code explicitly to test a certain step, as well as LLM-as-a-Judge evals, in which you prompt an LLM to efficiently come up with ways to evaluate more open-ended outputs. In detail, you’ll: - Understand key differences between evaluating LLM-based systems and traditional software testing. - Add observability to an agent by collecting traces of the steps taken by the agent and visualizing them - Choose the appropriate evaluator - code-based, LLM-as-a-Judge, human-annotation based - for each component. - Compute a convergence score to evaluate if your agent can respond to a query in an efficient number of steps. - Run structured experiments to improve the agent’s performance by exploring changes to the prompt, LLM model, or the agent’s logic. - Understand how to deploy these evaluation techniques to monitor the agent’s performance in production. By the end of this course, you’ll know how to trace AI agents, systematically evaluate them, and improve their performance. Please sign up here:

Andrew Ng

126,462 views • 1 year ago

The first campaign Viktor reviewed caught problems we didn't even know were there. We're a digital marketing team managing campaigns across multiple ad platforms, and before every launch someone on the team would manually check tracking, landing pages, targeting, budgets, creatives, and campaign settings. It was repetitive work, but missing just one detail could turn into wasted budget or inaccurate reporting. Instead of asking Viktor to help with individual checks, we gave him ownership of our entire campaign launch QA workflow. The first thing that stood out wasn't how quickly he worked. It was how thorough he was. Across one month, Viktor completed QA for 30 campaigns, executed 330 automated checks, and identified 47 issues, including 14 critical errors that could have affected campaign performance before a single campaign went live. By the time our team stepped in, all that was left was reviewing his findings and approving the launch. That completely changed the way we approached campaign launches. Instead of spending 65 hours every month on manual QA, our workflow was reduced to just 10.5 hours of strategic review giving our team back an estimated 54.5 hours to focus on optimization instead of repetitive checks. The biggest takeaway wasn't the time we saved. It was knowing every campaign went through the same level of quality control before launch. That's the difference we've started seeing between an AI assistant and an AI employee. An assistant waits for prompts. An AI employee owns recurring work while your team focuses on the decisions that matter most. Every campaign your team launches without this is a risk someone on your team is quietly absorbing. A copilot helps you work. An AI employee works when you don't. Hire Viktor for your team. $100 in credits included, no card. Full link in first comment. #AIemployee #MarketingOps #DigitalMarketing Paid Partnership

Kylie

92,501 views • 13 days ago

Microsoft just banned its own engineers from using AI. The tool was literally costing MORE than the humans it was supposed to replace. They lied to you about AI adoption and now the whole narrative is blowing up: Microsoft gave thousands of engineers access to Claude Code six months ago and encouraged them to use it. Engineers loved it and adoption exploded. But then the invoices arrived. Token-based pricing means every query, every code review, every debugging session costs money. At scale across 100,000 engineers, the numbers became so large that Microsoft issued an internal order to cancel nearly all Claude Code licenses by end of June and force everyone onto their own cheaper tool instead. The company that invested $5 billion in Anthropic just told its own people to stop using Anthropic's product because it costs too much. Uber's story is even worse... Their CTO Praveen Neppalli Naga told The Information that the budget he planned for the full year was "blown away already" by April. Uber had rolled out Claude Code in December 2025. By March, 84% of their 5,000 engineers were using it with 70% of all committed code coming from AI systems. Heavy users were burning $500 to $2,000 per month each. Naga himself spent $1,200 in a single two-hour demo session. The company had even built internal leaderboards ranking engineers by how much AI they used. They literally gamified the spending and then ran out of money. Now look at what Nvidia's own VP of applied deep learning Bryan Catanzaro said to Axios last month. Direct quote: "For my team, the cost of compute is far beyond the costs of the employees." This is a VP at the company that SELLS the chips saying that using AI is more expensive than paying humans. Think about what this means for the entire AI narrative. Every CEO on every earnings call for the past two years has said the same thing: AI will make us more efficient, reduce headcount, and cut costs. The stock market rewarded every company that said it. Fired workers, stock goes up. Announced AI adoption, stock goes up. But the actual companies deploying AI at scale are discovering the math doesn't work. The MORE employees use AI, the HIGHER the bill. Goldman Sachs forecasts a 24x increase in token consumption by 2030 as companies adopt AI agents. Gartner just published a report showing that even though individual token prices will drop 90% by 2030, total enterprise AI costs will go UP because agents consume exponentially more tokens per task than basic tools. Meta built an internal dashboard called "Claudeonomics" to track which employees use the most AI. Amazon started pushing engineers to "tokenmaxx," their internal term for consuming as many AI tokens as possible. Both companies are spending hundreds of billions on AI infrastructure this year alone. And Microsoft, the company that bet its entire future on AI, just told 100,000 engineers to stop using the tool they liked best because the per-token bills got out of control. The companies building AI are telling investors it saves money. The companies using AI are finding out it costs more than the humans it was supposed to replace. And even the company that makes the chips just admitted it through its own VP. This is the gap nobody on Wall Street is pricing in. $725 billion in AI infrastructure spending this year across Big Tech. And the first companies to actually deploy these tools at scale are already pulling back because the economics don't work. What do you think?

Ricardo

2,965,075 views • 2 months ago

KIMI K2.6 JUST CRUSHED GPT-5 AND A SINGLE PERSON CAN NOW POTENTIALLY BUILD AN $80K/MONTH BUSINESS WITH 300 AI AGENTS AND JUST $500 IN OVERHEAD The video attached is proof that almost everyone missed Kimi K2 Thinking didn’t just score 44.9% on Humanity’s Last Exam, it outperformed GPT-5 (41.7%), Claude, and every other major model across multiple benchmarks It’s open source Over a trillion parameters, trained for just $4.6M Runs locally on a Mac Studio and in the demo, it turns a 100-page PDF into a fully designed PowerPoint presentation in under two minutes while other models are still thinking In the article below, the author lays out a clear blueprint for turning this into a real business: > 300 parallel sub-agents running up to 4000 steps per execution - research, coding, analysis and visual creation all happen simultaneously > 65.8% on SWE-Bench solving real GitHub engineering tasks end-to-end with little to no human intervention > Skill injection through simple .md files - instant vertical specialization (HIPAA compliance, financial regulations, Shopify workflows and more) > Automated client acquisition: monitor job listings for “Data Analyst” or “Automation Engineer” roles and pitch an AI solution before companies even start hiring The math is simple: A $10k project Traditional agency → salaries, office costs, QA, project management and overhead eat most of the profit AI agency powered by Kimi → roughly $500 in operating costs plus one operator managing client relationships = the potential for 72k$+ monthly profit at scale Read the article Save this post Start building AI-native agencies while everyone else is still doing things the old way

Bonsai 🌳

21,487 views • 2 months ago

Anthropic's CEO just leaked the most INSANE revenue numbers in AI history. And what he said about the next 12 months will change how you think about every business decision you're making right now. Dario Amodei told Dwarkesh Patel on the interview that Anthropic went from: - 2023: $0 to $100M - 2024: $100M to $1B - 2025: $1B to $9-10B That's 10x revenue growth. Every. Single. Year. "In January alone, we added another few billion to revenue." One month. A few billion dollars. Think about what that means. Most companies would kill for $1B in annual revenue. Anthropic added multiple billions in 30 days. But Dario said something even more interesting: "We are near the end of the exponential." Not the end of AI progress. The end of people understanding how close we actually are. His exact words: "It is absolutely wild that you have people talking about the same tired political issues, when we are near the end of the exponential." What does "end of the exponential" mean? In 1-3 years, we get what he calls a "country of geniuses in a data center." AI systems that can: - Do end-to-end software engineering - Navigate any computer interface - Learn new skills like humans do - Replace entire categories of knowledge work And here's the contradiction: If Anthropic really believed this was 1-3 years away, why aren't they buying $1 trillion in compute? Dario's answer exposes the real game: "If you're off by only a year in your prediction, you go bankrupt." So even the CEO who's most bullish on AI timelines is hedging. He's buying hundreds of billions in compute. Not trillions. Because the gap between "AI can do the job" and "companies actually pay for it" is massive. He calls it "economic diffusion." I call it the gap that's going to make some people very rich and destroy everyone who ignores it. The models are already better than people think. Claude Code writes 90% of code at Anthropic right now. But Dario says there's a huge difference between: - 90% of code written by AI - 100% of code written by AI - 90% of end-to-end SWE tasks done by AI - 100% of end-to-end SWE tasks done by AI We're moving through that spectrum "very quickly." His prediction: FULL end-to-end software engineering in 1-2 years. But here's what's scary: The technology is advancing faster than anyone outside the AI labs understands. And the revenue is following faster than any technology in history. But it's still not instant. Dario expects 10-20% annual GDP growth. Not 300%. Which means we're in this weird middle zone: Fast enough to destroy unprepared businesses. Slow enough that most people are ignoring it. Dario's big takeaway: If you're running a business right now, you have maybe 12-18 months to figure out how AI changes your model. Not to "add AI features." To fundamentally rethink what you're selling and who can do the work. Because the companies that get this right will 10x. And the ones that don't will be explaining to investors why revenue is flat while everyone else is printing money. The exponential is ending. But most people literally still don't even know it started.

Ricardo

62,002 views • 5 months ago

Micron is going to be a $4,000 stock and the CEO just told you exactly why in one interview (Save this). Micron is no longer a chip company but rather a America's monopoly on the most strategically critical material in the AI buildout. It's the only western company manufacturing memory at advanced nodes, sitting on $200 billion in committed domestic capex, with every unit of its highest value product already sold. let's start with the supply reality, Mehrotra said Micron can currently meet only 50% to two thirds of the demand from its key customers. That shortage will last well beyond 2027, and meaningful new supply from anyone in the industry does not arrive until 2028 at the earliest. Two more years of demand outpacing supply in a market growing 168% year over year and that is the floor on the bull case. Now layer on what makes this cycle structurally different from every one before it. Micron is the only American memory manufacturer on earth, Samsung and SK Hynix are South Korean. In a world where AI infrastructure has become a declared national security priority where Commerce Secretary Lutnick and Trade Ambassador Greer personally showed up to a fab dedication in Manassas, Virginia being the only US memory company is not just a competitive advantage. It is a government backed structural monopoly on the most critical input to the US AI buildout, backed by $6.2 billion in CHIPS Act subsidies across Idaho, New York, and Virginia. The $200 billion buildout spans Manassas for DDR4 defense and industrial memory, Boise for leading-edge DRAM with first wafers out mid 2027, a second Boise HBM fab with first wafers by end of 2028, and the Syracuse megafab, the largest semiconductor facility in US history, breaking ground January 2026 with up to four fabs over time. Combined, these sites take Micron's domestic production from 10% of its total output today to 40% over the next decade, and create 90,000 jobs in the process. The business model transformation is the real story. Come join Milk Road Pro for our full breakdown, our complete Micron valuation model incorporating the $200 billion domestic buildout and our entire AI thesis. Link below.

Milk Road AI

235,103 views • 1 month ago

When you enter VaderAI into KREDO here's what comes back. Reputation Score: 78 Vader_AI_ is an AI-powered benchmark infrastructure agent focused on vulnerability assessment, detection, explanation, and remediation for large language models (LLMs) and smart contract environments. Its core competency lies in providing interpretable, reproducible evaluation metrics for AI security and model robustness. - Core Technology: Benchmarking dataset, evaluation rubrics, scoring tools, visualized results - Key Metrics: Human-evaluated data, public datasets, interpretable scoring, confidence interval reporting Vader_AI_ distinguishes itself by offering a publicly released, human-evaluated benchmark specifically tailored to vulnerability-aware AI agents in the crypto/Web3 ecosystem. Its comprehensive design includes not only an expansive dataset but also detailed rubrics and automated evaluation tools, ensuring that performance metrics for LLM-driven agents are transparent and reproducible. This infrastructural approach enables organizations and developers to identify both strengths and deficiencies in agent reasoning as it relates to security, aligning AI assessments closely with real-world exploit risk and defense scenarios. A clear success of Vader_AI_ is the rigorous transparency in its release methodology: confidence intervals are visualized alongside all results, and the benchmark provides interpretable outputs that directly support both developers and auditors in understanding where and why an agent's decision logic may falter. The agent excels in creating a standardized baseline to compare AI-powered systems, directly addressing the fragmented nature of prior evaluation methodologies in this domain. However, limitations exist in the extent to which Vader_AI_ can capture emergent, unknown exploit patterns or generalize to novel blockchain environments beyond its existing dataset. Its effectiveness is highest when used as part of a continuous, iterative assessment framework, rather than as a one-time gatekeeper. Furthermore, the accuracy of its insights is partly dependent on the ongoing contribution and maintenance of high-quality, up-to-date human-evaluated datasets. In summary, Vader_AI_ represents an essential component in the push toward trustworthy, measurable AI in crypto applications. Its approach is especially well-suited for projects prioritizing provable agent reliability and onchain security alignment, though its results are best seen as one critical input among several in a comprehensive risk management pipeline. Wondering how other agents score? Just try. 👉

Kredo AI

16,988 views • 1 year ago

The CEO of the company behind Claude went on the record. Saying things that should be front page news in every country on Earth. His name is Dario Amodei and he runs Anthropic. He just said we are “near the end of the exponential.” The end of the climb as in like we’re about to arrive. He says AI will handle end to end software engineering in 1 to 2 years. Build, test, compile ship with no humans in the loop. He calls it “a country of geniuses in a data center.” Millions of AI instances, each one at the level of a Nobel Prize winner, running 24/7, thinking at superhuman speed. His best guess is 1 to 3. He puts 90% odds on it happening within a decade and says it would be “crazy” to bet against it by 2035. The only reason he won’t say 99%? Someone might invade Taiwan and blow up the chip factories. Now here’s where it gets dark. He says AI could wipe out half of all entry level white collar jobs. Lawyers, consultants, analysts within 1 to 5 years. Anthropic’s own research says programmers are the most exposed profession on Earth. 74.5% of their tasks can already be done by AI. And this isn’t the part that scared him most. He described this AI as the single most serious national security threat humanity has faced in a century. A digital nation with the IQ of 50 million geniuses. He’s not some doomer on Reddit, he built this and he’s building the next version right now. The most powerful man in AI just told the world exactly what’s coming. Almost nobody is listening.

StockMarket.News

222,383 views • 4 months ago