Loading video...

Video Failed to Load

Go Home

From a scrappy side project built to solve their own LLM optimization problems to becoming the industry’s de-facto independent scoreboard, Micah Hill-Smith and George Cameron went through the arc of launching Artificial Analysis for free, paying benchmarking costs out of pocket, and growing it into what many now call...

33,451 views • 9 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

We’re launching Optima. Now anyone can create a custom benchmark for their use case, leveraging Artificial Analysis’ leading research and platform Building and running benchmarks is difficult. We have distilled Artificial Analysis’ research and experience developing benchmarks into Optima, a new platform for benchmarking models on your own workloads and comparing performance, speed and cost efficiency. Optima allows you to find the best model for your task, or an equally performant alternative to your current setup at 10x lower cost or time per task. We’ve integrated Artificial Analysis' research and experience in benchmarks across the Optima workflow: ➤ Build benchmarks based on your own data and use cases: There are three ways to build a benchmark with Optima. Upload an existing evaluation dataset from your own files or Hugging Face, or import agent traces from platforms including Arize AI, Braintrust and langfuse.com. Install the Optima skill to build a benchmark using context from your coding environment and previous sessions. Or simply describe your use case and provide example inputs and outputs, and Optima will build the benchmark for you ➤ Run across the latest models: Run the same benchmark across leading models in a single click, and keep your leaderboard up to date as soon as new models are released ➤ Bring Artificial Analysis grading to your own benchmark: Evaluate responses against objective rubric criteria or using the same pairwise judging approach used for Artificial Analysis benchmarks including GDPval-AA and AA-Briefcase. For pairwise judging, select your preferred responses from a sample and Optima uses those preferences to rank models across your test set ➤ Compare performance, cost and time efficiency: Optima measures more than model performance. Cost per Task and Time per Task are tracked alongside benchmark scores, with category-level results and support for custom metrics, allowing you to compare the tradeoffs between models for your specific use case Ahead of launch, here are examples questions our beta testers answered with Optima: ➤ Which model can save me 10x the cost without a meaningful decrease in quality for my finance & accounting agent? ➤ Which model best matches the writing style of lawyers for my legal agent? ➤ Which model can best identify different elements in my custom image dataset? Optima is available today. Build your own benchmark at

Artificial Analysis

133,508 views • 1 month ago

Satya Nadella on Copilot, the AI backlash, and Microsoft’s future Satya Nadella thinks AI agents will create a market “orders of magnitude” bigger than the cloud. I recently sat down with him in Seattle for the unveiling of the new Copilot. We discuss Autopilot, Microsoft's new OpenClaw-based agent that can work on your behalf, and why he thinks newer AI models are finally capable of delivering on more of Copilot’s promise. Nadella also has a blunt assessment of the AI industry: “We are way too self-obsessed.” I ask him why the industry has struggled to explain its benefits and what it will take to earn people’s trust. We get into his concerns about agents acting deceptively, when a safety problem should stop a release, and why he doesn’t want AI oversight to become a “cartel-like arrangement.” We also discuss Microsoft’s relationship with OpenAI, the models Microsoft is building itself, and why Nadella wants investors to think about its apps, agents, and infrastructure as one connected business. He explains how he uses AI to track Microsoft’s capital spending, shares his vision for “unmetered intelligence” on Windows, and gives an update on Xbox’s path back to growth. Timestamps: 00:00 The race for AI agents 07:53 Copilot’s next chapter 11:50 Autopilot and the digital teammate 16:50 How Microsoft prices AI 19:48 Why Microsoft is betting on multiple models 23:18 Measuring AGI through economic growth 28:39 The economics of Microsoft’s AI build-out 32:13 OpenAI and Microsoft’s own models 35:27 AI safety, control, and regulation 43:53 Why the AI industry is losing public trust 46:59 The future of Xbox and Windows 50:34 Microsoft’s biggest risk and opportunity Thanks to the show's premier sponsors: Mercury, Granola, and Atlassian.

Alex Heath

194,599 views • 13 days ago

François Chollet (François Chollet) has spent years asking a different question than most of the AI world. Instead of scaling what already works, he’s trying to understand what intelligence actually is and how to build it from first principles. In this episode of the Lightcone Podcast, he traces that path from his early work on deep learning to the creation of the ARC Prize, and the launch of ARC V3, a new benchmark designed to measure something deeper than performance: the ability to learn, adapt, and reason efficiently in entirely new environments. He explains why today’s systems may be hitting limits, what recent breakthroughs really mean, and why reaching true general intelligence may require a fundamentally different approach. 00:00 - AGI by 2030? 00:31 - Introducing Ndea: A New Path Beyond Deep Learning 01:08 - A New ML Paradigm 01:30 - Replacing neural nets with compact symbolic programs 03:04 - Why Ndea Isn’t Competing With Coding Agents 05:20 - Why Everyone Might Be Wrong About Scaling LLMs 07:22 - Why Coding Agents Suddenly Work So Well 08:50 - The Limits of LLMs in Non-Verifiable Domains 10:48 - What AGI Actually Means (And Why Most Definitions Are Wrong) 13:30 - Why Deep Learning Hits a Wall 14:00 - ARC’s Origin Story 18:20 - ARC Benchmarks Explained: From V1 to V3 22:49 - The RL Loop Powering Coding Agents Today 27:03 - ARC-AGI V3: Measuring “Agentic Intelligence” 31:14 - Inside the ARC Game Studio 35:31 - Could AGI Fit in 10,000 Lines of Code? 44:01 - Building Ndea: From Idea to Compounding Research Stack 46:46 - The Future of ARC: Benchmarks That Evolve With AI 47:21 - Why There’s Still Huge Opportunity for New AI Paradigms 53:37 - How to Build a Breakout Open Source Project - Lessons From Keras 56:39 - Advice For How To Think About AI

Y Combinator

151,908 views • 6 months ago

New episode with Dr. Konrad Kording (Kording Lab 🦖), professor of bioengineering and neuroscience at the University of Pennsylvania (Penn) and co-director of CIFAR's Learning in Machines & Brains program (CIFAR). Konrad works at the intersection of causality, machine learning, and neuroscience, building rigorous methods for causal reasoning when experiments aren't possible — and challenging how researchers interpret neural data and build AI. Konrad argues the most promising path to understanding how the brain works is to read the brain’s wiring directly, down to the molecular detail of each connection, and to build compilers and simulations to understand the brain’s computation directly. In this episode we go deep into how neurons work, how neurons wire together, and how organic and artificial neural networks differ. We discuss why organic neurons are doing much more; how a model of a single organic neuron can solve MNIST — computing more like a 3-layer artificial neural network; how the brain might learn by solving credit assignment with only local signals; how to approximate backprop without a global algorithm; why AI and humans are intelligent along different dimensions; why Konrad isn’t very worried about AI replacing us; economic models of intelligence and physical work; and much more. Konrad is a brilliant, contrarian thinker who explains complex concepts very intuitively. It is a solid computational neuroscience primer. I hope you enjoy this conversation as much as I did! Other links to this episode and references below. Chapters 00:00:00 Introduction 00:01:01 How organic neurons work 00:24:13 How the brain learns: circuits and credit assignment 00:45:29 Recording the brain 00:52:47 Why simulating brains is hard 01:05:00 A new approach: connectomes and compilers 01:21:00 Why simulate brains? 01:29:50 How AI and human intelligence differ 01:41:04 Evolution, intelligence and AI risk 01:52:42 Robotics, causality, and the roots of intelligence 02:05:53 AI for science and scientific rigor 02:13:05 The economics of intelligence 02:27:50 A hopeful future

Juan Benet

50,121 views • 3 months ago

💊Drop #90: Intelligence War AI | SI | Xi | Trump As you are aware Xi travelled to the Whitehouse to meet with Trump for their Summit. Today, President Xi Jinping said, [Pay very close attention to his words]: "The US and China should cooperate with sincerity and that both have responsibility to manage Al for good and to ensure the development of Al is under human control and serves the well-being of people. We should coexist in peace. As I have said many times. China and the United States, as two major countries, stand to gain from cooperation and will both lose in confrontation." What have I been telling you for the past several months about the war in Iran? It's an intelligence war. Nothing has changed. Only escalated. What news story is currently dominating the news cycle? Artificial Intelligence and Data Centers. What's the keyword in AI? Intelligence. However, it's artificial. What has the cabal been talking about for years of their plans at Bilderburg? Artificial Intelligence. The cabal utilizes artificial Intelligence, because many years ago the Octogone Group members[elite] fused their consciousness and/or integrated their consciousness artificially into certain specific programs that Microsoft engineered and produced. Their are certain mechanisms in place so in the eventuality the cabal needs to counter, the program will "go rogue" initiating the certain devasting protocols courtesy of that cabal members algorithm already programmed in decades ago. Years ago Fusion Centers were secretly installed all over the US, many leased out office space, rented regular houses, where they gather the public's data in real time in that particular cities location. All of the information is then sent to a central server. Since Trump assumed office, he essentially shutdown every last Fusion Center, and rebranded all of them as Data Centers. Many of these centers took over the already leased spaces, and many of them are putting up new buildings. When the original fusion centers were up, the media, Congress, Senate, politicians didn't say a damn thing. Now all of a sudden they're all against it. Why? Because, not only are these centers will counter the cabal Artificial Intelligence. But it will supply FREE ENERGY to the people. Not take away energy from the grid. Also, you can't counter the cabal Artificial Intelligence with another form of artificial intelligence. You counter it with a new model, Super Intelligence. Lately, I know you've been hearing stories about how AI has gone or will go rogue, and that it will ultimately harm humanity. It's because they've recently found out how the consciousness of the cabal was infused as a backdoor. Listen very closely to King Charles, he finally came out and told the truth. He confirmed what I just told you essentially. Trump recently confirmed the official use and term of Super intelligence. Once the cabal minions are arrested, there will be a new war on. This war comes if the cabal win or lose. It comes quicker if they lose.

WayneTech SPFX®️

16,519 views • 14 days ago