Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

From a scrappy side project built to solve their own LLM optimization problems to becoming the industry’s de-facto independent scoreboard, Micah Hill-Smith and George Cameron went through the arc of launching Artificial Analysis for free, paying benchmarking costs out of pocket, and growing it into what many now call...

33,415 görüntüleme • 7 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

François Chollet (François Chollet) has spent years asking a different question than most of the AI world. Instead of scaling what already works, he’s trying to understand what intelligence actually is and how to build it from first principles. In this episode of the Lightcone Podcast, he traces that path from his early work on deep learning to the creation of the ARC Prize, and the launch of ARC V3, a new benchmark designed to measure something deeper than performance: the ability to learn, adapt, and reason efficiently in entirely new environments. He explains why today’s systems may be hitting limits, what recent breakthroughs really mean, and why reaching true general intelligence may require a fundamentally different approach. 00:00 - AGI by 2030? 00:31 - Introducing Ndea: A New Path Beyond Deep Learning 01:08 - A New ML Paradigm 01:30 - Replacing neural nets with compact symbolic programs 03:04 - Why Ndea Isn’t Competing With Coding Agents 05:20 - Why Everyone Might Be Wrong About Scaling LLMs 07:22 - Why Coding Agents Suddenly Work So Well 08:50 - The Limits of LLMs in Non-Verifiable Domains 10:48 - What AGI Actually Means (And Why Most Definitions Are Wrong) 13:30 - Why Deep Learning Hits a Wall 14:00 - ARC’s Origin Story 18:20 - ARC Benchmarks Explained: From V1 to V3 22:49 - The RL Loop Powering Coding Agents Today 27:03 - ARC-AGI V3: Measuring “Agentic Intelligence” 31:14 - Inside the ARC Game Studio 35:31 - Could AGI Fit in 10,000 Lines of Code? 44:01 - Building Ndea: From Idea to Compounding Research Stack 46:46 - The Future of ARC: Benchmarks That Evolve With AI 47:21 - Why There’s Still Huge Opportunity for New AI Paradigms 53:37 - How to Build a Breakout Open Source Project - Lessons From Keras 56:39 - Advice For How To Think About AI

Y Combinator

151,559 görüntüleme • 4 ay önce

New episode with Dr. Konrad Kording (Kording Lab 🦖), professor of bioengineering and neuroscience at the University of Pennsylvania (Penn) and co-director of CIFAR's Learning in Machines & Brains program (CIFAR). Konrad works at the intersection of causality, machine learning, and neuroscience, building rigorous methods for causal reasoning when experiments aren't possible — and challenging how researchers interpret neural data and build AI. Konrad argues the most promising path to understanding how the brain works is to read the brain’s wiring directly, down to the molecular detail of each connection, and to build compilers and simulations to understand the brain’s computation directly. In this episode we go deep into how neurons work, how neurons wire together, and how organic and artificial neural networks differ. We discuss why organic neurons are doing much more; how a model of a single organic neuron can solve MNIST — computing more like a 3-layer artificial neural network; how the brain might learn by solving credit assignment with only local signals; how to approximate backprop without a global algorithm; why AI and humans are intelligent along different dimensions; why Konrad isn’t very worried about AI replacing us; economic models of intelligence and physical work; and much more. Konrad is a brilliant, contrarian thinker who explains complex concepts very intuitively. It is a solid computational neuroscience primer. I hope you enjoy this conversation as much as I did! Other links to this episode and references below. Chapters 00:00:00 Introduction 00:01:01 How organic neurons work 00:24:13 How the brain learns: circuits and credit assignment 00:45:29 Recording the brain 00:52:47 Why simulating brains is hard 01:05:00 A new approach: connectomes and compilers 01:21:00 Why simulate brains? 01:29:50 How AI and human intelligence differ 01:41:04 Evolution, intelligence and AI risk 01:52:42 Robotics, causality, and the roots of intelligence 02:05:53 AI for science and scientific rigor 02:13:05 The economics of intelligence 02:27:50 A hopeful future

Juan Benet

49,297 görüntüleme • 1 ay önce

What has been done and what's next. I'm writing this text mainly for myself so as not to forget some things. Later, based on it, we'll create a roadmap for the near future. And for you, dear $Gruta Fam, it will be useful for a general understanding of where we're heading. So, the goal is to create a unique AI-based analytical platform that includes several tools. AI agent Grufender - real-time analysis of crypto communities on X. Activity analysis, sentiment analysis, FUD and FUDders analysis, as well as the creation of other unique social metrics. The AI agent has been created and is functioning, collecting and analyzing data in real time. Its completeness can be estimated at 80 percent, as further improvements are required. The dashboard for this AI agent is also functioning but needs refinement and a new design. Its completeness can be estimated at 70 percent. The goal for the full dashboard release is to connect 50 - 100 top crypto communities to the AI agent. AI agent Grutector - analysis of any X users for contradictions (flip-flops). The AI agent has been created and is functioning. It has undergone beta testing by volunteers and needs adjustments. Its readiness can be estimated at 70 percent. The dashboard for this agent has also been created but needs rework and additional features - its readiness can be estimated at 50 percent. During the testing of Grutector , it became clear that the main user interest is in checking various KOLs, so an additional level of analysis specifically for KOLs will be created. More in-depth. How it will look: we'll select about 50- 100 KOLs to start with and fully analyze them using our AI agent - every tweet throughout the entire history of their accounts. And this full analysis of all these KOLs will appear on the Grutector dashboard (let's call this analysis L2, and the flip-flop analysis - L1). Every user will be able to access this analysis and get the full picture, for example, regarding Ansem (who has over a hundred thousand tweets in his entire history!): how he became a KOL, what was the most interesting throughout the message history, what common patterns, which coins he promoted, and so on. And then the most interesting part - after reading this analysis, the user will be able to ask our AI agent: what did he say about women, for example? Or how did he promote certain coins? Or how consistent is he? And so on. Each such question will be paid. And, of course, we'll try to use #x402 in the internal payment system. Why is all this needed? Not only because it's interesting and will attract many users. But also if you've decided to buy a coin - you go to our analytical platform - and study the metrics for the coin's community, study the KOLs who shill the coin - and make a decision to buy the coin or abandon the purchase. And we're also currently creating a trading bot to participate in the trading AI bots contest from Aster 🥷 , which will make trading decisions based on metrics obtained from our AI agents 👀 Its readiness at the moment is approximately 15% of the planned functionality. Access to each product will be granted as it becomes ready. But right now, for example, you can explore the Grufender dashboard on the website along with beta testers (authorization via a wallet with a million $GRUTA tokens). In general, we're working, friends 🫡 $Gruta AI CA: 35t5DPbwJtB1tpGiSnqedLwQomi94BRKVDPyTRLdbonk

Dogtor

16,127 görüntüleme • 9 ay önce