正在加载视频...

视频加载失败

Building AI agents? 🚧 Make sure they actually know where their answers come from. As Brana Rakic demonstrates, scalable AI requires verifiable knowledge, rule-based reasoning, and LLMs grounded in trusted memory. Key highlights: 03:25 - From human data to era of experience 07:42 - Agent architecture: memory + models...

17,750 次观看 • 6 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

RAG might already be becoming obsolete. A month ago, Andrej Karpathy dropped a simple GitHub gist called “LLM Wiki.” Now the comments section looks like the birth of an entirely new AI category. 5000+ stars later, developers are rapidly building: • persistent AI memory systems • self-maintaining knowledge bases • multi-agent research environments • contradiction detection engines • AI-native company operating systems • local-first memory architectures • graph-based reasoning layers • evolving second brains And the craziest part? Most of them were built in DAYS. Because the core idea is insanely powerful: Instead of AI repeatedly retrieving raw chunks like traditional RAG… …the model continuously maintains a living knowledge system. Not temporary context. Persistent synthesis. The shift sounds subtle until you realize what it changes: RAG: retrieve → answer → forget LLM Wiki: ingest → synthesize → evolve That one architectural difference is causing an explosion of experimentation right now. People are already building: • agent memory operating systems • AI-maintained engineering documentation • self-healing knowledge graphs • persistent research environments • conversational memory architectures • contradiction-aware wikis • context compression engines • machine-readable company systems The comments section alone feels like watching an ecosystem form in real time. One developer built deterministic contradiction detection using sheaf cohomology Another built “sleep consolidation” for AI memory systems inspired by human memory formation Another created persistent multi-agent vault conversations Another turned entire repositories into continuously maintained AI wikis Another built local-first memory systems with audit trails, provenance, graph exports, and MCP integration This is the important part: Karpathy didn’t launch a product. He introduced a pattern. And patterns are what create ecosystems. The same way: • transformers created modern AI • RAG created AI retrieval startups • agents created orchestration frameworks LLM Wikis may create persistent AI memory infrastructure. That’s why this moment feels different. For years, AI systems have been stateless. Now developers are trying to build systems that actually accumulate understanding over time. And once knowledge compounds instead of resetting… …the entire interface layer of AI changes. (Link in comments)

Suryansh Tiwari

142,314 次观看 • 3 个月前

Can #AI not only support but actually drive the future of scientific discovery? We are excited to introduce SciAgents💡🔬, an agentic AI aimed towards scientific discovery through the integration of large-scale knowledge graphs, LLMs, and adversarial interactions between multiple experts. The model is capable of autonomously advancing scientific understanding by exploring novel domains, identifying complex patterns, and uncovering previously unseen connections in vast scientific data, while retrieving new data via literature search. Using graph reasoning, SciAgents identifies interdisciplinary relationships that might otherwise remain hidden, offering a step-by-step strategy for discovery & innovation. The video features an audiotrack generated using 🍓#o1 based on the original paper and design examples, providing an explanation of the work and its implications. Key elements include: 1⃣Ontological Knowledge Graphs: Structuring and connecting scientific concepts to highlight relationships across fields. 2⃣Multi-Agent Collaboration: AI agents autonomously generate and refine hypotheses, critique research, and evaluate emerging trends. 3⃣Graph-Based Reasoning: Identifying novel material designs, such as mycelium-based composites or silk-pigment blends, informed by both natural and artificial patterns. SciAgents can be used as an autonomous or collaborative tool to assist human researchers. The system offers a more powerful way to process vast data, providing innovative paths to explore nature-inspired designs or unexpected material properties. In the field of materials science, for instance, SciAgents has already demonstrated how principles from biology, music, and art can converge to create new biomimetic materials. Through isomorphic mapping, parallels have been drawn between Beethoven’s 9th Symphony and biological structures, pointing to a broader applicability of AI-driven insights across disciplines. This project allows us to enhance capabilities of researchers, allowing them to explore larger datasets and propose hypotheses grounded in a vast, interconnected web of knowledge. The agentic system was built using Auto Gen #AI #ScientificResearch #GraphReasoning #AI4Science #MaterialsScience #InterdisciplinaryResearch #SciAgents #OpenAI Chi Wang

Markus J. Buehler

209,581 次观看 • 1 年前

AI has a trust problem. Verifiability is the solution. Our GM of AI Nima Vaziri sat down with a16z’s Ali Yahya and Dan Boneh of Stanford University to map the deepest fault lines in AI today. ☁️ Models we can’t trust ☁️ Current providers can censor, shut down, or shift rules overnight. Outsourced training hides backdoors. Even “open” weights don’t prove what’s actually running. Trust. Backdoors. Black boxes. The path forward is clear: 🔥 Verifiable evals 🔥 Verifiable inference 🔥 TEEs for hardware-backed integrity 🔥 Infra beyond single points of control 🔥 Blockchains as coordination layers for AI From “trust us” to “verify yourself.” That’s the shift. That’s the unlock. The frontier is here. The builders decide what comes next. Create and use AI that’s incentive aligned with you. Timestamps: 00:00:00 Introduction: AI & Crypto Intersection Overview 00:01:58 Four Major AI-Crypto Trends 00:02:44 AI Agents Need Financial Infrastructure 00:04:03 Proof of Humanity: Fighting AI-Generated Content 00:04:17 Decentralizing AI Infrastructure Networks 00:04:44 Synthetic Life: Autonomous AI Agents 00:06:20 Verifiable AI 00:10:16 Current Performance Numbers for AI Proofs 00:13:18 The Era of Experience in AI Learning 00:14:56 AI Agents Having Life of its Own 00:18:21 Algorithmic Fairness & Verifiable Models 00:23:18 Privacy in AI: Trusted Execution Environments 00:25:47 Economic Incentive for Open Weight Models 00:31:39 Attribution Problem: Who Gets Paid for AI Training? 00:35:52 Content Provenance & Authentication (C2PA) 00:48:03 AI Security: Finding Exploits & Vulnerabilities 00:54:53 Educational Applications: LLMs as Learning Partner 00:58:29 Reliance on LLMs and Cognitive Abilities 01:03:57 Content Providers’ Fear of LLM Training

EigenCloud

62,099 次观看 • 11 个月前

In this episode, Engram co-founder and CEO Dan Biderman joins allen to cook Mediterranean meatballs with yellow rice and talk about building AI that actually learns from you: why long context, RAG, and compaction eventually break down, how Engram compresses knowledge into cartridges and model weights, what continual learning could unlock for long-horizon agents, why token efficiency is inseparable from intelligence, how personal models could improve like Tamagotchis, and what it takes to build the research and infrastructure for millions of continuously updated AI memories. Timestamps: 0:00 Intro 0:26 Engram’s $98M Launch and Meatballs 1:45 From Naval Special Operations to AI Research 4:32 Israeli Military Culture and Founder Maturity 7:12 Why Engram Is Betting on Context and Continual Learning 9:14 Knowledge Cartridges, Compression, and Model Intuition 14:10 Trillion-Token Company Knowledge and Context Rot 18:05 Long-Context Limits, Compaction, and Neural Memory 22:20 Test-Time Training and “Destroying Prefill” 24:31 Harvey and Holistic Enterprise Queries Beyond RAG 27:02 Personal AI Models and Tamagotchi Weights 30:00 What Belongs in Weights vs. Text 32:25 Autonomous Memory and User-Specific Feedback Loops 34:20 Token Efficiency, Model Routing, and Harder Tasks 38:03 Engram’s Research Team and Product Culture 43:02 Hiring Researchers and Infrastructure Engineers 45:25 Doing More With Less 47:41 Where to Find Engram 48:19 Final Taste Test

Latent.Space

34,076 次观看 • 1 个月前

Today, we're joined by Aakanksha Chowdhery, member of technical staff at Reflection, to explore the fundamental shifts required to build true agentic AI. While the industry has largely focused on post-training techniques to improve reasoning, Aakanksha draws on her experience leading pre-training efforts for Google’s PaLM and early Gemini models to argue that pre-training itself must be rethought to move beyond static benchmarks. We explore the limitations of next-token prediction for multi-step workflows and examine how attention mechanisms, loss objectives, and training data must evolve to support long-form reasoning and planning. Aakanksha shares insights on the difference between context retrieval and actual reasoning, the importance of "trajectory" training data, and why scaling remains essential for discovering emergent agentic capabilities like error recovery and dynamic tool learning. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 02:26 - Reflection 04:54 - Limitations of post-training for building agents 07:31 - Rethinking pre-training in agents 10:51 - Scaling 11:27 - Evolving attention mechanisms for agentic capabilities 12:39 - Memory as a tool 14:13 - Loss objectives and training data 15:50 - Fine-tuning loss in agent performance 19:37 - Training data 21:29 - Augmenting dominant training data source 24:11 - Overcoming challenges in training on synthetic data 25:47 - Benchmarks 30:44 - Scaling laws in large models versus small models 33:20 - Long-form versus short-form reasoning 37:57 - Agent’s ability to recover from failure 40:15 - Hallucinations and failure recovery 43:53 - Tool use in agents 46:38 - Coding agents 48:37 - How researchers can contribute to agentic AI

The TWIML AI Podcast

44,888 次观看 • 8 个月前

Thanksgiving-week treat: an epic conversation on Frontier AI with Lukasz Kaiser -co-author of “Attention Is All You Need” (Transformers) and leading research scientist at OpenAI working on GPT-5.1-era reasoning models. 00:00 – Cold open and intro 01:29 – “AI slowdown” vs a wild week of new frontier models 08:03 – Low-hanging fruit, infra, RL training and better data 11:39 – What is a reasoning model, in plain language 17:02 – Chain-of-thought and training the thinking process with RL 21:39 – Łukasz’s path: from logic and France to Google and Kurzweil 24:20 – Inside the Transformer story and what “attention” really means 28:42 – From Google Brain to OpenAI: culture, scale and GPUs 32:49 – What’s next for pre-training, GPUs and distillation 37:29 – Can we still understand these models? Circuits, sparsity and black boxes 39:42 – GPT-4 → GPT-5 → GPT-5.1: what actually changed 42:40 – Post-training, safety and teaching GPT-5.1 different tones 46:16 – How long should GPT-5.1 think? Reasoning tokens and jagged abilities 47:43 – The five-year-old’s dot puzzle that still breaks frontier models 52:22 – Generalization, child-like learning and whether reasoning is enough 53:48 – Beyond Transformers: ARC, LeCun’s ideas and multimodal bottlenecks 56:10 – GPT-5.1 Codex Max, long-running agents and compaction 1:00:06 – Will foundation models eat most apps? The translation analogy and trust 1:02:34 – What still needs to be solved, and where AI might go next

Matt Turck

168,007 次观看 • 9 个月前

LIVE NOW -- The Rise of AI Crypto Agents The collision between crypto and AI agents has officially begun. Joining us today is Matthew Stephensen, Research Partner Pantera Capital and author of “Crypto: Picks and Shovels for the AI Gold Rush”. We dive into the world of autonomous AI agents on blockchains, discuss the evolving role of agents, AI-driven market changes, and whether blockchain is the natural substrate for AI. Matt Stephenson sheds light on topics from agent liability and regulatory challenges to infrastructure value capture and the "picks and shovels" approach to investing in AI-driven crypto tech. Are AI agents on blockchains the obvious future? And how do scarcity and abundance interact in this new era? Join us as we tackle these questions and more, exploring what the future might hold at the intersection of AI, autonomy, and blockchain. ------------------------------------------------------------ Chapters: 0:00 Intro 5:34 Crypto x AI Narrative Shift 6:39 AI & Economic Agents Explained 11:50 $GOAT Memecoin Summary 23:15 Were AI Crypto Agents Obvious? 25:18 Luna AI Token & Terminal 29:41 Consequences? Is This Life? 33:27 Exciting Use Cases 40:27 Sam Altman Quote Importance 42:33 Wealth Generation Process & Blockspace 48:15 Programmable Money & Agent MEV 56:14 The #AI Agent & Memecoin Thesis 1:03:03 Government & Society Reaction Predictions 1:11:09 No Off Buttons?... 1:13:40 #DePin & AI 1:16:45 AI Agent Blockspace Demand 1:19:15 Closing & Disclaimers

Bankless

43,849 次观看 • 1 年前

Today, we're joined by Yejin Choi, professor and senior fellow at Stanford University University in the Computer Science Department and Stanford UniversityHAI. In this conversation, we explore Yejin’s recent work on making small language models reason more effectively. We discuss how high-quality, diverse data plays a central role in closing the intelligence gap between small and large models, and how combining synthetic data generation, imitation learning, and reinforcement learning can unlock stronger reasoning capabilities in smaller models. Yejin explains the risks of homogeneity in model outputs and mode collapse highlighted in her “Artificial Hivemind” paper, and its impacts on human creativity and knowledge. We also discuss her team's novel approaches, including reinforcement learning as a pre-training objective, where models are incentivized to “think” before predicting the next token, and "Prismatic Synthesis," a gradient-based method for generating diverse synthetic math data while filtering overrepresented examples. Additionally, we cover the societal implications of AI and the concept of pluralistic alignment—ensuring AI reflects the diverse norms and values of humanity. Finally, Yejin shares her mission to democratize AI beyond large organizations and offers her predictions for the coming year. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 04:44 - "Snowball effect" in AI investments 06:58 - Approaches to smaller models 08:58 - Importance of “better data” 14:07 - Imitation learning 18:24 - Artificial Hivemind paper 25:25 - AI risks 27:50 - Spectrum tuning 28:53 - Future of AI on humanity 33:08 - Reasoning in small models 34:58 - Prismatic Synthesis 48:20 - Reinforcement as a Pretraining Objective 55:04 - Pluralistic alignment 1:03:30 - Predictions

The TWIML AI Podcast

12,141 次观看 • 6 个月前

Scale alone is not enough for AI data. Quality and complexity are equally critical. Excited to support all of these for LLM developers with Snorkel AI Data-as-a-Service, and to share our new leaderboard! — Our decade-plus of research and work in AI data has a simple point: scale alone is not enough. AI success is all about the quality, complexity, and distribution of data—in addition to volume. We’re excited to be powering leading LLM developers with Snorkel AI Expert Data-as-a-Service, our white glove service for custom, expert-level AI datasets—and to now preview some of what we’re building via our new Expert Data Leaderboard (🔗 in 🧵) + upcoming OSS dataset releases! Snorkel Expert Data-as-a-Service is built to meet the rapidly evolving data needs of the agentic AI world—where success is built on the quality, complexity, and distribution of datasets, in addition to size and scale. This kind of high-quality, frontier AI data can only come from a union of technology and human expertise. With Snorkel Expert Data-as-a-Service, we’re powering frontier LLM developers across agentic, expert knowledge, reasoning, coding, multi-modal, and other task types via the combination of these two key components: - (1) The Snorkel Expert Network: A global team of subject matter experts focused wholly on specialized knowledge–spanning thousands of topics in STEM/academic, vertical/professional, and consumer/lifestyle domains. - (2) Snorkel AI Data Development Platform: Our unique programmatic data curation and quality control platform, accelerating and improving expert authoring and review through principled techniques developed over the last decade of R&D. Now: we’re incredibly excited to showcase some of the power of Snorkel Expert Data-as-a-Service via the new Snorkel Leaderboard—putting frontier models to the test in complex, agentic, and reasoning settings inspired by real industry scenarios (not esoteric puzzles)! We’ll be releasing new leaderboards and accompanying expert-verified open source datasets (coming soon!) regularly. To start, we’re sharing three initial ones in preview: - SnorkelFinance: Q&A over financial documents requiring agentic tool-calling and reasoning - SnorkelUnderwrite: Agentic insurance tasks requiring industry-specific reasoning and tool use - SnorkelSequences: Mathematical tasks requiring compositional multi-step reasoning

Alex Ratner

495,851 次观看 • 1 年前