正在加载视频...

视频加载失败

OpenAI recently released its first open-weights model since GPT-2, entering a field led by DeepSeek and Alibaba's Qwen. Ankit () breaks down these top OSS models, including what sets them apart under the hood: mixture-of-experts, long-context training, and post-training techniques that shape reasoning and alignment—and how different design choices...

208,828 次观看 • 1 年前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

Thanksgiving-week treat: an epic conversation on Frontier AI with Lukasz Kaiser -co-author of “Attention Is All You Need” (Transformers) and leading research scientist at OpenAI working on GPT-5.1-era reasoning models. 00:00 – Cold open and intro 01:29 – “AI slowdown” vs a wild week of new frontier models 08:03 – Low-hanging fruit, infra, RL training and better data 11:39 – What is a reasoning model, in plain language 17:02 – Chain-of-thought and training the thinking process with RL 21:39 – Łukasz’s path: from logic and France to Google and Kurzweil 24:20 – Inside the Transformer story and what “attention” really means 28:42 – From Google Brain to OpenAI: culture, scale and GPUs 32:49 – What’s next for pre-training, GPUs and distillation 37:29 – Can we still understand these models? Circuits, sparsity and black boxes 39:42 – GPT-4 → GPT-5 → GPT-5.1: what actually changed 42:40 – Post-training, safety and teaching GPT-5.1 different tones 46:16 – How long should GPT-5.1 think? Reasoning tokens and jagged abilities 47:43 – The five-year-old’s dot puzzle that still breaks frontier models 52:22 – Generalization, child-like learning and whether reasoning is enough 53:48 – Beyond Transformers: ARC, LeCun’s ideas and multimodal bottlenecks 56:10 – GPT-5.1 Codex Max, long-running agents and compaction 1:00:06 – Will foundation models eat most apps? The translation analogy and trust 1:02:34 – What still needs to be solved, and where AI might go next

Matt Turck

168,007 次观看 • 9 个月前

My conversation with OpenAI co-founder Greg Brockman This is the most detailed first-person account of the 72 hours after Sam Altman was fired. We also go deep on what comes next: the global race to AGI, why ChatGPT stopped showing reasoning, how much of OpenAI's own code is now written by AI ("it's hard to know what percent is not"), and the untold story of how OpenAI actually started in 2015. 00:00:00 Introduction 00:00:49 Meeting Sam Altman and Starting OpenAI 00:02:40 Building the Founding Team 00:04:25 DeepMind's Lead Over OpenAI 00:04:54 Changing OpenAI to a For-Profit Model 00:06:05 Breakthrough Moments at OpenAI 00:08:22 What Dota 2 Meant for OpenAI 00:10:04 Reasoning Versus Prediction 00:11:59 Tensions Grow at OpenAI 00:15:44 Sam Altman's Firing 00:17:49 Greg Quits OpenAI 00:19:56 Sam Explores Deal with Microsoft's Satya 00:20:28 Petition for Altman's Return 00:23:43 Ilya Sutskever Leaves OpenAI 00:24:59 Lessons Learned after Sam Ousting 00:28:22 The Thing Ilya Said that Greg Can't Forget 00:32:22 Is AI Going Parabolic? 00:33:24 How Much of OpenAI's Code is Written by AI? 00:36:21 Do AI Chatbots Tell Us What We Want to Hear? 00:38:06 The Global AI Race to Reach AGI 00:38:40 What Happens if US Doesn't Reach AGI First? 00:39:49 Are Countries Stealing AI Advancements? 00:40:38 Why ChatGPT No Longer Shows Reasoning 00:41:47 The Finite Constraints of Compute 00:43:38 On Investing Early in Data Centers 00:46:31 The Future of Data Center Specialization 00:47:52 How to Decide Whose Queries to Serve 00:49:08 OpenAI on Consumer vs Enterprise Models 00:53:05 Data Centers in Space? 01:00:56 What Should AI Regulation Look Like? 01:04:33 The Future of AI-Powered Entrepreneurship 01:04:44 AI and Job Loss 01:07:15 The Skills Young People Should Invest In 01:11:30 What Does Success Look Like For You? Full episode on X below. Also find it on: • YouTube: • Spotify: • Apple:

Shane Parrish

450,952 次观看 • 4 个月前

Gemini 3, scaling laws and the 'finite data' era: my conversation with Sebastian Borgeaud, research engineer at Google DeepMind and a pre-training lead for Gemini 3 00:00 – Cold intro: “We’re ahead of schedule” + AI is now a system 00:58 – Oriol Vinyals's “secret recipe”: better pre- + post-training 02:09 – Why AI progress still isn’t slowing down 03:04 – Are models actually getting smarter? 04:36 – Two–three years out: what changes first? 06:34 – AI doing AI research: faster, not automated 07:45 – Frontier labs: same playbook or different bets? 10:19 – Post-transformers: will a disruption happen? 10:51 – DeepMind’s advantage: research × engineering × infra 12:26 – What a Gemini 3 pre-training lead actually does 13:59 – From Europe to Cambridge to DeepMind 18:06 – Why he left RL for real-world data 20:05 – From Gopher to Chinchilla to RETRO (and why it matters) 20:28 – “Research taste”: integrate or slow everyone down 23:00 – Fixes vs moonshots: how they balance the pipeline 24:37 – Research vs product pressure (and org structure) 26:24 – Gemini 3 under the hood: MoE in plain English 28:30 – Native multimodality: the hidden costs 30:03 – Scaling laws aren’t dead (but scale isn’t everything) 33:07 – Synthetic data: powerful, dangerous? 35:00 – Reasoning traces: what he can’t say (and why) 37:18 – Long context + attention: what’s next 38:40 – Retrieval vs RAG vs long context 41:49 – The real boss fight: evals (and contamination) 42:28 – Alignment: pre-training vs post-training 43:32 – Deep Think + agents + “vibe coding” 46:34 – Continual learning: updating models over time 49:35 – Advice for researchers + founders 53:35 – “No end in sight” for progress + closing

Matt Turck

51,484 次观看 • 9 个月前

We don't know what most microbial genes do. Can genomic language models help? there's only one way to find out! this is a 1 hour and 42 minute interview with an MIT professor (the famous Yunha Hwang) chatting about these questions, her work in solving them at Tatta Bio, and more. zoomer captions are back too Links in reply! Timestamps: 00:00:00 - Clips + sponsor roll from the wonderful LatchBio 00:02:07 – Introduction 00:02:23 – Why do microbial genomes matter 00:04:07 – Deep learning acceptance in metagenomics 00:05:25 – The case for genomic “context” over sequence matching 00:06:43 – OMG: the only ML-ready metagenomic dataset 00:09:27 – gLM2: A multimodal genomic language model 00:11:06 – What do you do with the output of genomic language models? 00:17:41 – How will OMG evolve? 00:20:26 – Why train on only microbial genomes, as opposed to all genomes? 00:22:58 – Do we need more sequences or more annotations? 00:23:54 – Is there a conserved microbial genome ‘language’? 00:28:11 – What non-obvious things can this genomic language model tell you? 00:33:08 – Semantic deduplication and evaluation 00:37:33 – How does benchmarking work for these types of models? 00:41:31 – Gaia: A genomic search engine 00:44:18 – Even ‘well-studied’ genomes are mostly unannotated 00:50:51 – Using agents on Gaia 00:54:53 – Will genomic language models reshape the tree of life? 00:59:18 – Current limitations of genomic language models 01:08:54 – Directed evolution as training data 01:12:35 – What is Tatta Bio? 01:19:02 – Building Google for genomic sequences (SeqHub) 01:25:46 – How to create communities around scientific OSS 01:29:06 – What’s the purpose in the centralization of the software? 01:35:37 – How will the way science is done change in 10 years?

owl

44,279 次观看 • 9 个月前

Today, we're joined by Aakanksha Chowdhery, member of technical staff at Reflection, to explore the fundamental shifts required to build true agentic AI. While the industry has largely focused on post-training techniques to improve reasoning, Aakanksha draws on her experience leading pre-training efforts for Google’s PaLM and early Gemini models to argue that pre-training itself must be rethought to move beyond static benchmarks. We explore the limitations of next-token prediction for multi-step workflows and examine how attention mechanisms, loss objectives, and training data must evolve to support long-form reasoning and planning. Aakanksha shares insights on the difference between context retrieval and actual reasoning, the importance of "trajectory" training data, and why scaling remains essential for discovering emergent agentic capabilities like error recovery and dynamic tool learning. 🗒️ For the full list of resources for this episode, visit the show notes page: 📖 CHAPTERS =============================== 00:00 - Introduction 02:26 - Reflection 04:54 - Limitations of post-training for building agents 07:31 - Rethinking pre-training in agents 10:51 - Scaling 11:27 - Evolving attention mechanisms for agentic capabilities 12:39 - Memory as a tool 14:13 - Loss objectives and training data 15:50 - Fine-tuning loss in agent performance 19:37 - Training data 21:29 - Augmenting dominant training data source 24:11 - Overcoming challenges in training on synthetic data 25:47 - Benchmarks 30:44 - Scaling laws in large models versus small models 33:20 - Long-form versus short-form reasoning 37:57 - Agent’s ability to recover from failure 40:15 - Hallucinations and failure recovery 43:53 - Tool use in agents 46:38 - Coding agents 48:37 - How researchers can contribute to agentic AI

The TWIML AI Podcast

45,470 次观看 • 9 个月前

No one: *insert lack of people here*: Me: Here's over four hours of CART 1996-02 onboard footage with timestamps if you want to find a driver that you like #INDYCAR Timestamps (Driver, Year, Track) 0:00 JPM '99 Long Beach 1:00 Andretti '97 Portland 1:43 Takagi '02 Motegi 2:08 JPM '00 Denver 2:57 Fernandez '99 Fontana 3:57 Gordon '96 Toronto 4:53 Andretti '98 Rio de Janeiro 6:41 Brack '02 Toronto 7:37 de Ferran '97 Fontana 8:59 da Matta '02 Monterrey 12:13 Gugelmin '01 Texas 12:40 Papis '00 Toronto 13:40 Bruno '01 Nazareth 14:10 de Ferran '96 Cleveland 15:06 Vasser '00 Michigan 15:37 Fernandez '98 Toronto 16:47 Dixon '01 Rockingham 17:29 Johansson '96 Vancouver 18:05 Tracy '02 Chicago 19:11 Rahal '96 Long Beach 19:59 Servia '02 Rockingham 20:43 Vasser '96 Toronto 22:16 Andretti '97 Fontana 23:08 Bruno '02 Denver 25:43 Pruett '98 Michigan 26:16 JPM '00 Road America 26:54 de Ferran '97 Gateway 27:26 Brack '02 Laguna Seca 29:33 Andretti '96 Milwaukee 31:33 Zanardi '98 Toronto 33:01 Papis '02 Fontana 33:53 Bruno '01 Toronto 35:15 Fittipaldi '96 Portland 36:12 Servia '02 Rockingham 36:40 Zanardi '98 Toronto 37:49 Andretti '96 Rio de Janeiro 38:34 Tracy '02 Vancouver 40:32 JPM '99 Homestead 41:11 Papis '01 Long Beach 41:39 Franchitti '02 Rockingham 42:11 Zanardi '98 Toronto 44:30 Gugelmin '01 Monterrey 45:10 Hearn '99 Milwaukee 46:15 Andretti '02 Monterrey 48:56 Brack '01 Chicago 49:26 da Matta '02 Road America 51:21 Pruett '98 Gateway 51:53 Tags '02 Toronto 52:34 Vasser '98 Motegi 53:06 Papis '02 Mid-Ohio 53:41 Hearn '98 Nazareth 54:18 da Matta '02 Monterrey 55:24 Vasser '00 Nazareth 56:25 Franchitti '02 Mexico City 57:40 Zanardi '98 Michigan 58:06 Franchitti '02 Portland 1:00:46 Vasser '00 Rio de Janeiro 1:01:34 Fittipaldi '98 Vancouver 1:02:42 Vasser '99 Nazareth 1:03:42 de Ferran '98 Denver 1:04:36 Fittipaldi '96 Mid-Ohio 1:05:29 Dixon '01 Vancouver 1:06:26 Rahal '98 Laguna Seca 1:08:10 Brack '02 Vancouver 1:09:16 Pruett '97 Rio de Janeiro 1:09:55 Rahal '98 Surfers Paradise 1:11:13 Minassian '01 Texas 1:12:48 Franchitti '02 Monterrey 1:14:27 Kanaan '99 Surfers Paradise 1:15:09 Vasser '96 Homestead 1:16:01 Franchitti '02 Mid-Ohio 1:17:16 Papis '00 Cleveland 1:18:18 Bruno '02 Road America 1:20:32 Andretti '99 Homestead 1:20:54 Bruno '01 Vancouver 1:22:07 Andretti '99 Motegi 1:23:11 JPM '00 Vancouver 1:23:41 Tracy '02 Laguna Seca 1:25:22 Minassian '01 Denver 1:26:23 Andretti '99 Gateway 1:26:49 Bruno '02 Denver 1:28:46 Andretti '99 Road America 1:30:04 Dixon '01 Surfers Paradise 1:30:37 Franchitti '02 Portland 1:32:48 Vasser '96 Vancouver 1:33:24 Andretti '98 Nazareth 1:34:00 Tracy '02 Toronto 1:34:56 JPM '99 Portland 1:36:45 Rahal '97 Homestead 1:37:24 Carpentier '02 Road America 1:39:02 da Matta '00 Nazareth 1:40:14 Zanardi '98 Toronto 1:40:59 Johansson '96 Portland 1:41:54 Brack '02 Long Beach 1:42:33 Boesel '97 Laguna Seca 1:43:28 Andretti '99 Homestead 1:44:26 Rahal '98 Denver 1:45:57 Johnstone '96 Portland 1:46:55 JPM '00 Chicago 1:47:35 da Matta '02 Laguna Seca 1:49:36 Brack '01 Chicago 1:50:44 Carpentier '02 Montreal 1:52:25 Gidley '01 Rockingham 1:53:01 Vasser '02 Mid-Ohio 1:54:23 Dixon '01 Houston 1:55:12 JPM '99 Milwaukee 1:56:15 Gugelmin '01 Mid-Ohio 1:57:13 Tracy '02 Miami 1:59:02 Pruett '98 Fontana 1:59:37 Andretti '99 Portland 2:00:32 Papis '00 Michigan 2:01:23 Andretti '96 Laguna Seca 2:03:21 Fernandez '99 Surfers Paradise 2:04:23 Brack '02 Fontana 2:05:24 Papis '00 Cleveland 2:06:08 da Matta '02 Laguna Seca 2:10:12 Ribiero '97 Michigan 2:10:59 Brack '02 Vancouver 2:12:26 Fernandez '99 Laguna Seca 2:13:49 Rahal '97 Fontana 2:14:27 Gidley '01 Road America 2:15:14 Fernandez '02 Milwaukee 2:16:26 Zanardi '98 Vancouver 2:17:54 Fittipaldi '97 Laguna Seca 2:18:58 Bell '02 Motegi 2:19:39 JPM '00 Laguna Seca 2:20:32 Pruett '98 Gateway 2:21:07 Franchitti '02 Toronto 2:23:08 JPM '99 Portland 2:24:08 Dixon '01 Lausitz 2:24:44 Vasser '02 Denver 2:25:55 Fernandez '01 Motegi 2:27:19 Tracy '02 Mexico City 2:28:20 Fittipaldi '99 Gateway 2:29:38 Carpentier '02 Mid-Ohio 2:30:13 Andretti '97 Surfers Paradise 2:31:40 Vasser '96 Nazareth 2:32:05 de Ferran '97 Long Beach 2:33:13 Vasser '99 Homestead 2:33:39 Kanaan '02 Cleveland 2:34:30 Fernandez '01 Chicago 2:35:58 Brack '02 Laguna Seca 2:37:21 Fernandez '01 Texas 2:37:49 JPM '99 Denver 2:38:44 de Ferran '97 Rio de Janeiro 2:39:28 Hearn '98 Road America 2:44:50 Brack '02 Vancouver 2:44:02 Andretti '98 Homestead 2:45:00 Tracy '02 Laguna Seca 2:46:16 Gidley '01 Chicago 2:47:12 Fittipaldi '96 Cleveland 2:48:18 JPM '00 Portland 2:49:57 Rahal '98 Motegi 2:50:28 Blundell '00 Long Beach 2:51:39 Franchitti '02 Rockingham 2:52:03 Andretti '99 Mid-Ohio 2:53:25 Tracy '02 Vancouver 2:54:33 Gugelmin '01 Portland 2:55:41 JPM '00 Denver 2:56:44 Franchitti '02 Motegi 2:57:25 Zanardi '98 Toronto 2:59:41 Franchitti '02 Rockingham 3:00:17 JPM '00 Laguna Seca 3:01:35 Tracy '02 Toronto 3:04:18 Brack '01 Michigan 3:04:51 da Matta '02 Monterrey 3:06:04 Jones '99 Rio de Janeiro 3:06:57 JPM '00 Houston 3:07:33 Fernandez '02 Monterrey 3:09:02 Bruno '01 Nazareth 3:10:20 Rahal '96 Toronto 3:11:23 da Matta '02 Road America 3:13:42 Pruett '96 Nazareth 3:14:40 JPM '00 Road America 3:15:23 Fittipaldi '99 Cleveland 3:16:31 Hearn '98 Rio de Janeiro 3:17:58 Andretti '02 Road America 3:18:54 Vasser '96 Milwaukee 3:19:37 Franchitti '02 Portland 3:20:34 Andretti '97 Toronto 3:21:26 Vasser '99 Milwaukee 3:22:06 Papis '00 Vancouver 3:22:48 Bruno '01 Laguna Seca 3:24:10 Rahal '97 Rio de Janeiro 3:25:22 Brack '02 Denver 3:26:15 Gordon '99 Chicago 3:27:03 Fittipaldi '96 Portland 3:27:41 Nakano '01 Motegi 3:28:34 Vasser '02 Denver 3:29:46 Andretti '99 Milwaukee 3:30:27 Pruett '98 Mid-Ohio 3:31:23 Papis '00 Milwaukee 3:31:55 Andretti '99 Denver 3:33:01 Bruno '01 Nazareth 3:33:34 Carpentier '02 Montreal 3:35:11 Andretti '99 Nazareth 3:36:01 JPM '00 Mid-Ohio 3:37:05 Rahal '98 Rio de Janeiro 3:37:48 Kanaan '02 Cleveland 3:39:01 Vasser '00 Motegi 3:40:10 Fittipaldi '02 Montreal 3:41:20 da Matta '00 Michigan 3:41:55 Zanardi '98 Toronto 3:42:57 Papis '00 Milwaukee 3:43:53 Bruno '02 Portland 3:44:48 da Matta '01 Fontana 3:45:42 Zanardi '98 Portland 3:46:42 Andretti '99 Nazareth 3:47:44 Carpentier '02 Mid-Ohio 3:48:55 Andretti '99 Fontana 3:49:42 Vasser '01 Houston 3:50:34 Rahal '98 Fontana 3:51:57 JPM '00 Portland 3:53:15 Andretti '99 Cleveland 3:54:30 Vasser '97 Homestead 3:55:30 Andretti '99 Cleveland 3:56:32 Servia '02 Rockingham 3:56:59 Fernandez '01 Surfers Paradise 3:58:27 Tracy '02 Fontana 3:59:08 JPM '99 Denver 4:00:16 Papis '00 Homestead 4:00:44 Franchitti '02 Laguna Seca 4:02:39 Bruno '01 Milwaukee 4:03:25 Carpentier '02 Vancouver 4:04:34 JPM '99 Milwaukee 4:05:04 Kanaan '02 Cleveland 4:05:55 Jones '99 Gateway 4:06:30 da Matta '02 Laguna Seca 4:07:38 Vasser '99 Motegi 4:08:13 Fontana '00 Long Beach 4:09:49 Bruno '01 Milwaukee 4:10:54 Carpentier '02 Cleveland 4:11:59 Andretti '99 Gateway 4:12:51 Franchitti '02 Road America 4:14:29 Zanardi '98 Toronto 4:15:51 da Matta '02 Fontana 4:16:51 Ribiero '97 Toronto 4:19:16 Nakano '00 Laguna Seca 4:20:12 Gidley '01 Fontana 4:20:42 Kanaan '99 Mid-Ohio 4:21:48 da Matta '01 Fontana 4:22:33 Andretti '02 Monterrey 4:24:22 Pruett '96 Toronto 4:26:40 Vasser '99 Michigan

Hickey

13,306 次观看 • 1 年前

Cursor Complete Guide for AI Coding... 1. The Basics, Composer, Cursor 2.0, Why use Cursor? 2. Multiple Agent Testing, Adding Database, Deploying to Vercel 3. Comparing the big 4: v0, Replit, Lovable, Cursor And more... with Senior Software Engineer Kehan Zhang TIME STAMPS --------------- 1. BASICS: 00:00 Introduction 01:01 Overview of Cursor and Its Features 01:47 Getting Started with Cursor 02:39 Understanding IDE and Vibe Coding 06:00 Cursor For Mobile Apps 10:26 Downloading and Installing Cursor 11:17 Creating and Managing Projects in Cursor 15:14 Building a Simple Game with Cursor 19:10 Advanced Features and Customization 40:28 Fixing Styling Rules 40:53 Redesigning the App 42:17 Exploring Cursor 2.0 Features 43:22 Setting Up the Project Structure 44:17 Adding and Testing Meme Templates 46:08 Debugging Text Issues 2. ADVANCED 49:46 Using Multiple Agents 01:10:40 Creating Custom Commands 01:14:15 Creating Commands in Settings Tab 01:15:11 Introduction to Instant DB 01:16:04 Setting Up Instant DB in Your Project 01:18:24 Building a Full Stack Application 01:19:04 Using the Agent to Plan and Build 01:26:06 Testing and Debugging the Application 01:53:02 Deploying the Application with Vercel 01:55:35 Setting Up the CLI 01:56:15 Understanding Command Line Interfaces (CLI) 01:57:32 Deploying Code to Vercel 01:58:07 Handling Environment Variables 01:58:44 Interacting with the Vercel Deployment 02:00:34 Exploring Cursor's Capabilities 3. COMPARING VIBE CODING TOOLS 02:09:48 Comparing Vibe Coding Tools 02:31:04 Final Thoughts and Recommendations

Riley Brown

65,392 次观看 • 10 个月前

qwen 3.8 max vs deepseek v4 flash 0731 vs kimi k3 vs gpt 5.6 sol – on rubik's cube and chess four frontier models built a rubik's cube stand and solved it, then built a chess board and played claude opus 5 on it the setup: Nous Research's hermes agent cli on OpenRouter tasks: 1. cube – build a 3d rubik's cube with a cli and a Three.js viewer, then solve an identical scrambled position on your own stand 2. chess – build a 3d chess stand, then play white against claude opus 5 as black, live, one move at a time. no engine, no solver, no opening book on either side. stockfish depth 14 grades every chess ply afterwards; neither player sees the score models: DeepSeek v4 flash 0731, OpenAI gpt-5.6 sol, Kimi.ai kimi k3, Qwen qwen 3.8 max gpt-5.6 sol and deepseek v4 flash solved their cubes – sol in 24 moves and seventeen seconds, deepseek in 32. qwen and kimi never got there, giving up at 96 and 207 moves then all four built chess stands and played white against claude opus 5 on them, and all four resigned: deepseek on move 13, sol on 19, kimi on 21, qwen holding out longest at 29 - build time, both stands #1 gpt-5.6 sol – 16m 43s #2 deepseek v4 flash – 97m 39s #3 kimi k3 – 166m 09s #4 qwen 3.8 max – 215m 08s - build attempts before a working stand #1 gpt-5.6 sol – 3 #2 qwen 3.8 max – 4 #3 kimi k3 – 4 #4 deepseek v4 flash – 5 - total tokens #1 gpt-5.6 sol – 6,713,754 #2 qwen 3.8 max – 17,272,507 #3 kimi k3 – 22,427,504 #4 deepseek v4 flash – 27,417,442 - total price #1 deepseek v4 flash – $0.557 #2 gpt-5.6 sol – $6.319 #3 qwen 3.8 max – $10.270 #4 kimi k3 – $16.667 observations: • deepseek v4 flash is the cheapest model here by a margin nobody else is near, and it got there while being the least efficient of the four. it burned 27.4m tokens – more than anyone, 5m more than kimi – and still finished both benchmarks for $0.557. that is $0.02 per million tokens against kimi's $0.74. it also needed the most passes to produce working stands, five, and that did not matter: all five deepseek passes together cost a thirtieth of kimi's two • so what deepseek cannot do is get it right the first time. what it can do is get it right the fifth time, for half a dollar. that is a different thing to be buying – not a good first draft, but the option to keep asking • gpt-5.6 sol is the opposite profile and the strongest of the four on pure efficiency. 16m 43s to build both stands, 6.7m tokens, three passes – under 40% of the next lowest token count and a quarter of deepseek's, on an eighth of qwen's clock. it also solved the cube fastest of anyone, 24 moves in seventeen seconds. sol is what you reach for when you want the answer now and can absorb $0.94 per million • sol's weakness is in what it does not check. its chess viewer deleted the capturing piece instead of the captured one, so pieces disappeared off the board mid-game – a defect the fifty-cent deepseek stand did not have. fast and terse turns out to be the same dial as fast and unverified • qwen 3.8 max is not the cheap open-weights option it gets treated as. $10.270 across the two benchmarks, second most expensive of the four, 18x deepseek, and by a distance the slowest – 215 minutes of build time, nearly thirteen times sol's. what the money buys is judgment: it played eighteen moves without a single error worth a hundredth of a pawn, then made exactly one bad move in the whole game, and averaged 44.6 centipawns lost across the longest game any of the four managed. it also could not solve a rubik's cube in 96 tries • kimi k3 is the one line with no reading that flatters it. most expensive at $16.667, last on the cube at 207 moves, last at chess at 478 centipawns lost per move. it is also the model that verified hardest – on the cube it wrote its own integrity check instead of trusting its output. that makes the result worse rather than better: the checking was real, and the reasoning underneath it still was not follow thehype. for 24/7 ai news, analysis and breakdowns

thehype.

84,777 次观看 • 1 个月前