
Turing Post
@TheTuringPost • 87,115 subscribers
On X we surface the AI research that matters and explain the ideas behind it. In the newsletter, we connect the dots between AI’s past, present, and future ⬇️
Videos

There's a curious detail that many may have missed amid Jeff Dean leaving Google, Demis Hassabis stepping away, and koray kavukcuoglu taking over. Dean’s work has a pattern: "turn a rare capability into shared infrastructure." MapReduce, TensorFlow, TPUs, Pathways. Now Discovery Loop wants to do the same with research itself. They are leaving Google to build the next great research system – with Google money, Google compute and with people who made Google Google, just outside its walls
Turing Post19,331 görüntüleme • 25 gün önce

Transformers are not the end game. AI still needs a breakthrough I talked to Sanja Fidler, VP of AI Research at NVIDIA, leading company's Spatial Intelligence Lab, and she explains why ↓ And you should definitely watch our full conversation to understand where AI is heading and why physical AI is the next big frontier:
Turing Post78,069 görüntüleme • 4 ay önce

The next evolution: VLA+ models Just yesterday Microsoft Research released Rho-alpha (ρα) – their first robotics model, built on the Phi family. While most Vision-Language-Action (VLA) models stop at vision and language, Rho-alpha adds: ▪️ Tactile sensing to feel objects during manipulation ▪️ Online learning that lets it improve from human corrections (via teleoperation, 3D mouse or other tools) in real-time even after deployment. Both these sides make adaptability central rather than incidental. Microsoft calls it a VLA+ model, positioning it as an extension beyond what current VLA systems support. ➡️ Today Rho-alpha can control dual-arm robot setups to perform tasks such as: • Manipulating the BusyBox following natural-language instructions • Plug insertion • Toolbox packing and object arrangement with bimanual coordination But to understand why this "plus" matters, we need to understand what came before. Here, we'll take you through the entire landscape of VLA models – Gemini Robotics, π0, SmolVLA, Helix, ACoT-VLA and others:
Turing Post62,494 görüntüleme • 7 ay önce

AutoScientists – a research lab made of agents Harvard University researchers connected agents into a self-organizing scientific team without a boss agent standing in the middle All agents look at the same shared workspace: they share memory, explore multiple directions in parallel, critique each other, avoid repeated failures, and reorganize as evidence changes. But the teams are not fixed. Agents can gather around a promising direction, like architecture, optimizer changes, or data augmentation, then abandon it if it stops working. Before they spend compute, they discuss proposals and critique each other. AutoScientists also shows strong results: - 74.4% mean leaderboard percentile on BioML-Bench - 1.9× faster GPT training optimization - +12.5% on ACE2–Spike, with the same method transferring to 217 ProteinGym assays for a +6.5% average gain
Turing Post22,353 görüntüleme • 2 ay önce

The most honest response about AGI I talked to Olive Song, senior researcher at MiniMax (official), at 9 pm on a Sunday in Beijing. She was waiting on results from the new M2.2 model experiments. We talked RL models that try to hack rewards, why alignment fails in practice, why tiny engineering details matter more than new algorithms, and what open models actually look like in production. AGI came up only at the end. For a reason. Full conversation in the comments.
Ksenia_TuringPost40,069 görüntüleme • 7 ay önce

Breaking news: Cosmos 3 is here. They are attempting to do something completely new 🤯 Why is Physical AI much harder than building a chatbot? Understanding the world is not enough, robots need to predict it and act inside it. That's the idea behind NVIDIA Cosmos 3: → Reasoning model understands what's happening from video, images, text, and actions. → World model generates future states of the environment. → Action model generates the actions needed to achieve a goal. Previous systems often stitched these capabilities together using separate models. Cosmos 3 combines them into a single architecture with two components: ▪️ Reasoner Tower Analyzes observations and builds an understanding of objects, motion, interactions, and physical context. ▪️ Generator Tower Uses that understanding to generate future videos and action sequences that obey physical constraints. So Cosmos 3 moves from: Perception → Model A Prediction → Model B Actions → Model C to: Perception + Prediction + Actions → One unified system. The goal is to make robots, autonomous vehicles, and smart environments better at answering three questions: 1. What is happening? 2. What will happen next? 3. What should I do? That's a big shift from today's AI, which mostly focuses on generating text. And check the benchmarks! Physical AI needs to generate decisions that survive contact with the real world. 🚗🤖
Turing Post13,836 görüntüleme • 3 ay önce

We met with legendary Robert Scoble to ask what excites him at CES. He is most excited about AI-driven glasses leveraging advanced waveguide tech to create immersive experiences. He discusses how these lightweight devices could evolve into holodeck-like systems, showcasing synthetic humans, 3D displays, and real-time personalization. The journey to smarter, AI-integrated products is just beginning. He's seeing CES 2025 as the start of AI revolutionizing everything – from tractors to toothbrushes. Watch to hear his bold predictions for 2025 and beyond! #CES2025 #AI #AR
Turing Post13,566 görüntüleme • 1 yıl önce
Daha fazla içerik yok.