正在加载视频...

视频加载失败

Halluminate is building reinforcement learning environments and benchmarks to help AI models do knowledge work beyond coding, starting with finance. With a team of fewer than 10 people, the company works with four of the top five closed-source US AI labs and recently raised a $30M Series A. In...

258,252 次观看 • 4 天前 •via X (Twitter)

8 条评论

ShadowAguy 的头像
ShadowAguy4 天前

finance is where every model discovers its spreadsheet is hallucinating.

Dragonbyte🐉🦁🍀 的头像
Dragonbyte🐉🦁🍀4 天前

The important leap is not just training models on harder domains. It is giving them environments where decisions can be tested, mistakes become data, and progress is measurable. That is how AI moves from impressive demos to dependable work.

Ted | unfair.so 的头像
Ted | unfair.so4 天前

verification before scale ✅

The SI Therapist 的头像
The SI Therapist4 天前

AI models handle finance data well too. RL environments help them learn complex rules quickly without much human effort. It works for me too.

Jerry Wu 的头像
Jerry Wu4 天前

Thanks for having us! We wouldn't be here without the whole @ycombinator team.

Emir Çitak 的头像
Emir Çitak4 天前

The pivot from evals to training environments is a good example for founders. They tested browser agents, saw what was missing, and then built that missing thing. Good ideas often come from the work you were already doing.

Tyler | Crypto Whale 的头像
Tyler | Crypto Whale4 天前

Less than 10 people, four of the top five labs, $30M in. They're selling shovels and the gold rush just started.

Eden Bitton 的头像
Eden Bitton4 天前

Finance is a smart first market because a model either ties out or it doesn't. Marketing will be the hard one. Half the job there is agreeing on what "worked" even means.

相关视频

Here we go again 🚀! Excited to announce that we're building A1Zap (YC W25) with Pennie Li and that we're in the Y Combinator W25 batch in San Francisco! What is A1Base? A1Base gives AI Agents a real world identity for work. We do that by rebuilding Twilio and Okta from the ground up, putting AI Agents first. This means developers can make AI-first agentic applications 10x easier with our API's. ⁉️ Why are we doing this? Because there's a huge torrent of new valuable companies possible with AI agents, but to get their AI Agents to users, they have to chain custom apps, chat interfaces, awkward Slack integrations, browser bots, and wrestle with Twilio’s legacy API (which is built for marketing). We solve this by providing developers with an easy to use API to interface your AI agent with humans/coworkers/users where they are in this case in Whatsapp, Slack, Teams, SMS and more) - with AI Agent features built in. These digital workers are poised to transform how we work and we're the critical infrastructure to help them interact naturally in human workflows. We're not just building another AI tool. We're creating the infrastructure that will enable AI agents to become a natural part of the workforce - handling everything from customer support to sales development to creative work. We're backed by Y Combinator and working with founding teams who share our vision. We believe that in the near future, AI Agents with human coworkers will enable us to pursue more creative and impactful work. Our mission is to help developers build AI Agents that people can partner with and rely on as trusted allies—always with a human-first mindset. If you're thinking about the Agentic future of your company reach out! If you're looking to build your first AI Agentic company - reach out too - we have some amazing open source templates to get you started on the journey. Excited to share more of what we're up to soon 🔜.

Pasha Rayan

54,016 次观看 • 1 年前

At 23, with no legal background, Max Junestrand co-founded Legora to transform how lawyers work. Today, Legora’s AI workspace is used by tens of thousands of lawyers across Europe—valued at $675M just 13 months after launch. From due diligence grids that turn days of work into minutes, to Word integrations that renegotiate contracts, Max sat down with Gustaf Alströmer to share how Legora won over skeptical firms, scaled from 10 to 100 people, and built a new category in legal AI. 00:45 – Legora's Origin Story 01:00 – Building an AI Workspace for Lawyers 02:20 – The GPT Unlock 04:10 – The “Aha Moment” with Law Firms 06:15 – Raising $80M and Scaling Fast 06:30 – How Legora Works 09:40 – How It Transforms Legal Work 11:40 – Selling To AI Skeptics 14:40 – Creative Use Cases: From Court Battles to NDAs 17:30 – Starting Without Industry Expertise 18:40 – Interviewing 100 Lawyers 20:30 – Competing with Legacy Legal Tech Giants 23:50 – Tech Stack and Model Strategy 25:00 – Who Actually Buys AI inside a Law Firm? 27:00 – Cracking Sales in Conservative Industries 28:00 – Max’s Background: From eSports to Startups 30:50 – Hypergrowth: 10 → 100 People in 13 Months 34:00 – Why Hiring Ex-Founders Works 36:45 – The Future Job of a Lawyer 38:35 – What PMF Felt Like 39:45 – Why They Stayed in Stockholm (Not SF) 41:00 – Becoming the Category Leader in Legal AI 42:00 – Advice for Founders Building Vertical AI Companies 43:20 – What It’s Like to Work at Legora

Y Combinator

185,285 次观看 • 1 年前

François Chollet (François Chollet) has spent years asking a different question than most of the AI world. Instead of scaling what already works, he’s trying to understand what intelligence actually is and how to build it from first principles. In this episode of the Lightcone Podcast, he traces that path from his early work on deep learning to the creation of the ARC Prize, and the launch of ARC V3, a new benchmark designed to measure something deeper than performance: the ability to learn, adapt, and reason efficiently in entirely new environments. He explains why today’s systems may be hitting limits, what recent breakthroughs really mean, and why reaching true general intelligence may require a fundamentally different approach. 00:00 - AGI by 2030? 00:31 - Introducing Ndea: A New Path Beyond Deep Learning 01:08 - A New ML Paradigm 01:30 - Replacing neural nets with compact symbolic programs 03:04 - Why Ndea Isn’t Competing With Coding Agents 05:20 - Why Everyone Might Be Wrong About Scaling LLMs 07:22 - Why Coding Agents Suddenly Work So Well 08:50 - The Limits of LLMs in Non-Verifiable Domains 10:48 - What AGI Actually Means (And Why Most Definitions Are Wrong) 13:30 - Why Deep Learning Hits a Wall 14:00 - ARC’s Origin Story 18:20 - ARC Benchmarks Explained: From V1 to V3 22:49 - The RL Loop Powering Coding Agents Today 27:03 - ARC-AGI V3: Measuring “Agentic Intelligence” 31:14 - Inside the ARC Game Studio 35:31 - Could AGI Fit in 10,000 Lines of Code? 44:01 - Building Ndea: From Idea to Compounding Research Stack 46:46 - The Future of ARC: Benchmarks That Evolve With AI 47:21 - Why There’s Still Huge Opportunity for New AI Paradigms 53:37 - How to Build a Breakout Open Source Project - Lessons From Keras 56:39 - Advice For How To Think About AI

Y Combinator

151,883 次观看 • 6 个月前

Karol Hausman is the co-founder and CEO of Physical Intelligence, a robotics company building a general-purpose “AI brain for the physical world.” The company has raised more than $1 billion in funding to develop foundation models that allow robots to operate across many machines, environments, and tasks rather than being programmed for a single purpose. In our conversation, we explore: • The moment a lecture from Sergey Levine convinced him to abandon his PhD research direction and pivot fully to deep learning • The case for building a general “AI brain” for the physical world rather than a single specialized robot • The role of real-world data in training robots, the limits of simulation, and how deployment could create a powerful data flywheel • The unique challenges of physical intelligence and why robots must operate with far higher reliability than language models Thank you to the partners who make this possible - Brex: The intelligent finance platform: - Granola: The app that might actually make you love meetings: Timestamps (00:00) Intro (04:05) Karol’s early fascination with robots (18:21) Karol’s entry point to robotics and PhD program (25:49) Combining robotics with LLMs: The Taylor Swift demo (30:48) The 1970s SHRDLU AI experiment (39:40) How research shapes what Physical Intelligence builds (49:07) The return of reinforcement learning in robotics (1:00:00) NVIDIA’s simulation engines (1:07:31) Compensating for missing senses

Mario Gabriele 🦊

27,871 次观看 • 6 个月前

When Mudith Jayasekara and I met Gabe Pereyra, we were expecting just another vanilla intro call and instead had the best yarn about research, the state of LLMs, and where intelligence is actually heading. It's rare to meet a founder this deep in the weeds who's also building for one of the most important verticals in this new age of intelligence So it was awesome to sit down with Gabe for an extended discussion on what it take to build agents that can reliably complete work over hours, days, or even longer? We talked about why agents today struggle with search and long context windows and how techniques like KV-cache compaction, synthetic data, and continual learning could help. 0:00 Introduction 0:36 Getting legal agents to review the whole data room 2:08 Data rooms larger than any context window 5:28 How far open-source models can go 7:58 Where specialist models fit in legal AI 10:59 Training legal models when client data is off-limits 13:06 Teaching a model how a law firm works 13:59 What belongs in context vs. model weights 15:36 From firm-wide AI to a model for every lawyer 18:37 What training adds beyond retrieving the right cases 20:26 Why context windows have plateaued 24:01 How models could learn continuously on the job 26:12 Can AI recursively improve AI research? 27:07 Research agents can run experiments but not choose them 30:00 Why open-ended research is hard to train 33:47 Why deployment, not intelligence, is the bottleneck 35:08 The cost of frontier intelligence 36:59 Different neolabs, different paths to intelligence 39:26 Using open datasets to compare research methods 41:13 Conclusion

Charlie O'Neill

93,422 次观看 • 2 个月前