Загрузка видео...

Не удалось загрузить видео

На главную

SITUATION EXPLAINED: How much are frontier labs actually spending on training data? .Sean Cai: "Frontier labs are spending about $10 to $15 billion per lab on data." "Really good long horizon tasks go up to $20,000 each. A complete browser-use version of SAP was rumored at $500,000." "Despite everybody...

334,625 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

Yesterday I interviewed Sean Cai about AI data. This is essentially a guide for founders on how to sell data and RL envs to AI labs. "I've never seen a data contract get turned down by a top lab, if it's good quality data, for budget reasons." 00:00 What areas of data are underserved? 02:10 For bio data, is it real-world or purely digital? 04:21 For cyber data, which subsets are most underserved? 05:50 What is the sales process like? 07:04 Why would a lab not renew or increase their purchase volume? 10:13 When a researcher is exploring a new direction, what's the first step? 11:35 In robotics data, what do you view as underserved? 13:12 What does the initial data delivery look like, what format? 13:53 Do labs have more sophisticated internal setups for running environments? 14:32 Are the non-frontier labs buying off-the-shelf data from Anthropic / OpenAI vendors? 16:11 Do Anthropic data vendors put expiry timeframes on the exclusivity? 16:42 Are purchase decisions researcher-led? 17:41 Decagon, Sierra, Ramp: what kinds of data are they buying? 19:06 Long-term, when do labs still need to buy external data vs train on user traces? 21:15 Will end-vendor benchmarks shift to performance per dollar? 22:04 How many labs are spending at the 1B+/yr data level? 23:53 Delta between Anthropic's stated $1B and your 10-20B/lab number? 26:05 What makes inference providers / neoclouds a good fit to acquire RL env cos?

Chris Barber (in SF)

130,324 просмотров • 2 месяцев назад

David Sacks says companies are trapped paying OpenAI & Anthropic because they can't figure out how to use open source models "I think enterprise CTOs would like to shift their token consumption to cheaper models for the obvious reason that it would be more efficient. They are seeing compute costs or token costs skyrocket right now, so everyone's trying to figure this out." "You also have the AI sovereignty issue that Alex Karp talked about. They're worried about giving up the secret sauce or the alpha in their business to a frontier lab that may one day be competing with them. "The problem is, I think in most cases, they don't have the technical ability to do it. Coinbase figured out how to do it. DoorDash figured out how to do it. They built a token routing system that allows them to send frontier tasks to frontier models and non frontier tasks to more mundane models. But I don't think your average enterprise has the technical capability to do that." "This is why the share of wallet of closed models, it actually increased. I think that open source went from 19% last year to 11% this year. So open source as a share of enterprise spending is actually decreasing." "I don't think that means usage is decreasing. I think usage is skyrocketing. It also may be the case that because the whole point of using an open model is you just pay for the compute costs, you don't have to pay a lab, so it may be that it's hard to measure that usage in terms of spend." "But nonetheless, anyone who's saying that these closed models are going to lose or are somehow losing, you're just not seeing it in the data."

dnap

110,207 просмотров • 13 дней назад