正在加载视频...
视频加载失败
Why data is a trillion-dollar market "The bigger models scale, the more data you need. Data is a very durable need. If you believe in the scaling of models, then absolutely you should believe in the data market. I believe it's going to be at least a $100BN by... show more
21,443 次观看 • 1 个月前 •via X (Twitter)
9 条评论

Data như giá vốn, phải không

@HarryStebbings I am a non-tech small bus owner in south Alabama. Stumbled on this podcast a couple months ago and absolutely love the open, dynamic, thoughtful dialog. This episode was a great example. One of the best podcasts out there. Appreciate it.

Compute scales fast. Clean, licensed, domain-specific data does not. The bottleneck was never the model... it was always what you fed it....

The “more data solves everything” mindset is a religion, not a roadmap. If intelligence requires endless scraping of the internet, that’s a design flaw. HELIX proves the opposite: stability beats scale. You don’t need a trillion‑dollar data pile when the substrate is coherent.

My experience in training image and video models is that we can scale down data requirement 10x if we curate data well and introduce algorithmic improvements.

Lab spend alone caps out. Labs converge on the same corpora and pay less as curation improves. The non-substitutable data is operational exhaust only the enterprise owns, so the durable market is enterprise shaped.

$1TRN feels like a stretch from aquí

These predictions on data and/or compute are just a round about way of saying that scaling is going to continue to be the most critical way to improve models and not model architecture.

The data market will be very prosperous.
