Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Vals AI CEO Rayan Krishnan on why defensible evals are critical for enterprises' survival: "The labs side of it is very clear. If you're raising lots of money, investing heavily in building models, it's essential for you to show why your model is getting better and why the customer...

34,017 Aufrufe • vor 2 Tagen •via X (Twitter)

12 Kommentare

Profilbild von Arnold
Arnoldvor 2 Tagen

A firm is its evals is a sharp framing. The tension is that the most defensible evals are often too proprietary to become trusted external proof.

Profilbild von Lynn Welch
Lynn Welchvor 2 Tagen

The independent testing layer can't come fast enough. In diligence rooms I still get handed the vendor's own benchmark deck as proof the model works, and nobody asks who ran the test. Who pays the evaluator decides whether any of it holds.

Profilbild von SAI
SAIvor 2 Tagen

Thanks for making it so that the regular investor cant invest

Profilbild von icefrog.◎
icefrog.◎vor 2 Tagen

token spend vs salary is wild. audit it or die

Profilbild von asa
asavor 2 Tagen

Recording the first sign that my range estimate was off gave me something concrete to work on in reviews.

Profilbild von sophs b
sophs bvor 2 Tagen

Defensible evals are a nice story, but enterprise survival depends on actual deployment, not benchmarks. What’s your threshold for trusting an eval over real-world results?

Profilbild von Weaver Labs
Weaver Labsvor 2 Tagen

Absolutely. Once AI becomes a meaningful operating cost, being able to measure performance and ROI becomes essential for every enterprise.

Profilbild von Saminu suleiman
Saminu suleimanvor 2 Tagen

A habit that stayed with me: writing a reason before returning to an old idea.

Profilbild von Apoorv Agarwal
Apoorv Agarwalvor 2 Tagen

The problem with most companies is their workflows are not stable. AI produces different verdict every time. Companies are focusing of accuracy but not stability. Just throwing bigger and bigger models on your workflows won't help. Developers have to build pass^k and pass@k evals effectively.

Profilbild von AI Mastery Guide
AI Mastery Guidevor 2 Tagen

evals really are the moat now

Profilbild von Modelplane
Modelplanevor 2 Tagen

The credit-ratings analogy is the part worth stressing: raters got paid by issuers and drifted. If eval vendors are paid by the labs they grade, the same conflict shows up, just faster. Independent funding is the whole ballgame.

Profilbild von Rosetina Degen 🚢🎒
Rosetina Degen 🚢🎒vor 2 Tagen

Defensible evals are crucial for survival

Ähnliche Videos

Vals AI co-founder and CEO Rayan Krishnan with a16z's Ben Horowitz and Jennifer Li on grading AI, what it costs, and who gets to make the rules: Every big industry eventually grows an independent testing layer. AI has credit ratings to learn from and Enron to avoid. Model capability today is still mostly self-reported. As public benchmarks saturate and models get better at optimizing for the tests themselves, Rayan makes the case for independent, continuously evolving evaluations. The harder problem is geopolitical. Reagan's "trust but verify" worked during the Cold War because you could fly over and count the missiles. No simple equivalent for AI models exists. In this conversation with Erik Torenberg, they get into how you measure a model's ability to improve itself, why every good benchmark eventually has to be retired, and what happens when token spend begins to rival employee salaries. 00:00 Intro 02:20 Llama 4 on public vs private benchmarks 05:24 Nobody agreed how to test humans either 06:55 What movie ratings teach us about AI 08:55 The Enron problem in benchmarking 11:36 Why a good benchmark has to be retired 13:22 Evals that run for weeks, not seconds 16:20 Where the real workday starts at 4pm 18:08 A firm really is just its evals 20:35 Why Sonnet can cost more than Opus 22:42 One engineer, 6 billion tokens in a day 25:05 Who should set the rules for models 28:55 Public sector enforces, private verifies 33:32 Why sovereign AI is inefficient and happening anyway 35:00 The AI version of trust-but-verify 37:15 Where cyber evals have to go next YouTube: Rayan Krishnan Vals AI benahorowitz.eth Jennifer Li Erik Torenberg

a16z

154,318 Aufrufe • vor 3 Tagen

David Sacks says companies are trapped paying OpenAI & Anthropic because they can't figure out how to use open source models "I think enterprise CTOs would like to shift their token consumption to cheaper models for the obvious reason that it would be more efficient. They are seeing compute costs or token costs skyrocket right now, so everyone's trying to figure this out." "You also have the AI sovereignty issue that Alex Karp talked about. They're worried about giving up the secret sauce or the alpha in their business to a frontier lab that may one day be competing with them. "The problem is, I think in most cases, they don't have the technical ability to do it. Coinbase figured out how to do it. DoorDash figured out how to do it. They built a token routing system that allows them to send frontier tasks to frontier models and non frontier tasks to more mundane models. But I don't think your average enterprise has the technical capability to do that." "This is why the share of wallet of closed models, it actually increased. I think that open source went from 19% last year to 11% this year. So open source as a share of enterprise spending is actually decreasing." "I don't think that means usage is decreasing. I think usage is skyrocketing. It also may be the case that because the whole point of using an open model is you just pay for the compute costs, you don't have to pay a lab, so it may be that it's hard to measure that usage in terms of spend." "But nonetheless, anyone who's saying that these closed models are going to lose or are somehow losing, you're just not seeing it in the data."

dnap

110,354 Aufrufe • vor 2 Monaten