Video wird geladen...
Video konnte nicht geladen werden
Vals AI CEO Rayan Krishnan on why defensible evals are critical for enterprises' survival: "The labs side of it is very clear. If you're raising lots of money, investing heavily in building models, it's essential for you to show why your model is getting better and why the customer... show more
34,017 Aufrufe • vor 2 Tagen •via X (Twitter)
12 Kommentare

A firm is its evals is a sharp framing. The tension is that the most defensible evals are often too proprietary to become trusted external proof.

The independent testing layer can't come fast enough. In diligence rooms I still get handed the vendor's own benchmark deck as proof the model works, and nobody asks who ran the test. Who pays the evaluator decides whether any of it holds.

Thanks for making it so that the regular investor cant invest

token spend vs salary is wild. audit it or die

Recording the first sign that my range estimate was off gave me something concrete to work on in reviews.

Defensible evals are a nice story, but enterprise survival depends on actual deployment, not benchmarks. What’s your threshold for trusting an eval over real-world results?

Absolutely. Once AI becomes a meaningful operating cost, being able to measure performance and ROI becomes essential for every enterprise.

A habit that stayed with me: writing a reason before returning to an old idea.

The problem with most companies is their workflows are not stable. AI produces different verdict every time. Companies are focusing of accuracy but not stability. Just throwing bigger and bigger models on your workflows won't help. Developers have to build pass^k and pass@k evals effectively.

evals really are the moat now

The credit-ratings analogy is the part worth stressing: raters got paid by issuers and drifted. If eval vendors are paid by the labs they grade, the same conflict shows up, just faster. Independent funding is the whole ballgame.

Defensible evals are crucial for survival
