Loading video...
Video Failed to Load
.christina kim says the frontier isn't benchmarks anymore. It's usage. Eval scores are saturated, but daily life isn’t. The real signal of progress is how many people use AI to get real things done. That’s how we’ll know we’re approaching AGI.
77,182 views • 1 year ago •via X (Twitter)
19 Comments

@christinahkim Yes! Usage over benchmarks! Now let's think about it: when thousands of users are begging to keep using 4o daily, and 5 can't replicate what made 4o special... are we calling that progress toward AGI or just really expensive product mismanagement?

@christinahkim @christinahkim so proud!!

@christinahkim Level of general usage has nothing to do with AGI you morons.

@christinahkim I agree, usage reflects true impact. Focus on practical applications will drive AI development forward meaningfully.

@christinahkim Use $PLTR as the leading indicator for getting "real things done." So far, so good.

@christinahkim it's all about AI in daily life, not just benchmarks

@christinahkim Exactly this. I've been thinking benchmarks are becoming vanity metrics - what matters is whether AI actually makes people's work better, not whether it scores 0.5% higher on some eval. The real test is: would you miss it if it was gone tomorrow?

In the case of n=1, what I care about with each ai model or product is “what can I easily do today, that I couldn’t do yesterday?” I appreciate incremental improvements but impressive advances do better. Good answers to that make it worthwhile to use new versions or products, otherwise i may as well stick with one UI already using a lot or the few UIs using etc.

@christinahkim Agree

@tylercowen @christinahkim

@christinahkim 🎯

@christinahkim One sign may be when AI that can rewrite all the banking systems COBOL ?

@christinahkim

@christinahkim Absolutely spot on. The gap between what AI can do in demos vs what people actually rely on daily is huge. Real progress will be measured by how seamlessly it integrates into workflows, not leaderboard scores

@christinahkim Widespread daily use is a compelling metric, but true AGI necessitates transformative impact, not just utility.

@christinahkim Focus on practical daily use

@christinahkim The true test is in the boring but important work where the rules are locked in the minds of entire teams. @autalyco is tackling accounts receivable in logistics and other ops heavy industry. A lot of work to do industries with messy data.

@christinahkim benchmarks for the deck, usage is real scoreboard

@christinahkim yes, the eval saturation is true and the end user hardly cares about the metric.

