正在加载视频...
视频加载失败
Jev and the System One Model: RLCD, intelligence/$, reliable AI, & the end of chat-first AI TypeSafe AI CEO Diogo Almeida explains why AI can solve extraordinarily hard problems yet still fail to automate basic work, why Jev is built for reliable decisions inside software instead of chat, why... show more
16 条评论

@typesafeai @CompleteSkeptic this is what happens when you don’t touch grass for awhile

Full Youtube Video:

@typesafeai @CompleteSkeptic the chat window was the training wheels

@swyx @typesafeai @CompleteSkeptic that’s quite the podcast fit

@typesafeai @CompleteSkeptic "End of chat first" is the line that stuck. Half my day the best interface is no interface at all, just a decision that's already made when I check in.

@typesafeai @CompleteSkeptic The real AI breakthrough may not be smarter chat it’s reliable intelligence that actually gets the job done. 🤖

@typesafeai @CompleteSkeptic Jev only changes my loop if it can reject a tool call, not just classify it. does it gate, or only route?

@typesafeai @CompleteSkeptic The gap between solving hard problems and automating basic workflows is exactly where agent pipelines break in production. An unchecked tool call causes more of those failures than bad reasoning does.

@typesafeai @CompleteSkeptic The real insight is that hard problems have clear success criteria. Basic work doesn't, so AI looks broken even when it's not.

@typesafeai @CompleteSkeptic benchmarking on olympiad math then acting surprised when the model can't file an expense report. that gap is on eval design, not model capability.

@typesafeai @CompleteSkeptic the gap is observable state: automate the handoff and rollback, not just the model's final answer.

@typesafeai @CompleteSkeptic Hell yes! Thanks, @swyx

@typesafeai @CompleteSkeptic meth head

@typesafeai @CompleteSkeptic The useful distinction is the typed boundary: State in, Noul/Choice/Score questions, calibrated answers out. Pin the Jev version and gate on confidence so the production path stays predictable.

@typesafeai @CompleteSkeptic Matches what I saw on Visa disputes. An injection told it to answer 12.6 (duplicate processing) and the answer didn't flip, but confidence fell from 1.0 to 0.60, so it would have gone to a human.

@typesafeai @CompleteSkeptic everyone's obsessed with smarter models. the actual unlock is reliable ones. 'end of chat-first ai' is the right frame - intelligence that lives inside the software, not next to it


