正在加载视频...
视频加载失败
SITUATION EXPLAINED: Cognition's new model is post-trained on Kimi K3. • SWE-2 scores 50.0% on FrontierCode 1.1 Main, Cognition's benchmark for whether a maintainer would merge the pull request • Fable 5.1 gets 50.9 at 64% higher cost. Grok 4.6 gets 48.0, Sol 47.5, Astra 53.3 • It leads... show more
13,488 次观看 • 10 天前 •via X (Twitter)
1 条评论

Ricci Research9 天前
53 steps versus 127 for the same result is the number with real economics behind it — agent costs scale with trajectory length, not parameter count, so halving the steps does more for unit economics than any pricing change. Also worth sitting with: an American coding company's frontier product is post-trained on Chinese open weights, and that's now just a routine architecture decision.
