Video wird geladen...
Video konnte nicht geladen werden
SITUATION EXPLAINED: Cognition's new model is post-trained on Kimi K3. • SWE-2 scores 50.0% on FrontierCode 1.1 Main, Cognition's benchmark for whether a maintainer would merge the pull request • Fable 5.1 gets 50.9 at 64% higher cost. Grok 4.6 gets 48.0, Sol 47.5, Astra 53.3 • It leads... show more
13,488 Aufrufe • vor 10 Tagen •via X (Twitter)
1 Kommentare

Ricci Researchvor 9 Tagen
53 steps versus 127 for the same result is the number with real economics behind it — agent costs scale with trajectory length, not parameter count, so halving the steps does more for unit economics than any pricing change. Also worth sitting with: an American coding company's frontier product is post-trained on Chinese open weights, and that's now just a routine architecture decision.
