Загрузка видео...
Не удалось загрузить видео
We have completed migrating all 11 million users on the Cline extension to our new SDK harness, built for open weights models. Task mistake rates have dropped from 6.34% → 0.62%. This was our biggest update ever, and almost nobody noticed since it was entirely under the hood.
16,837 просмотров • 13 дней назад •via X (Twitter)
Комментарии: 20

We used a special slow rollout process over the last month, which took some magic to work in VS Marketplace. More on the process here:

We’ve invested heavily into improving our harness with the Cline SDK. This marks a milestone where all Cline products are running off this new harness to get the best open weights performance.

Incredible work!

dropped from 6.34% to 0.62% and shipped it under the hood. that 5.72% difference is every user who stopped blaming themselves for the agent's mistakes

Awesome

Invisible to users, massive under the hood. Moving 11M users to a new SDK harness and reducing task mistake rates from 6.34% to 0.62% is an incredible engineering achievement. Congrats to the Cline team.

When Muse 1.3?

6% to 0.6% and nobody noticed. that's the only kind of update i trust.

@midudev que piensas de Cline

6.34% -> 0.62% task mistakes and nobody noticed, that's the whole story: the fix was the SDK harness, not a new model. swapping the foundation under 11M installs without breaking the IDE is the harder ship. the slow VS Marketplace rollout is the part nobody credits.

People still use cline?

Is memory ballooning issue solved now for long context conversations?

Anyone still using cline? Anyone still using VSCode?

A tenfold mistake-rate reduction shows harness design can outweigh model-brand differences.

10x fewer mistakes from changing the harness is the clearest proof that the runtime is part of the model.

Shipping that under the hood is a massive engineering win

Huge win 🔥 10x drop in mistakes and zero disruption.

Quiet harness swaps are the best kind of ship. We did one last month and nobody pinged us either. That 6.34% → 0.62% mistake drop is the changelog line I'd actually lead with.

That 10x drop is the tell — once tool-call validation + retry/repair live in the harness, open-weights stop looking "worse" and start looking under-scaffolded. Curious what share of the residual 0.62% is still harness edge cases (diff apply, truncated context) vs actual model reasoning misses.

Cline. which workflow still blows the budget?
