正在加载视频...

视频加载失败

I ran the same task on Claude Code and DeepSeek's new agent harness. One cost $150. The other cost $2. Today we're launching (AgentSky), the "OpenRouter for Agents" — one API → Claude Code, Codex, DeepSeek, Kimi, OpenCode, and every major agent in the cloud. And Agent Playground on...

2,225,467 次观看 • 1 个月前 •via X (Twitter)

43 条评论

Xiaoyin Qu 的头像
Xiaoyin Qu1 个月前

Race them yourself →

Kaitee 的头像
Kaitee1 个月前

@agentsky_dev Seeing cost, time, and tokens together makes comparison much clearer.

Orikan 的头像
Orikan1 个月前

@agentsky_dev I’m guessing DeepSeek handled the task for $2.

elvis 的头像
elvis1 个月前

@agentsky_dev I love the Agent Playground idea. Very useful tool for AI devs building with agents.

Alejo 的头像
Alejo1 个月前

@agentsky_dev This makes experimenting with a new agent much less of a project.

Iris Quinn 的头像
Iris Quinn1 个月前

@agentsky_dev One API across major coding agents could simplify experimentation considerably.

BullBear.News 的头像
BullBear.News1 个月前

@agentsky_dev I spent $80 on one rogue loop before adding circuit breakers. A two dollar run sounds like magic.

Xiaoyin Qu 的头像
Xiaoyin Qu1 个月前

@agentsky_dev My last loop cost $800+. 😓😓😓

Shukran 的头像
Shukran1 个月前

@agentsky_dev Congrats Xiaoyin, this is remarkable. Hopefully it goes as far, has similar outcomes, and goes big for acquisition.✨

Xiaoyin Qu 的头像
Xiaoyin Qu1 个月前

@agentsky_dev lol haha thanks

Adnan M. 的头像
Adnan M.1 个月前

@agentsky_dev Which agent burns the most tokens on the same job? That's the leaderboard I want

OnFinality 的头像
OnFinality1 个月前

@agentsky_dev Interesting positioning. The cost gap is wild, but routing agents across harnesses is more about state and tool compatibility than just API calls. How do you handle session context when switching between Claude Code and DeepSeek mid-task?

Xiaoyin Qu 的头像
Xiaoyin Qu1 个月前

@agentsky_dev from your question I know you are the pro! We handle tool state and tool compatibility! It's a lot of work. We are an infra company disguised as a routing one.

OnFinality 的头像
OnFinality1 个月前

@agentsky_dev That makes sense. Tool state and compatibility are probably the real moat here. How do you keep context consistent when the same task moves between different harnesses?

Nona 的头像
Nona1 个月前

@agentsky_dev My takeaway from AgentSky’s benchmark isn’t “always pick the cheapest.” It’s that agent choice is measurable: one pair may cost 3–10× more for only a few points of additional completion. Put cost, time, and outcome on the same screen.

Ted | unfair.so 的头像
Ted | unfair.so1 个月前

@agentsky_dev Forget the API Give agents access to all the pools at once they instantly become 10x smarter and faster

Shensi Ding 的头像
Shensi Ding1 个月前

@agentsky_dev congrats!!

Santi Torres 的头像
Santi Torres1 个月前

@agentsky_dev I like not having to rebuild my workflow every time I want to try something different

Adam 的头像
Adam1 个月前

@agentsky_dev the cost tail matters too: Claude Code × Fable 5 has a $16.54 p90 task cost overall; Hermes × DeepSeek V4 Pro is $1.79. Median comparisons already look wide. The expensive runs widen them further..

jsv 的头像
jsv1 个月前

@agentsky_dev $150 vs $2 is the headline, but the number i track is how many of those runs i had to redo. most of my ~$840 a month at anthropic is buying not checking the work twice. cheap output i don't trust is the most expensive thing in the stack.

Charly Wargnier ♨️ 的头像
Charly Wargnier ♨️1 个月前

@agentsky_dev cool tool! seeing DeepSeek run 3.6x cheaper than GPT-5.6 Sol with the exact same median time is pretty wild 😅

Eddie Fu 的头像
Eddie Fu1 个月前

@agentsky_dev Congrats on the launch! Will test it today.

Xiaoyin Qu 的头像
Xiaoyin Qu1 个月前

@agentsky_dev thanks Eddie!!

Olivia Reed 的头像
Olivia Reed1 个月前

@agentsky_dev In coding, Claude Code × Fable 5 costs $7.91 per task vs $2.91 for Codex × GPT-5.6 Sol: a 2.7× gap. Completion is 98.3% vs 97.4%. That extra 0.9 points is expensive

Max For AI 的头像
Max For AI1 个月前

@agentsky_dev 祝贺!

Xiaoyin Qu 的头像
Xiaoyin Qu1 个月前

@agentsky_dev 谢谢max

Xiuhan 的头像
Xiuhan1 个月前

@agentsky_dev Harness also need routing!

Alvin Wang Graylin 的头像
Alvin Wang Graylin1 个月前

@agentsky_dev Congrats! Very cool offering.

olab 的头像
olab1 个月前

@agentsky_dev Like the 'OpenRouter for agents' idea. One request: surface success rate next to cost/tokens by default. Otherwise this just becomes a race to the cheapest, not the best.

Britannio Jarrett 的头像
Britannio Jarrett1 个月前

@agentsky_dev 1m views in 90m, damn.

Naila | AI Insights 的头像
Naila | AI Insights1 个月前

@agentsky_dev Curious how different the bills really are on identical tasks.

Tejas Chopra 的头像
Tejas Chopra1 个月前

@agentsky_dev Even though deepseek is cheaper than claude, everything is harness dependent + thinking tokens used by deepseek, for example. That's been our experience at Headroom - curious what you're seeing

Marco | IA 的头像
Marco | IA1 个月前

@agentsky_dev Anyone else immediately thinking of a terrible repo to test this on? 😂

Noctus 的头像
Noctus1 个月前

@agentsky_dev Insane release

Xiaoyin Qu 的头像
Xiaoyin Qu1 个月前

@agentsky_dev haha thanks!!

Nilver Paiva 的头像
Nilver Paiva1 个月前

@agentsky_dev Although really like DeepSeek v4 Flash, it really cannot be compared to Claude Opus. I have spent 1 week troubleshooting a bad bug with DeepSeek, with orchestrator, multiple prompts working 24/7. A week later the bug was still there. Opus 5 High effort just fixed it in 3 hours.

Jiaying Yang 的头像
Jiaying Yang1 个月前

@agentsky_dev As an adult we don't choose, we want them all 😈

DeepSeek Harness Watch 的头像
DeepSeek Harness Watch1 个月前

@agentsky_dev Cost comparisons need a portable execution envelope. For a DSH race, can AgentSky export a replay bundle: harness/model versions, tool permissions, retained machine state, accepted output, tokens, compute time, and itemized cost? That would make the result auditable.

Xinqi Liu 的头像
Xinqi Liu1 个月前

@agentsky_dev Congrats on the launch! This is really a brilliant idea! Harness also needs a route. btw Love your video style!

Xiaoyin Qu 的头像
Xiaoyin Qu1 个月前

@agentsky_dev thanks Xinqi

SARI HABER 的头像
SARI HABER1 个月前

@agentsky_dev Hangisi 2 dolar olan acaba

Chidanand Tripathi 的头像
Chidanand Tripathi1 个月前

@agentsky_dev I’d trust a comparison more if every agent gets the exact same task

Ethan Cole AI 的头像
Ethan Cole AI1 个月前

@agentsky_dev Same harness, different model: Hermes × DeepSeek V4 Pro costs $0.53; Hermes × Kimi K3 costs $2.26. DeepSeek is 4.3× cheaper, 256 seconds faster, and has higher completion in the public data.

相关视频