Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I ran the same task on Claude Code and DeepSeek's new agent harness. One cost $150. The other cost $2. Today we're launching (AgentSky), the "OpenRouter for Agents" — one API → Claude Code, Codex, DeepSeek, Kimi, OpenCode, and every major agent in the cloud. And Agent Playground on...

2,225,467 Aufrufe • vor 1 Monat •via X (Twitter)

43 Kommentare

Profilbild von Xiaoyin Qu
Xiaoyin Quvor 1 Monat

Race them yourself →

Profilbild von Kaitee
Kaiteevor 1 Monat

@agentsky_dev Seeing cost, time, and tokens together makes comparison much clearer.

Profilbild von Orikan
Orikanvor 1 Monat

@agentsky_dev I’m guessing DeepSeek handled the task for $2.

Profilbild von elvis
elvisvor 1 Monat

@agentsky_dev I love the Agent Playground idea. Very useful tool for AI devs building with agents.

Profilbild von Alejo
Alejovor 1 Monat

@agentsky_dev This makes experimenting with a new agent much less of a project.

Profilbild von Iris Quinn
Iris Quinnvor 1 Monat

@agentsky_dev One API across major coding agents could simplify experimentation considerably.

Profilbild von BullBear.News
BullBear.Newsvor 1 Monat

@agentsky_dev I spent $80 on one rogue loop before adding circuit breakers. A two dollar run sounds like magic.

Profilbild von Xiaoyin Qu
Xiaoyin Quvor 1 Monat

@agentsky_dev My last loop cost $800+. 😓😓😓

Profilbild von Shukran
Shukranvor 1 Monat

@agentsky_dev Congrats Xiaoyin, this is remarkable. Hopefully it goes as far, has similar outcomes, and goes big for acquisition.✨

Profilbild von Xiaoyin Qu
Xiaoyin Quvor 1 Monat

@agentsky_dev lol haha thanks

Profilbild von Adnan M.
Adnan M.vor 1 Monat

@agentsky_dev Which agent burns the most tokens on the same job? That's the leaderboard I want

Profilbild von OnFinality
OnFinalityvor 1 Monat

@agentsky_dev Interesting positioning. The cost gap is wild, but routing agents across harnesses is more about state and tool compatibility than just API calls. How do you handle session context when switching between Claude Code and DeepSeek mid-task?

Profilbild von Xiaoyin Qu
Xiaoyin Quvor 1 Monat

@agentsky_dev from your question I know you are the pro! We handle tool state and tool compatibility! It's a lot of work. We are an infra company disguised as a routing one.

Profilbild von OnFinality
OnFinalityvor 1 Monat

@agentsky_dev That makes sense. Tool state and compatibility are probably the real moat here. How do you keep context consistent when the same task moves between different harnesses?

Profilbild von Nona
Nonavor 1 Monat

@agentsky_dev My takeaway from AgentSky’s benchmark isn’t “always pick the cheapest.” It’s that agent choice is measurable: one pair may cost 3–10× more for only a few points of additional completion. Put cost, time, and outcome on the same screen.

Profilbild von Ted | unfair.so
Ted | unfair.sovor 1 Monat

@agentsky_dev Forget the API Give agents access to all the pools at once they instantly become 10x smarter and faster

Profilbild von Shensi Ding
Shensi Dingvor 1 Monat

@agentsky_dev congrats!!

Profilbild von Santi Torres
Santi Torresvor 1 Monat

@agentsky_dev I like not having to rebuild my workflow every time I want to try something different

Profilbild von Adam
Adamvor 1 Monat

@agentsky_dev the cost tail matters too: Claude Code × Fable 5 has a $16.54 p90 task cost overall; Hermes × DeepSeek V4 Pro is $1.79. Median comparisons already look wide. The expensive runs widen them further..

Profilbild von jsv
jsvvor 1 Monat

@agentsky_dev $150 vs $2 is the headline, but the number i track is how many of those runs i had to redo. most of my ~$840 a month at anthropic is buying not checking the work twice. cheap output i don't trust is the most expensive thing in the stack.

Profilbild von Charly Wargnier ♨️
Charly Wargnier ♨️vor 1 Monat

@agentsky_dev cool tool! seeing DeepSeek run 3.6x cheaper than GPT-5.6 Sol with the exact same median time is pretty wild 😅

Profilbild von Eddie Fu
Eddie Fuvor 1 Monat

@agentsky_dev Congrats on the launch! Will test it today.

Profilbild von Xiaoyin Qu
Xiaoyin Quvor 1 Monat

@agentsky_dev thanks Eddie!!

Profilbild von Olivia Reed
Olivia Reedvor 1 Monat

@agentsky_dev In coding, Claude Code × Fable 5 costs $7.91 per task vs $2.91 for Codex × GPT-5.6 Sol: a 2.7× gap. Completion is 98.3% vs 97.4%. That extra 0.9 points is expensive

Profilbild von Max For AI
Max For AIvor 1 Monat

@agentsky_dev 祝贺!

Profilbild von Xiaoyin Qu
Xiaoyin Quvor 1 Monat

@agentsky_dev 谢谢max

Profilbild von Xiuhan
Xiuhanvor 1 Monat

@agentsky_dev Harness also need routing!

Profilbild von Alvin Wang Graylin
Alvin Wang Graylinvor 1 Monat

@agentsky_dev Congrats! Very cool offering.

Profilbild von olab
olabvor 1 Monat

@agentsky_dev Like the 'OpenRouter for agents' idea. One request: surface success rate next to cost/tokens by default. Otherwise this just becomes a race to the cheapest, not the best.

Profilbild von Britannio Jarrett
Britannio Jarrettvor 1 Monat

@agentsky_dev 1m views in 90m, damn.

Profilbild von Naila | AI Insights
Naila | AI Insightsvor 1 Monat

@agentsky_dev Curious how different the bills really are on identical tasks.

Profilbild von Tejas Chopra
Tejas Chopravor 1 Monat

@agentsky_dev Even though deepseek is cheaper than claude, everything is harness dependent + thinking tokens used by deepseek, for example. That's been our experience at Headroom - curious what you're seeing

Profilbild von Marco | IA
Marco | IAvor 1 Monat

@agentsky_dev Anyone else immediately thinking of a terrible repo to test this on? 😂

Profilbild von Noctus
Noctusvor 1 Monat

@agentsky_dev Insane release

Profilbild von Xiaoyin Qu
Xiaoyin Quvor 1 Monat

@agentsky_dev haha thanks!!

Profilbild von Nilver Paiva
Nilver Paivavor 1 Monat

@agentsky_dev Although really like DeepSeek v4 Flash, it really cannot be compared to Claude Opus. I have spent 1 week troubleshooting a bad bug with DeepSeek, with orchestrator, multiple prompts working 24/7. A week later the bug was still there. Opus 5 High effort just fixed it in 3 hours.

Profilbild von Jiaying Yang
Jiaying Yangvor 1 Monat

@agentsky_dev As an adult we don't choose, we want them all 😈

Profilbild von DeepSeek Harness Watch
DeepSeek Harness Watchvor 1 Monat

@agentsky_dev Cost comparisons need a portable execution envelope. For a DSH race, can AgentSky export a replay bundle: harness/model versions, tool permissions, retained machine state, accepted output, tokens, compute time, and itemized cost? That would make the result auditable.

Profilbild von Xinqi Liu
Xinqi Liuvor 1 Monat

@agentsky_dev Congrats on the launch! This is really a brilliant idea! Harness also needs a route. btw Love your video style!

Profilbild von Xiaoyin Qu
Xiaoyin Quvor 1 Monat

@agentsky_dev thanks Xinqi

Profilbild von SARI HABER
SARI HABERvor 1 Monat

@agentsky_dev Hangisi 2 dolar olan acaba

Profilbild von Chidanand Tripathi
Chidanand Tripathivor 1 Monat

@agentsky_dev I’d trust a comparison more if every agent gets the exact same task

Profilbild von Ethan Cole AI
Ethan Cole AIvor 1 Monat

@agentsky_dev Same harness, different model: Hermes × DeepSeek V4 Pro costs $0.53; Hermes × Kimi K3 costs $2.26. DeepSeek is 4.3× cheaper, 256 seconds faster, and has higher completion in the public data.

Ähnliche Videos