Loading video...
Video Failed to Load
Announcing Ship: an endpoint with the highest intelligence per dollar of any frontier model. Today, Ship makes using Opus and GPT 5.6 Sol 50% cheaper guaranteed by a quality SLA.
656,428 views • 2 months ago •via X (Twitter)
48 Comments

On every major benchmark, Ship models give similar or higher scores. We are so confident this extends to production traffic that we offer the industry’s first SLA on model quality, not just uptime. And 50% cheaper isn't just “up to 50%” or "50% on average", it's 50% less for every request. We take the risk that any individual request costs us more to execute.

Just change model="<original>" to model="ship-like/<original>". Same capability. Same behavior. Exactly half the price. You can see the evals and begin using it here:

Ship's closed alpha grew to trillions of tokens. Two of our favorite partners: @kilocode launched us as a stealth model to hundreds of thousands of users and reduced costs for developers @baz_scm who've channeled millions in savings to new features and higher quality

is a lab building best execution for intelligence, incubated by Martian, and Ship is just the beginning. The biggest questions in AI are economic. The answers are fundamentally ML. We're doing both. Ex-quant to understand the problem, ex-Anthropic and ex-DeepMind to solve it.

Ship is available now. Get an API key and try it: If you have private evals or production traffic, we’ll help you compare Ship directly against the models you use today.

Lower inference costs are nice. Maintaining the same behavior is the hard part. 👀

Congrats on the launch. Inference cost presents a ceiling on how many agent runs, retries, and verification passes a team can afford, so a fixed 50% cut that keeps the original model's quality and behavior goes a long way. Excited to try this out.

Let us know what you think!

Really cool!!

Can you elaborate in what the tradeoff is? I get you are offering lower prices while maintaining capabilities, so I’m assuming that you are doing that at lower latency. Any other tradeoff?

Our goal is to ensure absolute parity, so there's shouldn't really be a tradeoff. In practice, there might be some requests which miss the quality bar, but that's what the SLO and SLA are for! So we've removed as much of the risk as possible

Damn..Same Quality of GPT 5.6 SOL and Opus 4.8 at 50% cost ...

Please tell me this isn't just another benchmaxed model. My sanity can't take another one. 💀

We've tried to be very careful not to benchmax! We report a bunch of stats here, and I think behavioral equivalence is evidence against benchmaxing: But the best thing is to try it out!

okay, let me try it. I'll be dropping some feedback later on.

Really excited to see the conversation expand beyond raw intelligence. As agentic workflows become longer + more complex, speed and cost stop being secondary considerations for sure!

amazing news with the rising cost of intelligence with those models

could you at least provide an opt-out to retaining all our data?

developers read benchmarks founders read invoices

This is more ambitious than a model router. A router selects from a list of models. Ship searches across models, tools, harnesses, cascades, ensembles, programs, and other execution methods to find a cheaper way to satisfy the requested quality grade.

Wow that’s awesome 🙏

50% off is a no brainer.

insane stuff, how does it compare on latency? i would assume a meaningful impact since im guessing its not just heuristics that let you choose the approach that would work

We actually make this comparison in the blog post. On average we're (very slightly) faster! There are a lot of engineering tricks to make this possible, but the intuition is that you can start responding with faster models once you measure the other characteristics (quality, cost) carefully enough that they won't degrade

I so much like the JIT framing here. Would like Statistical non-inferiority instead of token-identity anyday.

Yeah, it's interesting to think about what it means for two AI models to "be the same". That's actually why @thesean_labs is named after

Cutting the bill in half without downgrading the model is a hell of a pitch.

this could be a big win 🫡

when people hear "cheaper endpoint" they picture a router grabbing whatever's cheapest right now. that's not what this is. ship hunts across models, tools, harnesses, cascades, ensembles and interventions to find the lowest-cost way to run your task while still matching your reference model's grade. the result: the same quality you get today, for a flat 50% less. point your calls at the ship-like version of what you already run and either pocket half the bill or plow it back into more quality on the same budget.

Very cool! well done team! Any possibilities to chose the models that would be used under the hood for compliance reasons?

For large-volume users, we do support this! Reach out to the Thesean team: [email protected]

This is 🔥

awesome! can't wait to test it properly!

DeepSWE bench please

Working on it!

The interesting claim here is not “cheaper AI.” It is a fixed 50% lower price while targeting the capabilities and behavior of the model your application already depends on. That is a much harder problem than routing every request to a cheaper model.

this is ridiculously cool and means the price of intelligence on tap just went down

this is amazing!

Here's what gets overlooked inference isn't free, and that cost quietly drains every AI product you build. Slash it by half and the math changes fast. Features you parked come back to life, a free tier finally pencils out, and you can open the doors to way more users. Running your product for a flat 50% less while keeping the exact same model quality matters more than people assume. You just point to the ship-like version of the model you're already running, and that's it.

Efficiency work usually happens before your model ever sees a real problem. This spends the optimization budget after, once it knows what you're solving. That's a genuinely different lever. Great work!

I'm very curious about how they managed to reduce the cost of Opus and GPT 5.6 Sol by 50% while surpassing all frontier models in intelligence per dollar?

it's awesome work

::.

Better performance at a lower cost is always welcome.

Congrats 👏, 50% cheaper is huge.

This is so good news

Now we test...

Let us know what you think! Very receptive to feedback rn
