Loading video...

Video Failed to Load

Go Home

Announcing Ship: an endpoint with the highest intelligence per dollar of any frontier model. Today, Ship makes using Opus and GPT 5.6 Sol 50% cheaper guaranteed by a quality SLA.

656,428 views • 2 months ago •via X (Twitter)

48 Comments

Martian's profile picture
Martian2 months ago

On every major benchmark, Ship models give similar or higher scores. We are so confident this extends to production traffic that we offer the industry’s first SLA on model quality, not just uptime. And 50% cheaper isn't just “up to 50%” or "50% on average", it's 50% less for every request. We take the risk that any individual request costs us more to execute.

Martian's profile picture
Martian2 months ago

Just change model="<original>" to model="ship-like/<original>". Same capability. Same behavior. Exactly half the price. You can see the evals and begin using it here:

Martian's profile picture
Martian2 months ago

Ship's closed alpha grew to trillions of tokens. Two of our favorite partners: @kilocode launched us as a stealth model to hundreds of thousands of users and reduced costs for developers @baz_scm who've channeled millions in savings to new features and higher quality

Martian's profile picture
Martian2 months ago

is a lab building best execution for intelligence, incubated by Martian, and Ship is just the beginning. The biggest questions in AI are economic. The answers are fundamentally ML. We're doing both. Ex-quant to understand the problem, ex-Anthropic and ex-DeepMind to solve it.

Martian's profile picture
Martian2 months ago

Ship is available now. Get an API key and try it: If you have private evals or production traffic, we’ll help you compare Ship directly against the models you use today.

Tiffany Fong's profile picture
Tiffany Fong2 months ago

Lower inference costs are nice. Maintaining the same behavior is the hard part. 👀

elvis's profile picture
elvis2 months ago

Congrats on the launch. Inference cost presents a ceiling on how many agent runs, retries, and verification passes a team can afford, so a fixed 50% cut that keeps the original model's quality and behavior goes a long way. Excited to try this out.

Martian's profile picture
Martian2 months ago

Let us know what you think!

Chubby♨️'s profile picture
Chubby♨️2 months ago

Really cool!!

Santiago's profile picture
Santiago2 months ago

Can you elaborate in what the tradeoff is? I get you are offering lower prices while maintaining capabilities, so I’m assuming that you are doing that at lower latency. Any other tradeoff?

Martian's profile picture
Martian2 months ago

Our goal is to ensure absolute parity, so there's shouldn't really be a tradeoff. In practice, there might be some requests which miss the quality bar, but that's what the SLO and SLA are for! So we've removed as much of the risk as possible

AshutoshShrivastava's profile picture
AshutoshShrivastava2 months ago

Damn..Same Quality of GPT 5.6 SOL and Opus 4.8 at 50% cost ...

OddSignal's profile picture
OddSignal2 months ago

Please tell me this isn't just another benchmaxed model. My sanity can't take another one. 💀

Martian's profile picture
Martian2 months ago

We've tried to be very careful not to benchmax! We report a bunch of stats here, and I think behavioral equivalence is evidence against benchmaxing: But the best thing is to try it out!

OddSignal's profile picture
OddSignal2 months ago

okay, let me try it. I'll be dropping some feedback later on.

Sarah Chieng's profile picture
Sarah Chieng2 months ago

Really excited to see the conversation expand beyond raw intelligence. As agentic workflows become longer + more complex, speed and cost stop being secondary considerations for sure!

Min Choi's profile picture
Min Choi2 months ago

amazing news with the rising cost of intelligence with those models

Sam L's profile picture
Sam L1 month ago

could you at least provide an opt-out to retaining all our data?

nabu's profile picture
nabu1 month ago

developers read benchmarks founders read invoices

Csaba Kissi's profile picture
Csaba Kissi2 months ago

This is more ambitious than a model router. A router selects from a list of models. Ship searches across models, tools, harnesses, cascades, ensembles, programs, and other execution methods to find a cheaper way to satisfy the requested quality grade.

WorldofAI's profile picture
WorldofAI2 months ago

Wow that’s awesome 🙏

Whale Insider's profile picture
Whale Insider2 months ago

50% off is a no brainer.

bricklerex's profile picture
bricklerex2 months ago

insane stuff, how does it compare on latency? i would assume a meaningful impact since im guessing its not just heuristics that let you choose the approach that would work

Martian's profile picture
Martian2 months ago

We actually make this comparison in the blog post. On average we're (very slightly) faster! There are a lot of engineering tricks to make this possible, but the intuition is that you can start responding with faster models once you measure the other characteristics (quality, cost) carefully enough that they won't degrade

Rohan Paul's profile picture
Rohan Paul2 months ago

I so much like the JIT framing here. Would like Statistical non-inferiority instead of token-identity anyday.

Martian's profile picture
Martian2 months ago

Yeah, it's interesting to think about what it means for two AI models to "be the same". That's actually why @thesean_labs is named after

Alex Prompter's profile picture
Alex Prompter2 months ago

Cutting the bill in half without downgrading the model is a hell of a pitch.

Khushi's profile picture
Khushi2 months ago

this could be a big win 🫡

Dhanush N's profile picture
Dhanush N2 months ago

when people hear "cheaper endpoint" they picture a router grabbing whatever's cheapest right now. that's not what this is. ship hunts across models, tools, harnesses, cascades, ensembles and interventions to find the lowest-cost way to run your task while still matching your reference model's grade. the result: the same quality you get today, for a flat 50% less. point your calls at the ship-like version of what you already run and either pocket half the bill or plow it back into more quality on the same budget.

Requesty's profile picture
Requesty2 months ago

Very cool! well done team! Any possibilities to chose the models that would be used under the hood for compliance reasons?

Martian's profile picture
Martian2 months ago

For large-volume users, we do support this! Reach out to the Thesean team: [email protected]

Nimrod Kor's profile picture
Nimrod Kor2 months ago

This is 🔥

Przemek Chojecki | PC's profile picture
Przemek Chojecki | PC2 months ago

awesome! can't wait to test it properly!

fabrico's profile picture
fabrico2 months ago

DeepSWE bench please

Martian's profile picture
Martian2 months ago

Working on it!

WebDeveloperMentor's profile picture
WebDeveloperMentor2 months ago

The interesting claim here is not “cheaper AI.” It is a fixed 50% lower price while targeting the capabilities and behavior of the model your application already depends on. That is a much harder problem than routing every request to a cheaper model.

Shak's profile picture
Shak2 months ago

this is ridiculously cool and means the price of intelligence on tap just went down

Shivay Lamba's profile picture
Shivay Lamba2 months ago

this is amazing!

Nandkishor's profile picture
Nandkishor2 months ago

Here's what gets overlooked inference isn't free, and that cost quietly drains every AI product you build. Slash it by half and the math changes fast. Features you parked come back to life, a free tier finally pencils out, and you can open the doors to way more users. Running your product for a flat 50% less while keeping the exact same model quality matters more than people assume. You just point to the ship-like version of the model you're already running, and that's it.

Tech with Mak's profile picture
Tech with Mak2 months ago

Efficiency work usually happens before your model ever sees a real problem. This spends the optimization budget after, once it knows what you're solving. That's a genuinely different lever. Great work!

Yanhua's profile picture
Yanhua2 months ago

I'm very curious about how they managed to reduce the cost of Opus and GPT 5.6 Sol by 50% while surpassing all frontier models in intelligence per dollar?

Jason Zhu's profile picture
Jason Zhu2 months ago

it's awesome work

Thibos_Nkutha's profile picture
Thibos_Nkutha2 months ago

::.

Aaliya's profile picture
Aaliya2 months ago

Better performance at a lower cost is always welcome.

Alex Nguyen's profile picture
Alex Nguyen2 months ago

Congrats 👏, 50% cheaper is huge.

Alvaro Cintas's profile picture
Alvaro Cintas2 months ago

This is so good news

Josh's profile picture
Josh2 months ago

Now we test...

Martian's profile picture
Martian2 months ago

Let us know what you think! Very receptive to feedback rn

Related Videos