Together AI's banner
Together AI's profile picture

Together AI

@togethercompute61,344 subscribers

Accelerate inference, model shaping, and pre-training on a research-optimized platform.

Shorts

Access Mixtral with the fastest inference performance anywhere! Up to 100 token/s for $0.0006/1K tokens — to our knowledge the fastest performance at the lowest price! Mixtral-8x7b-32kseqlen Mistral AI & DiscoLM-mixtral-8x7b-v2 are live on Together API!

Access Mixtral with the fastest inference performance anywhere! Up to 100 token/s for $0.0006/1K tokens — to our knowledge the fastest performance at the lowest price! Mixtral-8x7b-32kseqlen Mistral AI & DiscoLM-mixtral-8x7b-v2 are live on Together API!

924,940 просмотров

Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. A/B testing belongs at the endpoint, not in your app code. Same endpoint name, API, and keys for your clients. No feature flags, no hash-mod-100 in client code, no spreadsheet explaining what group A vs B means. Split a live endpoint's traffic into one control and up to 20 variants, each with a fixed percentage. Ramp with a single call. Delete the experiment and 100% of traffic returns to the control, with nothing left to unwind. Read the full walkthrough:

Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. A/B testing belongs at the endpoint, not in your app code. Same endpoint name, API, and keys for your clients. No feature flags, no hash-mod-100 in client code, no spreadsheet explaining what group A vs B means. Split a live endpoint's traffic into one control and up to 20 variants, each with a fixed percentage. Ramp with a single call. Delete the experiment and 100% of traffic returns to the control, with nothing left to unwind. Read the full walkthrough:

15,317 просмотров

We are thrilled to be a launch partner for Meta Llama 3. Experience Llama 3 now at up to 350 tokens per second for Llama 3 8B and up to 150 tokens per second for Llama 3 70B, running in full FP16 precision on the Together API! 🤯

We are thrilled to be a launch partner for Meta Llama 3. Experience Llama 3 now at up to 350 tokens per second for Llama 3 8B and up to 150 tokens per second for Llama 3 70B, running in full FP16 precision on the Together API! 🤯

88,229 просмотров

Videos

Больше нет контента для загрузки