Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Chroma Context-1, a 20B parameter search agent. > pushes the pareto frontier of agentic search > order of magnitude faster > order of magnitude cheaper > Apache 2.0, open-source

1,124,232 просмотров • 6 месяцев назад •via X (Twitter)

Комментарии: 40

Фото профиля Chroma
Chroma6 месяцев назад

Highly accurate search requires more than a single stage, the output of one search often informs the next. Frontier LLMs can approach these tasks through agentic search but long agentic search trajectories become cost and latency prohibitive. Context-1 solves these problems, delivering the best accuracy, speed, and cost across our generated benchmarks as well as public benchmarks like Browsecomp-Plus, SealQA, LongSealQA, and FRAMES.

Фото профиля Chroma
Chroma6 месяцев назад

Context-1 is trained on agentic search tasks using RL starting from gpt-oss-20b. Given a user query, it decomposes it into subqueries, searches iteratively, and returns a ranked set of supporting documents to a downstream reasoning model. It separates search from generation, each model performs the task it's best suited for.

Фото профиля Chroma
Chroma6 месяцев назад

One key technique the models harness leverages is self-editing context. As the agent searches, its context window fills with documents, many of which are irrelevant. Context-1 is trained to selectively prune its own context mid-search, freeing space for further exploration. As a result of being trained with this harness, Context-1 is far better than its base model at managing its own context window. This enables a 20B model with a 32k budget to outperform frontier models with larger context windows. The conservative token budget limits KV-cache invalidation costs, making mid-search pruning practical at serving time.

Фото профиля Chroma
Chroma6 месяцев назад

We trained Context-1 with SFT + RL using @thinkymachines Tinker on 8,000+ synthetic tasks across 4 domains: web, SEC filings, patent law, and email. Each task is multi-hop, requiring the agent to chain clues across documents. An extraction-based verification pipeline ensures label quality with high human-judge alignment. We're open-sourcing the full task generation codebase and weights of this model.

Фото профиля Chroma
Chroma6 месяцев назад

Context-1 matches or exceeds frontier models on retrieval performance. Results across our generated benchmarks and public evals (BrowseComp-Plus, SealQA, FRAMES, HotpotQA, HLE) demonstrate best in class performance. Search sub-agents are embarrassingly parallel. Running 4 parallel rollouts with rank fusion (Context-1 4x) for higher task performance is still cheaper than a single GPT-5 run.

Фото профиля Chroma
Chroma6 месяцев назад

We are releasing Context-1 as an open weights model, along with the full data generation pipeline used to train it. See our full report for a breakdown of our synthetic task generation, harness design, training methodology, along with an exhaustive evaluation of Context-1.

Фото профиля Max Rumpf
Max Rumpf6 месяцев назад

We published our research in December and told Chroma's CEO Jeff. 4 months later, Chroma republished it without citing it. We think this sets a pretty bad precedent:

Фото профиля Jackie.W
Jackie.W6 месяцев назад

This is the pattern I keep waiting for more teams to adopt — specialized smaller models as sub-agents instead of throwing frontier LLMs at every subtask. A 20B model that beats GPT-4.5 on search at a fraction of the cost is exactly what agent orchestration needs. The primary agent handles reasoning; the sub-agent handles retrieval. Each layer does what it's best at. The self-editing context trick (pruning irrelevant docs mid-search) is clever too. Context management is the real bottleneck in multi-hop agent tasks, not model capability.

Фото профиля 🔥 fire
🔥 fire6 месяцев назад

weights are here: ( link is inside the report, which you should read: )

Фото профиля Vaibhav (VB) Srivastav
Vaibhav (VB) Srivastav6 месяцев назад

gpt-oss 20b ftw! 🔥

Фото профиля Modal
Modal6 месяцев назад

Incredible! 💚

Фото профиля Nick Khami
Nick Khami6 месяцев назад

let's gooooo! cannot express enough how excited i am for this technology in general

Фото профиля Aamir
Aamir6 месяцев назад

this is super interesting. any chance you guys will release the code for the harness for the benchmarks?

Фото профиля Alessio Fanelli
Alessio Fanelli6 месяцев назад

I like the additional camera cuts to show that @kellyhongsn is not actually an AI model

Фото профиля dex
dex6 месяцев назад

very cool can I use this with claude code?

Фото профиля boris
boris6 месяцев назад

what’s easiest way to try this model ? super interesting

Фото профиля Alex Volkov
Alex Volkov6 месяцев назад

Wohoo! Congrats! I wish you guys would release this a bit earlier to include on @thursdai_pod , but maybe next week @jeffreyhuber !

Фото профиля Eoghan McCabe
Eoghan McCabe6 месяцев назад

That's awesome. Congrats

Фото профиля Henry Ventura
Henry Ventura6 месяцев назад

SO COOL! > agent edits its own context super pumped to see how you folks did that. I'm experimenting with that right now. it's been on the back of my mind for months

Фото профиля Sriraam
Sriraam6 месяцев назад

@willcb Omg this is so exciting and crazy that it’s open 🔥

Фото профиля Dima Persiianov
Dima Persiianov6 месяцев назад

How come 1x performs so much better than 4x on Legal / Email? Where does the variance come from, data or model stability?

Фото профиля Roy E. Bahat
Roy E. Bahat6 месяцев назад

#proudinvestor

Фото профиля jai
jai6 месяцев назад

so so cool, congrats!!

Фото профиля Calvin Chen
Calvin Chen6 месяцев назад

very cool @kellyhongsn !

Фото профиля Tejas Bhakta
Tejas Bhakta6 месяцев назад

awesome work!

Фото профиля Liran Tal
Liran Tal6 месяцев назад

I've been an early fan of Chroma from like 2 years back when RAG was all the rage Congrats team. Would love to see you innovate and release more often.

Фото профиля Sanjay Kariyappa
Sanjay Kariyappa6 месяцев назад

Very interesting work! We’ve also been thinking about model-driven context management in our recent paper SideQuest:

Фото профиля Nikhil Kulkarni
Nikhil Kulkarni6 месяцев назад

Another Kelly x Chroma research banger just dropped

Фото профиля Mitchell Troyanovsky
Mitchell Troyanovsky6 месяцев назад

Super cool

Фото профиля jedgar
jedgar6 месяцев назад

@kellyhongsn I noticed people have stopped asking if you're AI, so that's good, ha. :)

Фото профиля elbouz
elbouz6 месяцев назад

Also i have to say that im delighted there are no plot crimes in the release!! Congrats on such a dope release

Фото профиля Astasia Myers
Astasia Myers6 месяцев назад

Great job!

Фото профиля Dima Persiianov
Dima Persiianov6 месяцев назад

Thanks for the write up! Incredible reading

Фото профиля Shantanu
Shantanu6 месяцев назад

20B beating frontier on a narrow task was always inevitable, the interesting part is it's also 10x cheaper; generality has a real cost and most production tasks don't need it

Фото профиля Rudzinski Maciej (in SF 10/04-10/10)
Rudzinski Maciej (in SF 10/04-10/10)6 месяцев назад

Whole thread and flex yet no link to try it or to API?

Фото профиля Yaya Soumah
Yaya Soumah6 месяцев назад

@willcb Would be interesting to have this also as hosted service on top of your search related APIs

Фото профиля Sam Bhagwat
Sam Bhagwat6 месяцев назад

🚀🚀🚀

Фото профиля Nathan Benaich
Nathan Benaich6 месяцев назад

👀

Фото профиля SolomanNB
SolomanNB6 месяцев назад

Chroma is great, but I choose Qdrant.

Фото профиля Tristan Rhodes
Tristan Rhodes6 месяцев назад

Thank you for your gift!

Похожие видео