Video wird geladen...
Video konnte nicht geladen werden
Introducing Chroma Context-1, a 20B parameter search agent. > pushes the pareto frontier of agentic search > order of magnitude faster > order of magnitude cheaper > Apache 2.0, open-source
1,124,232 Aufrufe • vor 6 Monaten •via X (Twitter)
40 Kommentare

Highly accurate search requires more than a single stage, the output of one search often informs the next. Frontier LLMs can approach these tasks through agentic search but long agentic search trajectories become cost and latency prohibitive. Context-1 solves these problems, delivering the best accuracy, speed, and cost across our generated benchmarks as well as public benchmarks like Browsecomp-Plus, SealQA, LongSealQA, and FRAMES.

Context-1 is trained on agentic search tasks using RL starting from gpt-oss-20b. Given a user query, it decomposes it into subqueries, searches iteratively, and returns a ranked set of supporting documents to a downstream reasoning model. It separates search from generation, each model performs the task it's best suited for.

One key technique the models harness leverages is self-editing context. As the agent searches, its context window fills with documents, many of which are irrelevant. Context-1 is trained to selectively prune its own context mid-search, freeing space for further exploration. As a result of being trained with this harness, Context-1 is far better than its base model at managing its own context window. This enables a 20B model with a 32k budget to outperform frontier models with larger context windows. The conservative token budget limits KV-cache invalidation costs, making mid-search pruning practical at serving time.

We trained Context-1 with SFT + RL using @thinkymachines Tinker on 8,000+ synthetic tasks across 4 domains: web, SEC filings, patent law, and email. Each task is multi-hop, requiring the agent to chain clues across documents. An extraction-based verification pipeline ensures label quality with high human-judge alignment. We're open-sourcing the full task generation codebase and weights of this model.

Context-1 matches or exceeds frontier models on retrieval performance. Results across our generated benchmarks and public evals (BrowseComp-Plus, SealQA, FRAMES, HotpotQA, HLE) demonstrate best in class performance. Search sub-agents are embarrassingly parallel. Running 4 parallel rollouts with rank fusion (Context-1 4x) for higher task performance is still cheaper than a single GPT-5 run.

We are releasing Context-1 as an open weights model, along with the full data generation pipeline used to train it. See our full report for a breakdown of our synthetic task generation, harness design, training methodology, along with an exhaustive evaluation of Context-1.

We published our research in December and told Chroma's CEO Jeff. 4 months later, Chroma republished it without citing it. We think this sets a pretty bad precedent:

This is the pattern I keep waiting for more teams to adopt — specialized smaller models as sub-agents instead of throwing frontier LLMs at every subtask. A 20B model that beats GPT-4.5 on search at a fraction of the cost is exactly what agent orchestration needs. The primary agent handles reasoning; the sub-agent handles retrieval. Each layer does what it's best at. The self-editing context trick (pruning irrelevant docs mid-search) is clever too. Context management is the real bottleneck in multi-hop agent tasks, not model capability.

weights are here: ( link is inside the report, which you should read: )

gpt-oss 20b ftw! 🔥

Incredible! 💚

let's gooooo! cannot express enough how excited i am for this technology in general

this is super interesting. any chance you guys will release the code for the harness for the benchmarks?

I like the additional camera cuts to show that @kellyhongsn is not actually an AI model

very cool can I use this with claude code?

what’s easiest way to try this model ? super interesting

Wohoo! Congrats! I wish you guys would release this a bit earlier to include on @thursdai_pod , but maybe next week @jeffreyhuber !

That's awesome. Congrats

SO COOL! > agent edits its own context super pumped to see how you folks did that. I'm experimenting with that right now. it's been on the back of my mind for months

@willcb Omg this is so exciting and crazy that it’s open 🔥

How come 1x performs so much better than 4x on Legal / Email? Where does the variance come from, data or model stability?

#proudinvestor

so so cool, congrats!!

very cool @kellyhongsn !

awesome work!

I've been an early fan of Chroma from like 2 years back when RAG was all the rage Congrats team. Would love to see you innovate and release more often.

Very interesting work! We’ve also been thinking about model-driven context management in our recent paper SideQuest:

Another Kelly x Chroma research banger just dropped

Super cool

@kellyhongsn I noticed people have stopped asking if you're AI, so that's good, ha. :)

Also i have to say that im delighted there are no plot crimes in the release!! Congrats on such a dope release

Great job!

Thanks for the write up! Incredible reading

20B beating frontier on a narrow task was always inevitable, the interesting part is it's also 10x cheaper; generality has a real cost and most production tasks don't need it

Whole thread and flex yet no link to try it or to API?

@willcb Would be interesting to have this also as hosted service on top of your search related APIs

🚀🚀🚀

👀

Chroma is great, but I choose Qdrant.

Thank you for your gift!
