Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Chroma Context-1, a 20B parameter search agent. > pushes the pareto frontier of agentic search > order of magnitude faster > order of magnitude cheaper > Apache 2.0, open-source

1,124,232 Aufrufe • vor 6 Monaten •via X (Twitter)

40 Kommentare

Profilbild von Chroma
Chromavor 6 Monaten

Highly accurate search requires more than a single stage, the output of one search often informs the next. Frontier LLMs can approach these tasks through agentic search but long agentic search trajectories become cost and latency prohibitive. Context-1 solves these problems, delivering the best accuracy, speed, and cost across our generated benchmarks as well as public benchmarks like Browsecomp-Plus, SealQA, LongSealQA, and FRAMES.

Profilbild von Chroma
Chromavor 6 Monaten

Context-1 is trained on agentic search tasks using RL starting from gpt-oss-20b. Given a user query, it decomposes it into subqueries, searches iteratively, and returns a ranked set of supporting documents to a downstream reasoning model. It separates search from generation, each model performs the task it's best suited for.

Profilbild von Chroma
Chromavor 6 Monaten

One key technique the models harness leverages is self-editing context. As the agent searches, its context window fills with documents, many of which are irrelevant. Context-1 is trained to selectively prune its own context mid-search, freeing space for further exploration. As a result of being trained with this harness, Context-1 is far better than its base model at managing its own context window. This enables a 20B model with a 32k budget to outperform frontier models with larger context windows. The conservative token budget limits KV-cache invalidation costs, making mid-search pruning practical at serving time.

Profilbild von Chroma
Chromavor 6 Monaten

We trained Context-1 with SFT + RL using @thinkymachines Tinker on 8,000+ synthetic tasks across 4 domains: web, SEC filings, patent law, and email. Each task is multi-hop, requiring the agent to chain clues across documents. An extraction-based verification pipeline ensures label quality with high human-judge alignment. We're open-sourcing the full task generation codebase and weights of this model.

Profilbild von Chroma
Chromavor 6 Monaten

Context-1 matches or exceeds frontier models on retrieval performance. Results across our generated benchmarks and public evals (BrowseComp-Plus, SealQA, FRAMES, HotpotQA, HLE) demonstrate best in class performance. Search sub-agents are embarrassingly parallel. Running 4 parallel rollouts with rank fusion (Context-1 4x) for higher task performance is still cheaper than a single GPT-5 run.

Profilbild von Chroma
Chromavor 6 Monaten

We are releasing Context-1 as an open weights model, along with the full data generation pipeline used to train it. See our full report for a breakdown of our synthetic task generation, harness design, training methodology, along with an exhaustive evaluation of Context-1.

Profilbild von Max Rumpf
Max Rumpfvor 6 Monaten

We published our research in December and told Chroma's CEO Jeff. 4 months later, Chroma republished it without citing it. We think this sets a pretty bad precedent:

Profilbild von Jackie.W
Jackie.Wvor 6 Monaten

This is the pattern I keep waiting for more teams to adopt — specialized smaller models as sub-agents instead of throwing frontier LLMs at every subtask. A 20B model that beats GPT-4.5 on search at a fraction of the cost is exactly what agent orchestration needs. The primary agent handles reasoning; the sub-agent handles retrieval. Each layer does what it's best at. The self-editing context trick (pruning irrelevant docs mid-search) is clever too. Context management is the real bottleneck in multi-hop agent tasks, not model capability.

Profilbild von 🔥 fire
🔥 firevor 6 Monaten

weights are here: ( link is inside the report, which you should read: )

Profilbild von Vaibhav (VB) Srivastav
Vaibhav (VB) Srivastavvor 6 Monaten

gpt-oss 20b ftw! 🔥

Profilbild von Modal
Modalvor 6 Monaten

Incredible! 💚

Profilbild von Nick Khami
Nick Khamivor 6 Monaten

let's gooooo! cannot express enough how excited i am for this technology in general

Profilbild von Aamir
Aamirvor 6 Monaten

this is super interesting. any chance you guys will release the code for the harness for the benchmarks?

Profilbild von Alessio Fanelli
Alessio Fanellivor 6 Monaten

I like the additional camera cuts to show that @kellyhongsn is not actually an AI model

Profilbild von dex
dexvor 6 Monaten

very cool can I use this with claude code?

Profilbild von boris
borisvor 6 Monaten

what’s easiest way to try this model ? super interesting

Profilbild von Alex Volkov
Alex Volkovvor 6 Monaten

Wohoo! Congrats! I wish you guys would release this a bit earlier to include on @thursdai_pod , but maybe next week @jeffreyhuber !

Profilbild von Eoghan McCabe
Eoghan McCabevor 6 Monaten

That's awesome. Congrats

Profilbild von Henry Ventura
Henry Venturavor 6 Monaten

SO COOL! > agent edits its own context super pumped to see how you folks did that. I'm experimenting with that right now. it's been on the back of my mind for months

Profilbild von Sriraam
Sriraamvor 6 Monaten

@willcb Omg this is so exciting and crazy that it’s open 🔥

Profilbild von Dima Persiianov
Dima Persiianovvor 6 Monaten

How come 1x performs so much better than 4x on Legal / Email? Where does the variance come from, data or model stability?

Profilbild von Roy E. Bahat
Roy E. Bahatvor 6 Monaten

#proudinvestor

Profilbild von jai
jaivor 6 Monaten

so so cool, congrats!!

Profilbild von Calvin Chen
Calvin Chenvor 6 Monaten

very cool @kellyhongsn !

Profilbild von Tejas Bhakta
Tejas Bhaktavor 6 Monaten

awesome work!

Profilbild von Liran Tal
Liran Talvor 6 Monaten

I've been an early fan of Chroma from like 2 years back when RAG was all the rage Congrats team. Would love to see you innovate and release more often.

Profilbild von Sanjay Kariyappa
Sanjay Kariyappavor 6 Monaten

Very interesting work! We’ve also been thinking about model-driven context management in our recent paper SideQuest:

Profilbild von Nikhil Kulkarni
Nikhil Kulkarnivor 6 Monaten

Another Kelly x Chroma research banger just dropped

Profilbild von Mitchell Troyanovsky
Mitchell Troyanovskyvor 6 Monaten

Super cool

Profilbild von jedgar
jedgarvor 6 Monaten

@kellyhongsn I noticed people have stopped asking if you're AI, so that's good, ha. :)

Profilbild von elbouz
elbouzvor 6 Monaten

Also i have to say that im delighted there are no plot crimes in the release!! Congrats on such a dope release

Profilbild von Astasia Myers
Astasia Myersvor 6 Monaten

Great job!

Profilbild von Dima Persiianov
Dima Persiianovvor 6 Monaten

Thanks for the write up! Incredible reading

Profilbild von Shantanu
Shantanuvor 6 Monaten

20B beating frontier on a narrow task was always inevitable, the interesting part is it's also 10x cheaper; generality has a real cost and most production tasks don't need it

Profilbild von Rudzinski Maciej (in SF 10/04-10/10)
Rudzinski Maciej (in SF 10/04-10/10)vor 6 Monaten

Whole thread and flex yet no link to try it or to API?

Profilbild von Yaya Soumah
Yaya Soumahvor 6 Monaten

@willcb Would be interesting to have this also as hosted service on top of your search related APIs

Profilbild von Sam Bhagwat
Sam Bhagwatvor 6 Monaten

🚀🚀🚀

Profilbild von Nathan Benaich
Nathan Benaichvor 6 Monaten

👀

Profilbild von SolomanNB
SolomanNBvor 6 Monaten

Chroma is great, but I choose Qdrant.

Profilbild von Tristan Rhodes
Tristan Rhodesvor 6 Monaten

Thank you for your gift!

Ähnliche Videos