正在加载视频...

视频加载失败

Introducing Chroma Context-1, a 20B parameter search agent. > pushes the pareto frontier of agentic search > order of magnitude faster > order of magnitude cheaper > Apache 2.0, open-source

1,124,232 次观看 • 6 个月前 •via X (Twitter)

40 条评论

Chroma 的头像
Chroma6 个月前

Highly accurate search requires more than a single stage, the output of one search often informs the next. Frontier LLMs can approach these tasks through agentic search but long agentic search trajectories become cost and latency prohibitive. Context-1 solves these problems, delivering the best accuracy, speed, and cost across our generated benchmarks as well as public benchmarks like Browsecomp-Plus, SealQA, LongSealQA, and FRAMES.

Chroma 的头像
Chroma6 个月前

Context-1 is trained on agentic search tasks using RL starting from gpt-oss-20b. Given a user query, it decomposes it into subqueries, searches iteratively, and returns a ranked set of supporting documents to a downstream reasoning model. It separates search from generation, each model performs the task it's best suited for.

Chroma 的头像
Chroma6 个月前

One key technique the models harness leverages is self-editing context. As the agent searches, its context window fills with documents, many of which are irrelevant. Context-1 is trained to selectively prune its own context mid-search, freeing space for further exploration. As a result of being trained with this harness, Context-1 is far better than its base model at managing its own context window. This enables a 20B model with a 32k budget to outperform frontier models with larger context windows. The conservative token budget limits KV-cache invalidation costs, making mid-search pruning practical at serving time.

Chroma 的头像
Chroma6 个月前

We trained Context-1 with SFT + RL using @thinkymachines Tinker on 8,000+ synthetic tasks across 4 domains: web, SEC filings, patent law, and email. Each task is multi-hop, requiring the agent to chain clues across documents. An extraction-based verification pipeline ensures label quality with high human-judge alignment. We're open-sourcing the full task generation codebase and weights of this model.

Chroma 的头像
Chroma6 个月前

Context-1 matches or exceeds frontier models on retrieval performance. Results across our generated benchmarks and public evals (BrowseComp-Plus, SealQA, FRAMES, HotpotQA, HLE) demonstrate best in class performance. Search sub-agents are embarrassingly parallel. Running 4 parallel rollouts with rank fusion (Context-1 4x) for higher task performance is still cheaper than a single GPT-5 run.

Chroma 的头像
Chroma6 个月前

We are releasing Context-1 as an open weights model, along with the full data generation pipeline used to train it. See our full report for a breakdown of our synthetic task generation, harness design, training methodology, along with an exhaustive evaluation of Context-1.

Max Rumpf 的头像
Max Rumpf6 个月前

We published our research in December and told Chroma's CEO Jeff. 4 months later, Chroma republished it without citing it. We think this sets a pretty bad precedent:

Jackie.W 的头像
Jackie.W6 个月前

This is the pattern I keep waiting for more teams to adopt — specialized smaller models as sub-agents instead of throwing frontier LLMs at every subtask. A 20B model that beats GPT-4.5 on search at a fraction of the cost is exactly what agent orchestration needs. The primary agent handles reasoning; the sub-agent handles retrieval. Each layer does what it's best at. The self-editing context trick (pruning irrelevant docs mid-search) is clever too. Context management is the real bottleneck in multi-hop agent tasks, not model capability.

🔥 fire 的头像
🔥 fire6 个月前

weights are here: ( link is inside the report, which you should read: )

Vaibhav (VB) Srivastav 的头像
Vaibhav (VB) Srivastav6 个月前

gpt-oss 20b ftw! 🔥

Modal 的头像
Modal6 个月前

Incredible! 💚

Nick Khami 的头像
Nick Khami6 个月前

let's gooooo! cannot express enough how excited i am for this technology in general

Aamir 的头像
Aamir6 个月前

this is super interesting. any chance you guys will release the code for the harness for the benchmarks?

Alessio Fanelli 的头像
Alessio Fanelli6 个月前

I like the additional camera cuts to show that @kellyhongsn is not actually an AI model

dex 的头像
dex6 个月前

very cool can I use this with claude code?

boris 的头像
boris6 个月前

what’s easiest way to try this model ? super interesting

Alex Volkov 的头像
Alex Volkov6 个月前

Wohoo! Congrats! I wish you guys would release this a bit earlier to include on @thursdai_pod , but maybe next week @jeffreyhuber !

Eoghan McCabe 的头像
Eoghan McCabe6 个月前

That's awesome. Congrats

Henry Ventura 的头像
Henry Ventura6 个月前

SO COOL! > agent edits its own context super pumped to see how you folks did that. I'm experimenting with that right now. it's been on the back of my mind for months

Sriraam 的头像
Sriraam6 个月前

@willcb Omg this is so exciting and crazy that it’s open 🔥

Dima Persiianov 的头像
Dima Persiianov6 个月前

How come 1x performs so much better than 4x on Legal / Email? Where does the variance come from, data or model stability?

Roy E. Bahat 的头像
Roy E. Bahat6 个月前

#proudinvestor

jai 的头像
jai6 个月前

so so cool, congrats!!

Calvin Chen 的头像
Calvin Chen6 个月前

very cool @kellyhongsn !

Tejas Bhakta 的头像
Tejas Bhakta6 个月前

awesome work!

Liran Tal 的头像
Liran Tal6 个月前

I've been an early fan of Chroma from like 2 years back when RAG was all the rage Congrats team. Would love to see you innovate and release more often.

Sanjay Kariyappa 的头像
Sanjay Kariyappa6 个月前

Very interesting work! We’ve also been thinking about model-driven context management in our recent paper SideQuest:

Nikhil Kulkarni 的头像
Nikhil Kulkarni6 个月前

Another Kelly x Chroma research banger just dropped

Mitchell Troyanovsky 的头像
Mitchell Troyanovsky6 个月前

Super cool

jedgar 的头像
jedgar6 个月前

@kellyhongsn I noticed people have stopped asking if you're AI, so that's good, ha. :)

elbouz 的头像
elbouz6 个月前

Also i have to say that im delighted there are no plot crimes in the release!! Congrats on such a dope release

Astasia Myers 的头像
Astasia Myers6 个月前

Great job!

Dima Persiianov 的头像
Dima Persiianov6 个月前

Thanks for the write up! Incredible reading

Shantanu 的头像
Shantanu6 个月前

20B beating frontier on a narrow task was always inevitable, the interesting part is it's also 10x cheaper; generality has a real cost and most production tasks don't need it

Rudzinski Maciej (in SF 10/04-10/10) 的头像
Rudzinski Maciej (in SF 10/04-10/10)6 个月前

Whole thread and flex yet no link to try it or to API?

Yaya Soumah 的头像
Yaya Soumah6 个月前

@willcb Would be interesting to have this also as hosted service on top of your search related APIs

Sam Bhagwat 的头像
Sam Bhagwat6 个月前

🚀🚀🚀

Nathan Benaich 的头像
Nathan Benaich6 个月前

👀

SolomanNB 的头像
SolomanNB6 个月前

Chroma is great, but I choose Qdrant.

Tristan Rhodes 的头像
Tristan Rhodes6 个月前

Thank you for your gift!

相关视频