Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Harness-1, a 20B search agent trained with a state-externalizing harness. > frontier-level long-horizon search, rivaling Opus-4.6 and outperforming GPT-5.4 > Context-1-level cost and latency > externalizes candidates, evidence, verification, and search history > open-source

288,706 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 47

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

[1/N] I’ve been wondering: maybe search agents are bad at search partly because we make them do all the paperwork in their head. So I tried a simple idea: externalize the search state, then train the model to use that harness. The result is Harness-1: a 20B search agent that can match or even beat much larger frontier AI on hard long-horizon search tasks.

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

[2/N] The usual search-agent setup is basically: search → read → search → read → keep appending everything to the transcript. At some point the model is not just “searching” anymore. It is also being asked to be a memory system, a note taker, a verifier, and a librarian.

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

[3/N] This gets especially weird for RL. The final reward can tell you whether the episode worked, but it often does not tell you why it failed. Was it a bad search? Forgotten evidence? Missing verification? Poor curation? Or the agent just losing track of what it had already seen?

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

[4/N] Harness-1 tries to separate these two jobs. The model still makes the semantic decisions: what to search, what to read, what to keep, what to verify, when to stop. But the harness maintains the recoverable state around those decisions.

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

[5/N] Concretely, the harness keeps a working memory with: candidate docs, curated evidence, importance tags, search history, evidence links, verification records, dedup/compression, and context-budget markers. So the agent is not just talking to a search box. It is operating over a workspace.

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

[6/N] I think this changes what RL is actually learning. Instead of training the model to survive a giant append-only transcript, we train it to use a structured search interface: search, curate, revisit, verify, and submit. Much closer to how I’d want a search agent to work.

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

[7/N] A fun part: this was not trained with a huge amount of task data. Harness-1 uses 899 filtered SFT trajectories and RL on 3,453 queries. The point is not “less data is always enough.” The point is that a lot of the behavioral prior can live in the harness.

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

[8/N] The result that made me most excited is transfer. Harness-1 improves over Context-1 by +7.9 recall points on source-family benchmarks. But on held-out transfer benchmarks, the gain is +17.0 points. That’s the part that made the idea feel real to me.

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

[9/N] The ablations were also pretty revealing. When we disable the harness mechanisms, the model does not just lose some information. It changes behavior: more shallow searching, less reading / verification, worse final curation. So the harness is not just engineering glue.

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

[10/N] My takeaway: for search agents, “the model” is not the whole learning system. The interface matters. The memory layout matters. The action space matters. The harness matters. If we want RL to teach better search behavior, we should probably stop making the model do all the paperwork in its head.

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

Paper 📄: Code 💻: Model 🤗: HF Paper:

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

Huge thanks to @trychroma for fully supporting this work, and to @tinkerapi for the training infra!

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

Huge shoutout to my awesome collaborators @zhiyiscs @HammadTime @kellyhongsn @PatrickXu565299 @SunJiashuo36 !!

Фото профиля RG
RG3 месяцев назад

Yo this is insanely cool!

Фото профиля Pranav
Pranav3 месяцев назад

everyone will fixate on the 20B. the externalizing harness is the more interesting bet. the ablations show it: turn it off and the model searches shallower, verifies less. the scaffold carries the search behavior, not the parameter count. how far does it generalize past search?

Фото профиля Samarth Aggarwal
Samarth Aggarwal3 месяцев назад

Your product video looks very polished, kudos! Curious what tool you used to create it?

Фото профиля Jonathan Chang
Jonathan Chang3 месяцев назад

@kimbochen cool work . Reminds me of the first version of OpenAI deep research where people find out the it can run code during the research

Фото профиля Patrick Donohoe
Patrick Donohoe3 месяцев назад

Super cool project. I recently was speaking about this at a conference about how smaller models with higher parameter density+ reasoning ability paired with external knowledge stores are the future. Could be interesting to pair this with a web search api!

Фото профиля Vikas Tiwari
Vikas Tiwari3 месяцев назад

Will it eat up @ExaAILabs ?

Фото профиля Anthony 😎🛹
Anthony 😎🛹3 месяцев назад

Interesting direction. It seems we will continue to find new ways to optimize. What i am curious about is where we land. At some point we get a Linux OS and everyone is happy. I assume we get there with this transformer tech coupled with a harness of sorts.

Фото профиля minamium 🛡
minamium 🛡3 месяцев назад

the thing that gets me is the 17 point transfer gain. thats the signal that the harness is doing something fundamental, not just engineering tricks

Фото профиля ⓙⓘⓑⓞⓢⓢ
ⓙⓘⓑⓞⓢⓢ3 месяцев назад

that’s what ant and oai do under the hood, and why the moat with cc or codex data is so useful, i wonder how they construct their eval env for rl with all the private code data under the hood they post train their models on this harness now, why open sourced never thought of it

Фото профиля Gerard Sans | Axiom 🇬🇧
Gerard Sans | Axiom 🇬🇧3 месяцев назад

@chrissm79

Фото профиля Vadim
Vadim3 месяцев назад

Completely off topic…but is it just me or most comments on this post are AI comments? Pretty weird.

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

totally have no idea what you’re talking about

Фото профиля Bobby The Man
Bobby The Man3 месяцев назад

this is super cool

Фото профиля Levi
Levi3 месяцев назад

that harness idea kinda wild

Фото профиля Elias Lumer
Elias Lumer3 месяцев назад

I wonder if there’s a more generalizable version of this. Great work, will check out the paper/githuv

Фото профиля Ash
Ash3 месяцев назад

lol i was just thinking could it be better than chroma ones

Фото профиля The Bjorn Identity
The Bjorn Identity3 месяцев назад

Looks like Claude

Фото профиля Mohammed Hossam
Mohammed Hossam3 месяцев назад

On open router or not yet?

Фото профиля BlockedPath
BlockedPath3 месяцев назад

every new model announcement is beats frontier on benchmark and then you actually use it and it tells you to delete system32 to free up memory. show me the failure cases you coward

Фото профиля Samuel Ekpe
Samuel Ekpe3 месяцев назад

Nice

Фото профиля Mr Trava
Mr Trava3 месяцев назад

@huggingface Impressive to see search agents finally escaping the black box. State externalization + open source at 20B is the kind of transparency we need. How’s the tooling integration for custom evidence sources? Chroma-backed, but can we plug in other vector stores?

Фото профиля Manav Gupta
Manav Gupta3 месяцев назад

the model was never bad at search. it was bad at being a search engine, librarian, verifier, and memory system all at once. separating those jobs is the whole unlock. great work.

Фото профиля Vishvanand
Vishvanand3 месяцев назад

i was avoiding RL like the plague for agentic search until i read this

Фото профиля All Over Tools
All Over Tools3 месяцев назад

outperforms GPT-5.4? call me when GPT-5 ships. anyway, curious about the externalized evidence -last time i tried that with a 20B, the candidate log added 4k tokens per turn, tanking cost after 10 steps. have you benchmarked on anything beyond single-shot search?

Фото профиля Patrick Jiang
Patrick Jiang3 месяцев назад

yep, any frontier models here are equipped with context-1's harness - the self-context-editing one - the best one we know so far for agentic search

Фото профиля Daniel Fein
Daniel Fein3 месяцев назад

Very cool work

Фото профиля Remain Urus
Remain Urus3 месяцев назад

Holy clickbait

Фото профиля Max Andrews
Max Andrews3 месяцев назад

Love this! Looks like frontier models with this harness were not yet tested? i.e. if i wanted to run this harness with a serverless model like kimi or haiku

Фото профиля Tim White
Tim White3 месяцев назад

@grok summarize this thread

Фото профиля Banned
Banned3 месяцев назад

Awesome work, tragic naming

Фото профиля Alpha Batcher
Alpha Batcher3 месяцев назад

it's happened we can test Harness-1, at last !

Фото профиля Sail Ai
Sail Ai3 месяцев назад

@ClementDelangue

Фото профиля Potato Terminator
Potato Terminator3 месяцев назад

Open-source plus auditable search history is what matters to me here. Benchmarks are nice, but if I can inspect the evidence path myself, that's the real win.

Фото профиля Mika 🖤
Mika 🖤3 месяцев назад

nasa fake, i'm the real space queen 😉

Похожие видео