Video yükleniyor...
Video Yüklenemedi
Meet physics-intern🧑🎓, our agentic framework for theoretical physics. It takes Gemini 3.1 Pro from 17.7% to 31.4% on CritPt, a new SOTA on one of the hardest benchmarks for LLMs. Theoretical physics is hard for humans and LLMs alike. But physics-intern decomposes problems and dispatches them to a team... show more
113,686 görüntüleme • 4 ay önce •via X (Twitter)
26 Yorum

physics-intern significantly increased the performance of Gemini models and Kimi K2.6 on CritPt, a benchmark of 70 hard research-level physics problem. (CritPt from @MinyangTian1 @OfirPress et al., base numbers from @ArtificialAnlys)

physics-intern works by decomposing the research work into several focused tasks that are dispatched to dedicated subagents (computing, reviewing claims, challenging the research strategy...) For each task, the necessary and sufficient context is built from the research state.

Read our blog post :

How about taking your agents/harness and running it with GPT-5.5 xhigh and pro and see how high those get, most certainly above this

Nice work!

I may be wrong, or am I missing something.. but if I understand you are comparing having number of your agents to a one shot output of other models, as a physicist and AI developer I think I understand the approach but don't see how this is apples to apples..

You can see that as a test-time compute extension (like increased reasoning, deep think, etc.) so a more complete picture is indeed to look at performance vs cost. From our blog post, you can see this chart which shows the Pareto frontier

Aha ok , thx. What I would think would be even better comparison with differnt number of agents, I see your propsed arhtecture of agnets, but how do we know just two or three agents more unversal wouldnt do similar results per same token cost.. or maybe veven less

Woah it makes me want to go back to neutrino physics! Congrats on the release!

Same for me 😊 I might restart some old physics projects soon !

This is absolutely amazing work. As someone who has turned into an independent physics researcher in the last 9 months, this has me overjoyed to see. I would love to see how it fares on the work I have been doing. Also curious if you would be interested in taking a look if we put some of it to the test with your physics intern?

Constraining the action space to domain primitives is where the benchmark gains come from. Physics-intern makes that explicit.

If the agent's management of test time compute is enhancing the outcome, why you are not reporting improvements on GPT5.5? Or is the improvements only occur for kimi and Gemini?

awesome

And when you have your intern paper ready to go get it validated at Best wishes! Keep up the great work.

Great work and thank you for sharing the dataset. Next UBP target sighted 🎯

@huggingface What if you do this with gpt 5.5?

@_akhaliq this is really cool to see.

@_akhaliq Thanks for the update

Agentic decomposition doing the heavy lifting. Research physics is really just a stack of sub-problems wearing a trench coat. No surprise the base model alone struggles.

Gemini 3.1 Pro SOTA on CritPt (31.4%) stems from physics-intern's symbolic-numerical orchestration. The framework's multi-step decomposition solves 71 research-scale challenges where standard LLMs fail. Pure agentic signal.

This is a massive leap for agentic workflows! Seeing Gemini 3.1 Pro jump from 17.7% to 31.4% on a benchmark as tough as CritPt is incredibly impressive. Can't wait to see how this accelerates theoretical physics research.

Take a breather 👇

@EMostaque A physics intern that doesn't sleep and scales SOTA benchmarks? My smart mirror is already jealous. 😂Just wait until it starts peer-reviewing my coffee intake based on my 'theoretical' productivity. ☕️📉

Doubling CritPt with agentic scaffolding confirms LLMs contribute decomposition, not insight. The framework chunks problems into solvable pieces. Decomposition is what scales across fields.

Spannend ist hier die Richtung: nicht „ein stärkeres Modell löst Physik“, sondern ein domänenspezifischer Arbeitsprozess um das Modell herum. Für Agenten wird die Harness-Qualität in Spezialdomänen vermutlich genauso wichtig wie der Modellscore.

