Video wird geladen...
Video konnte nicht geladen werden
Most AI benchmarks test retrieval — can a model find the known answer? However, the hardest problems in science require discovery, can a system earn an answer nobody has yet? Meet TRACES 🧭 — the world's first benchmark for measuring discoverative AI: AI that can work through evidence, test... show more
1,709,123 Aufrufe • vor 1 Monat •via X (Twitter)
36 Kommentare

So what does TRACES actually measure? Not just the answer, but the path to the answer. Six capabilities that spell its name: *T — Tools: selecting, calling and correctly interpreting external tools *R — Repair: locating and correcting its own errors once feedback arrives *A — Alternatives: laying out competing hypotheses and keeping or discarding them as evidence accumulates *C — Coherence: holding state, constraints and logic intact across a long chain of work *E — Evidence: grounding every conclusion in observation, data, experiment or citation *S — Scope: stating the conditions under which a conclusion holds, and where it does not apply TRACES evaluates today whether the discovery process is rigorous and evidence-grounded, even when the final outcome is not yet known 🔍 📄 Paper:

Today, in full: the technical report — the definitions, the rubric, the registry of problems — spanning biomedicine, clinical translation and frontier-model engineering. Led by Apodex’s Lead Scientist @profshengwang, who brings together the technical framework behind this effort. Behind them: 10 STEM PhDs, two months, 561 industries surveyed to find the 423 problems worth doing this for — real, high-stakes questions that take human experts years to work through. Bring a problem. Bring a solver. Both doors are open. *Submit a Problem: *Submit a Solver:

Haha Big boss making the switch into the AI industry to innovate again is a probability worth learning from. Having AI hypothesize possible futures and then go verify the results is genuinely a forward-thinking idea — looking forward to seeing it come to fruition.

I think this is a really interesting shift from testing what AI knows to testing what AI can actually discover

Apodex is pushing boundaries of what AI can achieve.... This is exactly what the field has been missing.

Most AI benchmarks are glorified recall tests. If TRACES actually measures whether a model can reason through messy evidence and land on a verifiable conclusion, that’s a much cleaner bar. The real question is brutal: can it survive contact with uncertainty, or does it fold the second the answer key disappears? @tianqiao_chen is pointing at the right problem here

That is exactly what production agents need.

The R in Repair is the part I see fail most often in production. An agent that starts with a wrong hypothesis will defend it for 40 steps instead of going back and reconsidering. Scoring it separately changes how you interpret a failed run.

用10个STEM博士花两个月筛出的423个真实科学难题来测模型,这评测维度比普通benchmark扎实多了。

连解题路径都被拆解成可评估的步骤,答案对错反而不是唯一标准了。

Okay THIS is interesting. 👀 Finding an answer is one thing. Figuring out what the answer even is, testing it, being wrong, and still getting to something verifiable? That’s a completely different level of intelligence. 🧭

wow discoverative intelligence seems amazing and TRACES framework is cool great job

Ahh! Nice. Glad the tide is shifting from retrieval to discovery benchmarking in AI That’s where the most upside is.. especially in scientific research. Excited :)

六个维度里最难得的是Scope,能自己说清结论在什么条件下成立,这是真严谨。

This is absolutely brilliant! The Repair step is where agents often fail. If an agent starts with the wrong idea, it may keep trying to fix it instead of stopping and thinking again.

终于跳出 “背答案打分”,开始丈量 AI 真正做研究的本事。之前很多 AI 科研 Demo 更像运气抽奖,这个基准终于盯着推理过程而不只看最终结果了。能不能做到用 TRACES 去评测普通开源大模型的科研探索能力,而不只是闭源前沿模型?

This is the part creators often overlook. You don't need to know an idea will work before testing it. Build a hypothesis, publish, watch the response, and iterate. Real data beats sitting around trying to predict everything.

The no-answer-key approach is what makes this stand out

真正的科学突破从不在标准答案里,而在于发现能力。TRACES首次让AI像科学家一样发现——工具、修复纠错、假设、连贯、证据、成立边界。告别刷分,拥抱真实难题,这可能是AI从工具变成科学伙伴的关键一步!

okay this is actually a pretty cool way to benchmark AI

TRACES用Tools、Repair、Alternatives、Coherence、Evidence、Scope六大维度真正评测发现过程的严谨性,而不是只看最终答案,再用423个真实高风险科学难题来压测,这种对“无答案钥匙”探索能力的定义比传统benchmark扎实太多了。

This is an important distinction. AI is already great at accelerating execution, but research still depends heavily on insight, taste, and knowing what problem is actually worth solving. Helping researchers find better questions could be one of AI's biggest contributions.

这玩意儿终于来了,测真正会挖答案的AI,而不是只会背标准答案的刷题机器。期待看看实际跑出来啥效果。

this is lowkey insane congrats team

TRACES is an interesting step toward testing whether AI can actually discover new answers, not just retrieve known ones.

The shift from testing known answers to testing genuine discovery feels like a much harder benchmark to game.

真正拉开 AI 差距的,可能不是“知道多少”,而是“能不能发现没人知道的东西”。从答案检索器走向科学发现者,这才是下一场 AI 能力竞赛的分水岭。

Apodex launches TRACES, the world’s first discoverative AI benchmark, evaluating AI’s ability to solve scientific problems through evidence and hypothesis verification when no answer key exists.

Fascinating work. Most benchmarks reward recalling existing knowledge; TRACES moves toward real scientific‑discovery evaluation. Looking forward to seeing what models deliver here.

这个方向真的很有意思。以前我们更多是在测试 AI 能不能找到已知答案,但真正困难的问题,是能不能在未知中推理、验证假设,并一步步发现新的答案。TRACES 正在把 AI 评测带到一个更有意义的方向。🧭

This is the kind of benchmark research agents actually need: no answer key, just evidence, hypotheses, and verifiable conclusions.

This is a big one coming from the founder into reality, With the help of TRACES, AI can now work through evidences. It's a big one!!

This feels like a much more meaningful test for AI. Finding an answer is one thing, but being able to investigate, challenge hypotheses, and discover something genuinely new is a different level entirely. TRACES is a fascinating direction. 🧭

这个TRACES听起来超有意思,终于有测真正发现能力的基准了!

traces 这个思路不错,以前都是让AI找答案,不如让AI去发现答案,更加能体现AI的能力

The alternative hypothesis angle is 🔥
