
Martian
@withmartian • 3,786 subscribers
Understanding Intelligence. Measurement. Explanation. Application. That's how we're tackling AI interpretability: the greatest scientific problem of our age.
Shorts
Videos

We got 46% fewer errors than the single best LLM across the 16 most used benchmarks (TerminalBench, LiveCodeBench, etc). Here's how that's possible and what each model can achieve when used optimally (every benchmarks misses the majority of model capabilities) 👇 Interactive Site: Academic Paper:
Martian337,667 次观看 • 4 天前

Introducing Code Review Bench v0: The first independent code review benchmark. 200,000+ PRs. Unbiased. Fully OSS. Updated daily. Tool performance highlights 🧵👇 Featuring: Augment Code baz Claude CodeRabbit Cursor Google Gemini GitHub Graphite Greptile Kilo (acq. by Anaconda) OpenAI Developers Propel Qodo
Martian221,647 次观看 • 6 个月前

How good is Claude Code Review really, and is it worth $25+ per review? We scraped every OSS repo on GitHub that's using it to figure out how devs actually use it. Here's how it stacks up against 22 other tools: Featuring: Augment Code baz CodeAnt AI (YC W24) CodeRabbit Cognition cubic Cursor Google Gemini Greptile Kilo (acq. by Anaconda) Kody from Kodus Mesa Qodo
Martian49,789 次观看 • 5 个月前
没有更多内容可加载