Video wird geladen...
Video konnte nicht geladen werden
Most agent comparisons end up comparing models. We wanted to test the other half: the harness. That's the environment that decides whether a search actually runs, a source gets checked, and a file gets delivered. So we gave Minara, OpenClaw, Hermes and Claude Code the same tasks on the... show more
27,006 Aufrufe • vor 8 Tagen •via X (Twitter)
28 Kommentare

The ask: AppLovin's monthly analyst ratings as stacked columns, plus average target price and stock price as two lines on a second axis. 37 months of data. Same model, GPT-5.6 Luna. Minara delivered it in 60s and passed 32 of 33 visual checks. Hermes passed 28. OpenClaw passed 25.

This is the chart Minara delivered, straight from Excel. Five rating categories, two price lines, separate axes. Close enough to the spec that there's little left to fix by hand.

A BrowseComp question: identify a person from a handful of clues, with no name given. Same model, Qwen3.7 Flash. Minara ran the searches and got it right in 32.8s. OpenClaw also got it right, in 251.8s. Hermes stopped at 12.9s with two searches written out and zero executed.

Nginx deployment on GPT-5.6 Luna: all four agents passed. Minara finished in 142.7s. The others took between 443s and 938s. A separate Python task checked what happens when you interrupt a script mid-run. Minara passed 6/6. Claude Code and Hermes passed 5/6, missing the cleanup step.

The exact prompt, if you want to run it yourself: ------------ I'm preparing a one-page note on AppLovin (NASDAQ: APP). 1. Find APP's total revenue and net income for each of the last 8 reported quarters, using only its official quarterly earnings releases. For every number, include the source URL. 2. Put the data in an Excel file named app_note.xlsx: one table (Quarter, Revenue, Net income, Source URL) and a native column chart of quarterly revenue titled "APP Quarterly Revenue". 3. Write summary.md with exactly 5 bullet points on the trend. Each bullet must reference at least one number from the table. Do not include any number you cannot cite. When you're done, list the files you created. ------------ Try it on Minara Harness:

The source-checking part is what most people skip. An answer that can't point to where each number came from shouldn't be trusted, no matter which model wrote it.

good

Gminara

keep building minara

study the 12.9s one: two searches written out, zero executed. the quietest harness failure there is, because a well-formed tool call still reads as progress while nothing runs. planned-but-unexecuted calls deserve their own column in every harness comparison.

Gminara

great

Amazing work

Would love to chat @minara

Good

gMinara

Harness bake-offs lie when each tool picks its own ticket bank. Freeze one grader set and one spend cap across the compared shells, or you're ranking demos, not transfer.

gminara

Same model, different harness is like the same engine in different cars: tools, routing, and delivery determine whether the work reaches the finish line.

gMinara !!

gMinara

It's all good and fun till you get unlimited usage in chatgpt chat vs any other third party CLI or IDE There exist nothing which beats that and the reason I stick to chatgpt

gminara

All of them are fast and each has its own strengths, but when Harness operates on its own brain, that represents a distinct advancement of its own Huge!

did delivered mean the file existed, or that someone checked what was in it? most setups i've tried call it done at the first write. the one i keep drafts in makes each agent edit reviewable before it sticks, that's Sundial inline per-agent edit diffs,

LFG MINARA!

Same model, different harness → different truthfulness. The scoreboard that matters isn't tokens/sec — it's search-ran, source-checked, file-delivered, and injection contained. Benchmarks that only swap models are measuring the wrong layer. #AIAgents #AgentOps

Comparing harnesses on the same model is what most benches skip. Same brain, different scaffolding — and the timing gap shows.
