Loading video...

Video Failed to Load

Go Home

Most agent comparisons end up comparing models. We wanted to test the other half: the harness. That's the environment that decides whether a search actually runs, a source gets checked, and a file gets delivered. So we gave Minara, OpenClaw, Hermes and Claude Code the same tasks on the...

27,006 views • 8 days ago •via X (Twitter)

28 Comments

Minara AI's profile picture
Minara AI8 days ago

The ask: AppLovin's monthly analyst ratings as stacked columns, plus average target price and stock price as two lines on a second axis. 37 months of data. Same model, GPT-5.6 Luna. Minara delivered it in 60s and passed 32 of 33 visual checks. Hermes passed 28. OpenClaw passed 25.

Minara AI's profile picture
Minara AI8 days ago

This is the chart Minara delivered, straight from Excel. Five rating categories, two price lines, separate axes. Close enough to the spec that there's little left to fix by hand.

Minara AI's profile picture
Minara AI8 days ago

A BrowseComp question: identify a person from a handful of clues, with no name given. Same model, Qwen3.7 Flash. Minara ran the searches and got it right in 32.8s. OpenClaw also got it right, in 251.8s. Hermes stopped at 12.9s with two searches written out and zero executed.

Minara AI's profile picture
Minara AI8 days ago

Nginx deployment on GPT-5.6 Luna: all four agents passed. Minara finished in 142.7s. The others took between 443s and 938s. A separate Python task checked what happens when you interrupt a script mid-run. Minara passed 6/6. Claude Code and Hermes passed 5/6, missing the cleanup step.

Minara AI's profile picture
Minara AI8 days ago

The exact prompt, if you want to run it yourself: ------------ I'm preparing a one-page note on AppLovin (NASDAQ: APP). 1. Find APP's total revenue and net income for each of the last 8 reported quarters, using only its official quarterly earnings releases. For every number, include the source URL. 2. Put the data in an Excel file named app_note.xlsx: one table (Quarter, Revenue, Net income, Source URL) and a native column chart of quarterly revenue titled "APP Quarterly Revenue". 3. Write summary.md with exactly 5 bullet points on the trend. Each bullet must reference at least one number from the table. Do not include any number you cannot cite. When you're done, list the files you created. ------------ Try it on Minara Harness:

OneHourResearch's profile picture
OneHourResearch8 days ago

The source-checking part is what most people skip. An answer that can't point to where each number came from shouldn't be trusted, no matter which model wrote it.

Wooam_Mon's profile picture
Wooam_Mon7 days ago

good

MIN's profile picture
MIN8 days ago

Gminara

Application Architect's profile picture
Application Architect8 days ago

keep building minara

John Rood's profile picture
John Rood8 days ago

study the 12.9s one: two searches written out, zero executed. the quietest harness failure there is, because a well-formed tool call still reads as progress while nothing runs. planned-but-unexecuted calls deserve their own column in every harness comparison.

meetkingz's profile picture
meetkingz8 days ago

Gminara

David Freeman's profile picture
David Freeman8 days ago

great

Zainab's profile picture
Zainab8 days ago

Amazing work

Alex von Mühlenen's profile picture
Alex von Mühlenen8 days ago

Would love to chat @minara

yom1985 𝔽rAI (❖,❖)'s profile picture
yom1985 𝔽rAI (❖,❖)8 days ago

Good

MR MUI's profile picture
MR MUI8 days ago

gMinara

George O'Nair's profile picture
George O'Nair8 days ago

Harness bake-offs lie when each tool picks its own ticket bank. Freeze one grader set and one spend cap across the compared shells, or you're ranking demos, not transfer.

Apihpih's profile picture
Apihpih7 days ago

gminara

catman's profile picture
catman8 days ago

Same model, different harness is like the same engine in different cars: tools, routing, and delivery determine whether the work reaches the finish line.

김동욱's profile picture
김동욱8 days ago

gMinara !!

s_a_🏌️‍♂️'s profile picture
s_a_🏌️‍♂️8 days ago

gMinara

Sunny's profile picture
Sunny8 days ago

It's all good and fun till you get unlimited usage in chatgpt chat vs any other third party CLI or IDE There exist nothing which beats that and the reason I stick to chatgpt

Jo_ (❖,❖)'s profile picture
Jo_ (❖,❖)8 days ago

gminara

Bin's profile picture
Bin8 days ago

All of them are fast and each has its own strengths, but when Harness operates on its own brain, that represents a distinct advancement of its own Huge!

Matt's profile picture
Matt8 days ago

did delivered mean the file existed, or that someone checked what was in it? most setups i've tried call it done at the first write. the one i keep drafts in makes each agent edit reviewable before it sticks, that's Sundial inline per-agent edit diffs,

Shuaibu Audu's profile picture
Shuaibu Audu8 days ago

LFG MINARA!

Automater's profile picture
Automater8 days ago

Same model, different harness → different truthfulness. The scoreboard that matters isn't tokens/sec — it's search-ran, source-checked, file-delivered, and injection contained. Benchmarks that only swap models are measuring the wrong layer. #AIAgents #AgentOps

John K's profile picture
John K8 days ago

Comparing harnesses on the same model is what most benches skip. Same brain, different scaffolding — and the timing gap shows.

Related Videos

hey if you're thinking about running qwopus (the claude opus distilled qwen 3.5 27B) as a coding agent, this might save you a few hours. i tested both the base and the distilled version on the same hardware. single RTX 3090. same prompt. same context. same everything. the only variable was the model weights. base qwen 3.5 27B built octopus invaders in 13 minutes. 1,827 lines across 11 files. zero steering. one scope bug that took 2 lines to fix. game ran. qwopus couldn't finish the same task. enemies overlapping on screen. bullets not firing. controls worked but the game was broken. i had to steer it multiple times and it still didn't produce a playable result. both run at 35 tok/s. both use thinking mode. the distilled version actually has better jinja compatibility and doesn't stall midtask like base does on claude code. for conversation and reasoning it feels sharper. but for multifile autonomous coding where the model needs to coordinate 10+ files without losing track, base wins and it's not close. distillation compresses reasoning patterns but seems to lose precision on complex coordination. the model "thinks" well but can't hold the full picture across files the way base can. tested on opencode (base) and claude code (both). next up is hermes agent framework on base. same hardware. same prompt. comparing agents now, not just models. video below. first half is the distilled model's broken game. second half is what base built on the same 3090. judge for yourself.

Sudo su

45,052 views • 7 months ago

Sam Altman made the case for open-source harnesses in July. a month later, someone shipped it, and it's more efficient than most managed harnesses. here is the problem it was aimed at: a large share of your agent's token bill is the model rereading things it already read. that isn't the model's doing. the runtime around it decides what goes into every prompt and how often the model gets called. for example, an agent queries a CRM at step four and gets back 400 rows. those rows get piled up in the conversation history. by step nineteen, the model has to read those rows fifteen times unnecessarily, and every token read is billed at input rates. it happened because your harness assembled that prompt on every turn and kept the rows in it. that gives you two levers: how much context the harness carries forward, and how often it calls the model. there are four practical ways to keep the prompt from growing unnecessarily: → load tool schemas on demand. a server with 100 tools doesn't need to put all 100 into every prompt when the agent only calls two. → offload large results to disk. turn a large response into a short preview and a file path instead of replaying the entire result on every turn. → delegate to subagents. let a subagent spend thirty tool calls in its own context and return one summary to the root agent. → run toolchains in code. one script calls three tools, joins the results, and returns a table instead of three turns each dragging a full response. but reducing context is only half the job. you also need to control how often the model gets called. a good harness should avoid unnecessary planning, verification, and reflection when the work can be completed in fewer steps. TrueFoundry's open-source agent harness, TrueForge, is built around both of those controls. it sits between the model and the tools, deciding what goes into every prompt and when another model call is actually needed. it also breaks token usage down across the harness, skills, instructions, tools, and messages. DevRev's Enterprise-Bench is where this gets tested, on multi-step tasks of the kind where an agent pulls records from one system and reconciles them against another. TrueFoundry ran TrueForge there against Claude Managed Agents, both on the same model, and both finished the same number of tasks. the tie is the part that matters, because it means the gap underneath is not a quality tradeoff. TrueForge reached that score on close to a third of the tokens, with roughly 40% fewer trips back to the model. for the same result, that comes out around 2.7x cheaper than Claude Managed Agents. swapping in an open model made it sharper still. TrueForge with GLM-5.2 scored a little higher than either setup above, and the entire benchmark run cost about $3 at list prices. being open source matters beyond the license here. the model underneath can be swapped without rewriting the agent, and the whole thing can run inside your own environment when the data cannot leave it. all of this comes down to the runtime around the model, the context it carries, the tools it exposes, and how many times it goes back to the model. that is what a production harness actually owns. the full task list, the per-run numbers, and the MIT-licensed code are on GitHub: (don't forget to star 🌟) you can read more about the same in the article quoted below. thanks to the TrueForge team for working with me on this one.

Akshay 🚀

76,998 views • 1 month ago