Загрузка видео...

Не удалось загрузить видео

На главную

Most agent comparisons end up comparing models. We wanted to test the other half: the harness. That's the environment that decides whether a search actually runs, a source gets checked, and a file gets delivered. So we gave Minara, OpenClaw, Hermes and Claude Code the same tasks on the...

27,006 просмотров • 8 дней назад •via X (Twitter)

Комментарии: 28

Фото профиля Minara AI
Minara AI8 дней назад

The ask: AppLovin's monthly analyst ratings as stacked columns, plus average target price and stock price as two lines on a second axis. 37 months of data. Same model, GPT-5.6 Luna. Minara delivered it in 60s and passed 32 of 33 visual checks. Hermes passed 28. OpenClaw passed 25.

Фото профиля Minara AI
Minara AI8 дней назад

This is the chart Minara delivered, straight from Excel. Five rating categories, two price lines, separate axes. Close enough to the spec that there's little left to fix by hand.

Фото профиля Minara AI
Minara AI8 дней назад

A BrowseComp question: identify a person from a handful of clues, with no name given. Same model, Qwen3.7 Flash. Minara ran the searches and got it right in 32.8s. OpenClaw also got it right, in 251.8s. Hermes stopped at 12.9s with two searches written out and zero executed.

Фото профиля Minara AI
Minara AI8 дней назад

Nginx deployment on GPT-5.6 Luna: all four agents passed. Minara finished in 142.7s. The others took between 443s and 938s. A separate Python task checked what happens when you interrupt a script mid-run. Minara passed 6/6. Claude Code and Hermes passed 5/6, missing the cleanup step.

Фото профиля Minara AI
Minara AI8 дней назад

The exact prompt, if you want to run it yourself: ------------ I'm preparing a one-page note on AppLovin (NASDAQ: APP). 1. Find APP's total revenue and net income for each of the last 8 reported quarters, using only its official quarterly earnings releases. For every number, include the source URL. 2. Put the data in an Excel file named app_note.xlsx: one table (Quarter, Revenue, Net income, Source URL) and a native column chart of quarterly revenue titled "APP Quarterly Revenue". 3. Write summary.md with exactly 5 bullet points on the trend. Each bullet must reference at least one number from the table. Do not include any number you cannot cite. When you're done, list the files you created. ------------ Try it on Minara Harness:

Фото профиля OneHourResearch
OneHourResearch8 дней назад

The source-checking part is what most people skip. An answer that can't point to where each number came from shouldn't be trusted, no matter which model wrote it.

Фото профиля Wooam_Mon
Wooam_Mon7 дней назад

good

Фото профиля MIN
MIN8 дней назад

Gminara

Фото профиля Application Architect
Application Architect8 дней назад

keep building minara

Фото профиля John Rood
John Rood8 дней назад

study the 12.9s one: two searches written out, zero executed. the quietest harness failure there is, because a well-formed tool call still reads as progress while nothing runs. planned-but-unexecuted calls deserve their own column in every harness comparison.

Фото профиля meetkingz
meetkingz8 дней назад

Gminara

Фото профиля David Freeman
David Freeman8 дней назад

great

Фото профиля Zainab
Zainab8 дней назад

Amazing work

Фото профиля Alex von Mühlenen
Alex von Mühlenen8 дней назад

Would love to chat @minara

Фото профиля yom1985 𝔽rAI (❖,❖)
yom1985 𝔽rAI (❖,❖)8 дней назад

Good

Фото профиля MR MUI
MR MUI8 дней назад

gMinara

Фото профиля George O'Nair
George O'Nair8 дней назад

Harness bake-offs lie when each tool picks its own ticket bank. Freeze one grader set and one spend cap across the compared shells, or you're ranking demos, not transfer.

Фото профиля Apihpih
Apihpih7 дней назад

gminara

Фото профиля catman
catman8 дней назад

Same model, different harness is like the same engine in different cars: tools, routing, and delivery determine whether the work reaches the finish line.

Фото профиля 김동욱
김동욱8 дней назад

gMinara !!

Фото профиля s_a_🏌️‍♂️
s_a_🏌️‍♂️8 дней назад

gMinara

Фото профиля Sunny
Sunny8 дней назад

It's all good and fun till you get unlimited usage in chatgpt chat vs any other third party CLI or IDE There exist nothing which beats that and the reason I stick to chatgpt

Фото профиля Jo_ (❖,❖)
Jo_ (❖,❖)8 дней назад

gminara

Фото профиля Bin
Bin8 дней назад

All of them are fast and each has its own strengths, but when Harness operates on its own brain, that represents a distinct advancement of its own Huge!

Фото профиля Matt
Matt8 дней назад

did delivered mean the file existed, or that someone checked what was in it? most setups i've tried call it done at the first write. the one i keep drafts in makes each agent edit reviewable before it sticks, that's Sundial inline per-agent edit diffs,

Фото профиля Shuaibu Audu
Shuaibu Audu8 дней назад

LFG MINARA!

Фото профиля Automater
Automater8 дней назад

Same model, different harness → different truthfulness. The scoreboard that matters isn't tokens/sec — it's search-ran, source-checked, file-delivered, and injection contained. Benchmarks that only swap models are measuring the wrong layer. #AIAgents #AgentOps

Фото профиля John K
John K8 дней назад

Comparing harnesses on the same model is what most benches skip. Same brain, different scaffolding — and the timing gap shows.

Похожие видео

hey if you're thinking about running qwopus (the claude opus distilled qwen 3.5 27B) as a coding agent, this might save you a few hours. i tested both the base and the distilled version on the same hardware. single RTX 3090. same prompt. same context. same everything. the only variable was the model weights. base qwen 3.5 27B built octopus invaders in 13 minutes. 1,827 lines across 11 files. zero steering. one scope bug that took 2 lines to fix. game ran. qwopus couldn't finish the same task. enemies overlapping on screen. bullets not firing. controls worked but the game was broken. i had to steer it multiple times and it still didn't produce a playable result. both run at 35 tok/s. both use thinking mode. the distilled version actually has better jinja compatibility and doesn't stall midtask like base does on claude code. for conversation and reasoning it feels sharper. but for multifile autonomous coding where the model needs to coordinate 10+ files without losing track, base wins and it's not close. distillation compresses reasoning patterns but seems to lose precision on complex coordination. the model "thinks" well but can't hold the full picture across files the way base can. tested on opencode (base) and claude code (both). next up is hermes agent framework on base. same hardware. same prompt. comparing agents now, not just models. video below. first half is the distilled model's broken game. second half is what base built on the same 3090. judge for yourself.

Sudo su

45,052 просмотров • 7 месяцев назад

Sam Altman made the case for open-source harnesses in July. a month later, someone shipped it, and it's more efficient than most managed harnesses. here is the problem it was aimed at: a large share of your agent's token bill is the model rereading things it already read. that isn't the model's doing. the runtime around it decides what goes into every prompt and how often the model gets called. for example, an agent queries a CRM at step four and gets back 400 rows. those rows get piled up in the conversation history. by step nineteen, the model has to read those rows fifteen times unnecessarily, and every token read is billed at input rates. it happened because your harness assembled that prompt on every turn and kept the rows in it. that gives you two levers: how much context the harness carries forward, and how often it calls the model. there are four practical ways to keep the prompt from growing unnecessarily: → load tool schemas on demand. a server with 100 tools doesn't need to put all 100 into every prompt when the agent only calls two. → offload large results to disk. turn a large response into a short preview and a file path instead of replaying the entire result on every turn. → delegate to subagents. let a subagent spend thirty tool calls in its own context and return one summary to the root agent. → run toolchains in code. one script calls three tools, joins the results, and returns a table instead of three turns each dragging a full response. but reducing context is only half the job. you also need to control how often the model gets called. a good harness should avoid unnecessary planning, verification, and reflection when the work can be completed in fewer steps. TrueFoundry's open-source agent harness, TrueForge, is built around both of those controls. it sits between the model and the tools, deciding what goes into every prompt and when another model call is actually needed. it also breaks token usage down across the harness, skills, instructions, tools, and messages. DevRev's Enterprise-Bench is where this gets tested, on multi-step tasks of the kind where an agent pulls records from one system and reconciles them against another. TrueFoundry ran TrueForge there against Claude Managed Agents, both on the same model, and both finished the same number of tasks. the tie is the part that matters, because it means the gap underneath is not a quality tradeoff. TrueForge reached that score on close to a third of the tokens, with roughly 40% fewer trips back to the model. for the same result, that comes out around 2.7x cheaper than Claude Managed Agents. swapping in an open model made it sharper still. TrueForge with GLM-5.2 scored a little higher than either setup above, and the entire benchmark run cost about $3 at list prices. being open source matters beyond the license here. the model underneath can be swapped without rewriting the agent, and the whole thing can run inside your own environment when the data cannot leave it. all of this comes down to the runtime around the model, the context it carries, the tools it exposes, and how many times it goes back to the model. that is what a production harness actually owns. the full task list, the per-run numbers, and the MIT-licensed code are on GitHub: (don't forget to star 🌟) you can read more about the same in the article quoted below. thanks to the TrueForge team for working with me on this one.

Akshay 🚀

76,998 просмотров • 1 месяц назад