Загрузка видео...
Не удалось загрузить видео
tinyfish web agent just scored 90% on mind2web bench outperforming gemini by 21 points, openai by 29 and anthropic by 34 and we published every single run - all 300 tasks ran in parallel - in a public spreadsheet check out our runs, and try them yourself 👇
386,634 просмотров • 7 месяцев назад •via X (Twitter)
Комментарии: 55

our open-source cookbook is trending 500+ stars in 5 days tonnes of use cases you can check out, try for free, and fork ⭐️ it:

read more about our benchmark results: and here's the public spreadsheet of all the runs:

great job you guys! the results look solid

thank you, appreciate it

The parallel execution is pretty amazing . Most web agents step through one action at a time and take a lot of time .

yes exactly, this is how operations are done at scale

congrats for this release!!

thanks a lot!

The benchmark numbers look solid! Congrats.

Thank you so much!

Covering in the next edition of

Structured outputs allow for immediate integration into pipelines. Congratulation 👏🏽

exactly! nothing works in isolation

transparency like this raises the bar for everyone👌

Thanks how it should be done, thanks!

A big congrats to the entire team!

@Simantak242172 Thanks!

300 tasks in 8 minutes 😳

parallel execution 🙌

TinyFish just raised the bar.

The transparency here is what actually gets me. Most AI companies drop a chart and ask you to trust them. TinyFish literally linked a spreadsheet where you can watch the agent fail, see the reasoning, and verify the data yourself: In a world of over-promised AI, this is how you build trust with devs. What’s the first workflow you’d try to automate with 90% reliability?

what is the cameraperson with other major browser agent for mind2web?

Impressive leap forward! 89.9% on Mind2Web with full transparency on all 300 runs sets a new standard for reliable web agents. The hybrid LLM approach is exactly what the field needs to push past compounding errors. Huge congrats to the TinyFish team

Solving the problem of AI agents failing in real-world work environments is a true breakthrough.

Congrats on the launch! 🚀 90% on Mind2Web and all 300 runs public… that’s impressive.

There’s a real difference between a cool demo and something you can trust in production.

This is a serious benchmark result, especially with every run published publicly.

yes, gaming a bench for making headlines is too easy and meaningless. wanted to be true to the community and ourselves

TinyFish makes web data extraction reproducible and reliable.

thank you! 🙌

Love seeing benchmarks backed by full visibility. Open spreadsheet + parallel execution makes the result much more credible.

90% on Mind2Web is serious performance. The open spreadsheet makes it even more credible

Soo cool!

thank you!

let's gooo

yeaahh

90% on Mind2Web and fully reproducible runs? That’s how agent benchmarks should be done.

Okay, this is kinda crazy ngl. Huge congrats on the launch!

Thank you so much!

Huge milestone, 90 percent on Mind2Web with full transparency is impressive

yes, gaming the benchmarks is easy. we wanted to be true to the community and ourselves!

This is truly insane. Good job guys!

thank you! and do check out our open source cookbook:

👨🍳🔥

Can I use it to do e2e web testing?

TinyFish Web Agent scores 90% on the Mind2Web benchmark. A outperforms Gemini by 21 points, OpenAI by 29, and Anthropic by 34.

The real question isn’t “can it browse” It’s whether it can do the same task 10,000 times without drifting. That’s where infra actually starts.

a 100%

Sounds crazy 300 web tasks ran in parallel and completed in 8 minutes; it used to require an Infrastructure.

exactly! it's now available to EVERYONE in one api

Structured outputs enable immediate integration into pipelines.

exactly!!

This is incredible, really worth it

8kyAXig7XgEmnoF9GQCgXnDTDjDapHK6Q6r9vH6upump 4 month old OG

Good for the entire ecosystem.


