Loading video...

Video Failed to Load

Go Home

tinyfish web agent just scored 90% on mind2web bench outperforming gemini by 21 points, openai by 29 and anthropic by 34 and we published every single run - all 300 tasks ran in parallel - in a public spreadsheet check out our runs, and try them yourself 👇

386,634 views • 7 months ago •via X (Twitter)

55 Comments

TinyFish's profile picture
TinyFish7 months ago

our open-source cookbook is trending 500+ stars in 5 days tonnes of use cases you can check out, try for free, and fork ⭐️ it:

TinyFish's profile picture
TinyFish7 months ago

read more about our benchmark results: and here's the public spreadsheet of all the runs:

Shubham Saboo's profile picture
Shubham Saboo7 months ago

great job you guys! the results look solid

TinyFish's profile picture
TinyFish7 months ago

thank you, appreciate it

AshutoshShrivastava's profile picture
AshutoshShrivastava7 months ago

The parallel execution is pretty amazing . Most web agents step through one action at a time and take a lot of time .

TinyFish's profile picture
TinyFish7 months ago

yes exactly, this is how operations are done at scale

Francesco's profile picture
Francesco7 months ago

congrats for this release!!

TinyFish's profile picture
TinyFish7 months ago

thanks a lot!

Santiago's profile picture
Santiago7 months ago

The benchmark numbers look solid! Congrats.

TinyFish's profile picture
TinyFish7 months ago

Thank you so much!

Unwind AI's profile picture
Unwind AI7 months ago

Covering in the next edition of

Polanco | IA's profile picture
Polanco | IA7 months ago

Structured outputs allow for immediate integration into pipelines. Congratulation 👏🏽

TinyFish's profile picture
TinyFish7 months ago

exactly! nothing works in isolation

EyeingAI's profile picture
EyeingAI7 months ago

transparency like this raises the bar for everyone👌

TinyFish's profile picture
TinyFish7 months ago

Thanks how it should be done, thanks!

Simantak's profile picture
Simantak7 months ago

A big congrats to the entire team!

TinyFish's profile picture
TinyFish7 months ago

@Simantak242172 Thanks!

Gargi's profile picture
Gargi7 months ago

300 tasks in 8 minutes 😳

TinyFish's profile picture
TinyFish7 months ago

parallel execution 🙌

Chidanand Tripathi's profile picture
Chidanand Tripathi7 months ago

TinyFish just raised the bar.

Minimaldex's profile picture
Minimaldex7 months ago

The transparency here is what actually gets me. Most AI companies drop a chart and ask you to trust them. TinyFish literally linked a spreadsheet where you can watch the agent fail, see the reasoning, and verify the data yourself: In a world of over-promised AI, this is how you build trust with devs. What’s the first workflow you’d try to automate with 90% reliability?

nori's profile picture
nori7 months ago

what is the cameraperson with other major browser agent for mind2web?

Hasan's profile picture
Hasan7 months ago

Impressive leap forward! 89.9% on Mind2Web with full transparency on all 300 runs sets a new standard for reliable web agents. The hybrid LLM approach is exactly what the field needs to push past compounding errors. Huge congrats to the TinyFish team

Miguel Ángel | GptZone's profile picture
Miguel Ángel | GptZone7 months ago

Solving the problem of AI agents failing in real-world work environments is a true breakthrough.

Rony's profile picture
Rony7 months ago

Congrats on the launch! 🚀 90% on Mind2Web and all 300 runs public… that’s impressive.

Tech Fusionist | Kushal Gangil's profile picture
Tech Fusionist | Kushal Gangil7 months ago

There’s a real difference between a cool demo and something you can trust in production.

Markandey Sharma's profile picture
Markandey Sharma7 months ago

This is a serious benchmark result, especially with every run published publicly.

TinyFish's profile picture
TinyFish7 months ago

yes, gaming a bench for making headlines is too easy and meaningless. wanted to be true to the community and ourselves

SARAH's profile picture
SARAH7 months ago

TinyFish makes web data extraction reproducible and reliable.

TinyFish's profile picture
TinyFish7 months ago

thank you! 🙌

Manish Kumar Shah's profile picture
Manish Kumar Shah7 months ago

Love seeing benchmarks backed by full visibility. Open spreadsheet + parallel execution makes the result much more credible.

Future Coded's profile picture
Future Coded7 months ago

90% on Mind2Web is serious performance. The open spreadsheet makes it even more credible

Vishnu Urugonda's profile picture
Vishnu Urugonda7 months ago

Soo cool!

TinyFish's profile picture
TinyFish7 months ago

thank you!

Nelly;'s profile picture
Nelly;7 months ago

let's gooo

TinyFish's profile picture
TinyFish7 months ago

yeaahh

Parul Gautam's profile picture
Parul Gautam7 months ago

90% on Mind2Web and fully reproducible runs? That’s how agent benchmarks should be done.

Amit's profile picture
Amit7 months ago

Okay, this is kinda crazy ngl. Huge congrats on the launch!

TinyFish's profile picture
TinyFish7 months ago

Thank you so much!

Vipin Gautam (Viipin I Gautam)'s profile picture
Vipin Gautam (Viipin I Gautam)7 months ago

Huge milestone, 90 percent on Mind2Web with full transparency is impressive

TinyFish's profile picture
TinyFish7 months ago

yes, gaming the benchmarks is easy. we wanted to be true to the community and ourselves!

Valdo's profile picture
Valdo7 months ago

This is truly insane. Good job guys!

TinyFish's profile picture
TinyFish7 months ago

thank you! and do check out our open source cookbook:

Valdo's profile picture
Valdo7 months ago

👨‍🍳🔥

Jeremy Feng's profile picture
Jeremy Feng7 months ago

Can I use it to do e2e web testing?

Arpita Trisha's profile picture
Arpita Trisha7 months ago

TinyFish Web Agent scores 90% on the Mind2Web benchmark. A outperforms Gemini by 21 points, OpenAI by 29, and Anthropic by 34.

Ethan Pierce's profile picture
Ethan Pierce7 months ago

The real question isn’t “can it browse” It’s whether it can do the same task 10,000 times without drifting. That’s where infra actually starts.

TinyFish's profile picture
TinyFish7 months ago

a 100%

Atal's profile picture
Atal7 months ago

Sounds crazy 300 web tasks ran in parallel and completed in 8 minutes; it used to require an Infrastructure.

TinyFish's profile picture
TinyFish7 months ago

exactly! it's now available to EVERYONE in one api

SANI BULA's profile picture
SANI BULA7 months ago

Structured outputs enable immediate integration into pipelines.

TinyFish's profile picture
TinyFish7 months ago

exactly!!

Priyank Ahuja's profile picture
Priyank Ahuja7 months ago

This is incredible, really worth it

RatRace's profile picture
RatRace7 months ago

8kyAXig7XgEmnoF9GQCgXnDTDjDapHK6Q6r9vH6upump 4 month old OG

Sanchoy Hossain's profile picture
Sanchoy Hossain7 months ago

Good for the entire ecosystem.

Related Videos