Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

tinyfish web agent just scored 90% on mind2web bench outperforming gemini by 21 points, openai by 29 and anthropic by 34 and we published every single run - all 300 tasks ran in parallel - in a public spreadsheet check out our runs, and try them yourself 👇

386,634 Aufrufe • vor 7 Monaten •via X (Twitter)

55 Kommentare

Profilbild von TinyFish
TinyFishvor 7 Monaten

our open-source cookbook is trending 500+ stars in 5 days tonnes of use cases you can check out, try for free, and fork ⭐️ it:

Profilbild von TinyFish
TinyFishvor 7 Monaten

read more about our benchmark results: and here's the public spreadsheet of all the runs:

Profilbild von Shubham Saboo
Shubham Saboovor 7 Monaten

great job you guys! the results look solid

Profilbild von TinyFish
TinyFishvor 7 Monaten

thank you, appreciate it

Profilbild von AshutoshShrivastava
AshutoshShrivastavavor 7 Monaten

The parallel execution is pretty amazing . Most web agents step through one action at a time and take a lot of time .

Profilbild von TinyFish
TinyFishvor 7 Monaten

yes exactly, this is how operations are done at scale

Profilbild von Francesco
Francescovor 7 Monaten

congrats for this release!!

Profilbild von TinyFish
TinyFishvor 7 Monaten

thanks a lot!

Profilbild von Santiago
Santiagovor 7 Monaten

The benchmark numbers look solid! Congrats.

Profilbild von TinyFish
TinyFishvor 7 Monaten

Thank you so much!

Profilbild von Unwind AI
Unwind AIvor 7 Monaten

Covering in the next edition of

Profilbild von Polanco | IA
Polanco | IAvor 7 Monaten

Structured outputs allow for immediate integration into pipelines. Congratulation 👏🏽

Profilbild von TinyFish
TinyFishvor 7 Monaten

exactly! nothing works in isolation

Profilbild von EyeingAI
EyeingAIvor 7 Monaten

transparency like this raises the bar for everyone👌

Profilbild von TinyFish
TinyFishvor 7 Monaten

Thanks how it should be done, thanks!

Profilbild von Simantak
Simantakvor 7 Monaten

A big congrats to the entire team!

Profilbild von TinyFish
TinyFishvor 7 Monaten

@Simantak242172 Thanks!

Profilbild von Gargi
Gargivor 7 Monaten

300 tasks in 8 minutes 😳

Profilbild von TinyFish
TinyFishvor 7 Monaten

parallel execution 🙌

Profilbild von Chidanand Tripathi
Chidanand Tripathivor 7 Monaten

TinyFish just raised the bar.

Profilbild von Minimaldex
Minimaldexvor 7 Monaten

The transparency here is what actually gets me. Most AI companies drop a chart and ask you to trust them. TinyFish literally linked a spreadsheet where you can watch the agent fail, see the reasoning, and verify the data yourself: In a world of over-promised AI, this is how you build trust with devs. What’s the first workflow you’d try to automate with 90% reliability?

Profilbild von nori
norivor 7 Monaten

what is the cameraperson with other major browser agent for mind2web?

Profilbild von Hasan
Hasanvor 7 Monaten

Impressive leap forward! 89.9% on Mind2Web with full transparency on all 300 runs sets a new standard for reliable web agents. The hybrid LLM approach is exactly what the field needs to push past compounding errors. Huge congrats to the TinyFish team

Profilbild von Miguel Ángel | GptZone
Miguel Ángel | GptZonevor 7 Monaten

Solving the problem of AI agents failing in real-world work environments is a true breakthrough.

Profilbild von Rony
Ronyvor 7 Monaten

Congrats on the launch! 🚀 90% on Mind2Web and all 300 runs public… that’s impressive.

Profilbild von Tech Fusionist | Kushal Gangil
Tech Fusionist | Kushal Gangilvor 7 Monaten

There’s a real difference between a cool demo and something you can trust in production.

Profilbild von Markandey Sharma
Markandey Sharmavor 7 Monaten

This is a serious benchmark result, especially with every run published publicly.

Profilbild von TinyFish
TinyFishvor 7 Monaten

yes, gaming a bench for making headlines is too easy and meaningless. wanted to be true to the community and ourselves

Profilbild von SARAH
SARAHvor 7 Monaten

TinyFish makes web data extraction reproducible and reliable.

Profilbild von TinyFish
TinyFishvor 7 Monaten

thank you! 🙌

Profilbild von Manish Kumar Shah
Manish Kumar Shahvor 7 Monaten

Love seeing benchmarks backed by full visibility. Open spreadsheet + parallel execution makes the result much more credible.

Profilbild von Future Coded
Future Codedvor 7 Monaten

90% on Mind2Web is serious performance. The open spreadsheet makes it even more credible

Profilbild von Vishnu Urugonda
Vishnu Urugondavor 7 Monaten

Soo cool!

Profilbild von TinyFish
TinyFishvor 7 Monaten

thank you!

Profilbild von Nelly;
Nelly;vor 7 Monaten

let's gooo

Profilbild von TinyFish
TinyFishvor 7 Monaten

yeaahh

Profilbild von Parul Gautam
Parul Gautamvor 7 Monaten

90% on Mind2Web and fully reproducible runs? That’s how agent benchmarks should be done.

Profilbild von Amit
Amitvor 7 Monaten

Okay, this is kinda crazy ngl. Huge congrats on the launch!

Profilbild von TinyFish
TinyFishvor 7 Monaten

Thank you so much!

Profilbild von Vipin Gautam (Viipin I Gautam)
Vipin Gautam (Viipin I Gautam)vor 7 Monaten

Huge milestone, 90 percent on Mind2Web with full transparency is impressive

Profilbild von TinyFish
TinyFishvor 7 Monaten

yes, gaming the benchmarks is easy. we wanted to be true to the community and ourselves!

Profilbild von Valdo
Valdovor 7 Monaten

This is truly insane. Good job guys!

Profilbild von TinyFish
TinyFishvor 7 Monaten

thank you! and do check out our open source cookbook:

Profilbild von Valdo
Valdovor 7 Monaten

👨‍🍳🔥

Profilbild von Jeremy Feng
Jeremy Fengvor 7 Monaten

Can I use it to do e2e web testing?

Profilbild von Arpita Trisha
Arpita Trishavor 7 Monaten

TinyFish Web Agent scores 90% on the Mind2Web benchmark. A outperforms Gemini by 21 points, OpenAI by 29, and Anthropic by 34.

Profilbild von Ethan Pierce
Ethan Piercevor 7 Monaten

The real question isn’t “can it browse” It’s whether it can do the same task 10,000 times without drifting. That’s where infra actually starts.

Profilbild von TinyFish
TinyFishvor 7 Monaten

a 100%

Profilbild von Atal
Atalvor 7 Monaten

Sounds crazy 300 web tasks ran in parallel and completed in 8 minutes; it used to require an Infrastructure.

Profilbild von TinyFish
TinyFishvor 7 Monaten

exactly! it's now available to EVERYONE in one api

Profilbild von SANI BULA
SANI BULAvor 7 Monaten

Structured outputs enable immediate integration into pipelines.

Profilbild von TinyFish
TinyFishvor 7 Monaten

exactly!!

Profilbild von Priyank Ahuja
Priyank Ahujavor 7 Monaten

This is incredible, really worth it

Profilbild von RatRace
RatRacevor 7 Monaten

8kyAXig7XgEmnoF9GQCgXnDTDjDapHK6Q6r9vH6upump 4 month old OG

Profilbild von Sanchoy Hossain
Sanchoy Hossainvor 7 Monaten

Good for the entire ecosystem.

Ähnliche Videos