Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

tinyfish web agent just scored 90% on mind2web bench outperforming gemini by 21 points, openai by 29 and anthropic by 34 and we published every single run - all 300 tasks ran in parallel - in a public spreadsheet check out our runs, and try them yourself 👇

386,634 görüntüleme • 7 ay önce •via X (Twitter)

55 Yorum

TinyFish profil fotoğrafı
TinyFish7 ay önce

our open-source cookbook is trending 500+ stars in 5 days tonnes of use cases you can check out, try for free, and fork ⭐️ it:

TinyFish profil fotoğrafı
TinyFish7 ay önce

read more about our benchmark results: and here's the public spreadsheet of all the runs:

Shubham Saboo profil fotoğrafı
Shubham Saboo7 ay önce

great job you guys! the results look solid

TinyFish profil fotoğrafı
TinyFish7 ay önce

thank you, appreciate it

AshutoshShrivastava profil fotoğrafı
AshutoshShrivastava7 ay önce

The parallel execution is pretty amazing . Most web agents step through one action at a time and take a lot of time .

TinyFish profil fotoğrafı
TinyFish7 ay önce

yes exactly, this is how operations are done at scale

Francesco profil fotoğrafı
Francesco7 ay önce

congrats for this release!!

TinyFish profil fotoğrafı
TinyFish7 ay önce

thanks a lot!

Santiago profil fotoğrafı
Santiago7 ay önce

The benchmark numbers look solid! Congrats.

TinyFish profil fotoğrafı
TinyFish7 ay önce

Thank you so much!

Unwind AI profil fotoğrafı
Unwind AI7 ay önce

Covering in the next edition of

Polanco | IA profil fotoğrafı
Polanco | IA7 ay önce

Structured outputs allow for immediate integration into pipelines. Congratulation 👏🏽

TinyFish profil fotoğrafı
TinyFish7 ay önce

exactly! nothing works in isolation

EyeingAI profil fotoğrafı
EyeingAI7 ay önce

transparency like this raises the bar for everyone👌

TinyFish profil fotoğrafı
TinyFish7 ay önce

Thanks how it should be done, thanks!

Simantak profil fotoğrafı
Simantak7 ay önce

A big congrats to the entire team!

TinyFish profil fotoğrafı
TinyFish7 ay önce

@Simantak242172 Thanks!

Gargi profil fotoğrafı
Gargi7 ay önce

300 tasks in 8 minutes 😳

TinyFish profil fotoğrafı
TinyFish7 ay önce

parallel execution 🙌

Chidanand Tripathi profil fotoğrafı
Chidanand Tripathi7 ay önce

TinyFish just raised the bar.

Minimaldex profil fotoğrafı
Minimaldex7 ay önce

The transparency here is what actually gets me. Most AI companies drop a chart and ask you to trust them. TinyFish literally linked a spreadsheet where you can watch the agent fail, see the reasoning, and verify the data yourself: In a world of over-promised AI, this is how you build trust with devs. What’s the first workflow you’d try to automate with 90% reliability?

nori profil fotoğrafı
nori7 ay önce

what is the cameraperson with other major browser agent for mind2web?

Hasan profil fotoğrafı
Hasan7 ay önce

Impressive leap forward! 89.9% on Mind2Web with full transparency on all 300 runs sets a new standard for reliable web agents. The hybrid LLM approach is exactly what the field needs to push past compounding errors. Huge congrats to the TinyFish team

Miguel Ángel | GptZone profil fotoğrafı
Miguel Ángel | GptZone7 ay önce

Solving the problem of AI agents failing in real-world work environments is a true breakthrough.

Rony profil fotoğrafı
Rony7 ay önce

Congrats on the launch! 🚀 90% on Mind2Web and all 300 runs public… that’s impressive.

Tech Fusionist | Kushal Gangil profil fotoğrafı
Tech Fusionist | Kushal Gangil7 ay önce

There’s a real difference between a cool demo and something you can trust in production.

Markandey Sharma profil fotoğrafı
Markandey Sharma7 ay önce

This is a serious benchmark result, especially with every run published publicly.

TinyFish profil fotoğrafı
TinyFish7 ay önce

yes, gaming a bench for making headlines is too easy and meaningless. wanted to be true to the community and ourselves

SARAH profil fotoğrafı
SARAH7 ay önce

TinyFish makes web data extraction reproducible and reliable.

TinyFish profil fotoğrafı
TinyFish7 ay önce

thank you! 🙌

Manish Kumar Shah profil fotoğrafı
Manish Kumar Shah7 ay önce

Love seeing benchmarks backed by full visibility. Open spreadsheet + parallel execution makes the result much more credible.

Future Coded profil fotoğrafı
Future Coded7 ay önce

90% on Mind2Web is serious performance. The open spreadsheet makes it even more credible

Vishnu Urugonda profil fotoğrafı
Vishnu Urugonda7 ay önce

Soo cool!

TinyFish profil fotoğrafı
TinyFish7 ay önce

thank you!

Nelly; profil fotoğrafı
Nelly;7 ay önce

let's gooo

TinyFish profil fotoğrafı
TinyFish7 ay önce

yeaahh

Parul Gautam profil fotoğrafı
Parul Gautam7 ay önce

90% on Mind2Web and fully reproducible runs? That’s how agent benchmarks should be done.

Amit profil fotoğrafı
Amit7 ay önce

Okay, this is kinda crazy ngl. Huge congrats on the launch!

TinyFish profil fotoğrafı
TinyFish7 ay önce

Thank you so much!

Vipin Gautam (Viipin I Gautam) profil fotoğrafı
Vipin Gautam (Viipin I Gautam)7 ay önce

Huge milestone, 90 percent on Mind2Web with full transparency is impressive

TinyFish profil fotoğrafı
TinyFish7 ay önce

yes, gaming the benchmarks is easy. we wanted to be true to the community and ourselves!

Valdo profil fotoğrafı
Valdo7 ay önce

This is truly insane. Good job guys!

TinyFish profil fotoğrafı
TinyFish7 ay önce

thank you! and do check out our open source cookbook:

Valdo profil fotoğrafı
Valdo7 ay önce

👨‍🍳🔥

Jeremy Feng profil fotoğrafı
Jeremy Feng7 ay önce

Can I use it to do e2e web testing?

Arpita Trisha profil fotoğrafı
Arpita Trisha7 ay önce

TinyFish Web Agent scores 90% on the Mind2Web benchmark. A outperforms Gemini by 21 points, OpenAI by 29, and Anthropic by 34.

Ethan Pierce profil fotoğrafı
Ethan Pierce7 ay önce

The real question isn’t “can it browse” It’s whether it can do the same task 10,000 times without drifting. That’s where infra actually starts.

TinyFish profil fotoğrafı
TinyFish7 ay önce

a 100%

Atal profil fotoğrafı
Atal7 ay önce

Sounds crazy 300 web tasks ran in parallel and completed in 8 minutes; it used to require an Infrastructure.

TinyFish profil fotoğrafı
TinyFish7 ay önce

exactly! it's now available to EVERYONE in one api

SANI BULA profil fotoğrafı
SANI BULA7 ay önce

Structured outputs enable immediate integration into pipelines.

TinyFish profil fotoğrafı
TinyFish7 ay önce

exactly!!

Priyank Ahuja profil fotoğrafı
Priyank Ahuja7 ay önce

This is incredible, really worth it

RatRace profil fotoğrafı
RatRace7 ay önce

8kyAXig7XgEmnoF9GQCgXnDTDjDapHK6Q6r9vH6upump 4 month old OG

Sanchoy Hossain profil fotoğrafı
Sanchoy Hossain7 ay önce

Good for the entire ecosystem.

Benzer Videolar