Загрузка видео...

Не удалось загрузить видео

На главную

tinyfish web agent just scored 90% on mind2web bench outperforming gemini by 21 points, openai by 29 and anthropic by 34 and we published every single run - all 300 tasks ran in parallel - in a public spreadsheet check out our runs, and try them yourself 👇

386,634 просмотров • 7 месяцев назад •via X (Twitter)

Комментарии: 55

Фото профиля TinyFish
TinyFish7 месяцев назад

our open-source cookbook is trending 500+ stars in 5 days tonnes of use cases you can check out, try for free, and fork ⭐️ it:

Фото профиля TinyFish
TinyFish7 месяцев назад

read more about our benchmark results: and here's the public spreadsheet of all the runs:

Фото профиля Shubham Saboo
Shubham Saboo7 месяцев назад

great job you guys! the results look solid

Фото профиля TinyFish
TinyFish7 месяцев назад

thank you, appreciate it

Фото профиля AshutoshShrivastava
AshutoshShrivastava7 месяцев назад

The parallel execution is pretty amazing . Most web agents step through one action at a time and take a lot of time .

Фото профиля TinyFish
TinyFish7 месяцев назад

yes exactly, this is how operations are done at scale

Фото профиля Francesco
Francesco7 месяцев назад

congrats for this release!!

Фото профиля TinyFish
TinyFish7 месяцев назад

thanks a lot!

Фото профиля Santiago
Santiago7 месяцев назад

The benchmark numbers look solid! Congrats.

Фото профиля TinyFish
TinyFish7 месяцев назад

Thank you so much!

Фото профиля Unwind AI
Unwind AI7 месяцев назад

Covering in the next edition of

Фото профиля Polanco | IA
Polanco | IA7 месяцев назад

Structured outputs allow for immediate integration into pipelines. Congratulation 👏🏽

Фото профиля TinyFish
TinyFish7 месяцев назад

exactly! nothing works in isolation

Фото профиля EyeingAI
EyeingAI7 месяцев назад

transparency like this raises the bar for everyone👌

Фото профиля TinyFish
TinyFish7 месяцев назад

Thanks how it should be done, thanks!

Фото профиля Simantak
Simantak7 месяцев назад

A big congrats to the entire team!

Фото профиля TinyFish
TinyFish7 месяцев назад

@Simantak242172 Thanks!

Фото профиля Gargi
Gargi7 месяцев назад

300 tasks in 8 minutes 😳

Фото профиля TinyFish
TinyFish7 месяцев назад

parallel execution 🙌

Фото профиля Chidanand Tripathi
Chidanand Tripathi7 месяцев назад

TinyFish just raised the bar.

Фото профиля Minimaldex
Minimaldex7 месяцев назад

The transparency here is what actually gets me. Most AI companies drop a chart and ask you to trust them. TinyFish literally linked a spreadsheet where you can watch the agent fail, see the reasoning, and verify the data yourself: In a world of over-promised AI, this is how you build trust with devs. What’s the first workflow you’d try to automate with 90% reliability?

Фото профиля nori
nori7 месяцев назад

what is the cameraperson with other major browser agent for mind2web?

Фото профиля Hasan
Hasan7 месяцев назад

Impressive leap forward! 89.9% on Mind2Web with full transparency on all 300 runs sets a new standard for reliable web agents. The hybrid LLM approach is exactly what the field needs to push past compounding errors. Huge congrats to the TinyFish team

Фото профиля Miguel Ángel | GptZone
Miguel Ángel | GptZone7 месяцев назад

Solving the problem of AI agents failing in real-world work environments is a true breakthrough.

Фото профиля Rony
Rony7 месяцев назад

Congrats on the launch! 🚀 90% on Mind2Web and all 300 runs public… that’s impressive.

Фото профиля Tech Fusionist | Kushal Gangil
Tech Fusionist | Kushal Gangil7 месяцев назад

There’s a real difference between a cool demo and something you can trust in production.

Фото профиля Markandey Sharma
Markandey Sharma7 месяцев назад

This is a serious benchmark result, especially with every run published publicly.

Фото профиля TinyFish
TinyFish7 месяцев назад

yes, gaming a bench for making headlines is too easy and meaningless. wanted to be true to the community and ourselves

Фото профиля SARAH
SARAH7 месяцев назад

TinyFish makes web data extraction reproducible and reliable.

Фото профиля TinyFish
TinyFish7 месяцев назад

thank you! 🙌

Фото профиля Manish Kumar Shah
Manish Kumar Shah7 месяцев назад

Love seeing benchmarks backed by full visibility. Open spreadsheet + parallel execution makes the result much more credible.

Фото профиля Future Coded
Future Coded7 месяцев назад

90% on Mind2Web is serious performance. The open spreadsheet makes it even more credible

Фото профиля Vishnu Urugonda
Vishnu Urugonda7 месяцев назад

Soo cool!

Фото профиля TinyFish
TinyFish7 месяцев назад

thank you!

Фото профиля Nelly;
Nelly;7 месяцев назад

let's gooo

Фото профиля TinyFish
TinyFish7 месяцев назад

yeaahh

Фото профиля Parul Gautam
Parul Gautam7 месяцев назад

90% on Mind2Web and fully reproducible runs? That’s how agent benchmarks should be done.

Фото профиля Amit
Amit7 месяцев назад

Okay, this is kinda crazy ngl. Huge congrats on the launch!

Фото профиля TinyFish
TinyFish7 месяцев назад

Thank you so much!

Фото профиля Vipin Gautam (Viipin I Gautam)
Vipin Gautam (Viipin I Gautam)7 месяцев назад

Huge milestone, 90 percent on Mind2Web with full transparency is impressive

Фото профиля TinyFish
TinyFish7 месяцев назад

yes, gaming the benchmarks is easy. we wanted to be true to the community and ourselves!

Фото профиля Valdo
Valdo7 месяцев назад

This is truly insane. Good job guys!

Фото профиля TinyFish
TinyFish7 месяцев назад

thank you! and do check out our open source cookbook:

Фото профиля Valdo
Valdo7 месяцев назад

👨‍🍳🔥

Фото профиля Jeremy Feng
Jeremy Feng7 месяцев назад

Can I use it to do e2e web testing?

Фото профиля Arpita Trisha
Arpita Trisha7 месяцев назад

TinyFish Web Agent scores 90% on the Mind2Web benchmark. A outperforms Gemini by 21 points, OpenAI by 29, and Anthropic by 34.

Фото профиля Ethan Pierce
Ethan Pierce7 месяцев назад

The real question isn’t “can it browse” It’s whether it can do the same task 10,000 times without drifting. That’s where infra actually starts.

Фото профиля TinyFish
TinyFish7 месяцев назад

a 100%

Фото профиля Atal
Atal7 месяцев назад

Sounds crazy 300 web tasks ran in parallel and completed in 8 minutes; it used to require an Infrastructure.

Фото профиля TinyFish
TinyFish7 месяцев назад

exactly! it's now available to EVERYONE in one api

Фото профиля SANI BULA
SANI BULA7 месяцев назад

Structured outputs enable immediate integration into pipelines.

Фото профиля TinyFish
TinyFish7 месяцев назад

exactly!!

Фото профиля Priyank Ahuja
Priyank Ahuja7 месяцев назад

This is incredible, really worth it

Фото профиля RatRace
RatRace7 месяцев назад

8kyAXig7XgEmnoF9GQCgXnDTDjDapHK6Q6r9vH6upump 4 month old OG

Фото профиля Sanchoy Hossain
Sanchoy Hossain7 месяцев назад

Good for the entire ecosystem.

Похожие видео