正在加载视频...

视频加载失败

tinyfish web agent just scored 90% on mind2web bench outperforming gemini by 21 points, openai by 29 and anthropic by 34 and we published every single run - all 300 tasks ran in parallel - in a public spreadsheet check out our runs, and try them yourself 👇

386,634 次观看 • 7 个月前 •via X (Twitter)

55 条评论

TinyFish 的头像
TinyFish7 个月前

our open-source cookbook is trending 500+ stars in 5 days tonnes of use cases you can check out, try for free, and fork ⭐️ it:

TinyFish 的头像
TinyFish7 个月前

read more about our benchmark results: and here's the public spreadsheet of all the runs:

Shubham Saboo 的头像
Shubham Saboo7 个月前

great job you guys! the results look solid

TinyFish 的头像
TinyFish7 个月前

thank you, appreciate it

AshutoshShrivastava 的头像
AshutoshShrivastava7 个月前

The parallel execution is pretty amazing . Most web agents step through one action at a time and take a lot of time .

TinyFish 的头像
TinyFish7 个月前

yes exactly, this is how operations are done at scale

Francesco 的头像
Francesco7 个月前

congrats for this release!!

TinyFish 的头像
TinyFish7 个月前

thanks a lot!

Santiago 的头像
Santiago7 个月前

The benchmark numbers look solid! Congrats.

TinyFish 的头像
TinyFish7 个月前

Thank you so much!

Unwind AI 的头像
Unwind AI7 个月前

Covering in the next edition of

Polanco | IA 的头像
Polanco | IA7 个月前

Structured outputs allow for immediate integration into pipelines. Congratulation 👏🏽

TinyFish 的头像
TinyFish7 个月前

exactly! nothing works in isolation

EyeingAI 的头像
EyeingAI7 个月前

transparency like this raises the bar for everyone👌

TinyFish 的头像
TinyFish7 个月前

Thanks how it should be done, thanks!

Simantak 的头像
Simantak7 个月前

A big congrats to the entire team!

TinyFish 的头像
TinyFish7 个月前

@Simantak242172 Thanks!

Gargi 的头像
Gargi7 个月前

300 tasks in 8 minutes 😳

TinyFish 的头像
TinyFish7 个月前

parallel execution 🙌

Chidanand Tripathi 的头像
Chidanand Tripathi7 个月前

TinyFish just raised the bar.

Minimaldex 的头像
Minimaldex7 个月前

The transparency here is what actually gets me. Most AI companies drop a chart and ask you to trust them. TinyFish literally linked a spreadsheet where you can watch the agent fail, see the reasoning, and verify the data yourself: In a world of over-promised AI, this is how you build trust with devs. What’s the first workflow you’d try to automate with 90% reliability?

nori 的头像
nori7 个月前

what is the cameraperson with other major browser agent for mind2web?

Hasan 的头像
Hasan7 个月前

Impressive leap forward! 89.9% on Mind2Web with full transparency on all 300 runs sets a new standard for reliable web agents. The hybrid LLM approach is exactly what the field needs to push past compounding errors. Huge congrats to the TinyFish team

Miguel Ángel | GptZone 的头像
Miguel Ángel | GptZone7 个月前

Solving the problem of AI agents failing in real-world work environments is a true breakthrough.

Rony 的头像
Rony7 个月前

Congrats on the launch! 🚀 90% on Mind2Web and all 300 runs public… that’s impressive.

Tech Fusionist | Kushal Gangil 的头像
Tech Fusionist | Kushal Gangil7 个月前

There’s a real difference between a cool demo and something you can trust in production.

Markandey Sharma 的头像
Markandey Sharma7 个月前

This is a serious benchmark result, especially with every run published publicly.

TinyFish 的头像
TinyFish7 个月前

yes, gaming a bench for making headlines is too easy and meaningless. wanted to be true to the community and ourselves

SARAH 的头像
SARAH7 个月前

TinyFish makes web data extraction reproducible and reliable.

TinyFish 的头像
TinyFish7 个月前

thank you! 🙌

Manish Kumar Shah 的头像
Manish Kumar Shah7 个月前

Love seeing benchmarks backed by full visibility. Open spreadsheet + parallel execution makes the result much more credible.

Future Coded 的头像
Future Coded7 个月前

90% on Mind2Web is serious performance. The open spreadsheet makes it even more credible

Vishnu Urugonda 的头像
Vishnu Urugonda7 个月前

Soo cool!

TinyFish 的头像
TinyFish7 个月前

thank you!

Nelly; 的头像
Nelly;7 个月前

let's gooo

TinyFish 的头像
TinyFish7 个月前

yeaahh

Parul Gautam 的头像
Parul Gautam7 个月前

90% on Mind2Web and fully reproducible runs? That’s how agent benchmarks should be done.

Amit 的头像
Amit7 个月前

Okay, this is kinda crazy ngl. Huge congrats on the launch!

TinyFish 的头像
TinyFish7 个月前

Thank you so much!

Vipin Gautam (Viipin I Gautam) 的头像
Vipin Gautam (Viipin I Gautam)7 个月前

Huge milestone, 90 percent on Mind2Web with full transparency is impressive

TinyFish 的头像
TinyFish7 个月前

yes, gaming the benchmarks is easy. we wanted to be true to the community and ourselves!

Valdo 的头像
Valdo7 个月前

This is truly insane. Good job guys!

TinyFish 的头像
TinyFish7 个月前

thank you! and do check out our open source cookbook:

Valdo 的头像
Valdo7 个月前

👨‍🍳🔥

Jeremy Feng 的头像
Jeremy Feng7 个月前

Can I use it to do e2e web testing?

Arpita Trisha 的头像
Arpita Trisha7 个月前

TinyFish Web Agent scores 90% on the Mind2Web benchmark. A outperforms Gemini by 21 points, OpenAI by 29, and Anthropic by 34.

Ethan Pierce 的头像
Ethan Pierce7 个月前

The real question isn’t “can it browse” It’s whether it can do the same task 10,000 times without drifting. That’s where infra actually starts.

TinyFish 的头像
TinyFish7 个月前

a 100%

Atal 的头像
Atal7 个月前

Sounds crazy 300 web tasks ran in parallel and completed in 8 minutes; it used to require an Infrastructure.

TinyFish 的头像
TinyFish7 个月前

exactly! it's now available to EVERYONE in one api

SANI BULA 的头像
SANI BULA7 个月前

Structured outputs enable immediate integration into pipelines.

TinyFish 的头像
TinyFish7 个月前

exactly!!

Priyank Ahuja 的头像
Priyank Ahuja7 个月前

This is incredible, really worth it

RatRace 的头像
RatRace7 个月前

8kyAXig7XgEmnoF9GQCgXnDTDjDapHK6Q6r9vH6upump 4 month old OG

Sanchoy Hossain 的头像
Sanchoy Hossain7 个月前

Good for the entire ecosystem.

相关视频