Loading video...

Video Failed to Load

Go Home

Florence-2, the new vision foundation model by Microsoft, can now run 100% locally in your browser on WebGPU, thanks to Transformers.js! 🤗🤯 It supports tasks like image captioning, optical character recognition, object detection, and many more! 😍 WOW! Demo (+ source code) 👇

88,762 views • 2 years ago •via X (Twitter)

9 Comments

Xenova's profile picture
Xenova2 years ago

ONNX models: Source code: Demo:

nickmystic's profile picture
nickmystic2 years ago

amazing work!

Samuel Tallet's profile picture
Samuel Tallet2 years ago

It's awesome, thank you Xenova! Does the "Florence-2-large" model can also run on the client-side with WebGPU acceleration?

snats's profile picture
snats2 years ago

When is v3 coming out? I love this!

Aiflowly.com's profile picture
Aiflowly.com2 years ago

Amazing! We should add this to the list of our integrations.

Aaron Planell's profile picture
Aaron Planell2 years ago

@daviddincognit Maybe this can be interesting for you

PDS_B2BMGMT's profile picture
PDS_B2BMGMT2 years ago

I tried twice! I only got this

Thomas Hill's profile picture
Thomas Hill2 years ago

🔥

Gather Grove's profile picture
Gather Grove2 years ago

This actually works great even on my lame GPU, but it's accuracy is kinda random. The next version will be awesome.

Related Videos