Loading video...

Video Failed to Load

Go Home

..jesus open interpreter's first vision model, piloting my 8gb M1 macbook. 100% offline. this will be inside every computer in the world.

372,183 views • 2 years ago •via X (Twitter)

10 Comments

murat 🍥's profile picture
murat 🍥2 years ago

are there any plans for hierarchical vision models? i.e. if there are multiple play buttons on the screen, and i say "click play on youtube", it knows to first isolate youtube window for the next inference on the vision model?

killian's profile picture
killian2 years ago

actually yes, this exactly! soon, OI will be focused on the active window + let the LLM programmatically switch active windows. have been experimenting with the right way to do this. dramatically improves + speeds up vision model inference too, way less pixels.

Guli Moreno's profile picture
Guli Moreno2 years ago

This is so cool! Is this something like the Large Action Model proposed by @rabbit_hmi, where in a future state you could have it interact in the background with any app?

killian's profile picture
killian2 years ago

@rabbit_hmi thanks guli! yes, exactly what we're building—the 01 is an open-source Rabbit R1, this model is the fruit of that project.

Josh's profile picture
Josh2 years ago

Nice, y'all got icons working

killian's profile picture
killian2 years ago

yes! also you are a legend josh. would love to play with this model in the self-operating computer repo, building on the incredible advances you've made there. next OI should expose the model pretty cleanly, something like interpreter.point(screenshot_base64, query) -> coords

Roman Pshichenko's profile picture
Roman Pshichenko2 years ago

This is brilliant. I can see a integration testing framework built on top of this

killian's profile picture
killian2 years ago

yes!! great stuff in a similar vein at powered by OI

Happy's profile picture
Happy2 years ago

You're the best killian

killian's profile picture
killian2 years ago

thanks happy!

Related Videos