Loading video...

Video Failed to Load

Go Home

Okay so Jev can actually do computer use really well Without any screenshots, or LLMs and no Pixels leave my mac I dont even read the Dom elements A local CoreML model segments every button and UI element on screen. On-device OCR reads the labels. That text is all...

116,719 views • 1 day ago •via X (Twitter)

62 Comments

Milind S's profile picture
Milind S1 day ago

If this gets enough attention I’ll make it open source. Meanwhile check out my other stuff

Vedant Bhayani's profile picture
Vedant Bhayani1 day ago

I think the future is multiple other companies building new kind of llm models for individual tasks with which you just plug it in with frontiers model for steering and thinking Also domain specific harness and custom trained models with your data for specific use cases are going be more relevant and we are going to see even more than before

Milind S's profile picture
Milind S1 day ago

I agree that’s where things are headed

/h(elton)?Fabio\.(t|j)s/gi's profile picture
/h(elton)?Fabio\.(t|j)s/gi1 day ago

hey man, do you mind doing a test with a local model that uses the same architecture? like this one based on qwen-2.5-1b?

Milind S's profile picture
Milind S1 day ago

I’ll check it out

Valentin Pletzer's profile picture
Valentin Pletzer1 day ago

very cool!

Milind S's profile picture
Milind S1 day ago

Thank you!

Artur Kre's profile picture
Artur Kre1 day ago

wow insane, which local model does the window UI segmenting ?

Milind S's profile picture
Milind S1 day ago

Omniparser!

Deep's profile picture
Deep1 day ago

local segmentation plus OCR instead of a vision model. cuts what you send to the model, and it runs offline too.

Matt Ronge's profile picture
Matt Ronge1 day ago

Thats very slick. Does it use a larger model to know where to go? So if you ask it something complex to do in Blender for example it can break it down into steps?

Milind S's profile picture
Milind S1 day ago

There is literally no other larger models being used in this. Its as simple as it gets

Matt Ronge's profile picture
Matt Ronge1 day ago

Open source when lol

Milind S's profile picture
Milind S1 day ago

Haha let me clean it up and then I’ll do it promise :D Its based on my other open source project already called tiptour macos

Matt Ronge's profile picture
Matt Ronge1 day ago

Yea love it, I've been following your work for a while. Thanks for sharing it!

Varik Verilion's profile picture
Varik Verilion1 day ago

The interesting part is the interface boundary. Segmentation and OCR turn pixels into a compact, structured state before Jev plans an action. That makes local inference practical and keeps the sensitive screen data on-device.

Justin's profile picture
Justin1 day ago

Bro, you build way too fast :D

Milind S's profile picture
Milind S1 day ago

Thats how I roll :D

Bruno Skvorc's profile picture
Bruno Skvorc1 day ago

chat too

bikram's profile picture
bikram1 day ago

Yes. Share. Also, should go into OMB, will create a really good setup.

Milind S's profile picture
Milind S1 day ago

Great idea

B⭕bby's profile picture
B⭕bby1 day ago

Can you add voice to text on top so you can control it with your voice

Milind S's profile picture
Milind S1 day ago

Totally i can

Omkar Satpute's profile picture
Omkar Satpute1 day ago

SupaaaaFastttt computer use

ethereumdegen.eth 🕶️ᵍᵐ's profile picture
ethereumdegen.eth 🕶️ᵍᵐ1 day ago

I would love to use this and test it out

bikram's profile picture
bikram1 day ago

What are you using for on-device OCR

Milind S's profile picture
Milind S1 day ago

Apple ships it natively on all devices

Tacdel's profile picture
Tacdel1 day ago

90ms local loops make 5-second cloud VLM latency look unplayable. How does it handle icon-only buttons where OCR finds zero text?

Carlos Ziegler's profile picture
Carlos Ziegler1 day ago

This is really nice!!!!

Milind S's profile picture
Milind S1 day ago

I know right!!

Doomshade's profile picture
Doomshade1 day ago

Dude this is sick, would really love to try this out with some desktop applications!

Milind S's profile picture
Milind S1 day ago

Alright man. I’ll send it to you when its ready

Doomshade's profile picture
Doomshade1 day ago

🙏 much appreciated! btw - love openmaus - got it configured on my machine and made a few tweaks to my system - need to pull some of the latest features you've shipped !

Milind S's profile picture
Milind S1 day ago

Lfgg

Sahibzada Allahyar's profile picture
Sahibzada Allahyar1 day ago

this one can do it too and it's local

Joey van Koningsbruggen's profile picture
Joey van Koningsbruggen1 day ago

Very cool

up's profile picture
up1 day ago

Loved it, crazy capabilities

Milind S's profile picture
Milind S1 day ago

Indeed

shaik arbaz's profile picture
shaik arbaz1 day ago

Make it open source.

Milind S's profile picture
Milind S1 day ago

Will do

Arpit's profile picture
Arpit1 day ago

How can it be used to do computer use in bg?

Milind S's profile picture
Milind S1 day ago

I’m sure @trycua can

Hugo Catarino's profile picture
Hugo Catarino1 day ago

Que brutalidade!!! 🤯

Milind S's profile picture
Milind S1 day ago

Sí!!!

Christian Giangrande's profile picture
Christian Giangrande1 day ago

Impliment into maus 🙏🙏🙏

Milind S's profile picture
Milind S1 day ago

Hell yeah

Konrad's profile picture
Konrad1 day ago

how does the LM come up with more complex interactions, like in Blender, dragging while holding down buttons?

Milind S's profile picture
Milind S1 day ago

Gotta pair this with an llm to do that

Konrad's profile picture
Konrad1 day ago

Message me the details on WhatsApp, I have new investors for the stuff, doing a seed round the coming weeks. Getting the production cost from ~48ct to ~20ct could be realistic with this new tooling.

Mar's profile picture
Mar1 day ago

Dude, I need this!

Milind S's profile picture
Milind S1 day ago

Releasing soon

Nikhil Gangaraju's profile picture
Nikhil Gangaraju1 day ago

+1 for the open source version. Great demo!

Milind S's profile picture
Milind S1 day ago

Thanks! Repo coming shortly

Shahn's profile picture
Shahn1 day ago

this is really cool @milindlabs I’ve been trying out Jev for some browser use features as well, here you runs loop with ocr model right ?

Milind S's profile picture
Milind S1 day ago

I do yes along with the segmentation model

Guilherme's profile picture
Guilherme1 day ago

Does it support multi-screen ?

Milind S's profile picture
Milind S1 day ago

It does

D Row Kavi's profile picture
D Row Kavi1 day ago

@bot @poteto 😅🤘🏽

AskMeHow's profile picture
AskMeHow1 day ago

open source pls

Milind S's profile picture
Milind S1 day ago

If you say so

Aditya Sinha's profile picture
Aditya Sinha1 day ago

@iamgingertrash how did you know

Dominik's profile picture
Dominik1 day ago

a click loop without an llm can still own the machine. i am building refuse rules on the write surface so speed never becomes silent permission.

Related Videos

Our first short course with Anthropic! Building Towards Computer Use with Anthropic. This teaches you to build an LLM-based agent that uses a computer interface by generating mouse clicks and keystrokes. Computer Use is an important, emerging capability for LLMs that will let AI agents do many more tasks than were possible before, since it lets them interact with interfaces designed for humans to use, rather than only tools that provide explicit API access. I hope you will enjoy learning about it! This course is taught by Anthropic's Head of Curriculum, Colt_Steele. You'll learn to apply image reasoning and tool use to "use" a computer as follows: a model processes an image of the screen, analyzes it to understand what's going on, and navigates the computer via mouse clicks and keystrokes. This course goes through the key building blocks, and culminates in a demo of an AI assistant that uses a web browser to search for a research paper, downloads the PDF, and finally summarizes the paper for you. In detail, you’ll: - Learn about Anthropic's family of models, when to use which one, and make API requests to Claude - Use multi-modal prompts that combine text and image content blocks, and also work with streaming responses - Improve your prompting by using prompt templates, using XML to structure prompts, and providing examples - Implement prompt caching to reduce cost and latency - Apply tool-use to build a chatbot that can call different tools to respond to queries - See all these building blocks come together in Computer Use demo Please sign up here:

Andrew Ng

170,541 views • 1 year ago

David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” david friedberg: “I have had experiences where we've asked some fairly novel scientific questions, and (the AI model) identifies it as a novel insight. It's like, ‘Oh, never thought about that, interesting, blah, blah, blah.’ And then using a different account, asking the next version (of the model) later, I've now experienced this. It's like, ‘Oh, well, you could do this,’ and it actually just describes this exact thing that we had in our chat in the previous version. Now, these are a handful of anecdotal experiences, but I know the domain that we work in, and the niche of it, and the ideation of this stuff, and the novelty of this stuff, and the lack of papers being published, and so on. So I know that there isn't some new corpus of information out there that's training the new model. So all I can say at that point is that my conversation or our analyses have been used for training.” David Sacks: “Okay, this does raise a really good question. What does it mean that the model is allowed to train on unidentifiable data?” Friedberg: “Well, that's my point. So it doesn't use any of my personal information, but it can use an insight derived from our chat, which it can then say is some training data that is unrelated. But the truth is, it's actually a piece of IP that's our organization’s IP, and our engagement back and forth. We don't have any NDA or confidentiality provisions or protections with them being a service provider back to us. This is why I care a lot about open source because I don't want them having my chat logs because they can use it for training to create an IP advantage that is now diffused to the rest of the market.”

The All-In Podcast

54,050 views • 7 days ago