正在加载视频...

视频加载失败

Okay so Jev can actually do computer use really well Without any screenshots, or LLMs and no Pixels leave my mac I dont even read the Dom elements A local CoreML model segments every button and UI element on screen. On-device OCR reads the labels. That text is all...

116,719 次观看 • 1 天前 •via X (Twitter)

62 条评论

Milind S 的头像
Milind S1 天前

If this gets enough attention I’ll make it open source. Meanwhile check out my other stuff

Vedant Bhayani 的头像
Vedant Bhayani1 天前

I think the future is multiple other companies building new kind of llm models for individual tasks with which you just plug it in with frontiers model for steering and thinking Also domain specific harness and custom trained models with your data for specific use cases are going be more relevant and we are going to see even more than before

Milind S 的头像
Milind S1 天前

I agree that’s where things are headed

/h(elton)?Fabio\.(t|j)s/gi 的头像
/h(elton)?Fabio\.(t|j)s/gi1 天前

hey man, do you mind doing a test with a local model that uses the same architecture? like this one based on qwen-2.5-1b?

Milind S 的头像
Milind S1 天前

I’ll check it out

Valentin Pletzer 的头像
Valentin Pletzer1 天前

very cool!

Milind S 的头像
Milind S1 天前

Thank you!

Artur Kre 的头像
Artur Kre1 天前

wow insane, which local model does the window UI segmenting ?

Milind S 的头像
Milind S1 天前

Omniparser!

Deep 的头像
Deep1 天前

local segmentation plus OCR instead of a vision model. cuts what you send to the model, and it runs offline too.

Matt Ronge 的头像
Matt Ronge1 天前

Thats very slick. Does it use a larger model to know where to go? So if you ask it something complex to do in Blender for example it can break it down into steps?

Milind S 的头像
Milind S1 天前

There is literally no other larger models being used in this. Its as simple as it gets

Matt Ronge 的头像
Matt Ronge1 天前

Open source when lol

Milind S 的头像
Milind S1 天前

Haha let me clean it up and then I’ll do it promise :D Its based on my other open source project already called tiptour macos

Matt Ronge 的头像
Matt Ronge1 天前

Yea love it, I've been following your work for a while. Thanks for sharing it!

Varik Verilion 的头像
Varik Verilion1 天前

The interesting part is the interface boundary. Segmentation and OCR turn pixels into a compact, structured state before Jev plans an action. That makes local inference practical and keeps the sensitive screen data on-device.

Justin 的头像
Justin1 天前

Bro, you build way too fast :D

Milind S 的头像
Milind S1 天前

Thats how I roll :D

Bruno Skvorc 的头像
Bruno Skvorc1 天前

chat too

bikram 的头像
bikram1 天前

Yes. Share. Also, should go into OMB, will create a really good setup.

Milind S 的头像
Milind S1 天前

Great idea

B⭕bby 的头像
B⭕bby1 天前

Can you add voice to text on top so you can control it with your voice

Milind S 的头像
Milind S1 天前

Totally i can

Omkar Satpute 的头像
Omkar Satpute1 天前

SupaaaaFastttt computer use

ethereumdegen.eth 🕶️ᵍᵐ 的头像
ethereumdegen.eth 🕶️ᵍᵐ1 天前

I would love to use this and test it out

bikram 的头像
bikram1 天前

What are you using for on-device OCR

Milind S 的头像
Milind S1 天前

Apple ships it natively on all devices

Tacdel 的头像
Tacdel1 天前

90ms local loops make 5-second cloud VLM latency look unplayable. How does it handle icon-only buttons where OCR finds zero text?

Carlos Ziegler 的头像
Carlos Ziegler1 天前

This is really nice!!!!

Milind S 的头像
Milind S1 天前

I know right!!

Doomshade 的头像
Doomshade1 天前

Dude this is sick, would really love to try this out with some desktop applications!

Milind S 的头像
Milind S1 天前

Alright man. I’ll send it to you when its ready

Doomshade 的头像
Doomshade1 天前

🙏 much appreciated! btw - love openmaus - got it configured on my machine and made a few tweaks to my system - need to pull some of the latest features you've shipped !

Milind S 的头像
Milind S1 天前

Lfgg

Sahibzada Allahyar 的头像
Sahibzada Allahyar1 天前

this one can do it too and it's local

Joey van Koningsbruggen 的头像
Joey van Koningsbruggen1 天前

Very cool

up 的头像
up1 天前

Loved it, crazy capabilities

Milind S 的头像
Milind S1 天前

Indeed

shaik arbaz 的头像
shaik arbaz1 天前

Make it open source.

Milind S 的头像
Milind S1 天前

Will do

Arpit 的头像
Arpit1 天前

How can it be used to do computer use in bg?

Milind S 的头像
Milind S1 天前

I’m sure @trycua can

Hugo Catarino 的头像
Hugo Catarino1 天前

Que brutalidade!!! 🤯

Milind S 的头像
Milind S1 天前

Sí!!!

Christian Giangrande 的头像
Christian Giangrande1 天前

Impliment into maus 🙏🙏🙏

Milind S 的头像
Milind S1 天前

Hell yeah

Konrad 的头像
Konrad1 天前

how does the LM come up with more complex interactions, like in Blender, dragging while holding down buttons?

Milind S 的头像
Milind S1 天前

Gotta pair this with an llm to do that

Konrad 的头像
Konrad1 天前

Message me the details on WhatsApp, I have new investors for the stuff, doing a seed round the coming weeks. Getting the production cost from ~48ct to ~20ct could be realistic with this new tooling.

Mar 的头像
Mar1 天前

Dude, I need this!

Milind S 的头像
Milind S1 天前

Releasing soon

Nikhil Gangaraju 的头像
Nikhil Gangaraju1 天前

+1 for the open source version. Great demo!

Milind S 的头像
Milind S1 天前

Thanks! Repo coming shortly

Shahn 的头像
Shahn1 天前

this is really cool @milindlabs I’ve been trying out Jev for some browser use features as well, here you runs loop with ocr model right ?

Milind S 的头像
Milind S1 天前

I do yes along with the segmentation model

Guilherme 的头像
Guilherme1 天前

Does it support multi-screen ?

Milind S 的头像
Milind S1 天前

It does

D Row Kavi 的头像
D Row Kavi1 天前

@bot @poteto 😅🤘🏽

AskMeHow 的头像
AskMeHow1 天前

open source pls

Milind S 的头像
Milind S1 天前

If you say so

Aditya Sinha 的头像
Aditya Sinha1 天前

@iamgingertrash how did you know

Dominik 的头像
Dominik1 天前

a click loop without an llm can still own the machine. i am building refuse rules on the write surface so speed never becomes silent permission.

相关视频

Our first short course with Anthropic! Building Towards Computer Use with Anthropic. This teaches you to build an LLM-based agent that uses a computer interface by generating mouse clicks and keystrokes. Computer Use is an important, emerging capability for LLMs that will let AI agents do many more tasks than were possible before, since it lets them interact with interfaces designed for humans to use, rather than only tools that provide explicit API access. I hope you will enjoy learning about it! This course is taught by Anthropic's Head of Curriculum, Colt_Steele. You'll learn to apply image reasoning and tool use to "use" a computer as follows: a model processes an image of the screen, analyzes it to understand what's going on, and navigates the computer via mouse clicks and keystrokes. This course goes through the key building blocks, and culminates in a demo of an AI assistant that uses a web browser to search for a research paper, downloads the PDF, and finally summarizes the paper for you. In detail, you’ll: - Learn about Anthropic's family of models, when to use which one, and make API requests to Claude - Use multi-modal prompts that combine text and image content blocks, and also work with streaming responses - Improve your prompting by using prompt templates, using XML to structure prompts, and providing examples - Implement prompt caching to reduce cost and latency - Apply tool-use to build a chatbot that can call different tools to respond to queries - See all these building blocks come together in Computer Use demo Please sign up here:

Andrew Ng

170,541 次观看 • 1 年前

David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” david friedberg: “I have had experiences where we've asked some fairly novel scientific questions, and (the AI model) identifies it as a novel insight. It's like, ‘Oh, never thought about that, interesting, blah, blah, blah.’ And then using a different account, asking the next version (of the model) later, I've now experienced this. It's like, ‘Oh, well, you could do this,’ and it actually just describes this exact thing that we had in our chat in the previous version. Now, these are a handful of anecdotal experiences, but I know the domain that we work in, and the niche of it, and the ideation of this stuff, and the novelty of this stuff, and the lack of papers being published, and so on. So I know that there isn't some new corpus of information out there that's training the new model. So all I can say at that point is that my conversation or our analyses have been used for training.” David Sacks: “Okay, this does raise a really good question. What does it mean that the model is allowed to train on unidentifiable data?” Friedberg: “Well, that's my point. So it doesn't use any of my personal information, but it can use an insight derived from our chat, which it can then say is some training data that is unrelated. But the truth is, it's actually a piece of IP that's our organization’s IP, and our engagement back and forth. We don't have any NDA or confidentiality provisions or protections with them being a service provider back to us. This is why I care a lot about open source because I don't want them having my chat logs because they can use it for training to create an IP advantage that is now diffused to the rest of the market.”

The All-In Podcast

54,050 次观看 • 7 天前