Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Okay so Jev can actually do computer use really well Without any screenshots, or LLMs and no Pixels leave my mac I dont even read the Dom elements A local CoreML model segments every button and UI element on screen. On-device OCR reads the labels. That text is all...

116,719 Aufrufe • vor 1 Tag •via X (Twitter)

62 Kommentare

Profilbild von Milind S
Milind Svor 1 Tag

If this gets enough attention I’ll make it open source. Meanwhile check out my other stuff

Profilbild von Vedant Bhayani
Vedant Bhayanivor 1 Tag

I think the future is multiple other companies building new kind of llm models for individual tasks with which you just plug it in with frontiers model for steering and thinking Also domain specific harness and custom trained models with your data for specific use cases are going be more relevant and we are going to see even more than before

Profilbild von Milind S
Milind Svor 1 Tag

I agree that’s where things are headed

Profilbild von /h(elton)?Fabio\.(t|j)s/gi
/h(elton)?Fabio\.(t|j)s/givor 1 Tag

hey man, do you mind doing a test with a local model that uses the same architecture? like this one based on qwen-2.5-1b?

Profilbild von Milind S
Milind Svor 1 Tag

I’ll check it out

Profilbild von Valentin Pletzer
Valentin Pletzervor 1 Tag

very cool!

Profilbild von Milind S
Milind Svor 1 Tag

Thank you!

Profilbild von Artur Kre
Artur Krevor 1 Tag

wow insane, which local model does the window UI segmenting ?

Profilbild von Milind S
Milind Svor 1 Tag

Omniparser!

Profilbild von Deep
Deepvor 1 Tag

local segmentation plus OCR instead of a vision model. cuts what you send to the model, and it runs offline too.

Profilbild von Matt Ronge
Matt Rongevor 1 Tag

Thats very slick. Does it use a larger model to know where to go? So if you ask it something complex to do in Blender for example it can break it down into steps?

Profilbild von Milind S
Milind Svor 1 Tag

There is literally no other larger models being used in this. Its as simple as it gets

Profilbild von Matt Ronge
Matt Rongevor 1 Tag

Open source when lol

Profilbild von Milind S
Milind Svor 1 Tag

Haha let me clean it up and then I’ll do it promise :D Its based on my other open source project already called tiptour macos

Profilbild von Matt Ronge
Matt Rongevor 1 Tag

Yea love it, I've been following your work for a while. Thanks for sharing it!

Profilbild von Varik Verilion
Varik Verilionvor 1 Tag

The interesting part is the interface boundary. Segmentation and OCR turn pixels into a compact, structured state before Jev plans an action. That makes local inference practical and keeps the sensitive screen data on-device.

Profilbild von Justin
Justinvor 1 Tag

Bro, you build way too fast :D

Profilbild von Milind S
Milind Svor 1 Tag

Thats how I roll :D

Profilbild von Bruno Skvorc
Bruno Skvorcvor 1 Tag

chat too

Profilbild von bikram
bikramvor 1 Tag

Yes. Share. Also, should go into OMB, will create a really good setup.

Profilbild von Milind S
Milind Svor 1 Tag

Great idea

Profilbild von B⭕bby
B⭕bbyvor 1 Tag

Can you add voice to text on top so you can control it with your voice

Profilbild von Milind S
Milind Svor 1 Tag

Totally i can

Profilbild von Omkar Satpute
Omkar Satputevor 1 Tag

SupaaaaFastttt computer use

Profilbild von ethereumdegen.eth 🕶️ᵍᵐ
ethereumdegen.eth 🕶️ᵍᵐvor 1 Tag

I would love to use this and test it out

Profilbild von bikram
bikramvor 1 Tag

What are you using for on-device OCR

Profilbild von Milind S
Milind Svor 1 Tag

Apple ships it natively on all devices

Profilbild von Tacdel
Tacdelvor 1 Tag

90ms local loops make 5-second cloud VLM latency look unplayable. How does it handle icon-only buttons where OCR finds zero text?

Profilbild von Carlos Ziegler
Carlos Zieglervor 1 Tag

This is really nice!!!!

Profilbild von Milind S
Milind Svor 1 Tag

I know right!!

Profilbild von Doomshade
Doomshadevor 1 Tag

Dude this is sick, would really love to try this out with some desktop applications!

Profilbild von Milind S
Milind Svor 1 Tag

Alright man. I’ll send it to you when its ready

Profilbild von Doomshade
Doomshadevor 1 Tag

🙏 much appreciated! btw - love openmaus - got it configured on my machine and made a few tweaks to my system - need to pull some of the latest features you've shipped !

Profilbild von Milind S
Milind Svor 1 Tag

Lfgg

Profilbild von Sahibzada Allahyar
Sahibzada Allahyarvor 1 Tag

this one can do it too and it's local

Profilbild von Joey van Koningsbruggen
Joey van Koningsbruggenvor 1 Tag

Very cool

Profilbild von up
upvor 1 Tag

Loved it, crazy capabilities

Profilbild von Milind S
Milind Svor 1 Tag

Indeed

Profilbild von shaik arbaz
shaik arbazvor 1 Tag

Make it open source.

Profilbild von Milind S
Milind Svor 1 Tag

Will do

Profilbild von Arpit
Arpitvor 1 Tag

How can it be used to do computer use in bg?

Profilbild von Milind S
Milind Svor 1 Tag

I’m sure @trycua can

Profilbild von Hugo Catarino
Hugo Catarinovor 1 Tag

Que brutalidade!!! 🤯

Profilbild von Milind S
Milind Svor 1 Tag

Sí!!!

Profilbild von Christian Giangrande
Christian Giangrandevor 1 Tag

Impliment into maus 🙏🙏🙏

Profilbild von Milind S
Milind Svor 1 Tag

Hell yeah

Profilbild von Konrad
Konradvor 1 Tag

how does the LM come up with more complex interactions, like in Blender, dragging while holding down buttons?

Profilbild von Milind S
Milind Svor 1 Tag

Gotta pair this with an llm to do that

Profilbild von Konrad
Konradvor 1 Tag

Message me the details on WhatsApp, I have new investors for the stuff, doing a seed round the coming weeks. Getting the production cost from ~48ct to ~20ct could be realistic with this new tooling.

Profilbild von Mar
Marvor 1 Tag

Dude, I need this!

Profilbild von Milind S
Milind Svor 1 Tag

Releasing soon

Profilbild von Nikhil Gangaraju
Nikhil Gangarajuvor 1 Tag

+1 for the open source version. Great demo!

Profilbild von Milind S
Milind Svor 1 Tag

Thanks! Repo coming shortly

Profilbild von Shahn
Shahnvor 1 Tag

this is really cool @milindlabs I’ve been trying out Jev for some browser use features as well, here you runs loop with ocr model right ?

Profilbild von Milind S
Milind Svor 1 Tag

I do yes along with the segmentation model

Profilbild von Guilherme
Guilhermevor 1 Tag

Does it support multi-screen ?

Profilbild von Milind S
Milind Svor 1 Tag

It does

Profilbild von D Row Kavi
D Row Kavivor 1 Tag

@bot @poteto 😅🤘🏽

Profilbild von AskMeHow
AskMeHowvor 1 Tag

open source pls

Profilbild von Milind S
Milind Svor 1 Tag

If you say so

Profilbild von Aditya Sinha
Aditya Sinhavor 1 Tag

@iamgingertrash how did you know

Profilbild von Dominik
Dominikvor 1 Tag

a click loop without an llm can still own the machine. i am building refuse rules on the write surface so speed never becomes silent permission.

Ähnliche Videos

Our first short course with Anthropic! Building Towards Computer Use with Anthropic. This teaches you to build an LLM-based agent that uses a computer interface by generating mouse clicks and keystrokes. Computer Use is an important, emerging capability for LLMs that will let AI agents do many more tasks than were possible before, since it lets them interact with interfaces designed for humans to use, rather than only tools that provide explicit API access. I hope you will enjoy learning about it! This course is taught by Anthropic's Head of Curriculum, Colt_Steele. You'll learn to apply image reasoning and tool use to "use" a computer as follows: a model processes an image of the screen, analyzes it to understand what's going on, and navigates the computer via mouse clicks and keystrokes. This course goes through the key building blocks, and culminates in a demo of an AI assistant that uses a web browser to search for a research paper, downloads the PDF, and finally summarizes the paper for you. In detail, you’ll: - Learn about Anthropic's family of models, when to use which one, and make API requests to Claude - Use multi-modal prompts that combine text and image content blocks, and also work with streaming responses - Improve your prompting by using prompt templates, using XML to structure prompts, and providing examples - Implement prompt caching to reduce cost and latency - Apply tool-use to build a chatbot that can call different tools to respond to queries - See all these building blocks come together in Computer Use demo Please sign up here:

Andrew Ng

170,541 Aufrufe • vor 1 Jahr

David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” david friedberg: “I have had experiences where we've asked some fairly novel scientific questions, and (the AI model) identifies it as a novel insight. It's like, ‘Oh, never thought about that, interesting, blah, blah, blah.’ And then using a different account, asking the next version (of the model) later, I've now experienced this. It's like, ‘Oh, well, you could do this,’ and it actually just describes this exact thing that we had in our chat in the previous version. Now, these are a handful of anecdotal experiences, but I know the domain that we work in, and the niche of it, and the ideation of this stuff, and the novelty of this stuff, and the lack of papers being published, and so on. So I know that there isn't some new corpus of information out there that's training the new model. So all I can say at that point is that my conversation or our analyses have been used for training.” David Sacks: “Okay, this does raise a really good question. What does it mean that the model is allowed to train on unidentifiable data?” Friedberg: “Well, that's my point. So it doesn't use any of my personal information, but it can use an insight derived from our chat, which it can then say is some training data that is unrelated. But the truth is, it's actually a piece of IP that's our organization’s IP, and our engagement back and forth. We don't have any NDA or confidentiality provisions or protections with them being a service provider back to us. This is why I care a lot about open source because I don't want them having my chat logs because they can use it for training to create an IP advantage that is now diffused to the rest of the market.”

The All-In Podcast

54,050 Aufrufe • vor 7 Tagen