Загрузка видео...

Не удалось загрузить видео

На главную

Okay so Jev can actually do computer use really well Without any screenshots, or LLMs and no Pixels leave my mac I dont even read the Dom elements A local CoreML model segments every button and UI element on screen. On-device OCR reads the labels. That text is all...

116,719 просмотров • 1 день назад •via X (Twitter)

Комментарии: 62

Фото профиля Milind S
Milind S1 день назад

If this gets enough attention I’ll make it open source. Meanwhile check out my other stuff

Фото профиля Vedant Bhayani
Vedant Bhayani1 день назад

I think the future is multiple other companies building new kind of llm models for individual tasks with which you just plug it in with frontiers model for steering and thinking Also domain specific harness and custom trained models with your data for specific use cases are going be more relevant and we are going to see even more than before

Фото профиля Milind S
Milind S1 день назад

I agree that’s where things are headed

Фото профиля /h(elton)?Fabio\.(t|j)s/gi
/h(elton)?Fabio\.(t|j)s/gi1 день назад

hey man, do you mind doing a test with a local model that uses the same architecture? like this one based on qwen-2.5-1b?

Фото профиля Milind S
Milind S1 день назад

I’ll check it out

Фото профиля Valentin Pletzer
Valentin Pletzer1 день назад

very cool!

Фото профиля Milind S
Milind S1 день назад

Thank you!

Фото профиля Artur Kre
Artur Kre1 день назад

wow insane, which local model does the window UI segmenting ?

Фото профиля Milind S
Milind S1 день назад

Omniparser!

Фото профиля Deep
Deep1 день назад

local segmentation plus OCR instead of a vision model. cuts what you send to the model, and it runs offline too.

Фото профиля Matt Ronge
Matt Ronge1 день назад

Thats very slick. Does it use a larger model to know where to go? So if you ask it something complex to do in Blender for example it can break it down into steps?

Фото профиля Milind S
Milind S1 день назад

There is literally no other larger models being used in this. Its as simple as it gets

Фото профиля Matt Ronge
Matt Ronge1 день назад

Open source when lol

Фото профиля Milind S
Milind S1 день назад

Haha let me clean it up and then I’ll do it promise :D Its based on my other open source project already called tiptour macos

Фото профиля Matt Ronge
Matt Ronge1 день назад

Yea love it, I've been following your work for a while. Thanks for sharing it!

Фото профиля Varik Verilion
Varik Verilion1 день назад

The interesting part is the interface boundary. Segmentation and OCR turn pixels into a compact, structured state before Jev plans an action. That makes local inference practical and keeps the sensitive screen data on-device.

Фото профиля Justin
Justin1 день назад

Bro, you build way too fast :D

Фото профиля Milind S
Milind S1 день назад

Thats how I roll :D

Фото профиля Bruno Skvorc
Bruno Skvorc1 день назад

chat too

Фото профиля bikram
bikram1 день назад

Yes. Share. Also, should go into OMB, will create a really good setup.

Фото профиля Milind S
Milind S1 день назад

Great idea

Фото профиля B⭕bby
B⭕bby1 день назад

Can you add voice to text on top so you can control it with your voice

Фото профиля Milind S
Milind S1 день назад

Totally i can

Фото профиля Omkar Satpute
Omkar Satpute1 день назад

SupaaaaFastttt computer use

Фото профиля ethereumdegen.eth 🕶️ᵍᵐ
ethereumdegen.eth 🕶️ᵍᵐ1 день назад

I would love to use this and test it out

Фото профиля bikram
bikram1 день назад

What are you using for on-device OCR

Фото профиля Milind S
Milind S1 день назад

Apple ships it natively on all devices

Фото профиля Tacdel
Tacdel1 день назад

90ms local loops make 5-second cloud VLM latency look unplayable. How does it handle icon-only buttons where OCR finds zero text?

Фото профиля Carlos Ziegler
Carlos Ziegler1 день назад

This is really nice!!!!

Фото профиля Milind S
Milind S1 день назад

I know right!!

Фото профиля Doomshade
Doomshade1 день назад

Dude this is sick, would really love to try this out with some desktop applications!

Фото профиля Milind S
Milind S1 день назад

Alright man. I’ll send it to you when its ready

Фото профиля Doomshade
Doomshade1 день назад

🙏 much appreciated! btw - love openmaus - got it configured on my machine and made a few tweaks to my system - need to pull some of the latest features you've shipped !

Фото профиля Milind S
Milind S1 день назад

Lfgg

Фото профиля Sahibzada Allahyar
Sahibzada Allahyar1 день назад

this one can do it too and it's local

Фото профиля Joey van Koningsbruggen
Joey van Koningsbruggen1 день назад

Very cool

Фото профиля up
up1 день назад

Loved it, crazy capabilities

Фото профиля Milind S
Milind S1 день назад

Indeed

Фото профиля shaik arbaz
shaik arbaz1 день назад

Make it open source.

Фото профиля Milind S
Milind S1 день назад

Will do

Фото профиля Arpit
Arpit1 день назад

How can it be used to do computer use in bg?

Фото профиля Milind S
Milind S1 день назад

I’m sure @trycua can

Фото профиля Hugo Catarino
Hugo Catarino1 день назад

Que brutalidade!!! 🤯

Фото профиля Milind S
Milind S1 день назад

Sí!!!

Фото профиля Christian Giangrande
Christian Giangrande1 день назад

Impliment into maus 🙏🙏🙏

Фото профиля Milind S
Milind S1 день назад

Hell yeah

Фото профиля Konrad
Konrad1 день назад

how does the LM come up with more complex interactions, like in Blender, dragging while holding down buttons?

Фото профиля Milind S
Milind S1 день назад

Gotta pair this with an llm to do that

Фото профиля Konrad
Konrad1 день назад

Message me the details on WhatsApp, I have new investors for the stuff, doing a seed round the coming weeks. Getting the production cost from ~48ct to ~20ct could be realistic with this new tooling.

Фото профиля Mar
Mar1 день назад

Dude, I need this!

Фото профиля Milind S
Milind S1 день назад

Releasing soon

Фото профиля Nikhil Gangaraju
Nikhil Gangaraju1 день назад

+1 for the open source version. Great demo!

Фото профиля Milind S
Milind S1 день назад

Thanks! Repo coming shortly

Фото профиля Shahn
Shahn1 день назад

this is really cool @milindlabs I’ve been trying out Jev for some browser use features as well, here you runs loop with ocr model right ?

Фото профиля Milind S
Milind S1 день назад

I do yes along with the segmentation model

Фото профиля Guilherme
Guilherme1 день назад

Does it support multi-screen ?

Фото профиля Milind S
Milind S1 день назад

It does

Фото профиля D Row Kavi
D Row Kavi1 день назад

@bot @poteto 😅🤘🏽

Фото профиля AskMeHow
AskMeHow1 день назад

open source pls

Фото профиля Milind S
Milind S1 день назад

If you say so

Фото профиля Aditya Sinha
Aditya Sinha1 день назад

@iamgingertrash how did you know

Фото профиля Dominik
Dominik1 день назад

a click loop without an llm can still own the machine. i am building refuse rules on the write surface so speed never becomes silent permission.

Похожие видео

Our first short course with Anthropic! Building Towards Computer Use with Anthropic. This teaches you to build an LLM-based agent that uses a computer interface by generating mouse clicks and keystrokes. Computer Use is an important, emerging capability for LLMs that will let AI agents do many more tasks than were possible before, since it lets them interact with interfaces designed for humans to use, rather than only tools that provide explicit API access. I hope you will enjoy learning about it! This course is taught by Anthropic's Head of Curriculum, Colt_Steele. You'll learn to apply image reasoning and tool use to "use" a computer as follows: a model processes an image of the screen, analyzes it to understand what's going on, and navigates the computer via mouse clicks and keystrokes. This course goes through the key building blocks, and culminates in a demo of an AI assistant that uses a web browser to search for a research paper, downloads the PDF, and finally summarizes the paper for you. In detail, you’ll: - Learn about Anthropic's family of models, when to use which one, and make API requests to Claude - Use multi-modal prompts that combine text and image content blocks, and also work with streaming responses - Improve your prompting by using prompt templates, using XML to structure prompts, and providing examples - Implement prompt caching to reduce cost and latency - Apply tool-use to build a chatbot that can call different tools to respond to queries - See all these building blocks come together in Computer Use demo Please sign up here:

Andrew Ng

170,541 просмотров • 1 год назад

David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” david friedberg: “I have had experiences where we've asked some fairly novel scientific questions, and (the AI model) identifies it as a novel insight. It's like, ‘Oh, never thought about that, interesting, blah, blah, blah.’ And then using a different account, asking the next version (of the model) later, I've now experienced this. It's like, ‘Oh, well, you could do this,’ and it actually just describes this exact thing that we had in our chat in the previous version. Now, these are a handful of anecdotal experiences, but I know the domain that we work in, and the niche of it, and the ideation of this stuff, and the novelty of this stuff, and the lack of papers being published, and so on. So I know that there isn't some new corpus of information out there that's training the new model. So all I can say at that point is that my conversation or our analyses have been used for training.” David Sacks: “Okay, this does raise a really good question. What does it mean that the model is allowed to train on unidentifiable data?” Friedberg: “Well, that's my point. So it doesn't use any of my personal information, but it can use an insight derived from our chat, which it can then say is some training data that is unrelated. But the truth is, it's actually a piece of IP that's our organization’s IP, and our engagement back and forth. We don't have any NDA or confidentiality provisions or protections with them being a service provider back to us. This is why I care a lot about open source because I don't want them having my chat logs because they can use it for training to create an IP advantage that is now diffused to the rest of the market.”

The All-In Podcast

54,050 просмотров • 7 дней назад