ๆญฃๅจๅ ่ฝฝ่ง้ข...
่ง้ขๅ ่ฝฝๅคฑ่ดฅ
๐ข๐๐ฟ ๐๐ฒ๐ฎ๐บ ๐ฑ๐ถ๐๐ฐ๐ผ๐๐ฒ๐ฟ๐ฒ๐ฑ ๐๐ผ๐ ๐ฐ๐ฎ๐ป ๐๐๐ฒ ๐๐ฃ๐ง-๐ฐ๐ฉ๐ถ๐๐ถ๐ผ๐ป ๐๐ผ ๐ฐ๐ฟ๐ฒ๐ฎ๐๐ฒ ๐ฎ ๐๐ฒ๐น๐ณ-๐ผ๐ฝ๐ฒ๐ฟ๐ฎ๐๐ถ๐ป๐ด ๐ฐ๐ผ๐บ๐ฝ๐๐๐ฒ๐ฟ. By looking at the user interface, GPT-4 decides which series of click or type events are required to accomplish an objective. Here it is, writing a poem in Apple Notes.
470,500 ๆฌก่ง็ โข 2 ๅนดๅ โขvia X (Twitter)
9 ๆก่ฏ่ฎบ

This is insane! We hit our rate limit for 4V so fast :\ can't wait until the prod ready model is released

So for employers of remote workers, they can just fine tune on screen monitoring data they already collect and totally replace the worker without having to alter any existing business processes?

What happens when a multimodal model, by taking in a screenshot and user objective, predicts the next click (with X Y location) or typing event more accurately than a human? I am incredibly interested in this problem. We are working on it at HyperWrite.

how do you automate clicking?

GPT-4V decides on a window to click based on the objective and estimates x & y location in % which can be evaluated to pixels in Python. It does.. ok at this estimation.

This is SO cool! I need it

I need this in HyperWrite

RIP RPA?

@AdeptAILabs will probably release this on steroids
