Loading video...

Video Failed to Load

Go Home

We’re reimagining a 50-year-old interface - the mouse pointer - with AI. 🖱️ These experimental demos show how people can intuitively direct Gemini on their screens using motion, speech, and natural shorthand to get things done 🧵

1,684,389 views • 4 months ago •via X (Twitter)

38 Comments

Google DeepMind's profile picture
Google DeepMind4 months ago

With an AI-enabled pointer, help is always available where you’re working - without having to detour to additional apps. 📲 Point at a PDF and request bullet points for an email, hover over a table to ask for a pie chart, or highlight a recipe and simply say: "double these ingredients."

Google DeepMind's profile picture
Google DeepMind4 months ago

Current models require precise instructions, but our AI-enabled pointer removes that burden. 💡 By "seeing" what’s under your cursor, it instantly understands the specific word, image, or code block you need help with.

Google DeepMind's profile picture
Google DeepMind4 months ago

In the real world, we don't tend to speak in long paragraphs; we point and say: "fix this" or "move that". 💬 By combining gestures with speech, it lets you use natural shorthand to complete tasks.

Google DeepMind's profile picture
Google DeepMind4 months ago

For decades, your mouse only tracked where you were pointing. AI helps it understand what you're pointing at. 💭 This means a photo of a scribbled note could turn into an interactive to-do list, or a paused video frame can become a restaurant booking link.

Google DeepMind's profile picture
Google DeepMind4 months ago

These capabilities are guiding how we think about the next generation of interfaces. As we continue exploring what an AI-enabled mouse pointer would unlock, try our experiments in @GoogleAIStudio →

toi500's profile picture
toi5004 months ago

Why do you guys always use examples that target the general public but are genuinely useless in practice? Its always the same thing, shopping lists, calendar, travel plans... it is getting boring, man

Danny Limanseta's profile picture
Danny Limanseta4 months ago

Nice to see @FarzaTV’s clicky UX adopted by Google Gemini!

Ajeya's profile picture
Ajeya4 months ago

you should make it like clippy

kitze 🛠️ tinkerer.club's profile picture
kitze 🛠️ tinkerer.club4 months ago

this is farzatastic!

variousred's profile picture
variousred4 months ago

ok but what if there wasnt

leon's profile picture
leon4 months ago

the mouse feels like the wrong constraint here if ai is actually a new paradigm, build the hardware around that instead of forcing it through the old one

Eshan's profile picture
Eshan4 months ago

the cursor knowing what it's pointing at instead of just where it's pointing is genuinely the largest input paradigm shift since touch screens. every interaction model since 1968 has treated the pointer as a coordinate pair. x,y on a grid. the AI pointer treats it as a semantic reference. "this thing" instead of "pixel 847, 312." that's also why the voice plus gesture demo matters more than the visual demos. saying "fix this" while pointing at a code block transmits the same information as a 200 word prompt in about 0.8 seconds. the bandwidth of pointing plus two words exceeds the bandwidth of typing a paragraph. the mouse just became a pronoun

Francesco's profile picture
Francesco4 months ago

it's called @heyclicky

Mohit Mor's profile picture
Mohit Mor4 months ago

All this is super fun but costs a lot of tokens, if you ask AI about cool ideas it can give a lot of them but what about the cost? For every small thing if we start using llm tokens its going to cost a lot at current stage i believe

Noctre's profile picture
Noctre4 months ago

Sorry, but all of the examples you provided would've been done faster by just doing them manually, voice is an incredibly slow input method and prone to mistakes, not to mention that this makes each action a lot more expensive than if it was done with just regular input

Yann Kronberg's profile picture
Yann Kronberg4 months ago

Combining gesture, speech and visual context into one input signal is what @karpathy was pointing at last week. This is the first real hardware implementation of that thesis. How far does it go before it replaces the keyboard entirely?

Connor's profile picture
Connor4 months ago

this feels closer to the right ui for agents than pure voice, which is still a bit annoying. pointing is basically human attention made visible, which gives the model a much better prior over what the next action is supposed to touch

Yash Bhalgat's profile picture
Yash Bhalgat4 months ago

I wonder what happens to @heyclicky now 🫠

Nick Schmidt's profile picture
Nick Schmidt4 months ago

@tszzl Is this not @heyclicky

Yahia Bakour's profile picture
Yahia Bakour4 months ago

Clicky has been clicked @FarzaTV

ib's profile picture
ib4 months ago

Why is this better than Clicky?

Arai Chartz👽's profile picture
Arai Chartz👽4 months ago

Woah, with all this technology we're seeing nowadays, we're starting to look like what the future looks like in 1900s movies... Love this tech, hope it gets available for global purchase soon...

Murat's profile picture
Murat4 months ago

#keep4o #keepsonnet45 @DarioAmodei @sama you cant reach AGI if you keep removing LLM models

Abin Kumar's profile picture
Abin Kumar4 months ago

@mathemagic1an @heyclicky that you ?

Michał Piszczek's profile picture
Michał Piszczek4 months ago

The mouse pointer was the cheapest explicit-intent device ever shipped. Replacing it shifts inference cost to the model. Ambiguity becomes a runtime tax: cheaper than learning a UI, more expensive every session.

Maaz's profile picture
Maaz4 months ago

@ByteMohit see this

Darshan Savaliya (Aspora)'s profile picture
Darshan Savaliya (Aspora)4 months ago

Now everything is about how my model can grab right context without actually me describing it.

Mario Simic's profile picture
Mario Simic4 months ago

Open-source alternative already exists: AIPointer ⦿ - same point-and-ask concept, multi-provider, MIT, voice-enabled, ships today.

Ali Mehdi Mukadam's profile picture
Ali Mehdi Mukadam4 months ago

one step closer

Manoj's profile picture
Manoj4 months ago

Is it just me, or did anyone else instantly think of @heyclicky by @FarzaTV the second they saw this experiment?

Parsec 🪐's profile picture
Parsec 🪐4 months ago

Turning the humble mouse pointer unchanged for 50 years into a semantic pointer that actually understands what you're pointing at is a game changer. No more copy pasting or crafting perfect prompts; just point, gesture, speak naturally ("double this," "summarize for email," "fix that code block") and get things done in context. The demos feel like sci-fi becoming reality scribbled notes → actionable to dos, paused video → instant booking link. This is the kind of intuitive human AI collaboration we've been waiting for. Can't wait to play with the experiments in Google AI Studio. Huge step toward interfaces that finally meet us where we are! What do you all think ready to retire precise prompting? 🚀😊

Cybernorse's profile picture
Cybernorse4 months ago

AI-driven input surfaces new attack vectors motion, speech, shorthand injection. Hope security architecture was reimagined alongside the UX. Zero-trust for human-device interaction starts here. #Cybernorse

Stefan's profile picture
Stefan4 months ago

this is a step in the right direction trying things like this where the default for the past 50 years has been the same and untouched experiments like this lead to interesting discoveries

Muad'Deep - e/acc's profile picture
Muad'Deep - e/acc4 months ago

@tszzl Called it

Ngoc Anh Tran's profile picture
Ngoc Anh Tran4 months ago

This is kinda....useless.

The Calda's profile picture
The Calda4 months ago

Pointing at something and thinking "fix this" or "take notes on that" is how we already thought when using a computer... now the middle step of actually doing it is being removed

Tom Curonian's profile picture
Tom Curonian4 months ago

actually impressed. google made the cursor the scope signal for screen-aware agents and that's the move that matters. they otherwise burn tokens guessing what part of the screen you mean. the pointer answers it for free. Kudos @GoogleDeepMind

CityLad's profile picture
CityLad4 months ago

@FarzaTV is this not copycat of @heyclicky

Related Videos