Video wird geladen...
Video konnte nicht geladen werden
We’re reimagining a 50-year-old interface - the mouse pointer - with AI. 🖱️ These experimental demos show how people can intuitively direct Gemini on their screens using motion, speech, and natural shorthand to get things done 🧵
1,684,389 Aufrufe • vor 4 Monaten •via X (Twitter)
38 Kommentare

With an AI-enabled pointer, help is always available where you’re working - without having to detour to additional apps. 📲 Point at a PDF and request bullet points for an email, hover over a table to ask for a pie chart, or highlight a recipe and simply say: "double these ingredients."

Current models require precise instructions, but our AI-enabled pointer removes that burden. 💡 By "seeing" what’s under your cursor, it instantly understands the specific word, image, or code block you need help with.

In the real world, we don't tend to speak in long paragraphs; we point and say: "fix this" or "move that". 💬 By combining gestures with speech, it lets you use natural shorthand to complete tasks.

For decades, your mouse only tracked where you were pointing. AI helps it understand what you're pointing at. 💭 This means a photo of a scribbled note could turn into an interactive to-do list, or a paused video frame can become a restaurant booking link.

These capabilities are guiding how we think about the next generation of interfaces. As we continue exploring what an AI-enabled mouse pointer would unlock, try our experiments in @GoogleAIStudio →

Why do you guys always use examples that target the general public but are genuinely useless in practice? Its always the same thing, shopping lists, calendar, travel plans... it is getting boring, man

Nice to see @FarzaTV’s clicky UX adopted by Google Gemini!

you should make it like clippy

this is farzatastic!

ok but what if there wasnt

the mouse feels like the wrong constraint here if ai is actually a new paradigm, build the hardware around that instead of forcing it through the old one

the cursor knowing what it's pointing at instead of just where it's pointing is genuinely the largest input paradigm shift since touch screens. every interaction model since 1968 has treated the pointer as a coordinate pair. x,y on a grid. the AI pointer treats it as a semantic reference. "this thing" instead of "pixel 847, 312." that's also why the voice plus gesture demo matters more than the visual demos. saying "fix this" while pointing at a code block transmits the same information as a 200 word prompt in about 0.8 seconds. the bandwidth of pointing plus two words exceeds the bandwidth of typing a paragraph. the mouse just became a pronoun

it's called @heyclicky

All this is super fun but costs a lot of tokens, if you ask AI about cool ideas it can give a lot of them but what about the cost? For every small thing if we start using llm tokens its going to cost a lot at current stage i believe

Sorry, but all of the examples you provided would've been done faster by just doing them manually, voice is an incredibly slow input method and prone to mistakes, not to mention that this makes each action a lot more expensive than if it was done with just regular input

Combining gesture, speech and visual context into one input signal is what @karpathy was pointing at last week. This is the first real hardware implementation of that thesis. How far does it go before it replaces the keyboard entirely?

this feels closer to the right ui for agents than pure voice, which is still a bit annoying. pointing is basically human attention made visible, which gives the model a much better prior over what the next action is supposed to touch

I wonder what happens to @heyclicky now 🫠

@tszzl Is this not @heyclicky

Clicky has been clicked @FarzaTV

Why is this better than Clicky?

Woah, with all this technology we're seeing nowadays, we're starting to look like what the future looks like in 1900s movies... Love this tech, hope it gets available for global purchase soon...

#keep4o #keepsonnet45 @DarioAmodei @sama you cant reach AGI if you keep removing LLM models

@mathemagic1an @heyclicky that you ?

The mouse pointer was the cheapest explicit-intent device ever shipped. Replacing it shifts inference cost to the model. Ambiguity becomes a runtime tax: cheaper than learning a UI, more expensive every session.

@ByteMohit see this

Now everything is about how my model can grab right context without actually me describing it.

Open-source alternative already exists: AIPointer ⦿ - same point-and-ask concept, multi-provider, MIT, voice-enabled, ships today.

one step closer

Is it just me, or did anyone else instantly think of @heyclicky by @FarzaTV the second they saw this experiment?

Turning the humble mouse pointer unchanged for 50 years into a semantic pointer that actually understands what you're pointing at is a game changer. No more copy pasting or crafting perfect prompts; just point, gesture, speak naturally ("double this," "summarize for email," "fix that code block") and get things done in context. The demos feel like sci-fi becoming reality scribbled notes → actionable to dos, paused video → instant booking link. This is the kind of intuitive human AI collaboration we've been waiting for. Can't wait to play with the experiments in Google AI Studio. Huge step toward interfaces that finally meet us where we are! What do you all think ready to retire precise prompting? 🚀😊

AI-driven input surfaces new attack vectors motion, speech, shorthand injection. Hope security architecture was reimagined alongside the UX. Zero-trust for human-device interaction starts here. #Cybernorse

this is a step in the right direction trying things like this where the default for the past 50 years has been the same and untouched experiments like this lead to interesting discoveries

@tszzl Called it

This is kinda....useless.

Pointing at something and thinking "fix this" or "take notes on that" is how we already thought when using a computer... now the middle step of actually doing it is being removed

actually impressed. google made the cursor the scope signal for screen-aware agents and that's the move that matters. they otherwise burn tokens guessing what part of the screen you mean. the pointer answers it for free. Kudos @GoogleDeepMind

@FarzaTV is this not copycat of @heyclicky
