正在加载视频...

视频加载失败

We’re reimagining a 50-year-old interface - the mouse pointer - with AI. 🖱️ These experimental demos show how people can intuitively direct Gemini on their screens using motion, speech, and natural shorthand to get things done 🧵

1,684,389 次观看 • 4 个月前 •via X (Twitter)

38 条评论

Google DeepMind 的头像
Google DeepMind4 个月前

With an AI-enabled pointer, help is always available where you’re working - without having to detour to additional apps. 📲 Point at a PDF and request bullet points for an email, hover over a table to ask for a pie chart, or highlight a recipe and simply say: "double these ingredients."

Google DeepMind 的头像
Google DeepMind4 个月前

Current models require precise instructions, but our AI-enabled pointer removes that burden. 💡 By "seeing" what’s under your cursor, it instantly understands the specific word, image, or code block you need help with.

Google DeepMind 的头像
Google DeepMind4 个月前

In the real world, we don't tend to speak in long paragraphs; we point and say: "fix this" or "move that". 💬 By combining gestures with speech, it lets you use natural shorthand to complete tasks.

Google DeepMind 的头像
Google DeepMind4 个月前

For decades, your mouse only tracked where you were pointing. AI helps it understand what you're pointing at. 💭 This means a photo of a scribbled note could turn into an interactive to-do list, or a paused video frame can become a restaurant booking link.

Google DeepMind 的头像
Google DeepMind4 个月前

These capabilities are guiding how we think about the next generation of interfaces. As we continue exploring what an AI-enabled mouse pointer would unlock, try our experiments in @GoogleAIStudio →

toi500 的头像
toi5004 个月前

Why do you guys always use examples that target the general public but are genuinely useless in practice? Its always the same thing, shopping lists, calendar, travel plans... it is getting boring, man

Danny Limanseta 的头像
Danny Limanseta4 个月前

Nice to see @FarzaTV’s clicky UX adopted by Google Gemini!

Ajeya 的头像
Ajeya4 个月前

you should make it like clippy

kitze 🛠️ tinkerer.club 的头像
kitze 🛠️ tinkerer.club4 个月前

this is farzatastic!

variousred 的头像
variousred4 个月前

ok but what if there wasnt

leon 的头像
leon4 个月前

the mouse feels like the wrong constraint here if ai is actually a new paradigm, build the hardware around that instead of forcing it through the old one

Eshan 的头像
Eshan4 个月前

the cursor knowing what it's pointing at instead of just where it's pointing is genuinely the largest input paradigm shift since touch screens. every interaction model since 1968 has treated the pointer as a coordinate pair. x,y on a grid. the AI pointer treats it as a semantic reference. "this thing" instead of "pixel 847, 312." that's also why the voice plus gesture demo matters more than the visual demos. saying "fix this" while pointing at a code block transmits the same information as a 200 word prompt in about 0.8 seconds. the bandwidth of pointing plus two words exceeds the bandwidth of typing a paragraph. the mouse just became a pronoun

Francesco 的头像
Francesco4 个月前

it's called @heyclicky

Mohit Mor 的头像
Mohit Mor4 个月前

All this is super fun but costs a lot of tokens, if you ask AI about cool ideas it can give a lot of them but what about the cost? For every small thing if we start using llm tokens its going to cost a lot at current stage i believe

Noctre 的头像
Noctre4 个月前

Sorry, but all of the examples you provided would've been done faster by just doing them manually, voice is an incredibly slow input method and prone to mistakes, not to mention that this makes each action a lot more expensive than if it was done with just regular input

Yann Kronberg 的头像
Yann Kronberg4 个月前

Combining gesture, speech and visual context into one input signal is what @karpathy was pointing at last week. This is the first real hardware implementation of that thesis. How far does it go before it replaces the keyboard entirely?

Connor 的头像
Connor4 个月前

this feels closer to the right ui for agents than pure voice, which is still a bit annoying. pointing is basically human attention made visible, which gives the model a much better prior over what the next action is supposed to touch

Yash Bhalgat 的头像
Yash Bhalgat4 个月前

I wonder what happens to @heyclicky now 🫠

Nick Schmidt 的头像
Nick Schmidt4 个月前

@tszzl Is this not @heyclicky

Yahia Bakour 的头像
Yahia Bakour4 个月前

Clicky has been clicked @FarzaTV

ib 的头像
ib4 个月前

Why is this better than Clicky?

Arai Chartz👽 的头像
Arai Chartz👽4 个月前

Woah, with all this technology we're seeing nowadays, we're starting to look like what the future looks like in 1900s movies... Love this tech, hope it gets available for global purchase soon...

Murat 的头像
Murat4 个月前

#keep4o #keepsonnet45 @DarioAmodei @sama you cant reach AGI if you keep removing LLM models

Abin Kumar 的头像
Abin Kumar4 个月前

@mathemagic1an @heyclicky that you ?

Michał Piszczek 的头像
Michał Piszczek4 个月前

The mouse pointer was the cheapest explicit-intent device ever shipped. Replacing it shifts inference cost to the model. Ambiguity becomes a runtime tax: cheaper than learning a UI, more expensive every session.

Maaz 的头像
Maaz4 个月前

@ByteMohit see this

Darshan Savaliya (Aspora) 的头像
Darshan Savaliya (Aspora)4 个月前

Now everything is about how my model can grab right context without actually me describing it.

Mario Simic 的头像
Mario Simic4 个月前

Open-source alternative already exists: AIPointer ⦿ - same point-and-ask concept, multi-provider, MIT, voice-enabled, ships today.

Ali Mehdi Mukadam 的头像
Ali Mehdi Mukadam4 个月前

one step closer

Manoj 的头像
Manoj4 个月前

Is it just me, or did anyone else instantly think of @heyclicky by @FarzaTV the second they saw this experiment?

Parsec 🪐 的头像
Parsec 🪐4 个月前

Turning the humble mouse pointer unchanged for 50 years into a semantic pointer that actually understands what you're pointing at is a game changer. No more copy pasting or crafting perfect prompts; just point, gesture, speak naturally ("double this," "summarize for email," "fix that code block") and get things done in context. The demos feel like sci-fi becoming reality scribbled notes → actionable to dos, paused video → instant booking link. This is the kind of intuitive human AI collaboration we've been waiting for. Can't wait to play with the experiments in Google AI Studio. Huge step toward interfaces that finally meet us where we are! What do you all think ready to retire precise prompting? 🚀😊

Cybernorse 的头像
Cybernorse4 个月前

AI-driven input surfaces new attack vectors motion, speech, shorthand injection. Hope security architecture was reimagined alongside the UX. Zero-trust for human-device interaction starts here. #Cybernorse

Stefan 的头像
Stefan4 个月前

this is a step in the right direction trying things like this where the default for the past 50 years has been the same and untouched experiments like this lead to interesting discoveries

Muad'Deep - e/acc 的头像
Muad'Deep - e/acc4 个月前

@tszzl Called it

Ngoc Anh Tran 的头像
Ngoc Anh Tran4 个月前

This is kinda....useless.

The Calda 的头像
The Calda4 个月前

Pointing at something and thinking "fix this" or "take notes on that" is how we already thought when using a computer... now the middle step of actually doing it is being removed

Tom Curonian 的头像
Tom Curonian4 个月前

actually impressed. google made the cursor the scope signal for screen-aware agents and that's the move that matters. they otherwise burn tokens guessing what part of the screen you mean. the pointer answers it for free. Kudos @GoogleDeepMind

CityLad 的头像
CityLad4 个月前

@FarzaTV is this not copycat of @heyclicky

相关视频