正在加载视频...

视频加载失败

Using WebMCP tools, I can surface all sorts of functionality in my app. I can batch process, change themes and dark mode, check on higher level goals that I'm keeping track of. All of this functionality exists already! You can do it manually if you prefer. But you could...

20,831 次观看 • 6 天前 •via X (Twitter)

37 条评论

Samarth 的头像
Samarth6 天前

did the same for two apps that i now use frequently- BingeWatcher (movie and tv-show discovery with lineups and watchlists) and ROUGH//CUT (in browser video editor) :))

Sarah Drasner 的头像
Sarah Drasner6 天前

Nice!

Samarth 的头像
Samarth6 天前

thanks a lot!

Sarah Drasner 的头像
Sarah Drasner6 天前

A couple people asked me why I showed chat and not voice- the answer is simple: I was sick and I sound like a frog. 🐸 I'll record one when I am not making ribbit noises

Vishal Anton 的头像
Vishal Anton6 天前

This is cool.

Srinivasan K K 的头像
Srinivasan K K6 天前

yes, I did some experimentation earlier. What I understood is WebMCP for the frontend, MCP is for the backend but for AI agents.

Sarah Drasner 的头像
Sarah Drasner6 天前

Yes, you’re correct

Srinivasan K K 的头像
Srinivasan K K6 天前

👍

Abdul Wasey 的头像
Abdul Wasey6 天前

WebMCP makes the UI an agent API without rebuilding the backend.

Vinny 的头像
Vinny6 天前

There are some very creative use cases for it too:

Ritwik 的头像
Ritwik5 天前

WebMCP that still works as a normal UI is the right constraint. Agent-only surfaces rot the moment a human needs to take over.

Sarah Drasner 的头像
Sarah Drasner5 天前

Not necessarily- I think we’ll see more handoffs between experiences over time. There are certain batch actions I might have an agent do and then want to step in to augment individual values, for instance. However, just surfacing the existing functionality is just fine too!

Manu Schiller 的头像
Manu Schiller6 天前

Is WebMCP maybe a good fresh take on accessibility?

Nart Madi 的头像
Nart Madi6 天前

Hey Sarah, I'd love to show you what we're building for WebMCP @Senro_AI.

OneOrigine 的头像
OneOrigine6 天前

;🍎💯

Jarno 的头像
Jarno6 天前

Voice control for dev tools sounds great until you're debugging out loud in a coffee shop.

Sarah Drasner 的头像
Sarah Drasner6 天前

I’m not doing voice control here, I’m just typing. That’s sort of the point- it’s flexible to your situation ☺️

Vinicius Dallacqua 的头像
Vinicius Dallacqua6 天前

Do you have a link to this? I'd love to use it for benchmarking and metrics on an article I'm writting.

Sarah Drasner 的头像
Sarah Drasner6 天前

Yeah I’m getting some last features over the line and then I’ll open source it

Rohan Verma 的头像
Rohan Verma6 天前

What's the chat UI on the right?

Sarah Drasner 的头像
Sarah Drasner6 天前

I made it a part of the app. I embedded Gemini and surfaced a UI that can be voice or chat

rusa 的头像
rusa4 天前

the a11y angle is the sleeper. app actions as callable tools is basically a machine-readable menu, which is exactly what screen readers always wanted and rarely got reliably.

Sarah Drasner 的头像
Sarah Drasner4 天前

1000%

Sudhir Dudeja 的头像
Sudhir Dudeja6 天前

I made my app webmcp enable as well, Quiak question what are using apart from codex to test and it and does it require API key ?🤔

Max Villemure 的头像
Max Villemure6 天前

I really like the view, and the image on top. Great job

John Rood 的头像
John Rood6 天前

we run enough browser agents to feel this: a semantic tool surface beats teaching the model to click pixels every time. the accessibility path becomes an interface you can test, not a demo you hope works.

Ben Mo 的头像
Ben Mo6 天前

does one undo roll back the whole batch, or would it step through each task change?

Agrit Tiwari 的头像
Agrit Tiwari5 天前

This is good to do client side tool calling. Makes sense. Really like it so far. But is this Copilotkit for the sideChat? PS. tc.

Sarah Drasner 的头像
Sarah Drasner5 天前

Thanks! Just Gemini

Aivan Monceller 的头像
Aivan Monceller6 天前

Are you running an LLM on the browser? Or still calling a frontier in this case?

Sarah Drasner 的头像
Sarah Drasner6 天前

I've embedded Gemini and made it part of the app experience. You don't have to do it that way, though

Aivan Monceller 的头像
Aivan Monceller6 天前

What does "embedded Gemini" mean? Is the model actually running in the page and driving the existing features through WebMCP? What's the difference in practice between embedding Gemini in the app vs just calling the API?

Sarah Drasner 的头像
Sarah Drasner6 天前

Yeah it's running in the page and driving existing WebMCP tooling. I exposed some special tools to do things my app can't do as well like batch processing or checking against larger goals. No big difference here, and you can still bring your own/drive it another way. The app is for me, so I wanted it as part of my experience.

Saket Tawde 的头像
Saket Tawde6 天前

Woah! Imagine if the whole OS was like this?! Wait a second, wasn't there a ChromeOS?

LottieFiles 的头像
LottieFiles6 天前

The voice part is really interesting. Being able to just say what you want instead of clicking through menus would help a lot of people. Have you tested it with voice yet, or is that still the plan?

Sarah Drasner 的头像
Sarah Drasner6 天前

Yes! I almost did a voice one, and will probably still post it. I've been sick this week and my voice sounds like a frog 🐸 ribbit

Tony 的头像
Tony6 天前

The accessibility angle is what caught my attention. Being able to say what you want done could make a complicated app feel much more approachable.

相关视频