Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Using WebMCP tools, I can surface all sorts of functionality in my app. I can batch process, change themes and dark mode, check on higher level goals that I'm keeping track of. All of this functionality exists already! You can do it manually if you prefer. But you could...

20,831 Aufrufe • vor 6 Tagen •via X (Twitter)

37 Kommentare

Profilbild von Samarth
Samarthvor 6 Tagen

did the same for two apps that i now use frequently- BingeWatcher (movie and tv-show discovery with lineups and watchlists) and ROUGH//CUT (in browser video editor) :))

Profilbild von Sarah Drasner
Sarah Drasnervor 6 Tagen

Nice!

Profilbild von Samarth
Samarthvor 6 Tagen

thanks a lot!

Profilbild von Sarah Drasner
Sarah Drasnervor 6 Tagen

A couple people asked me why I showed chat and not voice- the answer is simple: I was sick and I sound like a frog. 🐸 I'll record one when I am not making ribbit noises

Profilbild von Vishal Anton
Vishal Antonvor 6 Tagen

This is cool.

Profilbild von Srinivasan K K
Srinivasan K Kvor 6 Tagen

yes, I did some experimentation earlier. What I understood is WebMCP for the frontend, MCP is for the backend but for AI agents.

Profilbild von Sarah Drasner
Sarah Drasnervor 6 Tagen

Yes, you’re correct

Profilbild von Srinivasan K K
Srinivasan K Kvor 6 Tagen

👍

Profilbild von Abdul Wasey
Abdul Waseyvor 6 Tagen

WebMCP makes the UI an agent API without rebuilding the backend.

Profilbild von Vinny
Vinnyvor 6 Tagen

There are some very creative use cases for it too:

Profilbild von Ritwik
Ritwikvor 6 Tagen

WebMCP that still works as a normal UI is the right constraint. Agent-only surfaces rot the moment a human needs to take over.

Profilbild von Sarah Drasner
Sarah Drasnervor 6 Tagen

Not necessarily- I think we’ll see more handoffs between experiences over time. There are certain batch actions I might have an agent do and then want to step in to augment individual values, for instance. However, just surfacing the existing functionality is just fine too!

Profilbild von Manu Schiller
Manu Schillervor 6 Tagen

Is WebMCP maybe a good fresh take on accessibility?

Profilbild von Nart Madi
Nart Madivor 6 Tagen

Hey Sarah, I'd love to show you what we're building for WebMCP @Senro_AI.

Profilbild von OneOrigine
OneOriginevor 6 Tagen

;🍎💯

Profilbild von Jarno
Jarnovor 6 Tagen

Voice control for dev tools sounds great until you're debugging out loud in a coffee shop.

Profilbild von Sarah Drasner
Sarah Drasnervor 6 Tagen

I’m not doing voice control here, I’m just typing. That’s sort of the point- it’s flexible to your situation ☺️

Profilbild von Vinicius Dallacqua
Vinicius Dallacquavor 6 Tagen

Do you have a link to this? I'd love to use it for benchmarking and metrics on an article I'm writting.

Profilbild von Sarah Drasner
Sarah Drasnervor 6 Tagen

Yeah I’m getting some last features over the line and then I’ll open source it

Profilbild von Rohan Verma
Rohan Vermavor 6 Tagen

What's the chat UI on the right?

Profilbild von Sarah Drasner
Sarah Drasnervor 6 Tagen

I made it a part of the app. I embedded Gemini and surfaced a UI that can be voice or chat

Profilbild von rusa
rusavor 4 Tagen

the a11y angle is the sleeper. app actions as callable tools is basically a machine-readable menu, which is exactly what screen readers always wanted and rarely got reliably.

Profilbild von Sarah Drasner
Sarah Drasnervor 4 Tagen

1000%

Profilbild von Sudhir Dudeja
Sudhir Dudejavor 6 Tagen

I made my app webmcp enable as well, Quiak question what are using apart from codex to test and it and does it require API key ?🤔

Profilbild von Max Villemure
Max Villemurevor 6 Tagen

I really like the view, and the image on top. Great job

Profilbild von John Rood
John Roodvor 6 Tagen

we run enough browser agents to feel this: a semantic tool surface beats teaching the model to click pixels every time. the accessibility path becomes an interface you can test, not a demo you hope works.

Profilbild von Ben Mo
Ben Movor 6 Tagen

does one undo roll back the whole batch, or would it step through each task change?

Profilbild von Agrit Tiwari
Agrit Tiwarivor 5 Tagen

This is good to do client side tool calling. Makes sense. Really like it so far. But is this Copilotkit for the sideChat? PS. tc.

Profilbild von Sarah Drasner
Sarah Drasnervor 5 Tagen

Thanks! Just Gemini

Profilbild von Aivan Monceller
Aivan Moncellervor 6 Tagen

Are you running an LLM on the browser? Or still calling a frontier in this case?

Profilbild von Sarah Drasner
Sarah Drasnervor 6 Tagen

I've embedded Gemini and made it part of the app experience. You don't have to do it that way, though

Profilbild von Aivan Monceller
Aivan Moncellervor 6 Tagen

What does "embedded Gemini" mean? Is the model actually running in the page and driving the existing features through WebMCP? What's the difference in practice between embedding Gemini in the app vs just calling the API?

Profilbild von Sarah Drasner
Sarah Drasnervor 6 Tagen

Yeah it's running in the page and driving existing WebMCP tooling. I exposed some special tools to do things my app can't do as well like batch processing or checking against larger goals. No big difference here, and you can still bring your own/drive it another way. The app is for me, so I wanted it as part of my experience.

Profilbild von Saket Tawde
Saket Tawdevor 6 Tagen

Woah! Imagine if the whole OS was like this?! Wait a second, wasn't there a ChromeOS?

Profilbild von LottieFiles
LottieFilesvor 6 Tagen

The voice part is really interesting. Being able to just say what you want instead of clicking through menus would help a lot of people. Have you tested it with voice yet, or is that still the plan?

Profilbild von Sarah Drasner
Sarah Drasnervor 6 Tagen

Yes! I almost did a voice one, and will probably still post it. I've been sick this week and my voice sounds like a frog 🐸 ribbit

Profilbild von Tony
Tonyvor 6 Tagen

The accessibility angle is what caught my attention. Being able to say what you want done could make a complicated app feel much more approachable.

Ähnliche Videos