Loading video...

Video Failed to Load

Go Home

I’ve been exploring Gemini 2.0’s new native audio output capability, which is available for early testers. I’m a developer at Google Creative Lab, and wanted to share one of my favorite experiments so far called ✨ VoiceCursor (🔊 sound on for video) Unlike traditional TTS, native audio lets you...

67,672 views • 1 year ago •via X (Twitter)

10 Comments

Trudy Painter's profile picture
Trudy Painter1 year ago

Gemini 2.0 native audio output is available in AI Studio for early testers. The prompt in this screencap is: Say this in an upbeat, happy tone: “You can steer a voice and … put emphasis on different words!” 🔗

Trudy Painter's profile picture
Trudy Painter1 year ago

✨Voice Cursor follows a similar prompting strategy. After you highlight a phrase, the Voice Cursor will ask the API for audio for the phrase in your selected voice and tone. (and you can edit the prompt sent to the Gemini API in the bottom box)

Trudy Painter's profile picture
Trudy Painter1 year ago

And for me, when the ✨Voice Cursor sits inside a familiar text editor, the highlight interaction feels fluid and comfortable. I’m excited about how native audio might enable new kinds of tools for how we write...

Trudy Painter's profile picture
Trudy Painter1 year ago

You can get the code to see how it works at Native audio output is available to early testers now, with a wider rollout expected next year. This voice cursor was built on top of Such a good repo Also - it’s super simple to change the tone prompt presets + how you make calls to the Gemini 2.0 API (see screenshot below).

jpa's profile picture
jpa1 year ago

so cool, trudy!

Codetard's profile picture
Codetard1 year ago

:)

Tom Bielecki's profile picture
Tom Bielecki1 year ago

@codexeditor audio as another annotation layer

Data & Analytics's profile picture
Data & Analytics1 year ago

@JeffDean @JeffDean, that native audio output sounds dope! Real game-changer for developers. How’s it stacking up against other tools you’ve tried?

steve ike's profile picture
steve ike1 year ago

This is really cool. Thanks for sharing, look forward to checking out the code and learning from your work.

𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺's profile picture
𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺1 year ago

Oh, the quality is remarkable! Thanks for sharing.

Related Videos

I asked Garry Tan how to use meta prompting to get better at AI: "My partners at YC Jared Friedman and Pete Koomen showed me how to do this. You can take almost anything that you do all the time and just drop it into a context window. And then say, “Here’s a bunch of inputs and outputs." And maybe you also add a bunch of notes. And then you tell it, “Write me a prompt that can act as an agent that takes this input and makes this output over here.” You can do this for almost any type of knowledge work. And you can even introspect. "What are things you notice that I did to convert this from the input to the output?”. And then you can just start using the prompt. Initially, it’s going to suck. Because it’s just not that smart yet. But what’s funny is now, I also use it to Iterate my writing. You can be very direct, "I would never say that", "Don’t say it like this", or "Oh, you used the long word there, use the short word". Just speak to it conversationally. And then when you're happy with the output, you can use that new output to make a new prompt. "Based on this conversation, give me a better initial prompt that incorporates all the things we talked about." And you can do this with literally everything. And in theory, there’s so much it applies to that people do day-to-day. You could use it for tweets. You could use it for editing podcasts. You can use it for pretty much everything. I have a folder of prompts that I use all the time. My YouTube prompt is on v27 or something. I'll go through this process with all the different max models. I'll use GPT 5.2 Pro. I’ll use Grok. I'll use Claude. Then, I’ll take all the outputs from all the models and put them into Claude and say "Here’s my prompt, here’s the output from four LLMs, including yourself. Rate each response and tell me what the pros and cons of each approach are." And I usually say "give it to me in numbered form". And then you can agree with one, disagree with two, tell it three is this or that. And then after that, you say given all of this, synthesize it."

The Peel

51,632 views • 7 months ago