Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

I’ve been exploring Gemini 2.0’s new native audio output capability, which is available for early testers. I’m a developer at Google Creative Lab, and wanted to share one of my favorite experiments so far called ✨ VoiceCursor (🔊 sound on for video) Unlike traditional TTS, native audio lets you...

67,672 görüntüleme • 1 yıl önce •via X (Twitter)

10 Yorum

Trudy Painter profil fotoğrafı
Trudy Painter1 yıl önce

Gemini 2.0 native audio output is available in AI Studio for early testers. The prompt in this screencap is: Say this in an upbeat, happy tone: “You can steer a voice and … put emphasis on different words!” 🔗

Trudy Painter profil fotoğrafı
Trudy Painter1 yıl önce

✨Voice Cursor follows a similar prompting strategy. After you highlight a phrase, the Voice Cursor will ask the API for audio for the phrase in your selected voice and tone. (and you can edit the prompt sent to the Gemini API in the bottom box)

Trudy Painter profil fotoğrafı
Trudy Painter1 yıl önce

And for me, when the ✨Voice Cursor sits inside a familiar text editor, the highlight interaction feels fluid and comfortable. I’m excited about how native audio might enable new kinds of tools for how we write...

Trudy Painter profil fotoğrafı
Trudy Painter1 yıl önce

You can get the code to see how it works at Native audio output is available to early testers now, with a wider rollout expected next year. This voice cursor was built on top of Such a good repo Also - it’s super simple to change the tone prompt presets + how you make calls to the Gemini 2.0 API (see screenshot below).

jpa profil fotoğrafı
jpa1 yıl önce

so cool, trudy!

Codetard profil fotoğrafı
Codetard1 yıl önce

:)

Tom Bielecki profil fotoğrafı
Tom Bielecki1 yıl önce

@codexeditor audio as another annotation layer

Data & Analytics profil fotoğrafı
Data & Analytics1 yıl önce

@JeffDean @JeffDean, that native audio output sounds dope! Real game-changer for developers. How’s it stacking up against other tools you’ve tried?

steve ike profil fotoğrafı
steve ike1 yıl önce

This is really cool. Thanks for sharing, look forward to checking out the code and learning from your work.

𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺 profil fotoğrafı
𝑫𝒂𝒏𝒊𝒆𝒍 𝑺𝒄𝒐𝒕𝒕 𝑴𝒂𝒕𝒕𝒉𝒆𝒘𝒔 🇦🇺1 yıl önce

Oh, the quality is remarkable! Thanks for sharing.

Benzer Videolar

I asked Garry Tan how to use meta prompting to get better at AI: "My partners at YC Jared Friedman and Pete Koomen showed me how to do this. You can take almost anything that you do all the time and just drop it into a context window. And then say, “Here’s a bunch of inputs and outputs." And maybe you also add a bunch of notes. And then you tell it, “Write me a prompt that can act as an agent that takes this input and makes this output over here.” You can do this for almost any type of knowledge work. And you can even introspect. "What are things you notice that I did to convert this from the input to the output?”. And then you can just start using the prompt. Initially, it’s going to suck. Because it’s just not that smart yet. But what’s funny is now, I also use it to Iterate my writing. You can be very direct, "I would never say that", "Don’t say it like this", or "Oh, you used the long word there, use the short word". Just speak to it conversationally. And then when you're happy with the output, you can use that new output to make a new prompt. "Based on this conversation, give me a better initial prompt that incorporates all the things we talked about." And you can do this with literally everything. And in theory, there’s so much it applies to that people do day-to-day. You could use it for tweets. You could use it for editing podcasts. You can use it for pretty much everything. I have a folder of prompts that I use all the time. My YouTube prompt is on v27 or something. I'll go through this process with all the different max models. I'll use GPT 5.2 Pro. I’ll use Grok. I'll use Claude. Then, I’ll take all the outputs from all the models and put them into Claude and say "Here’s my prompt, here’s the output from four LLMs, including yourself. Rate each response and tell me what the pros and cons of each approach are." And I usually say "give it to me in numbered form". And then you can agree with one, disagree with two, tell it three is this or that. And then after that, you say given all of this, synthesize it."

The Peel

51,632 görüntüleme • 7 ay önce