Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

“You can build a real relationship with someone using just your voice” Understanding the power of audio, JC built Playfriends into a 1-million-user voice platform where creators can monetise through voice alone. Now, with Beam, creators can receive virtual gift without even going live. Tune in to the discussion...

10,873 Aufrufe • vor 9 Monaten •via X (Twitter)

4 Kommentare

Profilbild von Nessa ♡
Nessa ♡vor 9 Monaten

@jc_playfriends @playfriends_gg @Beam_gifts Im host on playfriends for more than 1 year, and im so happy be part of this, for more things in the future, we deserve the best @playfriends_gg

Profilbild von aix
aixvor 9 Monaten

@jc_playfriends @playfriends_gg @Beam_gifts not the 6, 7 months-

Profilbild von imho
imhovor 9 Monaten

@jc_playfriends @playfriends_gg @Beam_gifts Voice only monetization is underrated in crypto Distribution and trust matter more than fancy UX

Profilbild von Bloopy 👑
Bloopy 👑vor 9 Monaten

@jc_playfriends @playfriends_gg @Beam_gifts Hey Sent u dm

Ähnliche Videos

Voice used to be AI’s forgotten modality - now it's having its big moment: rapid innovation, big funding rounds, major agentic applications My conversation with Neil Zeghidour, top AI researcher in the field (Google DeepMind, Meta, kyutai) and now CEO of Gradium This is a reference episode on all things voice AI 🔥 00:00 Intro 01:21 Voice AI’s big moment, and why we’re still early 03:34 Why voice lagged behind text/image/video 06:06 The convergence era: transformers for every modality 07:40 Beyond Her: always-on assistants, wake words, voice-first devices 11:01 Voice vs text: where voice fits (even for coding) 12:56 Neil’s origin story: from finance to machine learning, with help from Yann LeCun and Soumith Chintala 18:35 Neural codecs (SoundStream): compression as the unlock 22:30 Kyutai: open research, small elite teams, moving fast 31:32 Why big labs haven’t “won” voice AI4 34:01 On-device voice: where it works, why compact models matter 46:37 The last mile: real-world robustness, pronunciation, uptime 41:35 Benchmarking voice: why metrics fail, how they actually test 47:03 Cascades vs speech-to-speech: trade-offs + what’s next 54:05 Hardest frontier: noisy rooms, factories, multi-speaker chaos 1:00:50 New languages + dialects: what transfers, what doesn’t 1:02:54 Hardware & compute: why voice isn’t a 10,000-GPU game 1:07:27 What data do you need to train voice models 1:09:02 Deepfakes + privacy: why watermarking isn’t a solution 1:12:30 Voice + vision: multimodality, screen awareness, video+audio 1:14:43 Voice cloning vs voice design: where the market goes 1:16:32 Paris/Europe AI: talent density, underdog energy, what’s next

Matt Turck

22,995 Aufrufe • vor 7 Monaten

Voice AI can pass a Turing test. For about a minute. That's a generated clip, though. Have a human actually talk back and the number collapses to six or seven seconds, roughly where generated voice sat three years ago. One reason: real conversation isn't turn-based. Around 20% of the time more than one person is speaking, and laughter drives a lot of that overlap, since you laugh at a joke while it's still being told. We also adjust our pacing toward whoever we're talking to without noticing we're doing it. Voice models struggle with all of this. Full-duplex voice, where a model listens and speaks at the same time, is still extremely early. So an agent can know your joke is funny and still have to wait until you've finished before it laughs, by which point the timing has killed it. Miso Labs CEO & Co-Founder, Aoden Teo, describes a second consequence: agents get pushed toward almost "psychotically emotive" behavior. If they can only talk once you've stopped, they need some other way to show they were listening. You finish your sentence, and the thing goes "Hmm?". You've heard it. Underneath that sits an architecture problem. Voice models have to respond fast, which constrains how large they can be, and fast means something different here than it does in text. Working with an LLM like Claude, you care how quickly it finishes your code, more than how quickly it starts. Voice inverts that. Nobody needs 10 hours of audio generated in two seconds, because nobody can listen to 10 hours of audio in two seconds; what matters is reaction time. Most architectural decisions trade latency against throughput, and Aoden expects voice to keep moving away from LLM-style designs toward ones built around very low latency. Miso Labs is already pushing on it. Miso-1 got 3,000+ stars on GitHub and 5 million views on X, and they record data in their own LA studio because the internet doesn't contain every kind of audio a voice model might need. Nobody has released a podcast of someone reading millions and millions of email addresses, and people still want voice models that can read email addresses aloud, so teams end up generating some very strange training data themselves. The clip isn't the hard part. The hard part starts when you talk back. "The most emotive foundation models for voice" 🎙️Aoden Teo, CEO & Co-Founder, Miso Labs on Fondo.com START 1:03 Miso-1: 3K+ GitHub stars + 5M X views 1:59 Why emotiveness matters for games, UGC + interactive products 3:06 Measuring progress in voice AI with longer Turing tests 4:01 Why interactive conversation is harder than generating convincing clips 5:08 Full-duplex voice, interruptions + why laughter matters 6:04 Latency vs. throughput - and why voice differs from LLMs 7:09 Miso's LA recording studio + the challenge of voice training data 9:02 Talking teddy bears, UGC, anime + unexpected voice AI use cases 10:19 From serious chess player to math obsession to building Miso Labs 12:11 The surprise YC interview

David J Phillips

36,612 Aufrufe • vor 1 Monat