Sensitive content

This media may contain sensitive content.

Загрузка видео...

Не удалось загрузить видео

На главную

Sampl:e: - Listen with 🎧 ON: #Pornosexual Drone - my most complex audio layering with IRL voice #hypno and subliminal #goonaudio overlayed deliciously nasty Porn. Can you handle it? #pornaddict #goonharder #pornpiggy #goonencouragement #GoonCap #porncaps #PornIsLife #hypnoporn

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

#BaiLu on choosing to dub Li Peiyi with her own voice “I’d like to talk a bit about the dubbing. It was mainly because of the filming environment in Hengdian there were a lot of noise issues. When we first started shooting this drama, the plan was to use the original on-set audio entirely. But as filming continued, the environmental interference became too much. We had a huge amount of dialogue every day, and eventually the director told us, ‘Forget it either you dub it yourselves in post-production, or we’ll bring in professional voice actors.’ By the end, voice actress Qiao Shiyu had already recorded my lines three times. She worked incredibly hard! Later on, though, the director messaged me on WeChat and said, ‘Lulu, come to the editing room. Don’t just listen on your phone listen in surround sound and then make your choice.’ Since this character is somewhat similar to my role in The Legends (Zhao Yao), Teacher Qiao Shiyu had dubbed it very diligently three times. In the end, I spent an entire afternoon in the editing room, and the whole team decided together that my natural voice might suit Pei Yi better. My voice isn’t perfect, and my delivery isn’t the absolute best, but perhaps for Pei Yi who isn’t meant to sound overly polished or perfectly ‘pretty’ my vocal tone fit the character more. So we ultimately chose to use my own voice. As actors, we know that using original audio is a real test of our abilities. We’ll keep working hard, and for any areas where I didn’t do well enough, I hope everyone can be understanding.”

美

21,805 просмотров • 7 месяцев назад

Voice AI can pass a Turing test. For about a minute. That's a generated clip, though. Have a human actually talk back and the number collapses to six or seven seconds, roughly where generated voice sat three years ago. One reason: real conversation isn't turn-based. Around 20% of the time more than one person is speaking, and laughter drives a lot of that overlap, since you laugh at a joke while it's still being told. We also adjust our pacing toward whoever we're talking to without noticing we're doing it. Voice models struggle with all of this. Full-duplex voice, where a model listens and speaks at the same time, is still extremely early. So an agent can know your joke is funny and still have to wait until you've finished before it laughs, by which point the timing has killed it. Miso Labs CEO & Co-Founder, Aoden Teo, describes a second consequence: agents get pushed toward almost "psychotically emotive" behavior. If they can only talk once you've stopped, they need some other way to show they were listening. You finish your sentence, and the thing goes "Hmm?". You've heard it. Underneath that sits an architecture problem. Voice models have to respond fast, which constrains how large they can be, and fast means something different here than it does in text. Working with an LLM like Claude, you care how quickly it finishes your code, more than how quickly it starts. Voice inverts that. Nobody needs 10 hours of audio generated in two seconds, because nobody can listen to 10 hours of audio in two seconds; what matters is reaction time. Most architectural decisions trade latency against throughput, and Aoden expects voice to keep moving away from LLM-style designs toward ones built around very low latency. Miso Labs is already pushing on it. Miso-1 got 3,000+ stars on GitHub and 5 million views on X, and they record data in their own LA studio because the internet doesn't contain every kind of audio a voice model might need. Nobody has released a podcast of someone reading millions and millions of email addresses, and people still want voice models that can read email addresses aloud, so teams end up generating some very strange training data themselves. The clip isn't the hard part. The hard part starts when you talk back. "The most emotive foundation models for voice" 🎙️Aoden Teo, CEO & Co-Founder, Miso Labs on Fondo.com START 1:03 Miso-1: 3K+ GitHub stars + 5M X views 1:59 Why emotiveness matters for games, UGC + interactive products 3:06 Measuring progress in voice AI with longer Turing tests 4:01 Why interactive conversation is harder than generating convincing clips 5:08 Full-duplex voice, interruptions + why laughter matters 6:04 Latency vs. throughput - and why voice differs from LLMs 7:09 Miso's LA recording studio + the challenge of voice training data 9:02 Talking teddy bears, UGC, anime + unexpected voice AI use cases 10:19 From serious chess player to math obsession to building Miso Labs 12:11 The surprise YC interview

David J Phillips

36,612 просмотров • 22 дней назад

Learn to build conversational AI voice agents in "Building AI Voice Agents for Production", created in collaboration with LiveKit and RealAvatar, and taught by dsa (Co-founder & CEO of LiveKit), Shayne (Developer Advocate, LiveKit), and Nedelina Teneva (Head of AI at RealAvatar, an AI Fund portfolio company). Voice agents combine speech and reasoning capabilities to enable real-time conversations. They're already being used to support customer service, to improve accessibility in healthcare, for entertainment applications, and for talk therapy. In this course, you’ll learn to build voice agents that listen, reason, and respond naturally. You’ll follow the architecture used to create the "AI Andrew" Avatar, a collaborative project between and RealAvatar that responds to users in what sounds like my voice. You’ll build a voice agent from scratch and deploy it to the cloud, enabling support for many simultaneous users. What you’ll learn: - Understand the fundamentals of voice agents, including key components like speech-to-text (STT), text-to-speech (TTS), and LLMs, and how latency is introduced at each layer. - Explore voice agent architectures and the trade-offs between modular pipelines and speech-to-speech APIs. - Explore how platforms like LiveKit mitigate latency issues with optimized networking infrastructure and low-latency communication protocols. - Learn how to connect client devices to voice agents using WebRTC—and why it outperforms HTTP and WebSocket for low-latency audio streaming. - Incorporate voice activity detection (VAD), end-of-turn detection, and context management to detect turns, handle interruptions, and manage conversational flow. - Understand the trade-offs between latency, quality, and cost in an example in which you build a voice agent and change its voice. - Equip your agent with metrics to measure latency at each stage of the voice pipeline and learn the key levers you can pull to make your agent faster and more responsive. The voice agents built in this course also incorporate voice technology from , a supporting contributor to the project. By the end of this course, you'll have learned the components of an AI voice agent pipeline, combined them into a system with low-latency communication, and deployed them on cloud infrastructure so it scales to many users. I’m looking forward to seeing what voice agents you build from this course! Please sign up here:

Andrew Ng

87,868 просмотров • 1 год назад