正在加载视频...

视频加载失败

🚨 New model alert! Dialog by vibx — a leading text-to-speech model — now runs on GroqCloud™. That means natural-sounding speech with ultra-low latency, making real-time voice applications smoother and more responsive. Learn more & build fast — links in the comments!

47,474 次观看 • 1 年前 •via X (Twitter)

11 条评论

Groq Inc 的头像
Groq Inc1 年前

Learn more about our partnership → Build fast →

AssemblyAI 的头像
AssemblyAI1 年前

Announcing: Our most advanced speech-to-text model goes beyond accuracy to capture the real-world complexity of human conversation and deliver reliable, source-of-truth audio data. Explore Universal-2 updates 👇

Hatice Ozen 的头像
Hatice Ozen1 年前

@play_ht TURN UP THE VOLUME! 🗣️🗣️🗣️

PlayAI (formerly PlayHT) 的头像
PlayAI (formerly PlayHT)1 年前

🤝

Walgtech 👨🏻‍💻 的头像
Walgtech 👨🏻‍💻1 年前

@play_ht «Finally» 😅🔥

ZAZO 的头像
ZAZO1 年前

@play_ht 🔥🔥🔥🔥🔥🔥🔥🔥🔥🔥🔥🔥

Himujjal 的头像
Himujjal1 年前

@play_ht Model not found error. Fix that up

The Martian 的头像
The Martian1 年前

@play_ht Love the arabic accent

abdshomad 的头像
abdshomad1 年前

@play_ht Can it plays other languages? Like: Indonesian Language?

jeffscottworld (🛠️, 🤖, 🏡 ) 的头像
jeffscottworld (🛠️, 🤖, 🏡 )1 年前

@play_ht Do you guys have versions that sound conversational and not like she’s reading from a pamphlet?

aiandcivilization 的头像
aiandcivilization1 年前

@play_ht Would be nice to have a model with voice cloning like Zonos !

相关视频

Learn to build conversational AI voice agents in "Building AI Voice Agents for Production", created in collaboration with LiveKit and RealAvatar, and taught by dsa (Co-founder & CEO of LiveKit), Shayne (Developer Advocate, LiveKit), and Nedelina Teneva (Head of AI at RealAvatar, an AI Fund portfolio company). Voice agents combine speech and reasoning capabilities to enable real-time conversations. They're already being used to support customer service, to improve accessibility in healthcare, for entertainment applications, and for talk therapy. In this course, you’ll learn to build voice agents that listen, reason, and respond naturally. You’ll follow the architecture used to create the "AI Andrew" Avatar, a collaborative project between and RealAvatar that responds to users in what sounds like my voice. You’ll build a voice agent from scratch and deploy it to the cloud, enabling support for many simultaneous users. What you’ll learn: - Understand the fundamentals of voice agents, including key components like speech-to-text (STT), text-to-speech (TTS), and LLMs, and how latency is introduced at each layer. - Explore voice agent architectures and the trade-offs between modular pipelines and speech-to-speech APIs. - Explore how platforms like LiveKit mitigate latency issues with optimized networking infrastructure and low-latency communication protocols. - Learn how to connect client devices to voice agents using WebRTC—and why it outperforms HTTP and WebSocket for low-latency audio streaming. - Incorporate voice activity detection (VAD), end-of-turn detection, and context management to detect turns, handle interruptions, and manage conversational flow. - Understand the trade-offs between latency, quality, and cost in an example in which you build a voice agent and change its voice. - Equip your agent with metrics to measure latency at each stage of the voice pipeline and learn the key levers you can pull to make your agent faster and more responsive. The voice agents built in this course also incorporate voice technology from , a supporting contributor to the project. By the end of this course, you'll have learned the components of an AI voice agent pipeline, combined them into a system with low-latency communication, and deployed them on cloud infrastructure so it scales to many users. I’m looking forward to seeing what voice agents you build from this course! Please sign up here:

Andrew Ng

87,711 次观看 • 1 年前