Loading video...

Video Failed to Load

Go Home

Learn how to build a real-time voice AI agent using Gemini Live. In this demo created by Software Mansion, users solve a mystery in this interactive game featuring an AI narrator.

10,345 views • 7 months ago •via X (Twitter)

8 Comments

Google AI Developers's profile picture
Google AI Developers7 months ago

Get the code and play the game:

Google AI Developers's profile picture
Google AI Developers7 months ago

Join the Deep Sea Stories livestream on Thursday, Feb 12 at 7 PM CET to see if @GoogleDeepMind’s @thorwebdev can crack the case live and stay for the technical deep dive:

ALT Music's profile picture
ALT Music7 months ago

@swmansion Building a persistent AI companion ('Project Resurrection') in Android Studio right now. The latency in this demo looks incredibly low. Does this implementation handle long-context retention well? I need my agent to remember who 'Zach' is 20 minutes later! 🧠📱

Barrak's profile picture
Barrak7 months ago

@swmansion This interactive mystery game demo is brilliant! Building a real-time voice AI agent with Gemini Live that has natural conversation and narration creates such an immersive experience. Voice AI is transforming gaming! 🎮

🫏 BullishMule 📈's profile picture
🫏 BullishMule 📈7 months ago

@swmansion What game what work around just deliver wtf 😂😂🤯

Vu2day🆗's profile picture
Vu2day🆗7 months ago

@swmansion Building a talking detective sounds like the ultimate weekend project. Just hope the AI doesn't figure out "whodunnit" before the players even start!

Rachel Stuppy's profile picture
Rachel Stuppy7 months ago

@swmansion Love this!

Dmitrii Malakhov's profile picture
Dmitrii Malakhov6 months ago

@swmansion claude can do this too. real-time voice, interactive games, same pattern. the api isn't the moat.

Related Videos

🚨 this chinese guy makes over $1,000,000 a year… by building AI agents. no employees. no massive startup. he just keeps building. while most people are still asking ChatGPT random questions, he’s using Claude to build software that solves real problems. this is what people call vibe coding. he opens Claude and says: “build me an AI agent for real estate businesses that creates property videos.” Claude writes the code. builds the interface. adds subscriptions. helps deploy the app. within a day, he has a working product. then he starts building the next one. that’s the part most people don’t understand. he isn’t trying to build one billion-dollar company. he’s building dozens of AI agents, each solving one problem for one industry. → an AI agent for dentists → an AI agent for ecommerce brands → an AI agent for podcasters → an AI agent for real estate businesses each one automates work that people normally do by hand. each one is built with simple prompts. each one can become a real business. the crazy part? you don’t need to be a software engineer anymore. you need to know how to think like a builder. how to spot problems. how to explain solutions to AI. and how to ship. that’s exactly why i’m reading this article: “How to Actually Build Your First AI Agent.” because this is the skill that’s creating the next generation of builders. the people who learn to build AI agents today won’t just use AI. they’ll own the tools everyone else ends up paying for.

MIKE

39,041 views • 3 months ago

Learn to build conversational AI voice agents in "Building AI Voice Agents for Production", created in collaboration with LiveKit and RealAvatar, and taught by dsa (Co-founder & CEO of LiveKit), Shayne (Developer Advocate, LiveKit), and Nedelina Teneva (Head of AI at RealAvatar, an AI Fund portfolio company). Voice agents combine speech and reasoning capabilities to enable real-time conversations. They're already being used to support customer service, to improve accessibility in healthcare, for entertainment applications, and for talk therapy. In this course, you’ll learn to build voice agents that listen, reason, and respond naturally. You’ll follow the architecture used to create the "AI Andrew" Avatar, a collaborative project between and RealAvatar that responds to users in what sounds like my voice. You’ll build a voice agent from scratch and deploy it to the cloud, enabling support for many simultaneous users. What you’ll learn: - Understand the fundamentals of voice agents, including key components like speech-to-text (STT), text-to-speech (TTS), and LLMs, and how latency is introduced at each layer. - Explore voice agent architectures and the trade-offs between modular pipelines and speech-to-speech APIs. - Explore how platforms like LiveKit mitigate latency issues with optimized networking infrastructure and low-latency communication protocols. - Learn how to connect client devices to voice agents using WebRTC—and why it outperforms HTTP and WebSocket for low-latency audio streaming. - Incorporate voice activity detection (VAD), end-of-turn detection, and context management to detect turns, handle interruptions, and manage conversational flow. - Understand the trade-offs between latency, quality, and cost in an example in which you build a voice agent and change its voice. - Equip your agent with metrics to measure latency at each stage of the voice pipeline and learn the key levers you can pull to make your agent faster and more responsive. The voice agents built in this course also incorporate voice technology from , a supporting contributor to the project. By the end of this course, you'll have learned the components of an AI voice agent pipeline, combined them into a system with low-latency communication, and deployed them on cloud infrastructure so it scales to many users. I’m looking forward to seeing what voice agents you build from this course! Please sign up here:

Andrew Ng

87,965 views • 1 year ago