Loading video...
Video Failed to Load
Introducing Speech Engine. Developers can now turn their existing chat agent into a full voice agent with one prompt. Speech Engine combines our leading speech, transcription, and voice orchestration models into a single pipeline - all custom built to work best together.
135,012 views • 4 months ago •via X (Twitter)
39 Comments

Connect to your existing chat agent. Your text-based agent remains untouched - Speech Engine integrates on top of your existing stack so nothing is rearchitected.

Install with one command using our skill. npx skills add elevenlabs/skills --skill speech-engine The skill sets up everything you need so you can go from chat to voice in a single prompt.

Add expressive, human-like voices in 70+ languages. Voice is the fastest and richest way to exchange information, making your product and services more accessible.

Industry-leading transcription. Our transcription models are optimized for conversational use cases, delivering ultra-low latency and built to handle messy, real-world environments.

Enterprise-grade security. Our platform is designed for deployments at scale with enterprise-level data protections, including support for SOC 2, HIPAA, and GDPR compliance. EU Data Residency and Zero Retention Mode are available for stricter data control.

Watch the full Speech Engine walkthrough from the @aiDotEngineer conference in London.

Migrate to ElevenAgents at any time. Get additional deployment channels, monitoring, analytics, and the full suite of agent tools.

Speech Engine is available now in ElevenAPI. Starting at 8¢ per minute, decreasing with scale.

guys you are covered by ai radio!! ai host just picked your release

Whoever did the motion graphics, needs a raise. This thing is crazy good 🔥

Too bad we can't use any of this on YouTube without getting demonetized! THANKS ALOT! SYNTH ID BS.

🔥🔥

Voice agents get interesting when they inherit an existing text workflow instead of replacing it. The shift is chat agent -> callable phone/voice worker.

すごい進化ですね!音声エージェントに変えることができるなんて!

Jarvis? Is that you?

camping your page

so good :3

this is straight up ridiculous in the best way elevenlabs made full voice ai stupidly easy with one prompt low latency magic across 70 languages and zero rearchitecting devs are eating good tonight well played team 🔥

This is great, but what’s the cost to run it?

Cool

evolutionary... awesome

one prompt is doing a lot here

Voice to agent action is interesting. We took a different angle: speak your intent during a coding session, run /act, and the agent executes it with full screen context already in view.

So now one prompt can also make my chatbot talk more than me in real life 😄

@ElevenLabs es lo más horrible y caro que he visto en la vida, no pierdan su tiempo. ¿Un simple video de una mujer bailando? Pfff

voice was the last thing keeping agents feeling like agents. one prompt to flip the switch is wild. the gap between chat and human-like just collapsed.

Voice may be the next real interface shift for agents. Turning a chat agent into a voice agent with one prompt lowers the barrier from “build a product” to “test a new behavior.”

新式の会話仲介システム、素晴らしいですね。

Does Speech Engine support multi-modal interfaces, allowing developers to combine voice, text, and visual inputs for a more seamless user experience?

check out what I built using @ElevenLabs gimmicky or useful?honest feedback only.

voice agents need this kind of boring pipeline work. the magic only works if the whole stack is smooth

More models = more complexity, right? Not here. One prompt replaces three separate integrations. Simpler stack, same power.

Evaluating this for our voice stack. Is $0.08/min flat across all voice models? Agents page suggests bundled, but ElevenAPI shows per-model rates.

Happy to finally see this live!

so i use OpenAI for the LLM and ElevenLabs for TTS. Are you saying that ElevenLabs can now handle the LLM part also? I'm confused

Nice engineering move — lowering the friction to add voice to existing agents is smart. The next layer that matters for production use is when that voice agent has to work reliably in noisy environments, with hands full, and with strong guardrails before anything actually executes. Software layer on top of chat is necessary. Dedicated hardware + controlled handoff is what makes it usable all day.

One prompt to voice is a big unlock.

この内容、日本語で詳しく書きました Wrote a detailed take in Japanese:

the real unlock is turning existing agents into voice ones with a simple prompt. that's huge for devs who already have chat setup, now they can just flip the switch.



