
Soniox
@soniox_ai • 1,965 subscribers
Low-latency real-time speech-to-text, text-to-speech and translation APIs.
Shorts
Videos

Introducing Soniox TTS v2, our most powerful text-to-speech model yet. Soniox TTS v2 brings extraordinary voice quality, expressive control through audio tags, exceptional precision, high-fidelity voice cloning, more than 60 languages, natural language mixing, and low-latency streaming together in one model. Built from the ground up with a new model architecture, audio codec, and highly optimized inference engine, Soniox TTS v2 delivers frontier TTS at just $0.70 per generated hour. Available globally today as tts-rt-v2. Try it out: Turn the sound on and hear what’s possible. 🔊
Soniox666,014 views • 1 month ago

Soniox TTS lets you generate speech with distinct regional accents. The same English dialogue can be spoken with British, Australian, or Indian English voices, each with its own pronunciation, rhythm, and intonation. You can use accented voices from the Soniox voice library or preserve a speaker’s accent through voice cloning. For global products, this means users can hear voices that sound closer to the speech they hear around them every day. Learn more:
Soniox316,644 views • 1 month ago

Meet the new Soniox voice library. Soniox Text-to-Speech now includes more than 200 built-in voices, carefully curated for different applications, styles, accents, ages, and voice characteristics. Every voice works across 60+ languages, so you can choose the voice that fits your product and use that same voice globally while preserving its identity and character. You can now browse and filter voices by: • Gender • Age • Accent • Use case, such as conversational, narration, or social media • Style, such as warm, confident, friendly, calm, energetic, or deep This makes it much easier to find a voice for a specific application, for example: • Conversational + warm + middle-aged • Narration + deep + confident • Social media + young + energetic • Educational + calm + friendly We manually selected, reviewed, and tagged the voices so the metadata reflects how the voices actually sound and where they work well. We also redesigned the TTS Playground to make it easier to browse, filter, compare, and test voices with your own text and languages. Learn more:
Soniox126,350 views • 22 days ago

Make your HeyGen avatars multilingual and more expressive with Soniox TTS v2. • Choose from 200+ voices over 60+ languages. • Instantly clone your voice. • Switch languages naturally mid-sentence • Pronounce names and technical terms correctly. Generate speech with Soniox, lip-sync with HeyGen. 🍪
Soniox30,909 views • 6 days ago

Soniox TTS pronounces critical information precisely across all 60+ supported languages. Alphanumerics, email addresses, phone numbers, street addresses, IDs, and other structured information are spoken clearly and accurately, whether in German, Japanese, or any other supported language. This level of precision is critical in real-world applications, where getting a single character, number, or symbol wrong can matter, especially in healthcare, finance, customer support, and voice agents. Try it out yourself:
Soniox216,710 views • 1 month ago

We added Fish Audio's S2.1 Pro and Inworld AI's TTS-2 to Soniox Compare TTS. Run Soniox, Fish, and Inworld side by side on the same prompt and judge for yourself. All three sound natural. The gaps show up on names, numbers, emails, mixed-language text, and the expressiveness of audio tags, the critical details that matter most in production speech. Fish: about $0.75/hour. Inworld: $1.25/hour. Soniox: $0.70/hour. The video below shows the same prompt generating audio on all three models side by side. Compare TTS providers yourself. Don't trust benchmarks or marketing:
Soniox43,396 views • 15 days ago
No more content to load