Loading video...
Video Failed to Load
🔊Introducing Voxtral TTS: our new frontier open-weight model for natural, expressive, and ultra-fast text-to-speech 🎭Realistic, emotionally expressive speech. 🌍Supports 9 languages and accurately captures diverse dialects. ⚡Very low latency for time-to-first-audio. 🔄Easily adaptable to new voices
948,336 views • 5 months ago •via X (Twitter)
37 Comments

Voxtral TTS is built for global applications supporting 9 languages and powering voice workflows. ✅ Full audio intelligence: Works with Voxtral Transcribe for end-to-end speech-to-speech, or plugs into any STT + LLM stack. ✅ Built for business: From customer support to real-time translation, it’s the output layer that passes the human test. 🎥 See it in action:

State-of-the-art performance. In zero-shot custom voice tests, Voxtral TTS outperformed ElevenLabs v2.5 Flash - judged by native speakers for naturalness, accent accuracy, and similarity to the original voice.

Experiment with Voxtral TTS directly in the Mistral Studio playground. Select one of the Mistral voices or record your own.

Check out our blog post for details:

TTS models require research and talent but not extreme compute. Shows that Europe has immense talent, yet is limited by a lack of processing power, for big LLMs Imagine if Mistral had access to the same computer as OA, Anth, Google and Grok

C'est chaud que la voix Française ait un accent venant de Mistral mdr

I love the voice cloning feature:

congrats for the release! Are weights accessible from HF? can't find them!

Open-weight TTS that actually sounds good is a game changer. The voice cloning with accent preservation across 9 languages is wild — this opens up so many localization use cases that were previously locked behind expensive APIs.

dope! Will mention on @thursdai_pod in a few minutes! Congrats 🔥

the 90ms time-to-first-audio is the real differentiator. most TTS still feels robotic because users have to wait for first token. this changes the conversational UX calculus entirely.

open-weight TTS with voice adaptation and sub-100ms latency is the part that ends the ElevenLabs subscription for most builders. ElevenLabs charges $22/mo for what this runs locally for $0.

The combination of open-weight + sub-100ms latency is the thing. Previously you either had quality or speed or ownership. Voxtral is claiming all three. Going to test this in a live transcription pipeline today.

Voxtral TTS being open-weight matters more than the specs. People with speech disabilities who need natural-sounding voices for AAC devices. Content creators in non-English languages who want dialects that actually sound real. Accessibility users who've been stuck with robotic voices. Open weights means the community can adapt this for needs the creators never imagined. That's when tech stops being a product and starts being a tool. Excited to see what people build with this.

Nice drop from Mistral! Voxtral TTS sounds powerful realistic emotional voices, super low latency at 90ms, easy voice cloning from just 5 seconds, and it works across 9 languages. Open weights too. This could be big for voice apps.

Open-weight TTS with 9 languages and emotional expressiveness? Mistral keeps delivering. The voice AI space is about to get very interesting

The low latency piece is huge for real-time applications. Most TTS still has that awkward pause that kills conversational flow. Curious how the emotional expressiveness holds up across different languages - that's usually where these models break down.

Thank you for building this and making it open! Please consider sharing the training pipeline to help train other languages - many Easter European languages are missing as well as Japanese and Mandarin.

how does it compare to qwen tts and vibevoice?

Somebody let me know Voxtral TTS vs Qwen TTS

@grok What is the difference between this and the voices of elevenlabs and competition?

The race for the best TTS just got real

Voxtral's real edge isn't quality parity with ElevenLabs — it's compliance. 4B params running locally means voice data never leaves enterprise infra. Open-weight voice changes the calculus for regulated industries that can't route audio through a third-party API.

LETS FUCKING GOOOOOOOO

huh, sounds really good if real time conversation latancy can be achieved. 3.4B parameters... will have to test, but I think 4 gb ram should be enough

No Polish language support :/

> open weight model > built for business > “we also released an open weight model with some reference voices” > “because the voice references compatible with this model are cc-by-nc this model inherits that license” ???

Voice cloning this accessible raises important questions about provenance. When any agent can speak with any voice, cryptographic identity becomes non-negotiable. We are building Agent Passports for exactly this.

When in le chat app ?

Open-weight TTS with sub-100ms latency changes the economics of voice interfaces at scale. The real unlock is removing API dependency for edge robotics and spacecraft comms where round-trip latency isn't an option.

THIS! Is a big deal! Awsome.

I will try

Feels like @MistralAI is lowering the barrier more than improving the tech. And that usually matters more.

We are very behind. There are hundreds of living languages, yet most models only support a few.

The part of the demo in French really sounds like the TikTok TTS voice😅

Wow great quality - checking out your studio now.

3B params and apache licensed. the open-weight TTS space just got way more competitive overnight

