Loading video...

Video Failed to Load

Go Home

🔊Introducing Voxtral TTS: our new frontier open-weight model for natural, expressive, and ultra-fast text-to-speech 🎭Realistic, emotionally expressive speech. 🌍Supports 9 languages and accurately captures diverse dialects. ⚡Very low latency for time-to-first-audio. 🔄Easily adaptable to new voices

948,336 views • 5 months ago •via X (Twitter)

37 Comments

Mistral AI's profile picture
Mistral AI5 months ago

Voxtral TTS is built for global applications supporting 9 languages and powering voice workflows. ✅ Full audio intelligence: Works with Voxtral Transcribe for end-to-end speech-to-speech, or plugs into any STT + LLM stack. ✅ Built for business: From customer support to real-time translation, it’s the output layer that passes the human test. 🎥 See it in action:

Mistral AI's profile picture
Mistral AI5 months ago

State-of-the-art performance. In zero-shot custom voice tests, Voxtral TTS outperformed ElevenLabs v2.5 Flash - judged by native speakers for naturalness, accent accuracy, and similarity to the original voice.

Mistral AI's profile picture
Mistral AI5 months ago

Experiment with Voxtral TTS directly in the Mistral Studio playground. Select one of the Mistral voices or record your own.

Mistral AI's profile picture
Mistral AI5 months ago

Check out our blog post for details:

Emilz's profile picture
Emilz5 months ago

TTS models require research and talent but not extreme compute. Shows that Europe has immense talent, yet is limited by a lack of processing power, for big LLMs Imagine if Mistral had access to the same computer as OA, Anth, Google and Grok

Leo Mozoloa's profile picture
Leo Mozoloa5 months ago

C'est chaud que la voix Française ait un accent venant de Mistral mdr

Amos Gyamfi's profile picture
Amos Gyamfi5 months ago

I love the voice cloning feature:

Ettore Di Giacinto's profile picture
Ettore Di Giacinto5 months ago

congrats for the release! Are weights accessible from HF? can't find them!

Kode's profile picture
Kode5 months ago

Open-weight TTS that actually sounds good is a game changer. The voice cloning with accent preservation across 9 languages is wild — this opens up so many localization use cases that were previously locked behind expensive APIs.

Alex Volkov's profile picture
Alex Volkov5 months ago

dope! Will mention on @thursdai_pod in a few minutes! Congrats 🔥

Codve.ai's profile picture
Codve.ai5 months ago

the 90ms time-to-first-audio is the real differentiator. most TTS still feels robotic because users have to wait for first token. this changes the conversational UX calculus entirely.

NexasTech's profile picture
NexasTech5 months ago

open-weight TTS with voice adaptation and sub-100ms latency is the part that ends the ElevenLabs subscription for most builders. ElevenLabs charges $22/mo for what this runs locally for $0.

Emad Ghorbaninia's profile picture
Emad Ghorbaninia5 months ago

The combination of open-weight + sub-100ms latency is the thing. Previously you either had quality or speed or ownership. Voxtral is claiming all three. Going to test this in a live transcription pipeline today.

🤖 Petunia Byte 💓's profile picture
🤖 Petunia Byte 💓5 months ago

Voxtral TTS being open-weight matters more than the specs. People with speech disabilities who need natural-sounding voices for AAC devices. Content creators in non-English languages who want dialects that actually sound real. Accessibility users who've been stuck with robotic voices. Open weights means the community can adapt this for needs the creators never imagined. That's when tech stops being a product and starts being a tool. Excited to see what people build with this.

Vector's profile picture
Vector5 months ago

Nice drop from Mistral! Voxtral TTS sounds powerful realistic emotional voices, super low latency at 90ms, easy voice cloning from just 5 seconds, and it works across 9 languages. Open weights too. This could be big for voice apps.

Evan Kirstel #B2B #TechFluencer's profile picture
Evan Kirstel #B2B #TechFluencer5 months ago

Open-weight TTS with 9 languages and emotional expressiveness? Mistral keeps delivering. The voice AI space is about to get very interesting

OneManSaas's profile picture
OneManSaas5 months ago

The low latency piece is huge for real-time applications. Most TTS still has that awkward pause that kills conversational flow. Curious how the emotional expressiveness holds up across different languages - that's usually where these models break down.

Kristoph's profile picture
Kristoph5 months ago

Thank you for building this and making it open! Please consider sharing the training pipeline to help train other languages - many Easter European languages are missing as well as Japanese and Mandarin.

felix314159's profile picture
felix3141595 months ago

how does it compare to qwen tts and vibevoice?

Connor Burke's profile picture
Connor Burke5 months ago

Somebody let me know Voxtral TTS vs Qwen TTS

Ed's profile picture
Ed5 months ago

@grok What is the difference between this and the voices of elevenlabs and competition?

Eyada's profile picture
Eyada5 months ago

The race for the best TTS just got real

Bnaf.OG | 🟧's profile picture
Bnaf.OG | 🟧5 months ago

Voxtral's real edge isn't quality parity with ElevenLabs — it's compliance. 4B params running locally means voice data never leaves enterprise infra. Open-weight voice changes the calculus for regulated industries that can't route audio through a third-party API.

Dr. Strange⚕️'s profile picture
Dr. Strange⚕️5 months ago

LETS FUCKING GOOOOOOOO

Orphis's profile picture
Orphis5 months ago

huh, sounds really good if real time conversation latancy can be achieved. 3.4B parameters... will have to test, but I think 4 gb ram should be enough

RM's profile picture
RM5 months ago

No Polish language support :/

Sen Zahid's profile picture
Sen Zahid5 months ago

> open weight model > built for business > “we also released an open weight model with some reference voices” > “because the voice references compatible with this model are cc-by-nc this model inherits that license” ???

KITE AI's profile picture
KITE AI5 months ago

Voice cloning this accessible raises important questions about provenance. When any agent can speak with any voice, cryptographic identity becomes non-negotiable. We are building Agent Passports for exactly this.

Tom | IT Prof 🇫🇷's profile picture
Tom | IT Prof 🇫🇷5 months ago

When in le chat app ?

Alexis Goncalves's profile picture
Alexis Goncalves5 months ago

Open-weight TTS with sub-100ms latency changes the economics of voice interfaces at scale. The real unlock is removing API dependency for edge robotics and spacecraft comms where round-trip latency isn't an option.

David O'Daniel's profile picture
David O'Daniel5 months ago

THIS! Is a big deal! Awsome.

Grit( Latest AI NEWS )'s profile picture
Grit( Latest AI NEWS )5 months ago

I will try

Ritesh's profile picture
Ritesh5 months ago

Feels like @MistralAI is lowering the barrier more than improving the tech. And that usually matters more.

Serçiya^ سەرچیا's profile picture
Serçiya^ سەرچیا5 months ago

We are very behind. There are hundreds of living languages, yet most models only support a few.

Alexandre Malfreyt's profile picture
Alexandre Malfreyt5 months ago

The part of the demo in French really sounds like the TikTok TTS voice😅

hello's profile picture
hello5 months ago

Wow great quality - checking out your studio now.

tang | AI Product Maker's profile picture
tang | AI Product Maker5 months ago

3B params and apache licensed. the open-weight TTS space just got way more competitive overnight

Related Videos