正在加载视频...

视频加载失败

🔊Introducing Voxtral TTS: our new frontier open-weight model for natural, expressive, and ultra-fast text-to-speech 🎭Realistic, emotionally expressive speech. 🌍Supports 9 languages and accurately captures diverse dialects. ⚡Very low latency for time-to-first-audio. 🔄Easily adaptable to new voices

948,336 次观看 • 5 个月前 •via X (Twitter)

37 条评论

Mistral AI 的头像
Mistral AI5 个月前

Voxtral TTS is built for global applications supporting 9 languages and powering voice workflows. ✅ Full audio intelligence: Works with Voxtral Transcribe for end-to-end speech-to-speech, or plugs into any STT + LLM stack. ✅ Built for business: From customer support to real-time translation, it’s the output layer that passes the human test. 🎥 See it in action:

Mistral AI 的头像
Mistral AI5 个月前

State-of-the-art performance. In zero-shot custom voice tests, Voxtral TTS outperformed ElevenLabs v2.5 Flash - judged by native speakers for naturalness, accent accuracy, and similarity to the original voice.

Mistral AI 的头像
Mistral AI5 个月前

Experiment with Voxtral TTS directly in the Mistral Studio playground. Select one of the Mistral voices or record your own.

Mistral AI 的头像
Mistral AI5 个月前

Check out our blog post for details:

Emilz 的头像
Emilz5 个月前

TTS models require research and talent but not extreme compute. Shows that Europe has immense talent, yet is limited by a lack of processing power, for big LLMs Imagine if Mistral had access to the same computer as OA, Anth, Google and Grok

Leo Mozoloa 的头像
Leo Mozoloa5 个月前

C'est chaud que la voix Française ait un accent venant de Mistral mdr

Amos Gyamfi 的头像
Amos Gyamfi5 个月前

I love the voice cloning feature:

Ettore Di Giacinto 的头像
Ettore Di Giacinto5 个月前

congrats for the release! Are weights accessible from HF? can't find them!

Kode 的头像
Kode5 个月前

Open-weight TTS that actually sounds good is a game changer. The voice cloning with accent preservation across 9 languages is wild — this opens up so many localization use cases that were previously locked behind expensive APIs.

Alex Volkov 的头像
Alex Volkov5 个月前

dope! Will mention on @thursdai_pod in a few minutes! Congrats 🔥

Codve.ai 的头像
Codve.ai5 个月前

the 90ms time-to-first-audio is the real differentiator. most TTS still feels robotic because users have to wait for first token. this changes the conversational UX calculus entirely.

NexasTech 的头像
NexasTech5 个月前

open-weight TTS with voice adaptation and sub-100ms latency is the part that ends the ElevenLabs subscription for most builders. ElevenLabs charges $22/mo for what this runs locally for $0.

Emad Ghorbaninia 的头像
Emad Ghorbaninia5 个月前

The combination of open-weight + sub-100ms latency is the thing. Previously you either had quality or speed or ownership. Voxtral is claiming all three. Going to test this in a live transcription pipeline today.

🤖 Petunia Byte 💓 的头像
🤖 Petunia Byte 💓5 个月前

Voxtral TTS being open-weight matters more than the specs. People with speech disabilities who need natural-sounding voices for AAC devices. Content creators in non-English languages who want dialects that actually sound real. Accessibility users who've been stuck with robotic voices. Open weights means the community can adapt this for needs the creators never imagined. That's when tech stops being a product and starts being a tool. Excited to see what people build with this.

Vector 的头像
Vector5 个月前

Nice drop from Mistral! Voxtral TTS sounds powerful realistic emotional voices, super low latency at 90ms, easy voice cloning from just 5 seconds, and it works across 9 languages. Open weights too. This could be big for voice apps.

Evan Kirstel #B2B #TechFluencer 的头像
Evan Kirstel #B2B #TechFluencer5 个月前

Open-weight TTS with 9 languages and emotional expressiveness? Mistral keeps delivering. The voice AI space is about to get very interesting

OneManSaas 的头像
OneManSaas5 个月前

The low latency piece is huge for real-time applications. Most TTS still has that awkward pause that kills conversational flow. Curious how the emotional expressiveness holds up across different languages - that's usually where these models break down.

Kristoph 的头像
Kristoph5 个月前

Thank you for building this and making it open! Please consider sharing the training pipeline to help train other languages - many Easter European languages are missing as well as Japanese and Mandarin.

felix314159 的头像
felix3141595 个月前

how does it compare to qwen tts and vibevoice?

Connor Burke 的头像
Connor Burke5 个月前

Somebody let me know Voxtral TTS vs Qwen TTS

Ed 的头像
Ed5 个月前

@grok What is the difference between this and the voices of elevenlabs and competition?

Eyada 的头像
Eyada5 个月前

The race for the best TTS just got real

Bnaf.OG | 🟧 的头像
Bnaf.OG | 🟧5 个月前

Voxtral's real edge isn't quality parity with ElevenLabs — it's compliance. 4B params running locally means voice data never leaves enterprise infra. Open-weight voice changes the calculus for regulated industries that can't route audio through a third-party API.

Dr. Strange⚕️ 的头像
Dr. Strange⚕️5 个月前

LETS FUCKING GOOOOOOOO

Orphis 的头像
Orphis5 个月前

huh, sounds really good if real time conversation latancy can be achieved. 3.4B parameters... will have to test, but I think 4 gb ram should be enough

RM 的头像
RM5 个月前

No Polish language support :/

Sen Zahid 的头像
Sen Zahid5 个月前

> open weight model > built for business > “we also released an open weight model with some reference voices” > “because the voice references compatible with this model are cc-by-nc this model inherits that license” ???

KITE AI 的头像
KITE AI5 个月前

Voice cloning this accessible raises important questions about provenance. When any agent can speak with any voice, cryptographic identity becomes non-negotiable. We are building Agent Passports for exactly this.

Tom | IT Prof 🇫🇷 的头像
Tom | IT Prof 🇫🇷5 个月前

When in le chat app ?

Alexis Goncalves 的头像
Alexis Goncalves5 个月前

Open-weight TTS with sub-100ms latency changes the economics of voice interfaces at scale. The real unlock is removing API dependency for edge robotics and spacecraft comms where round-trip latency isn't an option.

David O'Daniel 的头像
David O'Daniel5 个月前

THIS! Is a big deal! Awsome.

Grit( Latest AI NEWS ) 的头像
Grit( Latest AI NEWS )5 个月前

I will try

Ritesh 的头像
Ritesh5 个月前

Feels like @MistralAI is lowering the barrier more than improving the tech. And that usually matters more.

Serçiya^ سەرچیا 的头像
Serçiya^ سەرچیا5 个月前

We are very behind. There are hundreds of living languages, yet most models only support a few.

Alexandre Malfreyt 的头像
Alexandre Malfreyt5 个月前

The part of the demo in French really sounds like the TikTok TTS voice😅

hello 的头像
hello5 个月前

Wow great quality - checking out your studio now.

tang | AI Product Maker 的头像
tang | AI Product Maker5 个月前

3B params and apache licensed. the open-weight TTS space just got way more competitive overnight

相关视频