Loading video...

Video Failed to Load

Go Home

HOLY CRAP, a new super tiny 1.6B param voice model just dropped that seems to.. outperform 11labs!? 😵‍💫 From Nari-labs, Dia is an Apache 2.0 voice model, that can generate laughs, sniffs and emotions, copy an existing voice and is effectively real time on larger GPUs:

525,803 views • 1 year ago •via X (Twitter)

10 Comments

Alex Volkov (Thursd/AI)'s profile picture
Alex Volkov (Thursd/AI)1 year ago

We'll of course cover this on the next @thursdai_pod (thanks for the heads up @rodrimora !)

Alex Volkov (Thursd/AI)'s profile picture
Alex Volkov (Thursd/AI)1 year ago

Here's the links: Demo Page: Github: HF Model: Try It:

Prince Canuma's profile picture
Prince Canuma1 year ago

Coming to MLX Audio this week 🔥 @lllucas and I are on it 🚀

Alex Volkov (Thursd/AI)'s profile picture
Alex Volkov (Thursd/AI)1 year ago

@lllucas can't wait to try this on my mac!

Prince Canuma's profile picture
Prince Canuma1 year ago

The best sniff emotion I have heard so far Other models feel like a sigh

Alex Volkov (Thursd/AI)'s profile picture
Alex Volkov (Thursd/AI)1 year ago

(sniff) the laughter is so dope as well, it comes in juuust before and you can hear the smile on the invisible's AI face as it starts to render the laughter

Jonathan's profile picture
Jonathan1 year ago

Intonation and tone are great, but quality is nowhere near elevenlabs.

Loic's profile picture
Loic1 year ago

Honestly voice generation feels like it has been stuck at the same level for at least a year, not real improvement

Yng_Pepe's profile picture
Yng_Pepe1 year ago

@turing_hamster

Michael Finney's profile picture
Michael Finney1 year ago

This on Pinokio yet @cocktailpeanut ?

Related Videos

VoxCPM 2 just dropped by OpenBMB Only 2B-param open-source TTS (Text-to-Speech) model built for production-grade multilingual voice work. Apache-2.0 license, Can run on only 8GB VRAM. • Eliminates the "robotic" feel of traditional TTS, delivering prosody and emotional depth suitable for high-stakes professional environments like filmmaking, gaming, animation, and audiobooks. • 30-language multilingual: no language tag needed, just type in a supported language and generate directly. • Voice design: create a brand-new voice from a text description alone, like age, tone, pace, or emotion. No reference audio required. Describe the desired voice characteristics (gender, age, tone, emotion, pace …) in Control Instruction, and VoxCPM2 will craft a unique voice from your description alone. • Controllable cloning: clone from a short clip, then steer delivery style without losing the speaker’s core voice. • Ultimate cloning: use reference audio + transcript for continuation-style cloning that keeps the tiny vocal details. • 48kHz output: takes 16kHz reference audio and produces studio-quality speech without an external upsampler. • Real-time ready: around 0.3 RTF on RTX 4090, even lower with Nano-VLLM. • Commercial use: Apache-2.0 licensed. Developer-Friendly Infrastructure: - Native Torch Inference: Direct support for PyTorch-based workflows. - Training Flexibility: Supports both full-parameter and LoRA fine-tuning for specific domain adaptation. - Production Readiness: Compatible with voxcpm-nanovllm for large-scale, high-concurrency deployment.

Rohan Paul

13,541 views • 3 months ago