Loading video...

Video Failed to Load

Go Home

Today, we’re releasing Octave: the first LLM built for text-to-speech. 🎨Design any voice with a prompt 🎬 Give acting instructions to control emotion and delivery (sarcasm, whispering, etc.) 🛠️Produce long-form content on our Creator Studio Unlike traditional TTS that just “reads” words aloud, Octave understands how meaning affects delivery...

394,020 views • 1 year ago •via X (Twitter)

12 Comments

Hume's profile picture
Hume1 year ago

🎨Voice Design Create any AI voice with a simple prompt. From "Southern ASMR meditation coach" to "film noir detective" — Octave instantly generates the voices you need for your content. Octave beats ElevenLabs Voice Design in a robust eval. Find out more👇

Hume's profile picture
Hume1 year ago

🎬Acting Instructions Octave is the first TTS system that can take natural language instructions to change emotional delivery and speaking style. Give directions like "sound sarcastic" or "whisper fearfully." For the first time, creators have total control.

Hume's profile picture
Hume1 year ago

🤔Context-Aware Expression Trained on 1000x more language than traditional TTS, Octave understands your script like a human actor, delivering realistic emotions, sarcasm, pace, word emphasis, and more. This allows it to understand plot twists, emotional cues, character traits, and how to combine them—it reads love letters tenderly and sports announcements energetically.

Hume's profile picture
Hume1 year ago

🛠️Tools for Creators and Developers Use our creator studio to edit and generate long-form content precisely with acting instructions. For developers, Octave is accessible through Python and TypeScript SDKs that handle authentication and provide typed interfaces for reliable integration.

Hume's profile picture
Hume1 year ago

Octave’s first-of-its-kind voice intelligence speaks for itself 😉 In a blind study, Octave outperformed ElevenLabs Voice Design: 🔊71.6% preferred Octave’s audio quality 🗣️51.7% found Octave more natural 🎯57.7% said Octave better matched voice descriptions

Hume's profile picture
Hume1 year ago

Create with Octave today: And the best part? Even with superior capabilities, Octave is cheaper than alternatives. Our blog shares more about Octave, our blind study, and what’s next:

Hume's profile picture
Hume1 year ago

Join our Discord to show off your projects, share product feedback, get technical support, or just hang with our empathic community. We are so excited to see what you create!

TurboScribe's profile picture
TurboScribe1 year ago

What REAL customers say about TurboScribe's unlimited AI transcription 👇 🇺🇸 "Not only does it transcribe with amazing accuracy, it also filters out a ton of the unnecessary noise associated with pauses in audio. Keep up the great work!" - Kevin

Moses Oh's profile picture
Moses Oh1 year ago

Wait … are all the voices in this video AI generated?

Hume's profile picture
Hume1 year ago

R u serious Moses

Rohan Paul's profile picture
Rohan Paul1 year ago

This is superb.. love it.. Characters finally speak in context without heavy manual adjustment.

Hume's profile picture
Hume1 year ago

thank you!! show us what you create :)

Related Videos

VoxCPM 2 just dropped by OpenBMB Only 2B-param open-source TTS (Text-to-Speech) model built for production-grade multilingual voice work. Apache-2.0 license, Can run on only 8GB VRAM. • Eliminates the "robotic" feel of traditional TTS, delivering prosody and emotional depth suitable for high-stakes professional environments like filmmaking, gaming, animation, and audiobooks. • 30-language multilingual: no language tag needed, just type in a supported language and generate directly. • Voice design: create a brand-new voice from a text description alone, like age, tone, pace, or emotion. No reference audio required. Describe the desired voice characteristics (gender, age, tone, emotion, pace …) in Control Instruction, and VoxCPM2 will craft a unique voice from your description alone. • Controllable cloning: clone from a short clip, then steer delivery style without losing the speaker’s core voice. • Ultimate cloning: use reference audio + transcript for continuation-style cloning that keeps the tiny vocal details. • 48kHz output: takes 16kHz reference audio and produces studio-quality speech without an external upsampler. • Real-time ready: around 0.3 RTF on RTX 4090, even lower with Nano-VLLM. • Commercial use: Apache-2.0 licensed. Developer-Friendly Infrastructure: - Native Torch Inference: Direct support for PyTorch-based workflows. - Training Flexibility: Supports both full-parameter and LoRA fine-tuning for specific domain adaptation. - Production Readiness: Compatible with voxcpm-nanovllm for large-scale, high-concurrency deployment.

Rohan Paul

13,541 views • 4 months ago

Brain signals and LLM embeddings converge for predicting every spoken or heard word. Beautiful research from Google AI They compared human brain activity during real conversations with internal embeddings from a speech-to-text LLM. Measured electrode signals in speech and language-related brain regions and matched them to the model’s word-level features. 🤖Key Highlights → Brain activity aligns linearly with LLM embeddings for real-life spoken conversations. → Sequence of comprehension: first speech sounds, then word meaning. → Sequence of production: planned meaning, then articulation, then hearing one’s own voice. → Consistent predictive coding (pre-onset anticipation, post-onset surprise) mirrors LLM next-word prediction. → Lower-tier auditory regions still show partial sensitivity to semantic information. 🤖 Model-Brain Alignment They observed a clear sequence: during comprehension, auditory cortex (superior temporal gyrus) showed strong correlation with speech embeddings, then language embeddings aligned with Broca’s area. During production, Broca’s area correlated with language embeddings before articulation, followed by motor cortex signals matching speech embeddings. This suggests that next-word prediction and higher-level meaning representation in the model parallel the brain’s approach. ⚙ So the study revealed a shared computational principle of predicting words in context. Even though the Transformer-based LLM processes words in parallel layers, the human brain processes them serially yet mirrors similar statistical regularities. This supports a “soft hierarchy” where both lower-level acoustic processing and higher-level semantic processing partially overlap in the brain.

Rohan Paul

15,213 views • 1 year ago