Loading video...

Video Failed to Load

Go Home

Introducing Octave 2: our next-generation multilingual text-to-speech model What’s new: - Fluent in 11+ languages - 40% faster (<200ms latency⁠⁠) & 50% cheaper than Octave 1 - Multi-speaker conversation - More reliable pronunciation - New voice conversion & phoneme editing capabilities For the month of October, we’re offering 50%...

7,077,598 views • 11 months ago •via X (Twitter)

48 Comments

Hume AI's profile picture
Hume AI11 months ago

Our new models speak over 11 languages, including English, Japanese 🇯🇵, Korean 🇰🇷, Spanish 🇪🇸, French 🇫🇷, Portuguese 🇵🇹, Italian 🇮🇹, German 🇩🇪, Russian 🇷🇺, Hindi 🇮🇳, and Arabic 🇦🇪. They’re faster (<200ms), more efficient, and more accurate. We are also rolling out a conversationally tuned model, EVI 4 mini, that is powered by the intelligence of any frontier model like Claude 4.5 Sonnet or GPT-5. Try our live translation demo & EVI 4 mini chat, built on a few voice samples and a prompt:

Hume AI's profile picture
Hume AI11 months ago

We'll be releasing more evaluations, more languages, and access to voice conversion and phoneme editing soon. In the meantime, we’re excited to see what you create! Read more about product &amp; research in our blog ↓

Hume AI's profile picture
Hume AI11 months ago

We’re also HIRING across roles! Apply below on our careers page

Umesh's profile picture
Umesh11 months ago

Congratulations 👏 It's amazing!

Hume AI's profile picture
Hume AI11 months ago

thanks for the support umesh!

proper's profile picture
proper11 months ago

2.0!! congrats on the launch!

Kol Tregaskes's profile picture
Kol Tregaskes11 months ago

Nice guys, thanks for the early access!

Hume AI's profile picture
Hume AI11 months ago

would love to hear what you think Kol!

Alan Cowen's profile picture
Alan Cowen11 months ago

日本語に最適な音声AI

Hume AI's profile picture
Hume AI11 months ago

The land of @BIG_CLAPPY!!

mori's profile picture
mori11 months ago

LFG!! ⭐️

mori's profile picture
mori11 months ago

can’t wait to see what ppl build w thisss

Hume AI's profile picture
Hume AI11 months ago

show us what you build!

Alec Freudenstein's profile picture
Alec Freudenstein11 months ago

INCREDIBLE WORK TEAM!!!! SO PROUD OF EVERYONE WHO PUT SO MUCH TIME INTO THIS!!!! NOBODY SHIPS LIKE HUME!!!!!!!!

Janet Ho's profile picture
Janet Ho11 months ago

Can’t wait to go to Japan with this

Linus ✦ Ekenstam's profile picture
Linus ✦ Ekenstam11 months ago

Guys, thanks for giving me early access to this, been sharing with the family last couple of days, and it’s been a blast. Extremely bullish on the trajectory you are on, keep shipping, this upgrade is spot on. 🫡

Hume AI's profile picture
Hume AI11 months ago

thanks Linus!!

Lucas | IA's profile picture
Lucas | IA11 months ago

SO good

Hume AI's profile picture
Hume AI11 months ago

thank you 😊

Alvaro Cintas's profile picture
Alvaro Cintas11 months ago

SO good!

Miguel Ángel | GptZone's profile picture
Miguel Ángel | GptZone11 months ago

I am just trying it right now and it seems spectacular!!

Robert Scoble's profile picture
Robert Scoble11 months ago

Great getting to talk to your CEO about this launch:

Hume AI's profile picture
Hume AI11 months ago

awesome conversation!

Abdul Shakoor's profile picture
Abdul Shakoor11 months ago

octave 2 feels like a serious upgrade.

Hume AI's profile picture
Hume AI11 months ago

so glad you think so!

cate bligh's profile picture
cate bligh11 months ago

Great video! I'm looking forward to playing with this more!

Hume AI's profile picture
Hume AI11 months ago

love the videos you create with us cate!

Dr. Apolotary's profile picture
Dr. Apolotary11 months ago

Why is he talking with a strong English accent in Japanese

Hume AI's profile picture
Hume AI11 months ago

because he’s English! our model preserves accents based on the language of the base voice

Adam's profile picture
Adam11 months ago

Unbelievable! I'll talk about it soon 👀

Jack's profile picture
Jack11 months ago

So amazing

Nelly;'s profile picture
Nelly;11 months ago

no excuses now.. I can actually sound fluent in 11 languages😂

Qing Liu's profile picture
Qing Liu11 months ago

Soooo proud 🥺 This is really amazing!!

@Czr305's profile picture
@Czr30511 months ago

What about for EVI?

Hume AI's profile picture
Hume AI11 months ago

we have EVI4 mini also in preview

@Czr305's profile picture
@Czr30511 months ago

How do I gain access ?

Eleftheria Batsou's profile picture
Eleftheria Batsou11 months ago

Let me check that!!! Excited for all these features

ayn's profile picture
ayn11 months ago

Congratulations HUME!! 🫶🏻

cova's profile picture
cova11 months ago

Congratulations!! 🫶

WISe Social's profile picture
WISe Social11 months ago

Exciting advancements with Octave 2. Its multi-speaker feature could really enhance AI agent interactions. Imagine seamless transactions through secure channels using satellite connectivity, perfect for a world where speed and reliability are key.

Luke Miler's profile picture
Luke Miler11 months ago

sweet! will add to shortcut

Afiz ⚡️'s profile picture
Afiz ⚡️11 months ago

multilingual text-to-speech. I would definitely want to give it a try.

Jameson McSmirnoff 🌊🛥's profile picture
Jameson McSmirnoff 🌊🛥11 months ago

Deutsch ist euer Endgegner!

Roni Rahman's profile picture
Roni Rahman11 months ago

Looks really good!

MetaDJ's profile picture
MetaDJ11 months ago

Awesome! 🔥

Jussi's profile picture
Jussi11 months ago

Will it support audio to audio? I am looking for a audio model to change the vocals of a song I generated on Suno to my own voice

Jaden Tripp's profile picture
Jaden Tripp11 months ago

IPA input?

Rafael Estrela | IA's profile picture
Rafael Estrela | IA11 months ago

Incrivel !

Related Videos

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 views • 1 year ago

Voice used to be AI’s forgotten modality - now it's having its big moment: rapid innovation, big funding rounds, major agentic applications My conversation with Neil Zeghidour, top AI researcher in the field (Google DeepMind, Meta, kyutai) and now CEO of Gradium This is a reference episode on all things voice AI 🔥 00:00 Intro 01:21 Voice AI’s big moment, and why we’re still early 03:34 Why voice lagged behind text/image/video 06:06 The convergence era: transformers for every modality 07:40 Beyond Her: always-on assistants, wake words, voice-first devices 11:01 Voice vs text: where voice fits (even for coding) 12:56 Neil’s origin story: from finance to machine learning, with help from Yann LeCun and Soumith Chintala 18:35 Neural codecs (SoundStream): compression as the unlock 22:30 Kyutai: open research, small elite teams, moving fast 31:32 Why big labs haven’t “won” voice AI4 34:01 On-device voice: where it works, why compact models matter 46:37 The last mile: real-world robustness, pronunciation, uptime 41:35 Benchmarking voice: why metrics fail, how they actually test 47:03 Cascades vs speech-to-speech: trade-offs + what’s next 54:05 Hardest frontier: noisy rooms, factories, multi-speaker chaos 1:00:50 New languages + dialects: what transfers, what doesn’t 1:02:54 Hardware & compute: why voice isn’t a 10,000-GPU game 1:07:27 What data do you need to train voice models 1:09:02 Deepfakes + privacy: why watermarking isn’t a solution 1:12:30 Voice + vision: multimodality, screen awareness, video+audio 1:14:43 Voice cloning vs voice design: where the market goes 1:16:32 Paris/Europe AI: talent density, underdog energy, what’s next

Matt Turck

22,980 views • 6 months ago

Open science is how we continue to push technology forward and today at Meta FAIR we’re sharing eight new AI research artifacts including new models, datasets and code to inspire innovation in the community. More in the video from Joelle Pineau. This work is another important step towards our goal of achieving Advanced Machine Intelligence (AMI). What we’re releasing: • Meta Spirit LM: An open source language model for seamless speech and text integration. • Meta Segment Anything Model 2.1: An updated checkpoint with improved results on visually similar objects, small objects and occlusion handling. Plus a new developer suite to make it easier for developers to build with SAM 2. • Layer Skip: Inference code and fine-tuned checkpoints demonstrating a new method for enhancing LLM performance. • SALSA: New code to enable researchers to benchmark AI-based attacks in support of validating security for post-quantum cryptography. • Meta Lingua: A lightweight and self-contained codebase designed to train language models at scale. • Meta Open Materials: New open source models and the largest dataset of its kind to accelerate AI-driven discovery of new inorganic materials. • MEXMA: A new research paper and code for our novel pre-trained cross-lingual sentence encoder with coverage across 80 languages. • Self-Taught Evaluator: a new method for generating synthetic preference data to train reward models without relying on human annotations. Access to state-of-the-art AI creates opportunities for everyone. We’re excited to share this work and look forward to seeing the community innovation that results from it. Details and access to everything released by FAIR today ➡️

AI at Meta

150,477 views • 1 year ago