正在加载视频...

视频加载失败

Introducing Octave 2: our next-generation multilingual text-to-speech model What’s new: - Fluent in 11+ languages - 40% faster (<200ms latency⁠⁠) & 50% cheaper than Octave 1 - Multi-speaker conversation - More reliable pronunciation - New voice conversion & phoneme editing capabilities For the month of October, we’re offering 50%...

7,077,598 次观看 • 11 个月前 •via X (Twitter)

48 条评论

Hume AI 的头像
Hume AI11 个月前

Our new models speak over 11 languages, including English, Japanese 🇯🇵, Korean 🇰🇷, Spanish 🇪🇸, French 🇫🇷, Portuguese 🇵🇹, Italian 🇮🇹, German 🇩🇪, Russian 🇷🇺, Hindi 🇮🇳, and Arabic 🇦🇪. They’re faster (<200ms), more efficient, and more accurate. We are also rolling out a conversationally tuned model, EVI 4 mini, that is powered by the intelligence of any frontier model like Claude 4.5 Sonnet or GPT-5. Try our live translation demo & EVI 4 mini chat, built on a few voice samples and a prompt:

Hume AI 的头像
Hume AI11 个月前

We'll be releasing more evaluations, more languages, and access to voice conversion and phoneme editing soon. In the meantime, we’re excited to see what you create! Read more about product &amp; research in our blog ↓

Hume AI 的头像
Hume AI11 个月前

We’re also HIRING across roles! Apply below on our careers page

Umesh 的头像
Umesh11 个月前

Congratulations 👏 It's amazing!

Hume AI 的头像
Hume AI11 个月前

thanks for the support umesh!

proper 的头像
proper11 个月前

2.0!! congrats on the launch!

Kol Tregaskes 的头像
Kol Tregaskes11 个月前

Nice guys, thanks for the early access!

Hume AI 的头像
Hume AI11 个月前

would love to hear what you think Kol!

Alan Cowen 的头像
Alan Cowen11 个月前

日本語に最適な音声AI

Hume AI 的头像
Hume AI11 个月前

The land of @BIG_CLAPPY!!

mori 的头像
mori11 个月前

LFG!! ⭐️

mori 的头像
mori11 个月前

can’t wait to see what ppl build w thisss

Hume AI 的头像
Hume AI11 个月前

show us what you build!

Alec Freudenstein 的头像
Alec Freudenstein11 个月前

INCREDIBLE WORK TEAM!!!! SO PROUD OF EVERYONE WHO PUT SO MUCH TIME INTO THIS!!!! NOBODY SHIPS LIKE HUME!!!!!!!!

Janet Ho 的头像
Janet Ho11 个月前

Can’t wait to go to Japan with this

Linus ✦ Ekenstam 的头像
Linus ✦ Ekenstam11 个月前

Guys, thanks for giving me early access to this, been sharing with the family last couple of days, and it’s been a blast. Extremely bullish on the trajectory you are on, keep shipping, this upgrade is spot on. 🫡

Hume AI 的头像
Hume AI11 个月前

thanks Linus!!

Lucas | IA 的头像
Lucas | IA11 个月前

SO good

Hume AI 的头像
Hume AI11 个月前

thank you 😊

Alvaro Cintas 的头像
Alvaro Cintas11 个月前

SO good!

Miguel Ángel | GptZone 的头像
Miguel Ángel | GptZone11 个月前

I am just trying it right now and it seems spectacular!!

Robert Scoble 的头像
Robert Scoble11 个月前

Great getting to talk to your CEO about this launch:

Hume AI 的头像
Hume AI11 个月前

awesome conversation!

Abdul Shakoor 的头像
Abdul Shakoor11 个月前

octave 2 feels like a serious upgrade.

Hume AI 的头像
Hume AI11 个月前

so glad you think so!

cate bligh 的头像
cate bligh11 个月前

Great video! I'm looking forward to playing with this more!

Hume AI 的头像
Hume AI11 个月前

love the videos you create with us cate!

Dr. Apolotary 的头像
Dr. Apolotary11 个月前

Why is he talking with a strong English accent in Japanese

Hume AI 的头像
Hume AI11 个月前

because he’s English! our model preserves accents based on the language of the base voice

Adam 的头像
Adam11 个月前

Unbelievable! I'll talk about it soon 👀

Jack 的头像
Jack11 个月前

So amazing

Nelly; 的头像
Nelly;11 个月前

no excuses now.. I can actually sound fluent in 11 languages😂

Qing Liu 的头像
Qing Liu11 个月前

Soooo proud 🥺 This is really amazing!!

@Czr305 的头像
@Czr30511 个月前

What about for EVI?

Hume AI 的头像
Hume AI11 个月前

we have EVI4 mini also in preview

@Czr305 的头像
@Czr30511 个月前

How do I gain access ?

Eleftheria Batsou 的头像
Eleftheria Batsou11 个月前

Let me check that!!! Excited for all these features

ayn 的头像
ayn11 个月前

Congratulations HUME!! 🫶🏻

cova 的头像
cova11 个月前

Congratulations!! 🫶

WISe Social 的头像
WISe Social11 个月前

Exciting advancements with Octave 2. Its multi-speaker feature could really enhance AI agent interactions. Imagine seamless transactions through secure channels using satellite connectivity, perfect for a world where speed and reliability are key.

Luke Miler 的头像
Luke Miler11 个月前

sweet! will add to shortcut

Afiz ⚡️ 的头像
Afiz ⚡️11 个月前

multilingual text-to-speech. I would definitely want to give it a try.

Jameson McSmirnoff 🌊🛥 的头像
Jameson McSmirnoff 🌊🛥11 个月前

Deutsch ist euer Endgegner!

Roni Rahman 的头像
Roni Rahman11 个月前

Looks really good!

MetaDJ 的头像
MetaDJ11 个月前

Awesome! 🔥

Jussi 的头像
Jussi11 个月前

Will it support audio to audio? I am looking for a audio model to change the vocals of a song I generated on Suno to my own voice

Jaden Tripp 的头像
Jaden Tripp11 个月前

IPA input?

Rafael Estrela | IA 的头像
Rafael Estrela | IA11 个月前

Incrivel !

相关视频

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 次观看 • 1 年前

Voice used to be AI’s forgotten modality - now it's having its big moment: rapid innovation, big funding rounds, major agentic applications My conversation with Neil Zeghidour, top AI researcher in the field (Google DeepMind, Meta, kyutai) and now CEO of Gradium This is a reference episode on all things voice AI 🔥 00:00 Intro 01:21 Voice AI’s big moment, and why we’re still early 03:34 Why voice lagged behind text/image/video 06:06 The convergence era: transformers for every modality 07:40 Beyond Her: always-on assistants, wake words, voice-first devices 11:01 Voice vs text: where voice fits (even for coding) 12:56 Neil’s origin story: from finance to machine learning, with help from Yann LeCun and Soumith Chintala 18:35 Neural codecs (SoundStream): compression as the unlock 22:30 Kyutai: open research, small elite teams, moving fast 31:32 Why big labs haven’t “won” voice AI4 34:01 On-device voice: where it works, why compact models matter 46:37 The last mile: real-world robustness, pronunciation, uptime 41:35 Benchmarking voice: why metrics fail, how they actually test 47:03 Cascades vs speech-to-speech: trade-offs + what’s next 54:05 Hardest frontier: noisy rooms, factories, multi-speaker chaos 1:00:50 New languages + dialects: what transfers, what doesn’t 1:02:54 Hardware & compute: why voice isn’t a 10,000-GPU game 1:07:27 What data do you need to train voice models 1:09:02 Deepfakes + privacy: why watermarking isn’t a solution 1:12:30 Voice + vision: multimodality, screen awareness, video+audio 1:14:43 Voice cloning vs voice design: where the market goes 1:16:32 Paris/Europe AI: talent density, underdog energy, what’s next

Matt Turck

22,980 次观看 • 6 个月前

Open science is how we continue to push technology forward and today at Meta FAIR we’re sharing eight new AI research artifacts including new models, datasets and code to inspire innovation in the community. More in the video from Joelle Pineau. This work is another important step towards our goal of achieving Advanced Machine Intelligence (AMI). What we’re releasing: • Meta Spirit LM: An open source language model for seamless speech and text integration. • Meta Segment Anything Model 2.1: An updated checkpoint with improved results on visually similar objects, small objects and occlusion handling. Plus a new developer suite to make it easier for developers to build with SAM 2. • Layer Skip: Inference code and fine-tuned checkpoints demonstrating a new method for enhancing LLM performance. • SALSA: New code to enable researchers to benchmark AI-based attacks in support of validating security for post-quantum cryptography. • Meta Lingua: A lightweight and self-contained codebase designed to train language models at scale. • Meta Open Materials: New open source models and the largest dataset of its kind to accelerate AI-driven discovery of new inorganic materials. • MEXMA: A new research paper and code for our novel pre-trained cross-lingual sentence encoder with coverage across 80 languages. • Self-Taught Evaluator: a new method for generating synthetic preference data to train reward models without relying on human annotations. Access to state-of-the-art AI creates opportunities for everyone. We’re excited to share this work and look forward to seeing the community innovation that results from it. Details and access to everything released by FAIR today ➡️

AI at Meta

150,477 次观看 • 1 年前