Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing Octave 2: our next-generation multilingual text-to-speech model What’s new: - Fluent in 11+ languages - 40% faster (<200ms latency⁠⁠) & 50% cheaper than Octave 1 - Multi-speaker conversation - More reliable pronunciation - New voice conversion & phoneme editing capabilities For the month of October, we’re offering 50%...

7,077,598 görüntüleme • 11 ay önce •via X (Twitter)

48 Yorum

Hume AI profil fotoğrafı
Hume AI11 ay önce

Our new models speak over 11 languages, including English, Japanese 🇯🇵, Korean 🇰🇷, Spanish 🇪🇸, French 🇫🇷, Portuguese 🇵🇹, Italian 🇮🇹, German 🇩🇪, Russian 🇷🇺, Hindi 🇮🇳, and Arabic 🇦🇪. They’re faster (<200ms), more efficient, and more accurate. We are also rolling out a conversationally tuned model, EVI 4 mini, that is powered by the intelligence of any frontier model like Claude 4.5 Sonnet or GPT-5. Try our live translation demo & EVI 4 mini chat, built on a few voice samples and a prompt:

Hume AI profil fotoğrafı
Hume AI11 ay önce

We'll be releasing more evaluations, more languages, and access to voice conversion and phoneme editing soon. In the meantime, we’re excited to see what you create! Read more about product &amp; research in our blog ↓

Hume AI profil fotoğrafı
Hume AI11 ay önce

We’re also HIRING across roles! Apply below on our careers page

Umesh profil fotoğrafı
Umesh11 ay önce

Congratulations 👏 It's amazing!

Hume AI profil fotoğrafı
Hume AI11 ay önce

thanks for the support umesh!

proper profil fotoğrafı
proper11 ay önce

2.0!! congrats on the launch!

Kol Tregaskes profil fotoğrafı
Kol Tregaskes11 ay önce

Nice guys, thanks for the early access!

Hume AI profil fotoğrafı
Hume AI11 ay önce

would love to hear what you think Kol!

Alan Cowen profil fotoğrafı
Alan Cowen11 ay önce

日本語に最適な音声AI

Hume AI profil fotoğrafı
Hume AI11 ay önce

The land of @BIG_CLAPPY!!

mori profil fotoğrafı
mori11 ay önce

LFG!! ⭐️

mori profil fotoğrafı
mori11 ay önce

can’t wait to see what ppl build w thisss

Hume AI profil fotoğrafı
Hume AI11 ay önce

show us what you build!

Alec Freudenstein profil fotoğrafı
Alec Freudenstein11 ay önce

INCREDIBLE WORK TEAM!!!! SO PROUD OF EVERYONE WHO PUT SO MUCH TIME INTO THIS!!!! NOBODY SHIPS LIKE HUME!!!!!!!!

Janet Ho profil fotoğrafı
Janet Ho11 ay önce

Can’t wait to go to Japan with this

Linus ✦ Ekenstam profil fotoğrafı
Linus ✦ Ekenstam11 ay önce

Guys, thanks for giving me early access to this, been sharing with the family last couple of days, and it’s been a blast. Extremely bullish on the trajectory you are on, keep shipping, this upgrade is spot on. 🫡

Hume AI profil fotoğrafı
Hume AI11 ay önce

thanks Linus!!

Lucas | IA profil fotoğrafı
Lucas | IA11 ay önce

SO good

Hume AI profil fotoğrafı
Hume AI11 ay önce

thank you 😊

Alvaro Cintas profil fotoğrafı
Alvaro Cintas11 ay önce

SO good!

Miguel Ángel | GptZone profil fotoğrafı
Miguel Ángel | GptZone11 ay önce

I am just trying it right now and it seems spectacular!!

Robert Scoble profil fotoğrafı
Robert Scoble11 ay önce

Great getting to talk to your CEO about this launch:

Hume AI profil fotoğrafı
Hume AI11 ay önce

awesome conversation!

Abdul Shakoor profil fotoğrafı
Abdul Shakoor11 ay önce

octave 2 feels like a serious upgrade.

Hume AI profil fotoğrafı
Hume AI11 ay önce

so glad you think so!

cate bligh profil fotoğrafı
cate bligh11 ay önce

Great video! I'm looking forward to playing with this more!

Hume AI profil fotoğrafı
Hume AI11 ay önce

love the videos you create with us cate!

Dr. Apolotary profil fotoğrafı
Dr. Apolotary11 ay önce

Why is he talking with a strong English accent in Japanese

Hume AI profil fotoğrafı
Hume AI11 ay önce

because he’s English! our model preserves accents based on the language of the base voice

Adam profil fotoğrafı
Adam11 ay önce

Unbelievable! I'll talk about it soon 👀

Jack profil fotoğrafı
Jack11 ay önce

So amazing

Nelly; profil fotoğrafı
Nelly;11 ay önce

no excuses now.. I can actually sound fluent in 11 languages😂

Qing Liu profil fotoğrafı
Qing Liu11 ay önce

Soooo proud 🥺 This is really amazing!!

@Czr305 profil fotoğrafı
@Czr30511 ay önce

What about for EVI?

Hume AI profil fotoğrafı
Hume AI11 ay önce

we have EVI4 mini also in preview

@Czr305 profil fotoğrafı
@Czr30511 ay önce

How do I gain access ?

Eleftheria Batsou profil fotoğrafı
Eleftheria Batsou11 ay önce

Let me check that!!! Excited for all these features

ayn profil fotoğrafı
ayn11 ay önce

Congratulations HUME!! 🫶🏻

cova profil fotoğrafı
cova11 ay önce

Congratulations!! 🫶

WISe Social profil fotoğrafı
WISe Social11 ay önce

Exciting advancements with Octave 2. Its multi-speaker feature could really enhance AI agent interactions. Imagine seamless transactions through secure channels using satellite connectivity, perfect for a world where speed and reliability are key.

Luke Miler profil fotoğrafı
Luke Miler11 ay önce

sweet! will add to shortcut

Afiz ⚡️ profil fotoğrafı
Afiz ⚡️11 ay önce

multilingual text-to-speech. I would definitely want to give it a try.

Jameson McSmirnoff 🌊🛥 profil fotoğrafı
Jameson McSmirnoff 🌊🛥11 ay önce

Deutsch ist euer Endgegner!

Roni Rahman profil fotoğrafı
Roni Rahman11 ay önce

Looks really good!

MetaDJ profil fotoğrafı
MetaDJ11 ay önce

Awesome! 🔥

Jussi profil fotoğrafı
Jussi11 ay önce

Will it support audio to audio? I am looking for a audio model to change the vocals of a song I generated on Suno to my own voice

Jaden Tripp profil fotoğrafı
Jaden Tripp11 ay önce

IPA input?

Rafael Estrela | IA profil fotoğrafı
Rafael Estrela | IA11 ay önce

Incrivel !

Benzer Videolar

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 görüntüleme • 1 yıl önce

Voice used to be AI’s forgotten modality - now it's having its big moment: rapid innovation, big funding rounds, major agentic applications My conversation with Neil Zeghidour, top AI researcher in the field (Google DeepMind, Meta, kyutai) and now CEO of Gradium This is a reference episode on all things voice AI 🔥 00:00 Intro 01:21 Voice AI’s big moment, and why we’re still early 03:34 Why voice lagged behind text/image/video 06:06 The convergence era: transformers for every modality 07:40 Beyond Her: always-on assistants, wake words, voice-first devices 11:01 Voice vs text: where voice fits (even for coding) 12:56 Neil’s origin story: from finance to machine learning, with help from Yann LeCun and Soumith Chintala 18:35 Neural codecs (SoundStream): compression as the unlock 22:30 Kyutai: open research, small elite teams, moving fast 31:32 Why big labs haven’t “won” voice AI4 34:01 On-device voice: where it works, why compact models matter 46:37 The last mile: real-world robustness, pronunciation, uptime 41:35 Benchmarking voice: why metrics fail, how they actually test 47:03 Cascades vs speech-to-speech: trade-offs + what’s next 54:05 Hardest frontier: noisy rooms, factories, multi-speaker chaos 1:00:50 New languages + dialects: what transfers, what doesn’t 1:02:54 Hardware & compute: why voice isn’t a 10,000-GPU game 1:07:27 What data do you need to train voice models 1:09:02 Deepfakes + privacy: why watermarking isn’t a solution 1:12:30 Voice + vision: multimodality, screen awareness, video+audio 1:14:43 Voice cloning vs voice design: where the market goes 1:16:32 Paris/Europe AI: talent density, underdog energy, what’s next

Matt Turck

22,980 görüntüleme • 6 ay önce

Open science is how we continue to push technology forward and today at Meta FAIR we’re sharing eight new AI research artifacts including new models, datasets and code to inspire innovation in the community. More in the video from Joelle Pineau. This work is another important step towards our goal of achieving Advanced Machine Intelligence (AMI). What we’re releasing: • Meta Spirit LM: An open source language model for seamless speech and text integration. • Meta Segment Anything Model 2.1: An updated checkpoint with improved results on visually similar objects, small objects and occlusion handling. Plus a new developer suite to make it easier for developers to build with SAM 2. • Layer Skip: Inference code and fine-tuned checkpoints demonstrating a new method for enhancing LLM performance. • SALSA: New code to enable researchers to benchmark AI-based attacks in support of validating security for post-quantum cryptography. • Meta Lingua: A lightweight and self-contained codebase designed to train language models at scale. • Meta Open Materials: New open source models and the largest dataset of its kind to accelerate AI-driven discovery of new inorganic materials. • MEXMA: A new research paper and code for our novel pre-trained cross-lingual sentence encoder with coverage across 80 languages. • Self-Taught Evaluator: a new method for generating synthetic preference data to train reward models without relying on human annotations. Access to state-of-the-art AI creates opportunities for everyone. We’re excited to share this work and look forward to seeing the community innovation that results from it. Details and access to everything released by FAIR today ➡️

AI at Meta

150,477 görüntüleme • 1 yıl önce