Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Octave 2: our next-generation multilingual text-to-speech model What’s new: - Fluent in 11+ languages - 40% faster (<200ms latency⁠⁠) & 50% cheaper than Octave 1 - Multi-speaker conversation - More reliable pronunciation - New voice conversion & phoneme editing capabilities For the month of October, we’re offering 50%...

7,077,598 Aufrufe • vor 11 Monaten •via X (Twitter)

48 Kommentare

Profilbild von Hume AI
Hume AIvor 11 Monaten

Our new models speak over 11 languages, including English, Japanese 🇯🇵, Korean 🇰🇷, Spanish 🇪🇸, French 🇫🇷, Portuguese 🇵🇹, Italian 🇮🇹, German 🇩🇪, Russian 🇷🇺, Hindi 🇮🇳, and Arabic 🇦🇪. They’re faster (<200ms), more efficient, and more accurate. We are also rolling out a conversationally tuned model, EVI 4 mini, that is powered by the intelligence of any frontier model like Claude 4.5 Sonnet or GPT-5. Try our live translation demo & EVI 4 mini chat, built on a few voice samples and a prompt:

Profilbild von Hume AI
Hume AIvor 11 Monaten

We'll be releasing more evaluations, more languages, and access to voice conversion and phoneme editing soon. In the meantime, we’re excited to see what you create! Read more about product &amp; research in our blog ↓

Profilbild von Hume AI
Hume AIvor 11 Monaten

We’re also HIRING across roles! Apply below on our careers page

Profilbild von Umesh
Umeshvor 11 Monaten

Congratulations 👏 It's amazing!

Profilbild von Hume AI
Hume AIvor 11 Monaten

thanks for the support umesh!

Profilbild von proper
propervor 11 Monaten

2.0!! congrats on the launch!

Profilbild von Kol Tregaskes
Kol Tregaskesvor 11 Monaten

Nice guys, thanks for the early access!

Profilbild von Hume AI
Hume AIvor 11 Monaten

would love to hear what you think Kol!

Profilbild von Alan Cowen
Alan Cowenvor 11 Monaten

日本語に最適な音声AI

Profilbild von Hume AI
Hume AIvor 11 Monaten

The land of @BIG_CLAPPY!!

Profilbild von mori
morivor 11 Monaten

LFG!! ⭐️

Profilbild von mori
morivor 11 Monaten

can’t wait to see what ppl build w thisss

Profilbild von Hume AI
Hume AIvor 11 Monaten

show us what you build!

Profilbild von Alec Freudenstein
Alec Freudensteinvor 11 Monaten

INCREDIBLE WORK TEAM!!!! SO PROUD OF EVERYONE WHO PUT SO MUCH TIME INTO THIS!!!! NOBODY SHIPS LIKE HUME!!!!!!!!

Profilbild von Janet Ho
Janet Hovor 11 Monaten

Can’t wait to go to Japan with this

Profilbild von Linus ✦ Ekenstam
Linus ✦ Ekenstamvor 11 Monaten

Guys, thanks for giving me early access to this, been sharing with the family last couple of days, and it’s been a blast. Extremely bullish on the trajectory you are on, keep shipping, this upgrade is spot on. 🫡

Profilbild von Hume AI
Hume AIvor 11 Monaten

thanks Linus!!

Profilbild von Lucas | IA
Lucas | IAvor 11 Monaten

SO good

Profilbild von Hume AI
Hume AIvor 11 Monaten

thank you 😊

Profilbild von Alvaro Cintas
Alvaro Cintasvor 11 Monaten

SO good!

Profilbild von Miguel Ángel | GptZone
Miguel Ángel | GptZonevor 11 Monaten

I am just trying it right now and it seems spectacular!!

Profilbild von Robert Scoble
Robert Scoblevor 11 Monaten

Great getting to talk to your CEO about this launch:

Profilbild von Hume AI
Hume AIvor 11 Monaten

awesome conversation!

Profilbild von Abdul Shakoor
Abdul Shakoorvor 11 Monaten

octave 2 feels like a serious upgrade.

Profilbild von Hume AI
Hume AIvor 11 Monaten

so glad you think so!

Profilbild von cate bligh
cate blighvor 11 Monaten

Great video! I'm looking forward to playing with this more!

Profilbild von Hume AI
Hume AIvor 11 Monaten

love the videos you create with us cate!

Profilbild von Dr. Apolotary
Dr. Apolotaryvor 11 Monaten

Why is he talking with a strong English accent in Japanese

Profilbild von Hume AI
Hume AIvor 11 Monaten

because he’s English! our model preserves accents based on the language of the base voice

Profilbild von Adam
Adamvor 11 Monaten

Unbelievable! I'll talk about it soon 👀

Profilbild von Jack
Jackvor 11 Monaten

So amazing

Profilbild von Nelly;
Nelly;vor 11 Monaten

no excuses now.. I can actually sound fluent in 11 languages😂

Profilbild von Qing Liu
Qing Liuvor 11 Monaten

Soooo proud 🥺 This is really amazing!!

Profilbild von @Czr305
@Czr305vor 11 Monaten

What about for EVI?

Profilbild von Hume AI
Hume AIvor 11 Monaten

we have EVI4 mini also in preview

Profilbild von @Czr305
@Czr305vor 11 Monaten

How do I gain access ?

Profilbild von Eleftheria Batsou
Eleftheria Batsouvor 11 Monaten

Let me check that!!! Excited for all these features

Profilbild von ayn
aynvor 11 Monaten

Congratulations HUME!! 🫶🏻

Profilbild von cova
covavor 11 Monaten

Congratulations!! 🫶

Profilbild von WISe Social
WISe Socialvor 11 Monaten

Exciting advancements with Octave 2. Its multi-speaker feature could really enhance AI agent interactions. Imagine seamless transactions through secure channels using satellite connectivity, perfect for a world where speed and reliability are key.

Profilbild von Luke Miler
Luke Milervor 11 Monaten

sweet! will add to shortcut

Profilbild von Afiz ⚡️
Afiz ⚡️vor 11 Monaten

multilingual text-to-speech. I would definitely want to give it a try.

Profilbild von Jameson McSmirnoff 🌊🛥
Jameson McSmirnoff 🌊🛥vor 11 Monaten

Deutsch ist euer Endgegner!

Profilbild von Roni Rahman
Roni Rahmanvor 11 Monaten

Looks really good!

Profilbild von MetaDJ
MetaDJvor 11 Monaten

Awesome! 🔥

Profilbild von Jussi
Jussivor 11 Monaten

Will it support audio to audio? I am looking for a audio model to change the vocals of a song I generated on Suno to my own voice

Profilbild von Jaden Tripp
Jaden Trippvor 11 Monaten

IPA input?

Profilbild von Rafael Estrela | IA
Rafael Estrela | IAvor 11 Monaten

Incrivel !

Ähnliche Videos

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 Aufrufe • vor 1 Jahr

Voice used to be AI’s forgotten modality - now it's having its big moment: rapid innovation, big funding rounds, major agentic applications My conversation with Neil Zeghidour, top AI researcher in the field (Google DeepMind, Meta, kyutai) and now CEO of Gradium This is a reference episode on all things voice AI 🔥 00:00 Intro 01:21 Voice AI’s big moment, and why we’re still early 03:34 Why voice lagged behind text/image/video 06:06 The convergence era: transformers for every modality 07:40 Beyond Her: always-on assistants, wake words, voice-first devices 11:01 Voice vs text: where voice fits (even for coding) 12:56 Neil’s origin story: from finance to machine learning, with help from Yann LeCun and Soumith Chintala 18:35 Neural codecs (SoundStream): compression as the unlock 22:30 Kyutai: open research, small elite teams, moving fast 31:32 Why big labs haven’t “won” voice AI4 34:01 On-device voice: where it works, why compact models matter 46:37 The last mile: real-world robustness, pronunciation, uptime 41:35 Benchmarking voice: why metrics fail, how they actually test 47:03 Cascades vs speech-to-speech: trade-offs + what’s next 54:05 Hardest frontier: noisy rooms, factories, multi-speaker chaos 1:00:50 New languages + dialects: what transfers, what doesn’t 1:02:54 Hardware & compute: why voice isn’t a 10,000-GPU game 1:07:27 What data do you need to train voice models 1:09:02 Deepfakes + privacy: why watermarking isn’t a solution 1:12:30 Voice + vision: multimodality, screen awareness, video+audio 1:14:43 Voice cloning vs voice design: where the market goes 1:16:32 Paris/Europe AI: talent density, underdog energy, what’s next

Matt Turck

22,980 Aufrufe • vor 6 Monaten

Open science is how we continue to push technology forward and today at Meta FAIR we’re sharing eight new AI research artifacts including new models, datasets and code to inspire innovation in the community. More in the video from Joelle Pineau. This work is another important step towards our goal of achieving Advanced Machine Intelligence (AMI). What we’re releasing: • Meta Spirit LM: An open source language model for seamless speech and text integration. • Meta Segment Anything Model 2.1: An updated checkpoint with improved results on visually similar objects, small objects and occlusion handling. Plus a new developer suite to make it easier for developers to build with SAM 2. • Layer Skip: Inference code and fine-tuned checkpoints demonstrating a new method for enhancing LLM performance. • SALSA: New code to enable researchers to benchmark AI-based attacks in support of validating security for post-quantum cryptography. • Meta Lingua: A lightweight and self-contained codebase designed to train language models at scale. • Meta Open Materials: New open source models and the largest dataset of its kind to accelerate AI-driven discovery of new inorganic materials. • MEXMA: A new research paper and code for our novel pre-trained cross-lingual sentence encoder with coverage across 80 languages. • Self-Taught Evaluator: a new method for generating synthetic preference data to train reward models without relying on human annotations. Access to state-of-the-art AI creates opportunities for everyone. We’re excited to share this work and look forward to seeing the community innovation that results from it. Details and access to everything released by FAIR today ➡️

AI at Meta

150,477 Aufrufe • vor 1 Jahr