Загрузка видео...

Не удалось загрузить видео

На главную

Introducing Octave 2: our next-generation multilingual text-to-speech model What’s new: - Fluent in 11+ languages - 40% faster (<200ms latency⁠⁠) & 50% cheaper than Octave 1 - Multi-speaker conversation - More reliable pronunciation - New voice conversion & phoneme editing capabilities For the month of October, we’re offering 50%...

7,077,598 просмотров • 11 месяцев назад •via X (Twitter)

Комментарии: 48

Фото профиля Hume AI
Hume AI11 месяцев назад

Our new models speak over 11 languages, including English, Japanese 🇯🇵, Korean 🇰🇷, Spanish 🇪🇸, French 🇫🇷, Portuguese 🇵🇹, Italian 🇮🇹, German 🇩🇪, Russian 🇷🇺, Hindi 🇮🇳, and Arabic 🇦🇪. They’re faster (<200ms), more efficient, and more accurate. We are also rolling out a conversationally tuned model, EVI 4 mini, that is powered by the intelligence of any frontier model like Claude 4.5 Sonnet or GPT-5. Try our live translation demo & EVI 4 mini chat, built on a few voice samples and a prompt:

Фото профиля Hume AI
Hume AI11 месяцев назад

We'll be releasing more evaluations, more languages, and access to voice conversion and phoneme editing soon. In the meantime, we’re excited to see what you create! Read more about product &amp; research in our blog ↓

Фото профиля Hume AI
Hume AI11 месяцев назад

We’re also HIRING across roles! Apply below on our careers page

Фото профиля Umesh
Umesh11 месяцев назад

Congratulations 👏 It's amazing!

Фото профиля Hume AI
Hume AI11 месяцев назад

thanks for the support umesh!

Фото профиля proper
proper11 месяцев назад

2.0!! congrats on the launch!

Фото профиля Kol Tregaskes
Kol Tregaskes11 месяцев назад

Nice guys, thanks for the early access!

Фото профиля Hume AI
Hume AI11 месяцев назад

would love to hear what you think Kol!

Фото профиля Alan Cowen
Alan Cowen11 месяцев назад

日本語に最適な音声AI

Фото профиля Hume AI
Hume AI11 месяцев назад

The land of @BIG_CLAPPY!!

Фото профиля mori
mori11 месяцев назад

LFG!! ⭐️

Фото профиля mori
mori11 месяцев назад

can’t wait to see what ppl build w thisss

Фото профиля Hume AI
Hume AI11 месяцев назад

show us what you build!

Фото профиля Alec Freudenstein
Alec Freudenstein11 месяцев назад

INCREDIBLE WORK TEAM!!!! SO PROUD OF EVERYONE WHO PUT SO MUCH TIME INTO THIS!!!! NOBODY SHIPS LIKE HUME!!!!!!!!

Фото профиля Janet Ho
Janet Ho11 месяцев назад

Can’t wait to go to Japan with this

Фото профиля Linus ✦ Ekenstam
Linus ✦ Ekenstam11 месяцев назад

Guys, thanks for giving me early access to this, been sharing with the family last couple of days, and it’s been a blast. Extremely bullish on the trajectory you are on, keep shipping, this upgrade is spot on. 🫡

Фото профиля Hume AI
Hume AI11 месяцев назад

thanks Linus!!

Фото профиля Lucas | IA
Lucas | IA11 месяцев назад

SO good

Фото профиля Hume AI
Hume AI11 месяцев назад

thank you 😊

Фото профиля Alvaro Cintas
Alvaro Cintas11 месяцев назад

SO good!

Фото профиля Miguel Ángel | GptZone
Miguel Ángel | GptZone11 месяцев назад

I am just trying it right now and it seems spectacular!!

Фото профиля Robert Scoble
Robert Scoble11 месяцев назад

Great getting to talk to your CEO about this launch:

Фото профиля Hume AI
Hume AI11 месяцев назад

awesome conversation!

Фото профиля Abdul Shakoor
Abdul Shakoor11 месяцев назад

octave 2 feels like a serious upgrade.

Фото профиля Hume AI
Hume AI11 месяцев назад

so glad you think so!

Фото профиля cate bligh
cate bligh11 месяцев назад

Great video! I'm looking forward to playing with this more!

Фото профиля Hume AI
Hume AI11 месяцев назад

love the videos you create with us cate!

Фото профиля Dr. Apolotary
Dr. Apolotary11 месяцев назад

Why is he talking with a strong English accent in Japanese

Фото профиля Hume AI
Hume AI11 месяцев назад

because he’s English! our model preserves accents based on the language of the base voice

Фото профиля Adam
Adam11 месяцев назад

Unbelievable! I'll talk about it soon 👀

Фото профиля Jack
Jack11 месяцев назад

So amazing

Фото профиля Nelly;
Nelly;11 месяцев назад

no excuses now.. I can actually sound fluent in 11 languages😂

Фото профиля Qing Liu
Qing Liu11 месяцев назад

Soooo proud 🥺 This is really amazing!!

Фото профиля @Czr305
@Czr30511 месяцев назад

What about for EVI?

Фото профиля Hume AI
Hume AI11 месяцев назад

we have EVI4 mini also in preview

Фото профиля @Czr305
@Czr30511 месяцев назад

How do I gain access ?

Фото профиля Eleftheria Batsou
Eleftheria Batsou11 месяцев назад

Let me check that!!! Excited for all these features

Фото профиля ayn
ayn11 месяцев назад

Congratulations HUME!! 🫶🏻

Фото профиля cova
cova11 месяцев назад

Congratulations!! 🫶

Фото профиля WISe Social
WISe Social11 месяцев назад

Exciting advancements with Octave 2. Its multi-speaker feature could really enhance AI agent interactions. Imagine seamless transactions through secure channels using satellite connectivity, perfect for a world where speed and reliability are key.

Фото профиля Luke Miler
Luke Miler11 месяцев назад

sweet! will add to shortcut

Фото профиля Afiz ⚡️
Afiz ⚡️11 месяцев назад

multilingual text-to-speech. I would definitely want to give it a try.

Фото профиля Jameson McSmirnoff 🌊🛥
Jameson McSmirnoff 🌊🛥11 месяцев назад

Deutsch ist euer Endgegner!

Фото профиля Roni Rahman
Roni Rahman11 месяцев назад

Looks really good!

Фото профиля MetaDJ
MetaDJ11 месяцев назад

Awesome! 🔥

Фото профиля Jussi
Jussi11 месяцев назад

Will it support audio to audio? I am looking for a audio model to change the vocals of a song I generated on Suno to my own voice

Фото профиля Jaden Tripp
Jaden Tripp11 месяцев назад

IPA input?

Фото профиля Rafael Estrela | IA
Rafael Estrela | IA11 месяцев назад

Incrivel !

Похожие видео

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 просмотров • 1 год назад

Voice used to be AI’s forgotten modality - now it's having its big moment: rapid innovation, big funding rounds, major agentic applications My conversation with Neil Zeghidour, top AI researcher in the field (Google DeepMind, Meta, kyutai) and now CEO of Gradium This is a reference episode on all things voice AI 🔥 00:00 Intro 01:21 Voice AI’s big moment, and why we’re still early 03:34 Why voice lagged behind text/image/video 06:06 The convergence era: transformers for every modality 07:40 Beyond Her: always-on assistants, wake words, voice-first devices 11:01 Voice vs text: where voice fits (even for coding) 12:56 Neil’s origin story: from finance to machine learning, with help from Yann LeCun and Soumith Chintala 18:35 Neural codecs (SoundStream): compression as the unlock 22:30 Kyutai: open research, small elite teams, moving fast 31:32 Why big labs haven’t “won” voice AI4 34:01 On-device voice: where it works, why compact models matter 46:37 The last mile: real-world robustness, pronunciation, uptime 41:35 Benchmarking voice: why metrics fail, how they actually test 47:03 Cascades vs speech-to-speech: trade-offs + what’s next 54:05 Hardest frontier: noisy rooms, factories, multi-speaker chaos 1:00:50 New languages + dialects: what transfers, what doesn’t 1:02:54 Hardware & compute: why voice isn’t a 10,000-GPU game 1:07:27 What data do you need to train voice models 1:09:02 Deepfakes + privacy: why watermarking isn’t a solution 1:12:30 Voice + vision: multimodality, screen awareness, video+audio 1:14:43 Voice cloning vs voice design: where the market goes 1:16:32 Paris/Europe AI: talent density, underdog energy, what’s next

Matt Turck

22,980 просмотров • 6 месяцев назад

Open science is how we continue to push technology forward and today at Meta FAIR we’re sharing eight new AI research artifacts including new models, datasets and code to inspire innovation in the community. More in the video from Joelle Pineau. This work is another important step towards our goal of achieving Advanced Machine Intelligence (AMI). What we’re releasing: • Meta Spirit LM: An open source language model for seamless speech and text integration. • Meta Segment Anything Model 2.1: An updated checkpoint with improved results on visually similar objects, small objects and occlusion handling. Plus a new developer suite to make it easier for developers to build with SAM 2. • Layer Skip: Inference code and fine-tuned checkpoints demonstrating a new method for enhancing LLM performance. • SALSA: New code to enable researchers to benchmark AI-based attacks in support of validating security for post-quantum cryptography. • Meta Lingua: A lightweight and self-contained codebase designed to train language models at scale. • Meta Open Materials: New open source models and the largest dataset of its kind to accelerate AI-driven discovery of new inorganic materials. • MEXMA: A new research paper and code for our novel pre-trained cross-lingual sentence encoder with coverage across 80 languages. • Self-Taught Evaluator: a new method for generating synthetic preference data to train reward models without relying on human annotations. Access to state-of-the-art AI creates opportunities for everyone. We’re excited to share this work and look forward to seeing the community innovation that results from it. Details and access to everything released by FAIR today ➡️

AI at Meta

150,477 просмотров • 1 год назад