Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Two Realtime API updates: - You can now build speech-to-speech experiences with five new voices—which are much more expressive and steerable. 🤣🤫🤪 - We're lowering the price by using prompt caching. Cached text inputs are discounted 50% and cached audio inputs are discounted 80%. 📉

243,432 görüntüleme • 1 yıl önce •via X (Twitter)

10 Yorum

Vaibhav (VB) Srivastav profil fotoğrafı
Vaibhav (VB) Srivastav1 yıl önce

Congratulations - our cascaded approach has had that for quite a long while 🤗

Degen Wave 🌊 profil fotoğrafı
Degen Wave 🌊1 yıl önce

Can we get a voice that is at least close to being as good as Sky was plz?

Mckay Wrigley profil fotoğrafı
Mckay Wrigley1 yıl önce

This one is super appreciated!

Alex Volkov (Thursd/AI) profil fotoğrafı
Alex Volkov (Thursd/AI)1 yıl önce

Can we get some clarification / confirmation of the following? - are output tokens are turned into input tokens on next turn? - Is the caching automatic or do we need to do anything? - can we get an example calculation of a 3 turn conversation and transparent pricing? 🙏

Daniel Nguyen ⚡ profil fotoğrafı
Daniel Nguyen ⚡1 yıl önce

Thank you. Can you add sample audio so we can add voice preview in our app? Similar to mp3 files on this page

• profil fotoğrafı
1 yıl önce

“Sorry, that’s against my guidelines” is what needs to be fixed most

Haider. profil fotoğrafı
Haider.1 yıl önce

ok, but what's so funny?

Everett World profil fotoğrafı
Everett World1 yıl önce

Nice, but Sol is still the best voice!

Jan profil fotoğrafı
Jan1 yıl önce

we need an updated text to speech api : (

Kol Tregaskes profil fotoğrafı
Kol Tregaskes1 yıl önce

Liking the new voices.

Benzer Videolar

NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A Unified Audio-Text LLM That Preserves the Text Intelligence of Its Backbone Most unified audio models pay a text tax. Add audio output, and reasoning benchmarks drop — even when the only new output is speech. NVIDIA just released one that doesn't. Audex (Nemotron-Labs-Audex-30B-A3B) is a 30B MoE with 3B active parameters, built on the text-only Nemotron-Cascade-2-30B-A3B backbone. Audio inputs are projected into the text embedding space. Text tokens and quantized audio tokens are then generated the same way, inside one MoE decoder — no thinker–talker split, no stacked cascade. Here's what's actually interesting: → One model, audio in and out: understanding, ASR, translation, TTS, text-to-audio, speech-to-speech → Text holds vs its own backbone: IMO AnswerBench 81.1 vs 79.3, MMLU-Redux 86.4 vs 86.3 → Beats text-only Qwen3.5-35B-A3B on several tasks: LiveCodeBench v6 85.3 vs 74.6, IFBench 77.8 vs 70.2 → The usual tax, for contrast: Qwen3-Omni-30B-A3B-Thinking drops to 60.4 on HMMT vs 71.4 for its text backbone → 6.82 WER on OpenASR, ahead of Step-Audio-R1.1-33B (7.91) and Qwen3-Omni-Thinking (8.00) → Two codecs: X-Codec2 for speech (50 tok/s, FSQ, 65,536 codebook), X-Codec for general audio (200 tok/s, 4 flattened RVQ layers) — the only strong open model generating general audio beyond speech Full analysis: Paper: Model weights: NVIDIA AI NVIDIA Wei Ping

Marktechpost AI

28,301 görüntüleme • 2 ay önce