Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Wow! New Speech to Speech model - Fish Agent v0.1 3B by Fish Audio 🔥 > Trained on 700K hours of multilingual audio > Continue-pretrained version of Qwen-2.5-3B-Instruct for 200B audio & text tokens > Zero-shot voice cloning > Text + audio input/ Audio output > Ultra-fast inference w/...

67,003 görüntüleme • 1 yıl önce •via X (Twitter)

10 Yorum

Vaibhav (VB) Srivastav profil fotoğrafı
Vaibhav (VB) Srivastav1 yıl önce

Check out the model here:

Vaibhav (VB) Srivastav profil fotoğrafı
Vaibhav (VB) Srivastav1 yıl önce

They also put together a pretty dope space to try out the model:

Juan Pablo Gallego profil fotoğrafı
Juan Pablo Gallego1 yıl önce

@huggingface @FishAudio It was clearly trained on TTS voice. Sounds very robotic

William Payne profil fotoğrafı
William Payne1 yıl önce

@FishAudio How does it compare to Mochi?

Ishtiaq Rahman profil fotoğrafı
Ishtiaq Rahman1 yıl önce

@FishAudio Unfortunately, 3+ seconds for non-RAG questions is slow as a mule. OpenAI's speech-to-speech is much, much faster. Naah, this ain't it.

Houdini profil fotoğrafı
Houdini1 yıl önce

@FishAudio Youre always early on these bangers. Nice work.

Houdini profil fotoğrafı
Houdini1 yıl önce

@FishAudio @cocktailpeanut

bl4nk profil fotoğrafı
bl4nk1 yıl önce

@FishAudio Pls drop google colab code

Sadi Moodi profil fotoğrafı
Sadi Moodi1 yıl önce

@FishAudio no published code to use, so its useless

Sahil 🇮🇳/acc profil fotoğrafı
Sahil 🇮🇳/acc1 yıl önce

@FishAudio why do these guys don't do hindi?

Benzer Videolar

NVIDIA Releases Audex (Nemotron-Labs-Audex-30B-A3B): A Unified Audio-Text LLM That Preserves the Text Intelligence of Its Backbone Most unified audio models pay a text tax. Add audio output, and reasoning benchmarks drop — even when the only new output is speech. NVIDIA just released one that doesn't. Audex (Nemotron-Labs-Audex-30B-A3B) is a 30B MoE with 3B active parameters, built on the text-only Nemotron-Cascade-2-30B-A3B backbone. Audio inputs are projected into the text embedding space. Text tokens and quantized audio tokens are then generated the same way, inside one MoE decoder — no thinker–talker split, no stacked cascade. Here's what's actually interesting: → One model, audio in and out: understanding, ASR, translation, TTS, text-to-audio, speech-to-speech → Text holds vs its own backbone: IMO AnswerBench 81.1 vs 79.3, MMLU-Redux 86.4 vs 86.3 → Beats text-only Qwen3.5-35B-A3B on several tasks: LiveCodeBench v6 85.3 vs 74.6, IFBench 77.8 vs 70.2 → The usual tax, for contrast: Qwen3-Omni-30B-A3B-Thinking drops to 60.4 on HMMT vs 71.4 for its text backbone → 6.82 WER on OpenASR, ahead of Step-Audio-R1.1-33B (7.91) and Qwen3-Omni-Thinking (8.00) → Two codecs: X-Codec2 for speech (50 tok/s, FSQ, 65,536 codebook), X-Codec for general audio (200 tok/s, 4 flattened RVQ layers) — the only strong open model generating general audio beyond speech Full analysis: Paper: Model weights: NVIDIA AI NVIDIA Wei Ping

Marktechpost AI

28,301 görüntüleme • 2 ay önce