Video yükleniyor...
Video Yüklenemedi
Wow! New Speech to Speech model - Fish Agent v0.1 3B by Fish Audio 🔥 > Trained on 700K hours of multilingual audio > Continue-pretrained version of Qwen-2.5-3B-Instruct for 200B audio & text tokens > Zero-shot voice cloning > Text + audio input/ Audio output > Ultra-fast inference w/... show more
67,003 görüntüleme • 1 yıl önce •via X (Twitter)
10 Yorum

Vaibhav (VB) Srivastav1 yıl önce
Check out the model here:

Vaibhav (VB) Srivastav1 yıl önce
They also put together a pretty dope space to try out the model:

Juan Pablo Gallego1 yıl önce
@huggingface @FishAudio It was clearly trained on TTS voice. Sounds very robotic

William Payne1 yıl önce
@FishAudio How does it compare to Mochi?

Ishtiaq Rahman1 yıl önce
@FishAudio Unfortunately, 3+ seconds for non-RAG questions is slow as a mule. OpenAI's speech-to-speech is much, much faster. Naah, this ain't it.

Houdini1 yıl önce
@FishAudio Youre always early on these bangers. Nice work.

Houdini1 yıl önce
@FishAudio @cocktailpeanut

bl4nk1 yıl önce
@FishAudio Pls drop google colab code

Sadi Moodi1 yıl önce
@FishAudio no published code to use, so its useless

Sahil 🇮🇳/acc1 yıl önce
@FishAudio why do these guys don't do hindi?
