正在加载视频...
视频加载失败
OuteTTS 1.0: Upgrades in Quality, Cloning, and 20 Languages - Best Performance: Generate audio around 42 seconds in a single run (approximately 8,192 tokens). It is recomended not to near the limits of this windows when generating. Usually, the best results are up to 7,000 tokens. - Context Reduction... show more
25,779 次观看 • 1 年前 •via X (Twitter)
6 条评论

Our speech-to-text models are the most accurate on the market with top rankings across industry benchmarks. - The highest accuracy rates—up to 95% - Up to 30% fewer hallucinations than other leaders - Low latency—63 minutes converts in 35 seconds Try via API for free today 👇

Weird that no model yet caught up to elevenlabs in overall cosistency

"OuteTTS 1.0 is redefining TTS innovation! High-quality multilingual speech synthesis across 20 languages, improved speaker cloning, and generating 42 seconds of audio (up to 8,192 tokens) are major milestones. Staying around 7,000 tokens ensures the best quality. This sets a new benchmark for adaptable, expressive TTS systems. 🚀 #TTS #Innovation"

ai talk faster than me

Just checked this out on Huggingface where it had a non commercial license. Am I missing something

list of languages and more

