Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

🚀Introducing Aero-1-Audio lmms-lab ⚡️Lightweight 1.5B model 💡Outperforms SOTA models like Whisper, Qwen-2-Audio & commercial services from ElevenLabs - Blog: - Code: - Demo Gradio: . Thanks to AK!

17,982 görüntüleme • 1 yıl önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

I coded a Speech-to-Text model from scratch. 𝐇𝐞𝐫𝐞 𝐢𝐬 𝐭𝐡𝐞 𝐛𝐥𝐨𝐠 𝐟𝐨𝐫 𝐭𝐡𝐞 𝐬𝐚𝐦𝐞: No APIs. No pre-trained models. Just PyTorch, an A100 GPU, and hours of debugging. This started months ago. I wanted to understand how machines hear. Not surface-level understanding. I wanted to build the whole thing myself. So I built it piece by piece: autoencoders, VAEs, VQ-VAEs, Residual Vector Quantization, and CTC loss. Each one took days to get right. Trained for 3 hours on 13,100 audio clips. Got complete garbage. Changed the tokenizer from BPE to character-level. Rechecked everything. Asked AVB who built STT models before. His answer: these models are tricky to train and need days of compute, not hours. Cut the dataset to 200 clips. After 2 hours, actual words appeared. Overfitted? Absolutely. But watching noise turn into recognizable English was satisfying. I have made a blog about this as well so you can learn about the same and my process - Audio fundamentals and waveform representation - Why attention breaks on raw audio - Convolutional downsampling - Transformer encoder with positional encoding - Vector Quantization, straight-through estimator, and RVQ - CTC loss and greedy decoding - Full training loop with VQ loss warmup - What went wrong and what finally worked Resources: - Blog: - Code: More Resoures CTC loss AVB videos SoundStream Paper LJ speech dataset wav2vec paper RVQ blog Next up: I've already trained two TTS architectures from scratch. Video post about those coming soon. But first, I'm dropping a visual breakdown of Vision Transformers, covering how they work and how to fine-tune them. Follow me Mayank Pratap Singh you're into audio deep learning. Repost so others can find this

Mayank Pratap Singh

51,382 görüntüleme • 5 ay önce

This is THE moment of Physical AI! We are officially announcing Cosmos 3: Omnimodal World Models for Physical AI 🚀 - Cosmos 3 is an omnimodal world model: within a unified architecture, it can understand and generate language, images, video, audio, and actions. - It is not just a VLM, not just a video generator, not just an audio-visual generative model, and not just a physics simulator / world-action model. It can understand images and videos, generate images, videos, and audio, simulate future worlds, predict actions, and generate robot policies—enabling models to truly begin to “touch the world.” - Cosmos 3 is the #1 open-weight reasoner / T2I / I2V / robot policy across many benchmarks. Huge thanks to every teammate who fought side by side on this journey—from architecture, data, training, infra, serving, and evaluation to post-training. Every part of this project carries an incredible amount of hard work. This was my first time leading a project as Tech Lead, and I feel truly fortunate. The future of Physical AI needs models that can not only “see” and “describe” the world, but also “imagine,” “simulate,” and “act”—and eventually close the loop with the real world. I hope Cosmos 3 can become an important starting point for this direction, and I’m excited to push Physical AI into its next stage together with the open-source community. Welcome to the era of Physical AI. HuggingFace: Project Website: Code:

Max Zhaoshuo Li 李赵硕

1,078,824 görüntüleme • 3 ay önce

Sarvam Beats GPT-4o: India’s New AI Model Claims Top Spot in Indic Speech Sarvam AI, an Indian startup, recently launched Sarvam Audio, a speech recognition model that claims superior performance over GPT-4o Transcribe on Indic language benchmarks. This development highlights India's push for AI sovereignty in handling local linguistic nuances. Sarvam Audio supports 22 Indian languages from the Eighth Schedule, plus Indian English, with strong handling of code-mixing like Hindi-English blends. It features built-in speaker diarization for up to eight speakers and processes long-form audio such as podcasts or meetings. Trained on the IndicVoices dataset 12,000 hours from over 16,000 speakers across 208 districts it captures real-world noise and spontaneous speech. The model reportedly outperforms GPT-4o Transcribe and Gemini 3 Flash in transcription accuracy (lower Word Error Rate) on IndicVoices benchmarks for unnormalized, normalized, and code-mixed speech. Sarvam attributes this to specialization on Indian accents and patterns, unlike global models trained on Western data. Detailed public benchmarks are pending independent verification. Key Applications 🔴 Call centers and logistics for multilingual transcription. 🔴 Banking, fintech, and e-commerce for customer interactions. 🔴 Podcasts, meetings, and lectures via API for real-time or batch processing. ​ 🔴 This B2B-focused tool aligns with India's IndiaAI Mission, backed by government GPU access for sovereign LLMs. Credit : AIM Networks.

Augadh

43,429 görüntüleme • 7 ay önce

BREAKING: Inside The ElevenLabs Summit The future of voice-first interfaces. CEO, Mati Staniszewski (@matiii) Head of Growth, Luke Harries (Luke Harries) Series A Lead, Bryan Kim (Bryan Kim) of a16z Klarna CEO, Sebastian Siemiatkowski (Sebastian Siemiatkowski) AIUC CEO, Rune Kvist (Rune Kvist) Founded in just 2022, has scaled to $300M ARR & raised $781M in total funding across five funding rounds, most recently closing a $500M Series D led by Sequoia Capital at an $11 billion valuation. ElevenLabs began with a breakthrough human-like text-to-speech model & has since expanded into speech-to-text, dubbing, sound effects, music, & conversational AI. Today, it combines these models with integrations & enterprise-grade infrastructure to power production-ready platforms for businesses, creators, & developers: - ElevenAgents enables enterprises to deploy voice and chat agents at scale, with customers including Deutsche Telekom, Revolut, Square, & the Ukrainian government using it for support, commerce, and citizen services. - ElevenCreative allows brands like Duolingo, NVIDIA, & TIME to generate and localize high-quality audio in 70+ languages. - ElevenAPI provides low-latency voice infrastructure for developers, powering platforms from Meta & Epic Games to Salesforce, MasterClass, & Harvey—reaching more than one billion users globally. Leaving the Summit, after seeing the product depth, long-term vision, and caliber of operators & partners around them, it’s clear ElevenLabs has built one of the strongest & fastest-growing companies in AI today. Really impressive.

Molly O’Shea

143,008 görüntüleme • 6 ay önce