Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing Meta Omnilingual Automatic Speech Recognition (ASR), a suite of models providing ASR capabilities for over 1,600 languages, including 500 low-coverage languages never before served by any ASR system. While most ASR systems focus on a limited set of languages that are well-represented on the internet, this release marks...

544,998 Aufrufe • vor 9 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,026 Aufrufe • vor 4 Monaten

LINGUISTIC IMPERIALISM Linguistic imperialism is the process whereby dominant powers impose their language on those they colonise, suppressing indigenous languages and thus marginalising their speakers and sustaining power inequalities. Indigenous languages in colonial Africa were frowned upon, while colonial languages were made mandatory. In Anglophone Africa, policies were all written in English; media and broadcasting used English; and there were many systemic policies forbidding the use of indigenous languages. For example, students in many Kenyan schools were punished for using any other languages, with most of these punishments involving shaming tactics such as the wearing of bones from dead animals as chains around the neck. This also happened to the indigenous people who were forced to speak Arabic and change their names in places like Sudan. This warped the consciousness of individuals, leading to a loss of appreciation for indigenous languages and cultures - and promoted the adoption of the coloniser’s worldview, values, systems and structures. UNESCO's 1953 report, The Use of Vernacular Languages in Education, indicated that around 40% of the global population received education in an unfamiliar language. Sub-Saharan Africa, which is home to nearly 30% of the world’s languages, still uses the colonisers’ languages as national languages. Consequently, the education and values instilled remain those of the colonisers. Many languages are at risk of extinction, with only a few speakers left. The loss of these languages is equivalent to the loss of African heritage and culture.

African Stream

20,321 Aufrufe • vor 1 Jahr

U.N. official wants to decolonize AI and train new models with third-world "systems of knowledge" Deputy Secretary-General Amina J. Mohammed of Nigeria: "Colonial conquest brought catastrophe and genocide, in many cases on a continental scale. Entire communities were destroyed. Indigenous people were driven far from their lands. The assault was also on memory itself. Languages were suppressed. Knowledge gained through millennia of civilization, including in the Mayan codices, was deliberately put to the flame. Many traditions rooted in land and community and memory were lost." "These are living traditions carrying a wisdom that can help the region shape technological change in its own image towards its own vision of a good life, to pursue development on its own terms. And that starts with ensuring the people of this region can see themselves in the technology. More than 800 indigenous peoples live here, representing 60 million people. Many of their languages and systems of knowledge are sustained through oral traditions when AI systems are trained largely on written material that has been gathered elsewhere. That knowledge can be distorted or made invisible. Right now, Chile's National Center for Artificial Intelligence and 30 partner institutions are developing a LatAm GPT, an open-source model trained on regional data. Its first version is in Spanish and Portuguese, but researchers are also developing tools for other languages -- with the consent of the indigenous communities concerned. And that is what an inclusive reset can look like. Technology built with a fuller account of the people that it is meant to serve."

Breitbart News

45,032 Aufrufe • vor 9 Tagen