Video wird geladen...
Video konnte nicht geladen werden
Introducing Voxtral WebGPU: Real-time speech transcription entirely in your browser. This demo runs Voxtral-Mini-4B, a powerful streaming ASR model from Mistral AI, locally on WebGPU. The model supports 13 languages and is capable of <500 ms latency. Fully private. Zero cost.
94,623 Aufrufe • vor 6 Monaten •via X (Twitter)
28 Kommentare

Powered by 🤗 Transformers.js v4 (preview). I'm excited to see what the community builds with it! Try it out yourself! Demo + source code 👇

@MistralAI this is fire, it switches languages just like that. Awesome!

@MistralAI Thanks! 🙏 I know right, the model is super impressive!

@MistralAI You have no idea how pump I am 😂 I feel like i have to add this somewhere to any of my free apps. 😂 Man the change from spanish to english and english to spanish right on the spot it's really insane.

@MistralAI

@MistralAI Which OS/browser are you running? We recommend Chrome (or a Chromium-based browser), and most of our tests are on Mac/Windows.

@MistralAI it could be that the lack of a standard graphics architecture on linux makes webgpu harder.

@MistralAI lets goooo!

@MistralAI You called it 😂

@MistralAI been following you long enough to know excited to build with this!

@MistralAI It’s good. Been running voxtral mini for a while with

@MistralAI Wohh, this is so good. It was able to correctly transcribe the name of my street, which is an Indigenous word in Portuguese. Congratulation!

@MistralAI I feel we're getting closer and closer to a world where advanced voice AI is readily available to everyone. 高度な音声AIが誰にでもすぐ使える世界に、どんどん近づいていると感じます。

@MistralAI Cannot wait till we drop our AI Performance results in our benchmark test 🥰 i can see the vision on this and will be curious to how Voxtral will compare to other loads Also can't wait to try them out

@MistralAI is this multithreaded? what kind of cpu/gpu usage does it have?

@MistralAI can i use it on a second window in youtube?

🚨 AI Tech Update Real-time speech transcription is now possible directly in your browser with Voxtral WebGPU. • Runs Voxtral‑Mini‑4B locally using WebGPU • Developed by Mistral AI • Supports 13 languages • <500 ms latency for near real-time transcription • Fully private processing happens on your device • Zero cost, no cloud required A glimpse into the future of on-device AI and private speech recognition. #AI #SpeechRecognition #WebGPU #MistralAI

@MistralAI 🎙️ Hey Insta fam!

@MistralAI Does this run on Oculus Quest chrome browser?

@MistralAI C++/GGML based Voxtral Mini 4B Realtime released! 16.39x (0.0610 RTF) realtime offline, 170ms TTFT streaming on RTX 5090.

@MistralAI Amazing, can’t wait to integrate this.

@MistralAI wow, must try this

@MistralAI

@MistralAI this should be able to transcribe youtube videos right ?

@MistralAI Argh, no support for Danish - my test got transmogrified to Dutch 😂 Which actually makes sense as the best guess in the lack of Danish support

@MistralAI yea but new granite by ibm ...

@MistralAI Been waiting for something like this. Running a 4B model locally with sub-500ms latency is wild—WebGPU keeps unlocking new possibilities.

@Scobleizer @MistralAI lol the fact that this runs entirely in browser with zero backend is kind of insane, been waiting for local speech-to-text that actually works 😭
