
Xenova
@xenovacom • 20,591 subscribers
Bringing the power of machine learning to the web. Currently working on Transformers.js (@huggingface 🤗)
Shorts
Videos

I gave Fable 5 one job: write custom WebGPU kernels for Gemma 4 inference. It climbed to 84 tok/s, then hit a wall, insisting further optimization was impossible. Hours later, Anthropic rolled back invisible LLM development safeguards, and it hit 255 tok/s. The next day, access to Fable 5 was suspended globally.
Xenova1,167,502 Aufrufe • vor 2 Monaten

Before Fable 5 was shut down, it pushed Gemma 4 to 255 tok/s on WebGPU. Some didn't believe it was real. Today we're releasing the demo and kernels it wrote for you to see yourself. Run it locally in your browser. Agentic kernel optimization is the future of on-device inference
Xenova484,504 Aufrufe • vor 2 Monaten

Bonsai 27B just changed the local LLM game forever. 1-bit quantization shrinks it from 54GB to just 3.8GB (-93%), while retaining 90% of its intelligence. That's insane. With custom WebGPU kernels written by Fable 5 and GPT 5.6 Sol, the model now runs locally in your browser!
Xenova210,659 Aufrufe • vor 1 Monat

While we eagerly await Fable 5's return, our agentic WebGPU kernel optimization framework kept running. Opus 4.8 picked up where Fable left off, pushing Liquid AI's new LFM2.5 230M to an unbelievable 1,400 tok/s... running locally in your browser. Don't blink or you'll miss it.
Xenova174,867 Aufrufe • vor 2 Monaten

WTF?! This changes image generation forever! 🤯 PrismML just released Binary and Ternary Bonsai Image 4B! That's right, 1-bit diffusion models are here. Only ~3GB in size (FLUX.2 Klein 4B is 16GB). The most shocking part? It can run 100% locally in your browser. Try it now! 👇
Xenova183,996 Aufrufe • vor 3 Monaten

NEW: OpenAI releases Privacy Filter, their first open model of 2026! 🤗 Apache-2.0! It's a bidirectional token-classification adaptation of GPT-OSS, trained to mask personally identifiable information (PII) in text. At only 1.5B params, it can even run locally in your browser!
Xenova220,052 Aufrufe • vor 4 Monaten

Behold... GPT-OSS (20B) running 100% locally in your browser on WebGPU. This shouldn't be possible — but with Transformers.js v4 and ONNX Runtime Web, it is! A new class of AI apps is emerging. Zero-install, infinite distribution. Simply visit a website and run models locally.
Xenova311,512 Aufrufe • vor 6 Monaten

Chrome's new `window.ai` feature is going to change the web forever! 🤯 It allows you to run Gemini Nano, a powerful 3.25B parameter LLM, 100% locally in your browser! We've also added experimental support to 🤗 Transformers.js, making it super easy to use! 😍 Check it out! 👇
Xenova581,694 Aufrufe • vor 2 Jahren

NEW: Mistral AI releases Mistral 3, a family of multimodal models, including three start-of-the-art dense models (3B, 8B, and 14B) and Mistral Large 3 (675B, 41B active). All Apache 2.0! 🤗 Surprisingly, the 3B is small enough to run 100% locally in your browser on WebGPU! 🤯
Xenova225,277 Aufrufe • vor 9 Monaten

NEW: Alibaba just released Qwen 3.5 Small — a family of powerful multimodal models available in a range of sizes (0.8B, 2B, 4B, and 9B parameters). Perfect for on-device applications! They can even run 100% locally in your browser on WebGPU, powered by Transformers.js! 🤯
Xenova102,592 Aufrufe • vor 6 Monaten

Okay, this is actually insane... You can now run LFM2.5-1.2B-Thinking (a 1.2B parameter LLM from @LiquidAI) at over 200 tokens per second directly in your browser on WebGPU! 🤯 Zero install. Fully private. Blazingly fast. Powered by Transformers.js and ONNX Runtime Web
Xenova103,001 Aufrufe • vor 6 Monaten

Introducing Voxtral WebGPU: Real-time speech transcription entirely in your browser. This demo runs Voxtral-Mini-4B, a powerful streaming ASR model from Mistral AI, locally on WebGPU. The model supports 13 languages and is capable of <500 ms latency. Fully private. Zero cost.
Xenova94,356 Aufrufe • vor 5 Monaten

It's finally possible: real-time in-browser speech recognition with OpenAI Whisper! 🤯 The model runs fully on-device using Transformers.js and ONNX Runtime Web, and supports multilingual transcription across 100 different languages! 🔥 Check out the demo (+ source code)! 👇
Xenova262,819 Aufrufe • vor 2 Jahren

RF-DETR, the state-of-the-art model series for real-time object detection, can now run 100% locally in your browser on WebGPU with 🤗 Transformers.js v4! The models are Apache-2.0 licensed, making them a perfect fit for both personal and commercial applications. Try the demo 👇
Xenova77,179 Aufrufe • vor 6 Monaten

BOOM! 💥 Today I added WebGPU support for Andrej Karpathy's nanochat models, meaning they can run 100% locally in your browser (no server)! The d32 version runs at over 50 tps on my M4 Max 🚀 Pretty wild that you can now deploy AI applications using just a single index.html file 😅
Xenova95,481 Aufrufe • vor 10 Monaten