ๆญฃๅœจๅŠ ่ฝฝ่ง†้ข‘...

่ง†้ข‘ๅŠ ่ฝฝๅคฑ่ดฅ

End to End Speech models are on fire - LLAMA-OMNI 8B - Apache licensed! ๐Ÿ”ฅ > Speech Encoder - Whisper Large v3 > LLM backbone - Llama 3.1 8B Instruct > Speech Decoder - HuBERT (UnitY) > Simultaneously generate Speech + Text > Less than 250 ms latency >...

47,921 ๆฌก่ง‚็œ‹ โ€ข 1 ๅนดๅ‰ โ€ขvia X (Twitter)

10 ๆก่ฏ„่ฎบ

Vaibhav (VB) Srivastav ็š„ๅคดๅƒ
Vaibhav (VB) Srivastav1 ๅนดๅ‰

Model checkpoint:

Vaibhav (VB) Srivastav ็š„ๅคดๅƒ
Vaibhav (VB) Srivastav1 ๅนดๅ‰

Github repo:

Qingkai Fang ็š„ๅคดๅƒ
Qingkai Fang1 ๅนดๅ‰

Thanks for sharing our work!

Vaibhav (VB) Srivastav ็š„ๅคดๅƒ
Vaibhav (VB) Srivastav1 ๅนดๅ‰

๐Ÿ”ฅ

Tommy D. Rossi ็š„ๅคดๅƒ
Tommy D. Rossi1 ๅนดๅ‰

I wouldn't call this end to end, let's keep that term for single multi modal models that do everything by themselves

ThisAndThat ็š„ๅคดๅƒ
ThisAndThat1 ๅนดๅ‰

less than 250ms latency on what?

Vaibhav (VB) Srivastav ็š„ๅคดๅƒ
Vaibhav (VB) Srivastav1 ๅนดๅ‰

Time to first audio chunk according to their GH.

Waifuology ็š„ๅคดๅƒ
Waifuology1 ๅนดๅ‰

License looks good, but the voice quality isn't really there yet.

Hiro ็š„ๅคดๅƒ
Hiro1 ๅนดๅ‰

Do you know what are supported languages?

Trying my best :-) ็š„ๅคดๅƒ
Trying my best :-)1 ๅนดๅ‰

Can it detect emotion?

็›ธๅ…ณ่ง†้ข‘

๐Ÿ”ฅHOLY SMOKES! $TAO holders! ๐Ÿš€ SUBNET 19 (VISION) ON BITTENSOR IS ABSOLUTELY CRUSHING IT! In my 5+ years covering crypto and AI, this is one of the most impressive implementations I've seen. The combination of scale, performance, and decentralization is absolutely next level! ๐Ÿš€ @namoray_dev @Corcel_X ๐Ÿ’จ INSANE Speed Performance: - Llama 3.1 8B: 196.18 tokens/s with +107.23% advantage - Llama 3.1 70B: 124.96 tokens/s with +154.96% advantage - Llama 3.2 3B: 166.69 tokens/s with +21.66% advantage ๐Ÿ”ฅ Top Tier Model Integration: - Meta-Llama-3-70B & 8B Instruct - FLUX.1-schnell for Text-to-Image - ProteusV0.4-Lightning (Text & Image) - Multiple model variations for redundancy ๐Ÿ”ฅ What Makes This INSANE: - Complete decentralization - No single point of failure - Multiple model choices for redundancy - Real-time performance tracking - Transparent incentive structure The incentive distribution curve shows a healthy network with: - Strong rewards for top performers - Fair distribution across all participants - Clear path for growth and improvement - Sustainable economic model What's truly MIND-BLOWING is how they've managed to: 1. Scale to millions of operations 2. Maintain high quality across multiple tasks 3. Create a fair, competitive marketplace 4. Build in redundancy and reliability 5. Achieve true decentralization This isn't just another subnet - this is the future of decentralized AI inference happening RIGHT NOW! ๐Ÿ”ฅ 1. MASSIVE Scale & Adoption: - We're seeing 7M+ tokens being processed - 14K+ processing steps being executed - Multiple AI models running simultaneously - Incredible miner participation across the network 2. Revolutionary Task Distribution: - Llama 3.1 70B leading with 20% weighting - Avatar Generation at 15% - Perfectly balanced task distribution for optimal network performance - Multiple specialized tasks including Text-to-Image and Image-to-Image processing 3. Elite Performance Metrics: - Top miners hitting 0.00775 incentive rates - Consistent performance across the network - Impressive scaling from top to bottom performers - Strong incentive curve maintaining network quality ๐Ÿ“ˆ Network Performance: - Consistent upward trend in tokens/s - Quality scores maintaining high levels (>0.9) - Steady improvement in miner performance - Rock-solid network reliability โšก Platform Highlights: - Permissionless, serverless architecture - Global network of Always-On GPUs - Instant API access - Full decentralization - Multi-model support with seamless switching What makes this TRULY SPECIAL is the consistent upward trajectory in both speed and quality, while maintaining a decentralized architecture. The performance advantages over industry standards (+154.96% for 70B!) are absolutely mind-blowing! ๐Ÿš€ This isn't just another AI subnet - it's a glimpse into the future of decentralized AI inference! The combination of speed, reliability, and model variety makes this one of the most impressive implementations in the space! ๐Ÿ”ฅ ๐Ÿ“ฝ Watch Now on YouTube and TikTok: Source ๐Ÿ”—

Andy ฯ„ฯ„

11,616 ๆฌก่ง‚็œ‹ โ€ข 1 ๅนดๅ‰

Chinese robotics company Astribot released their latest World-Action Model (WAM), Lumo-2. Technical breakdown: - based on a frozen ๐Ÿฅถ Qwen-3.5 4B VLM - trained in 3 progressive stages: 1. Action is aligned with latent world dynamics (an abstract representation of action). Real-world actions are anchored to physical constraints, while the latent space is guided to focus on motion-relevant changes. This bidirectional relationship makes the model physically grounded -> critical for a world model. 2. Action is aligned with vision and language. Reusing the vision backbone and action encoder from the frozen VLM, the authors add a custom vocabulary (for new actions), a semantic module, an action decoder, and an action projector. This aligns the (new) action representations with the (existing) vision-language semantic space. Most importantly: it builds a direct mapping from natural-language instructions to motor execution. 3. End-to-end training on language, video, and robot data. Only the new modules (everything outside the frozen backbone) are trained end-to-end across temporal reasoning, physical understanding, long-horizon, and dexterous manipulation. At the end of the day, Lumo-2 is not the best on benchmarks, but that's not the point. What's genuinely new: - a way to combine latent world modeling and action generation through progressive alignment - a physically-grounded latent dynamics space - it lifts performance on unseen objects using un-annotated human egocentric video + Vision Pro captures, no special transfer algorithm needed Why it matters: - the whole model is thin trainable adapters (semantic module, action decoder/projector) on a frozen 4B backbone (cheap) - that scale is suited for real-time embedded inference (~2.71ร— decode speedup, no accuracy loss) - its real moat is long-horizon execution, where the added temporal memory pays off far more than on any other task As a result, this robot can now make your latte (5x sped up video):

Lรฉo

32,173 ๆฌก่ง‚็œ‹ โ€ข 23 ๅคฉๅ‰

QVAC SDK 0.15.0 is live. This release adds multiple prompts batching, brings a native AMD GPU backend to the stack, moves more vision encoders onto mobile GPUs, and adds a second local coding-agent integration. Main highlights: - Prompt batching for the LLM addon. Batch multiple prompts into one job and process them concurrently, with each answer returned the moment its generation finishes. - Native AMD GPU backend. A first-class HIP/ROCm backend in @qvac/vla-ggml, auto-selected over Vulkan with clean fallback when ROCm is absent. - A second local coding agent. OpenClaw joins OpenCode for local, cloud-free agent workflows. AGENTS - OpenCode plugin update (@qvac/opencode-plugin). Aligned with the current SDK, CLI, and AI SDK provider packages. A fresh install runs OpenCode against managed local QVAC models out of the box, from the default qvac/qwen3.5-9b, with no manual qvac serve setup. - OpenClaw plugin (@qvac/openclaw-plugin). A second coding-agent integration alongside OpenCode. A fresh setup installs the plugin, creates a local qvac provider through onboarding, and runs a QVAC model through OpenClaw๐Ÿฆž's local service path. LANGUAGE MODELS - Prompt batching (LLM addon). Batch multiple prompts in one job and run them concurrently, each answer returns the moment its generation finishes, no waiting on the others. - Reasoning-context trimming on hybrid + recurrent models (@qvac/llm-llamacpp). remove_thinking_from_context now works beyond pure-attention models. Same JS API, no throw. VOICE AND SPEECH - Transcription (transcription-parakeet 0.9.0). More robust CPU fallback on GPU failure and a faster Vulkan backend on Pixel 9. - Text-to-speech features (tts-ggml 0.4.0). Adds LavaSR for noise removal and adjustable output frequency up to 48 kHz, plus Japanese via Chatterbox. - Text-to-speech fixes (tts-ggml 0.4.1). CPU fallback on GPU failure, a q8_0 KV crash fix on Metal with Chatterbox. VISION - Qwen3.5 vision encoder on GPU (Android). Image encoder moves onto the phone GPU, with a smarter tile-grid preprocessor and default image-token caps, for flagship Android: Vulkan on Mali (Pixel 9 Pro) and OpenCL on Adreno 830 (Galaxy S25). - Gemma-4 vision encoder on GPU (Android). Vision encoder runs on the phone GPU instead of CPU, same flagship Android targets. PLATFORM AND PERFORMANCE - AMD GPU backend (@qvac/vla-ggml). Native HIP/ROCm backend, auto-selected over Vulkan with clean fallback when ROCm is absent (Linux x64 only). Comes with ~23% faster than Vulkan, ~14% faster than PyTorch-ROCm, parity preserved. Unified code style. A cleaner, more consistent, easier-to-contribute codebase. Let's build. npm install @qvac/sdk

QVAC

29,259,075 ๆฌก่ง‚็œ‹ โ€ข 1 ไธชๆœˆๅ‰

Google just proved that bigger isn't always better. Their 308M parameter model is outperforming models 2x its size. Google just released ๐—˜๐—บ๐—ฏ๐—ฒ๐—ฑ๐—ฑ๐—ถ๐—ป๐—ด๐—š๐—ฒ๐—บ๐—บ๐—ฎ, and it's proving that lightweight embedding models can punch way above their weight class. At just 308M parameters (578MB), it's the new state-of-the-art for models under 500M parameters across MTEB multilingual, English, and code benchmarks. But the really impressive part is that it ranks 8th overall on MTEB(Multilingual, v2) - that's ๐Ÿญ๐Ÿณ ๐—ฝ๐—น๐—ฎ๐—ฐ๐—ฒ๐˜€ above the second-best sub-500M model, and it's delivering performance ๐—ฐ๐—ผ๐—บ๐—ฝ๐—ฎ๐—ฟ๐—ฎ๐—ฏ๐—น๐—ฒ ๐˜๐—ผ ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น๐˜€ ๐—ป๐—ฒ๐—ฎ๐—ฟ๐—น๐˜† ๐—ฑ๐—ผ๐˜‚๐—ฏ๐—น๐—ฒ ๐—ถ๐˜๐˜€ ๐˜€๐—ถ๐˜‡๐—ฒ. There are three key parts of their training recipe that sets it apart: ๐Ÿญ. ๐—˜๐—ป๐—ฐ๐—ผ๐—ฑ๐—ฒ๐—ฟ-๐——๐—ฒ๐—ฐ๐—ผ๐—ฑ๐—ฒ๐—ฟ ๐—œ๐—ป๐—ถ๐˜๐—ถ๐—ฎ๐—น๐—ถ๐˜‡๐—ฎ๐˜๐—ถ๐—ผ๐—ป Instead of starting from a decoder-only Gemma 3 model, they first adapted it to encoder-decoder, then used just the encoder. By basing EmbeddingGemma off an LLM that already has world and language understanding, it gives it a stronger starting point. ๐Ÿฎ. ๐—ง๐—ต๐—ฟ๐—ฒ๐—ฒ-๐—Ÿ๐—ผ๐˜€๐˜€ ๐—ง๐—ฟ๐—ฎ๐—ถ๐—ป๐—ถ๐—ป๐—ด They combine three different loss functions, instead of just having one: โ€ข Contrastive loss (NCE) with in-batch negatives and hardness weighting โ€ข Spread-out regularization to ensure embeddings utilize the full space (for quantization and ANN retrieval) โ€ข Embedding matching distillation from Gemini Embedding - not just learning from relevance scores, but directly aligning the embedding space with the teacher model ๐Ÿฏ. ๐— ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ฆ๐—ผ๐˜‚๐—ฝ๐—ถ๐—ป๐—ด Rather than just averaging checkpoints from the same training run, they use optimization techniques to find multiple specialized training mixtures. Each mixture creates an "expert" model in different domains, and averaging all their parameters creates a final model that's actually better than individual models. Extras: โ€ข Matryoshka embeddings supporting 768, 512, 256, and 128 dimensions โ€ข Quantization-aware training - maintains quality even at int4 precision โ€ข 100+ languages from Gemma 3 pretraining โ€ข Exceptional performance on low-resource languages (check their XTREME-UP results) Is it the absolute best embedding model? No - Gemini Embedding still leads overall. But that's not really the point. EmbeddingGemma proves you can achieve state-of-the-art performance in a small package that's actually deployable on-device, in low-latency applications, and in resource-constrained environments. This makes good embeddings accessible for use cases that I'm seeing more and more: offline applications, privacy-sensitive deployments, and high-throughput scenarios where inference cost actually matters. Full paper: Shoutout to the EmbeddingGemma team at Google DeepMind for this awesome open source work ๐Ÿ’™ and to Daniel Williams for helping me with this video! ๐Ÿซถ

Victoria Slocum

21,610 ๆฌก่ง‚็œ‹ โ€ข 9 ไธชๆœˆๅ‰