Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Marvis-TTS-v0.2 is here 🚀 A local first TTS model capable of realtime performance even on older iPhones that Lucas Newman and I built. What’s new: ✨ Blazing fast — 100M (tiny) & 250M parameter models 🌍 Multilingual — English, French, German 🎭 Enhanced voice cloning — More natural &...

118,492 Aufrufe • vor 10 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,486 Aufrufe • vor 4 Monaten

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 Aufrufe • vor 3 Monaten

What if your voice AI could interrupt you the moment it figured out your question - sometimes even before you finished asking it? Last week, I sat down with Neil, CEO of Gradium and co-founder of Kyutai , to talk about the future of speech-to-speech models and why he believes today's cascaded voice systems will soon look "archaic and brittle." Some highlights from our conversation: 🎯 How Kyutai built Moshi—a full duplex conversational AI with "negative latency"—in 6 months with just 4-6 people (while big tech teams had 10-20x the resources) 🧠 Why speech-to-speech models lose intelligence compared to their text counterparts (and what's being done about it) 📱 Pocket TTS: The first voice cloning model that runs on your phone's CPU—not GPU, CPU 🤖 Why robotics and spatial audio represent the next frontier (hint: current voice systems completely break in these environments) 👶 The efficiency gap: Babies learn to speak fluently from <5,000 hours of audio. Current models train on millions of hours. We're doing something wrong. My favorite vision from Neil? The first truly contrarian AI that interrupts you mid-sentence to tell you why you're wrong. Not just more natural conversation—but actually useful for testing ideas and playing devil's advocate. Full episode and detailed blog post linked in the comments 👇 What's your take - will speech-to-speech replace cascaded systems, or will modularity keep cascaded architectures dominant even as naturalness improves?

Brooke Hopkins

13,000 Aufrufe • vor 6 Monaten

🐆 Rapid-MLX v0.12 is here. We’ve officially evolved from a simple chat app into a full-fledged, on-device AI studio for Apple Silicon! 🖥️✨ We didn't just push the MLX inference engine to its limits and expand support for a massive lineup of local open-source models—we are alpha-launching the highly anticipated Desktop Version. (A huge shoutout to the IoTeX community for grinding through the closed beta with us. Your feedback was incredible and helped shape this beast.) Here are the game-changing features you can run on your Mac right now, 100% free and 100% offline 👇 🚀 Blazing Fast Local LLMs Run anything from 4B up to Qwen3.5-122B completely offline. No guessing games—we recommend models matched perfectly to your Mac's actual RAM. Rich chat includes syntax highlighting, markdown tables, and honest tok/s metrics. 🎨 Local Image Generation A brand new Images tab to render directly on your machine. Pick a model (FLUX.2-klein, Z-Image-Turbo), prompt, and refine. Everything lands in a visual filmstrip. 👁️ Vision & Live Web Tools Attach an image and chat about it with local vision. Need real-time data? Our built-in web tools (weather, search, page-fetch) run mid-answer with strict, transparent privacy controls. 🤖 Plug-and-Play Coding Agents Wire up Claude Code, Codex, Cline, or Continue in seconds. One copy-paste from the Launch tab spins up a local OpenAI/Anthropic-compatible endpoint. 🔒 Private by Design Everything runs on-device. Signed, notarized, and entirely local. Your data stays yours. Turn your Mac into an AI powerhouse today. ⚡️

raullen

34,130 Aufrufe • vor 25 Tagen

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 Aufrufe • vor 1 Jahr

One week ago, we launched Typeahead. The product is already meaningfully better. We’ve shipped new features, Typeahead 2.0 is almost ready, and we’re already planning 3.0. 3.0 is going to be big. That pace is part of the point. Sam Asante and I started a new company together because we wanted to see how far we could push local AI software. Typeahead is the first thing we shipped. It is a local AI writing app for Mac. You type, suggestions appear inline, and it learns how you actually write. It works offline. You pay once. $79 and you own it. The product is intentionally simple. The thesis behind it is bigger. We think local models are going to create a new class of software. Fast, private and offline. By default. Personal without being creepy. Useful without turning everything into a subscription. Most AI products today assume the model lives in the cloud. That will not be the only path. The machines we already own are getting powerful enough. The models are getting small enough. And the experience can start to feel less like chatting with a remote service and more like using software that belongs on your computer. We have already built a few fully local experiments together. Typeahead felt right because the value is immediate. A few people have asked me how this fits with Crazy Egg. Crazy Egg is still my main work. This is a focused company with Sam, built around a thesis we both believe in. I’ve learned that the people I work with shape the work more than almost anything else. Sam was already on my short list of people I wanted to build with. Typeahead is the first public proof of the thesis. Watch the video. Get it here:

Hiten Shah

18,627 Aufrufe • vor 2 Monaten

Voice AI turn taking is a solved problem. The single most common complaint about voice AI, today, is that agents interrupt too often. But the voice agents I build for myself now respond quickly and interrupt me less often than the people I talk to every day. (I actually measured this.) Mark Backman made a Pipecat AI PR two weeks ago that was the last piece of the puzzle for turn taking so good that I no longer ever think about it. The approach combines three layers of processing: 1. Voice activity detection, with a short (200ms) trigger. 2. A native audio turn detection model that's small, fast, and runs on CPU. This model captures audio nuances like inflection and filler sounds that don't get transcribed. 3. A prompt mixin for the conversation LLM that decides turn completion based on conversation context. None of these are new. We've been using VAD for a long time. We trained the first version of the Pipecat Smart Turn native audio model in December 2024. And we've been experimenting with prompt-based large model turn detection (sometimes called "selective refusal") for more than a year. Now, the Smart Turn model and the SOTA LLMs we're using in voice agents have both gotten so good that using them together feels like we've finally "solved" turn detection. Mark also figured out how to elegantly apply a "single-token tagging" technique to this problem. We sometimes use single-token tagging in place of tool calling, when we need a near-zero latency programmatic trigger. Mark's Pipecat mixin defines three single-token characters and prompts the LLM to output exactly one of them at the beginning of every response. - ✓ means the agent should respond normally (immediately) - ○ is a "short incomplete" - the agent should wait 5 seconds - ◐ is a "long incomplete" - the agent should wait 10 seconds The wait times, and the details of the prompt, are configurable, of course. Watch the video to see me talk to an agent that handles all my various pauses and inflections, plus phrases like "let me think," pretty much the way a person would handle them, in terms of response latency. Also, in the second half of the video, I ask the agent to adjust its response pattern because I'm going to tell it a phone number. This kind of "in-context" adjustment of response wait times is really useful. The LLM in the video is GTP-4.1. We've tested the prompt and single-token adherance with GPT-4.1, Gemini 2.5 Flash, Anthropic Claude Sonnet 4.5, and AWS Nova 2 Pro. Note that older models in all these families (and, in general, smaller open weights models) aren't able to reliably output these single-token tags. But the new models we're using these days are pretty amazing.

kwindla

27,016 Aufrufe • vor 6 Monaten

First impressions on Muse Glimmer! It's incredibly fast for a dense model, currently running an average of 208tps with a max of 274tps on a single 5090 with their DFLASH config. Comparatively, though, both using Open Code, Qwopus Coder (with thinking off) produced a much better shark survival game than the one I got from Glimmer. Meta's new dense model is currently just lacking some HTML canvas taste, but this is something that can be added via SFT as long as the model is stable and capable from a back-end programming perspective. And it seems to be, without a doubt. The big kicker here is that I ran this at extra high thinking, and it did not take long at all to run. Our current local leader, Qwen 27B 3.6, has a tendency to overthink, but with glimmer, that is not the case. Right now, my recommendation for general local programming (Apps, Games, Websites, Visual Tools) in this class is still Qwopus Coder with thinking disabled, or Qwopus Fusion with thinking enabled. Of course Shark Survival is a very basic domain-specific test, but I find that the result scales very well across many domains. If we're going to be shipping apps generated entirely locally, visual taste is somewhat of a bare minimum requirement, solely in my opinion, and Qwen's models in this class offer significantly more at the moment. That's actually why I initially started getting into finetuning with Qwen 3.5, they were the first base that was able to do really good front-end with some opus-trace fine-tuning. Qwen 3.6 has taste even in the base model, and we know Qwen 3.8 is going to blow us all away! Regardless, this looks like a very tempting new base model. As a first offering from Meta in this class for a long time, I am incredibly impressed and elated to have it. We now finally have a proper Single GPU frontier race, instead of us just begging Qwen for more releases. Single GPU open frontier model race is a VERY good thing. Please keep pushing Meta

Kyle Hessling

17,485 Aufrufe • vor 25 Tagen

Today we're announcing $280M in Series B funding at a $2B valuation, led by our long-time investor and partner, Menlo Ventures. When we announced our Series A last May, most conversations I had about voice began with someone explaining why they doubted it. I don't have those conversations anymore. People tell me instead how much time they save talking instead of typing and what they want us to build next. In a little over a year, voice has gone from something people were questioning to something they rely on, and it happened faster than we expected. That shift is why we've raised a new round of funding. This funding represents a deeper investment in our products, lab, models, and our team. Alongside the funding, we're announcing a preview of our first proprietary speech model, Canto. Canto is a 2B parameter speech model trained for the places people actually talk: loud rooms, windy streets, a toddler in the background, a second language mixed into the first. Models like this will change how we all interact with devices, and Canto is the first in a rapid line of them. Each one will be larger and more capable than the last, and each should improve Flow's dictation in a way you can feel the day it ships. Larger speech models also open the door to the vision we dreamed up 5 years ago: using your voice as the primary way you interact with devices. That still has to be invented. A voice interface people rely on all day, in every app, doesn't exist anywhere yet. It's time for that to change. To Sahaj Garg, our team, and every person who gave Wispr Flow a chance - this one's for you.

Tanay Kothari

613,494 Aufrufe • vor 18 Tagen

Sundar Pichai just confirmed that nobody on earth is using Google’s best model. Including him. Pichai: “The models we all use the most is maybe like a few months behind the maximum capability we can deliver.” Finished intelligence is sitting in reserve, built and working, waiting for the economics to catch up. What reaches you is not the edge of research. It is the edge of what can be served to billions of people without losing money on every request. Pichai: “for each generation, we feel like we’ve been able to get the Pro model at like, I don’t know, 80-90% of Ultra’s capability.” Google stopped shipping Ultra. The capability existed. Serving it did not. Pichai: “But what we’ve been able to do is to go to the next generation and make the next generation’s Pro as good as the previous generation’s Ultra.” The frontier gets built, held back, compressed, and released a generation later at a price that works. Every efficiency gain drains part of that backlog at once, which is why progress keeps arriving in jumps that feel larger than the research behind them. You are not watching discovery. You are watching a queue clear. Benchmarks stopped tracking any of this, because they measure the ceiling and almost nobody works at the ceiling. Fridman: “benchmarks are less and less capable of capturing the intelligence of models, the effectiveness of models.” A model that is slightly less capable and dramatically faster wins nearly every real task, because latency decides what you are willing to ask in the first place. Fridman: “you could argue Gemini Flash is much more impactful than Pro. Just because of the latency, it’s super intelligent already.” Wait thirty seconds for an answer and you ask a few questions a day. Get it back instantly and you rebuild your workflow around it. Value is created at the point of use, never at the top of a leaderboard. Intelligence stopped being the scarce input. Delivery became the constraint, and delivery is an engineering problem, which is the one category of problem humans have never lost. Electricity was real for decades before it reached a kitchen. The invention finishes early. Distribution takes longer, and it always gets built. Scaling did not stall. What closed is the gap between what exists and what reaches you. The best model in the world has already been built. What is left is a cost curve, and cost curves only run one direction.

Dustin

125,137 Aufrufe • vor 26 Tagen

Here’s my written & video review of the new 2026 Tesla Model Y Performance after driving it for a week. This is the best-value new Tesla you can buy. Crazy performance, no real drawbacks, and all for just $57,490. Let’s dive in. Price: I haven’t seen many others mention this: every option on the Model Y Performance is included at no extra cost in the US (except FSD). So if you spec a Model Y Premium AWD with an upgraded paint color, tow package, white interior, and upgraded 20" wheels, a fully loaded Model Y Premium AWD ends up only about $2,500 less expensive than a fully loaded Model Y Performance. Ride Quality: I thought I might feel worse ride quality vs my Premium AWD Model Y with 20" wheels, but I struggled to find any real difference, despite the larger 21" wheels, firmer suspension setting, and 0.6" lower ride height on the Performance trim. A true testament to Tesla's engineering magic on this thing. Exterior Design: Unlike the previous Model Y Performance, Tesla made some exterior design tweaks with a new front and rear fascia to spice things up a bit. The result, in my opinion, is the best-looking Model Y trim you can buy. It definitely has a more aggressive presence in person, even if it's subtle. The carbon-fiber spoiler boosts high-speed stability and cuts aerodynamic drag by 10%. The new 21’’ Arachnid 2.0 wheels look fantastic in person, one of my favorite designs ever from Tesla. Staggered wheel and tire fitment provides better grip and steering. The beefier 275mm rear tires (255mm in the front) give the vehicle a better stance from behind. Vehicle-to-Load (V2L): For the first time on a Model Y in North America, you can now plug in anything you want to the exterior charge port with an adapter, even a campsite! It provides up to 2.4 kW of power (120V at 20A) from two household outlets. It's a great feature. Interior: It’s what you know and love, but with a few changes that elevate the ownership experience. New with the Performance is a larger 16" center screen (vs. 15.4" on non-Performance models), with thinner bezels and higher resolution. It’s not a huge difference on paper, but you definitely notice it in daily use. The carbon-fiber décor on the door cards and dash is a nice touch, though I would like to see it extended to the center console. The new performance seats are the best seats of any Tesla I've ever experienced. While they retain aggressive bolstering in the torso area, the bottom seat cushion has less aggressive bolstering than on the Model 3 Performance seats, making it easier to get in and out. It also doesn’t squeeze your thighs too tightly. The powered thigh extenders add comfort on longer drives, especially for taller people who want extra support. And of course, they’re heated and ventilated. The headrests also feel more comfortable than the ones in my Model Y. I want these seats. Unlike the old Model Y Performance, the new one has no Track Mode. Why? Because nobody used it lol. No point in putting engineering resources into something people won’t use. The refreshed Model 3 Performance still has it, though. Cabin Quietness: Despite the thinner-profile tires, there is no noticeable difference vs my Premium Model Y. Decibel reading results at highway speeds were visually the same compared to my 2026 Model Y Premium (65-66). Driving Impressions: It’s amazing. Sharp, precise, and agile. The vehicle feels stable at all times. Acceleration is blistering (3.3s 0–60 mph), with plenty of punch even at higher speeds. The tires offer good grip, and cornering is fantastic for an SUV. The suspension setup is great. More steering wheel feedback would be nice, though. Cruising around traffic is a joy. The brakes are much improved over the previous Model Y Performance and are far better suited for spirited driving. There are three acceleration modes: Chill, Standard, and Insane. Just stay in Insane. You’d be insane not to lol. The car also lets you switch between two ride and handling modes: Standard and Sport. The difference isn’t huge, but Standard is better if you’ve got passengers. FSD: I unfortunately wasn’t able to get FSD V14 on this car. It had V13.2.9, so I didn’t use it much. But in the little time I did, it was smooth and comfortable. It didn’t bother me because V14 will perform just as well here as it does on my 2026 Model Y Premium AWD. Conclusion: You won’t find another new SUV today that offers this level of performance for the price. Back in 2022, when Tesla couldn’t build Model Ys fast enough, a fully loaded Model Y Performance cost over $90,440. Today, the refreshed and far more capable 2026 Model Y Performance is just $57,490 fully loaded, and you can simply subscribe to FSD for $99/month. The 2026 Model Y Performance delivers utility, great performance, comfort, tech, self-driving and everything else people love about the Model Y. It’s a no-brainer purchase. I want one badly, but I'll need to show restraint, as I’m saving up for a house lol.

Sawyer Merritt

224,054 Aufrufe • vor 9 Monaten

Today, we’re excited to announce our $50M Series B, led by Greenfield Partners (formerly TPG Capital), with participation from Lightspeed and Notable Capital. 🚀 At PatronusAI, we develop simulations and evals to train and improve AI. The first phase of AI was built on static benchmarks, but that era is over now. As agents are used to solve longer and longer tasks, they need to practice in dynamic, living worlds to get better. Simulations are the critical infrastructure powering this next phase. As a company, we’re behind the most influential research and products in AI evaluation, like FinanceBench, Lynx, and Percival. And things have moved at the speed of light since. ⚡ We partner with the world's leading frontier AI labs and enterprises, and our revenue has grown more than 15x over the past year. Additionally, today, we’re introducing a preview of the first Digital World Model for AI agent training and simulation: Patronus-DWM. Digital World Models are language diffusion world models that predict realistic environment behaviors and steer agent actions across digital workflows. Just as physical world models predict how objects move through space, we’re developing the equivalent for the digital world: predicting how agents act in digital workflows, then using that to scale the creation of high-quality training data for LLMs. Digital World Models help us push the frontier of ultra long horizon workflows, and unlock a new class of self-improving RL environments. This is our scalable approach to simulating all of the world’s intelligence. The round was also joined by Datadog, Inc., Samsung Ventures, Gokul Rajaram, Factorial Capital, and a large cohort of amazing AI leaders and researchers across Anthropic, OpenAI, Google DeepMind, NVIDIA, Recursive, and more. ✨ It has been the ride of a lifetime. But we’re just getting started. The best is yet to come. "Do not go gentle into that good night, Rage, rage against the dying of the light" - Dylan Thomas (1954)

Anand Kannappan

42,227 Aufrufe • vor 2 Monaten