Sensitive content

This media may contain sensitive content.

Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Serena | The First Descendant Models: Madruga3D Audio: Evilaudio Voice: DelaliciousVA #R34 #nsfw #r34nsfw #3dartnsfw #rule34 #Serena #serenatfd #TheFirstDescendant

67,943 Aufrufe • vor 6 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

What if your voice AI could interrupt you the moment it figured out your question - sometimes even before you finished asking it? Last week, I sat down with Neil, CEO of Gradium and co-founder of Kyutai , to talk about the future of speech-to-speech models and why he believes today's cascaded voice systems will soon look "archaic and brittle." Some highlights from our conversation: 🎯 How Kyutai built Moshi—a full duplex conversational AI with "negative latency"—in 6 months with just 4-6 people (while big tech teams had 10-20x the resources) 🧠 Why speech-to-speech models lose intelligence compared to their text counterparts (and what's being done about it) 📱 Pocket TTS: The first voice cloning model that runs on your phone's CPU—not GPU, CPU 🤖 Why robotics and spatial audio represent the next frontier (hint: current voice systems completely break in these environments) 👶 The efficiency gap: Babies learn to speak fluently from <5,000 hours of audio. Current models train on millions of hours. We're doing something wrong. My favorite vision from Neil? The first truly contrarian AI that interrupts you mid-sentence to tell you why you're wrong. Not just more natural conversation—but actually useful for testing ideas and playing devil's advocate. Full episode and detailed blog post linked in the comments 👇 What's your take - will speech-to-speech replace cascaded systems, or will modularity keep cascaded architectures dominant even as naturalness improves?

Brooke Hopkins

13,000 Aufrufe • vor 6 Monaten

Voice used to be AI’s forgotten modality - now it's having its big moment: rapid innovation, big funding rounds, major agentic applications My conversation with Neil Zeghidour, top AI researcher in the field (Google DeepMind, Meta, kyutai) and now CEO of Gradium This is a reference episode on all things voice AI 🔥 00:00 Intro 01:21 Voice AI’s big moment, and why we’re still early 03:34 Why voice lagged behind text/image/video 06:06 The convergence era: transformers for every modality 07:40 Beyond Her: always-on assistants, wake words, voice-first devices 11:01 Voice vs text: where voice fits (even for coding) 12:56 Neil’s origin story: from finance to machine learning, with help from Yann LeCun and Soumith Chintala 18:35 Neural codecs (SoundStream): compression as the unlock 22:30 Kyutai: open research, small elite teams, moving fast 31:32 Why big labs haven’t “won” voice AI4 34:01 On-device voice: where it works, why compact models matter 46:37 The last mile: real-world robustness, pronunciation, uptime 41:35 Benchmarking voice: why metrics fail, how they actually test 47:03 Cascades vs speech-to-speech: trade-offs + what’s next 54:05 Hardest frontier: noisy rooms, factories, multi-speaker chaos 1:00:50 New languages + dialects: what transfers, what doesn’t 1:02:54 Hardware & compute: why voice isn’t a 10,000-GPU game 1:07:27 What data do you need to train voice models 1:09:02 Deepfakes + privacy: why watermarking isn’t a solution 1:12:30 Voice + vision: multimodality, screen awareness, video+audio 1:14:43 Voice cloning vs voice design: where the market goes 1:16:32 Paris/Europe AI: talent density, underdog energy, what’s next

Matt Turck

22,980 Aufrufe • vor 6 Monaten

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 Aufrufe • vor 3 Monaten

Voice AI turn taking is a solved problem. The single most common complaint about voice AI, today, is that agents interrupt too often. But the voice agents I build for myself now respond quickly and interrupt me less often than the people I talk to every day. (I actually measured this.) Mark Backman made a Pipecat AI PR two weeks ago that was the last piece of the puzzle for turn taking so good that I no longer ever think about it. The approach combines three layers of processing: 1. Voice activity detection, with a short (200ms) trigger. 2. A native audio turn detection model that's small, fast, and runs on CPU. This model captures audio nuances like inflection and filler sounds that don't get transcribed. 3. A prompt mixin for the conversation LLM that decides turn completion based on conversation context. None of these are new. We've been using VAD for a long time. We trained the first version of the Pipecat Smart Turn native audio model in December 2024. And we've been experimenting with prompt-based large model turn detection (sometimes called "selective refusal") for more than a year. Now, the Smart Turn model and the SOTA LLMs we're using in voice agents have both gotten so good that using them together feels like we've finally "solved" turn detection. Mark also figured out how to elegantly apply a "single-token tagging" technique to this problem. We sometimes use single-token tagging in place of tool calling, when we need a near-zero latency programmatic trigger. Mark's Pipecat mixin defines three single-token characters and prompts the LLM to output exactly one of them at the beginning of every response. - ✓ means the agent should respond normally (immediately) - ○ is a "short incomplete" - the agent should wait 5 seconds - ◐ is a "long incomplete" - the agent should wait 10 seconds The wait times, and the details of the prompt, are configurable, of course. Watch the video to see me talk to an agent that handles all my various pauses and inflections, plus phrases like "let me think," pretty much the way a person would handle them, in terms of response latency. Also, in the second half of the video, I ask the agent to adjust its response pattern because I'm going to tell it a phone number. This kind of "in-context" adjustment of response wait times is really useful. The LLM in the video is GTP-4.1. We've tested the prompt and single-token adherance with GPT-4.1, Gemini 2.5 Flash, Anthropic Claude Sonnet 4.5, and AWS Nova 2 Pro. Note that older models in all these families (and, in general, smaller open weights models) aren't able to reliably output these single-token tags. But the new models we're using these days are pretty amazing.

kwindla

27,016 Aufrufe • vor 6 Monaten

🚨 JUST IN: MICROSOFT just open sourced a VOICE AI THAT TRANSCRIBES 60 MINUTES OF AUDIO in a single pass. 100% FREE. It knows who spoke. It knows when they spoke. It knows exactly what they said. All in one shot. No chunking. No context loss. It's called VibeVoice. Not a transcription tool. Not a basic speech to text wrapper. A frontier voice AI family with ASR, TTS, and real time streaming. All open source. All free. Here's what it actually does 👇 VibeVoice ASR - Speech Recognition: → Processes 60 minutes of continuous audio in a single pass → Never slices audio into chunks so global context is never lost → Identifies WHO spoke, WHEN they spoke and WHAT they said simultaneously → Supports customized hotwords for domain specific accuracy → Works in 50+ languages natively → Already adopted by Hugging Face Transformers library → Already being built on by the open source community BY PEOPLE WHO HAD NO IDEA THIS LEVEL OF ACCURACY WAS ALREADY FREE. VibeVoice TTS - Text to Speech: → Generates up to 90 minutes of speech in a single pass → Supports up to 4 distinct speakers in one conversation → Natural turn taking and speaker consistency throughout → Expressive speech that captures emotional nuances → Supports English, Chinese and multiple other languages VibeVoice Realtime - Streaming TTS: → Only 300 millisecond first audible latency → Streams text input in real time → 0.5B parameters so it actually deploys anywhere → Robust long form generation up to 10 minutes → Lightweight enough for production use today The core innovation nobody is talking about: Most voice AI models slice long audio into short chunks. Every time they slice, they lose context. Speaker tracking breaks. Semantic coherence breaks. Accuracy drops. VibeVoice uses continuous speech tokenizers running at an ultra low frame rate of 7.5 Hz. This preserves audio fidelity while dramatically boosting computational efficiency. The entire 60 minutes stays in context. Nothing gets lost. Nobody gets misidentified. The numbers: → VibeVoice ASR 7B - available now on Hugging Face → VibeVoice Realtime 0.5B - try it on Colab right now → 50+ supported languages → 11 distinct English voice styles → 9 multilingual speaker voices → Already integrated into Hugging Face Transformers → Finetuning code now available The wildest part? A voice powered input method called Vibing just built itself on top of VibeVoice ASR. Available on macOS and Windows right now. The open source community is already shipping products on top of this. 100% Open Source. Free to use. Free to fine tune. Free to build on. 🔖 Save this before your competitors find it first. 👇

Kanika

221,558 Aufrufe • vor 5 Monaten

BREAKING: Inside The ElevenLabs Summit The future of voice-first interfaces. CEO, Mati Staniszewski (@matiii) Head of Growth, Luke Harries (Luke Harries) Series A Lead, Bryan Kim (Bryan Kim) of a16z Klarna CEO, Sebastian Siemiatkowski (Sebastian Siemiatkowski) AIUC CEO, Rune Kvist (Rune Kvist) Founded in just 2022, has scaled to $300M ARR & raised $781M in total funding across five funding rounds, most recently closing a $500M Series D led by Sequoia Capital at an $11 billion valuation. ElevenLabs began with a breakthrough human-like text-to-speech model & has since expanded into speech-to-text, dubbing, sound effects, music, & conversational AI. Today, it combines these models with integrations & enterprise-grade infrastructure to power production-ready platforms for businesses, creators, & developers: - ElevenAgents enables enterprises to deploy voice and chat agents at scale, with customers including Deutsche Telekom, Revolut, Square, & the Ukrainian government using it for support, commerce, and citizen services. - ElevenCreative allows brands like Duolingo, NVIDIA, & TIME to generate and localize high-quality audio in 70+ languages. - ElevenAPI provides low-latency voice infrastructure for developers, powering platforms from Meta & Epic Games to Salesforce, MasterClass, & Harvey—reaching more than one billion users globally. Leaving the Summit, after seeing the product depth, long-term vision, and caliber of operators & partners around them, it’s clear ElevenLabs has built one of the strongest & fastest-growing companies in AI today. Really impressive.

Molly O’Shea

143,008 Aufrufe • vor 6 Monaten

✨ I made a 100% AI video of a single night in 🇩🇪 Berlin's underground techno scene Doing an AI shoot is very much like a real life shoot, you have to prepare locations, settings, outfits and models, you can do the same: 1) First I generated photos with the [ 🇩🇪 Berlin nightlife ] photo pack on they're all non-existing AI people except me and my gf 2) Then I put the pics I liked in Kling AI to make them into 10 second videos (with Professional Mode), soon I'll have API access and it'll be one-click inside Photo AI to turn a photo into a 10-sec video, it then takes about 5 minutes to make a video. It's not always great, especially dancing is hard to get right! 3) Then I wanted to have a German narrator talking about Berlin's techno nights in a poetic way like a documentary, so I asked ChatGPT 4o: "write in the style of a cult German novelist about the Berlin techno scene" I then copied that text into Eleven Labs and selected a German AI voice 4) I then collected all the videos, the music and the audio narrator and edited them together in Final Cut Pro adding music from HÖR and adding audio effects 5) Then I added English captions with CapCut Time to make it: 3 hours for 1 minute of video I hope you like it! My dream is to have this entire pipeline in Photo AI at some point. Connecting to Kling AI's video API is the first step to that. Imagine just writing a prompt and you end up with a video like this without all the work Credits: Music: Asquith - Let Me (Rave Mix) [ASQ004] Barbax - Si Vis Pacem Para Bellum (Original Mix) Mixed by Ellen Allien @ HÖR Video: 100% AI characters and video by Photo AI Edited by me Narrator voice: 100% written by ChatGPT 4o 100% narrated by Eleven Labs AI voice

@levelsio

676,094 Aufrufe • vor 2 Jahren

🚨FOLLOW-UP: The Erika Kirk Audio is rREAL, and the context is even more DAMNING than we Initially Thought. Let's clear up the disinformation first. The audio file from the DOJ has no year listed, only April 20th. Anyone telling you they know the exact year is lying or speculating. The facts are what matter, and the facts are in the recording. Now, let's connect this bombshell audio to the other bombshell revelations from Candace Owens. Candace has already exposed Erika's deep ties to Epstein's network in both Romania and New York City. The most crucial piece of information? Erika was the real estate point person managing the apartment building where Jeffrey Epstein housed his underage victims. Candace revealed that multiple eyewitnesses from Next Model Management remembered Erika Kirk taking meetings at their office. Her job? To find housing for the young Eastern European models they were "packing like sardines" into apartments. She was the specific contact for a white building on the Upper East Side around 68th Street. Lo and behold, that building matches the description of 301 East 66th Street—a building owned by Mark Epstein, Jeffrey's brother—that newly unsealed records revealed was used to secretly house underage victims. The building functioned as a "logistical hub for Epstein's world," where girls were groomed and manipulated before being trafficked to his townhouse. This isn't a coincidence. This is the missing piece of the puzzle. If Erika Kirk was the real estate point person for Epstein's victim housing, her role on that DOJ audio call makes perfect, horrifying sense. She wasn't just some associate. She was the scheduler and procurer. The audio where she says, "me and him are gonna put a schedule together for you and your sister" is a direct reflection of her day-to-day job managing the logistics of Epstein's trafficking operation. I know some people are trying to muddy the waters by claiming it's Haley Robson. It's not. After months of investigative research, I have listened to more of Erika Kirk's speeches than probably anyone else. I have become an expert on her voice. Her vocal cadence, her tone, her distinct speech patterns—it's a perfect match. It's not even close. This is why Erika Kirk needs to respond. The silence is deafening. She must either deny that it is her voice on that DOJ audio recording or admit to it. Until she does, we have no choice but to accept this as the smoking gun proof that she was a key player in Jeffrey Epstein's child sex trafficking ring. CALL TO ACTION: Please do me a favor and downvote these paid shills and bots attacking me by attempting to community note my post below. This is an information war and they don''t want people to know the truth. If they are successful in community noting my post it will kill it's reach in the algorithm, That means the people who need to be woken up and see it them most will never get tha chance. I need your help. We must all stand together. Thanks for your attention to this mater.🙏 The context is everything. The dots are connected. What more proof do you need? Share this so they can't ignore it. 👇

Project Constitution

696,996 Aufrufe • vor 6 Monaten

Last Friday on Pi Day, we held AI Dev 25, a new conference for AI Developers. Tickets had (unfortunately) sold out shortly after we announced their availability, but I came away energized by the day of coding and technical discussions with fellow AI Builders! Let me share here my observations from the event. I'd decided to start AI Dev because while there're great academic AI conferences that disseminate research work (such as NeurIPS, ICML and ICLR) and also great meetings held by individual companies, often focused on each company's product offerings, there were few vendor-neutral conferences for AI developers. With the wide range of AI tools now available, there is a rich set of opportunities for developers to build new things (and to share ideas on how to build things!), but also a need for a neutral forum that helps developers do so. Based on an informal poll, about half the attendees had traveled to San Francisco from outside the Bay Area for this meeting, including many who had come from overseas. I was thrilled by the enthusiasm to be part of this AI Builder community. To everyone who came, thank you! Other aspects of the event that struck me: - First, agentic AI continues to be a strong theme. The topic attendees most wanted to hear about (based on free text responses to our in-person survey at the start of the event) was agents! - Google's Paige Bailey talked about embedding AI in everything and using a wide range of models to do so. I also particularly enjoyed her demos of Astra and Deep Research agents. - Meta's Amit Sangani talked compellingly as usual about open models. Specifically, he described developers fine-tuning smaller models on specific data, resulting in superior performance than with large general purpose models. While there're still many companies using fine-tuning that should really just be prompting, I'm also seeing continued growth of fine-tuning in applications that are reaching scale and that are becoming valuable. - Many speakers also spoke about the importance of being pragmatic about what problems we are solving, as opposed to buying into the AGI hype. For example, Nebius' Roman Chernin put it simply: Focusing on solving real problems is important! - Lastly, I was excited to hear continued enthusiasm for the Voice Stack. Justin Uberti gave a talk about OpenAI’s realtime audio API to a packed room, with many people pulling out laptops to try things out themselves in code! has a strong “Learner First” mentality; our foremost goal is always to help learners. I was thrilled that a few attendees told me they enjoyed how technical the sessions were, and said they learned many things that they're sure they will use. (In fact, I, too, came away with a few ideas from the sessions!) I was also struck that, both during the talks and at the technical demo booths, the rooms were packed with attendees who were highly engaged throughout the whole day. I'm glad that we were able to have a meeting filled with technical and engineering discussions. I'm delighted that AI Dev 25 went off so well, and am grateful to all the attendees, volunteers, speakers, sponsors, partners, and team members that made the event possible. I regretted only that the physical size of the event space prevented us from admitting more attendees this time. There is something magical about bringing people together physically to share ideas, make friends, and to learn from and help each other. I hope we'll be able to bring even more people together in the future. [Original text: ]

Andrew Ng

47,037 Aufrufe • vor 1 Jahr