Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Introducing VIVA 2.5, with our most advanced models yet. Krisp's Voice Isolation has been running in production for 2 years, across 1B+ mins of voice AI conversations a month, improving WER by focusing on primary speaker only. As more teams ran voice isolation in front of their STT, we...

25,958,617 Aufrufe • vor 1 Monat •via X (Twitter)

34 Kommentare

Profilbild von Alistair Ward
Alistair Wardvor 1 Monat

@krispHQ Looks like shit pal

Profilbild von .
.vor 1 Monat

@krispHQ Genuinely, what's the benefit of this technology?

Profilbild von simon
simonvor 1 Monat

@krispHQ I fucking hate ai

Profilbild von Kennedy Leflar🌹🕊✝️
Kennedy Leflar🌹🕊✝️vor 1 Monat

@krispHQ Still looks fake, the uncanny valley vibe will never leave AI slop

Profilbild von Tiny
Tinyvor 1 Monat

@krispHQ I dont. I'm old school i can still think for myself. Thats what a brain is for 😉

Profilbild von villager
villagervor 1 Monat

@krispHQ who needs 2 years of data to figure that out lol

Profilbild von Yori
Yorivor 23 Tagen

@krispHQ its so HD that it went from looking convincing back to looking uncanny lmao

Profilbild von coolman 6666-mar
coolman 6666-marvor 1 Monat

@krispHQ give me back my water

Profilbild von Berty
Bertyvor 1 Monat

@krispHQ I didn’t know Vauxhall fitted a 2.5 in the viva 🤷‍♂️😀

Profilbild von asdghmejfn
asdghmejfnvor 22 Tagen

@krispHQ Let’s get rid of this Ai slop and make em lose the money the spent to corrupt our minds and importantly our kids minds personally I don’t want my kid to use Ai to learn as it legit makes em dumb typing everything in and nog learning anymore

Profilbild von Menmaro
Menmarovor 21 Tagen

@krispHQ Looks and sounds like AI slop

Profilbild von Charlotte
Charlottevor 1 Monat

@krispHQ super intéressant de voir l'évolution de l'isolation vocale en prod. vous mesurez aussi la latence ajoutée par le traitement ?

Profilbild von Militant Boomer 🇵🇸 #FreePalestine
Militant Boomer 🇵🇸 #FreePalestinevor 22 Tagen

@krispHQ why don't u work on something that actually helps humanity instead of perfecting make-believe "customer service agents" that end up referring to actual human beings to answer the simplest questions —but then they don't even know their own business's customer service hours #Fail

Profilbild von A Buffalo
A Buffalovor 1 Monat

@krispHQ We don’t want any of this

Profilbild von JORDAN
JORDANvor 1 Monat

@krispHQ This should make voice agents feel much more reliable.

Profilbild von Franchaela Fingerton
Franchaela Fingertonvor 1 Monat

@krispHQ This shit is the reason glaciers are melting and leaves are falling in August. Let actors work ffs

Profilbild von DB
DBvor 1 Monat

@krispHQ Yeah, I had a Vauxhall Viva in the 70s, it wasn't that big though 😕

Profilbild von OwlishNick
OwlishNickvor 20 Tagen

@krispHQ Hey A.I engineer here and I worked on Viva and I can confirm it was trained on massive amounts of porn especially gay porn. I said "Davit why do you need all this for model training?" And he said it was very necessary to have terabytes of porn sent to him personally.

Profilbild von Lisa O'Brien
Lisa O'Brienvor 26 Tagen

@krispHQ I love him. ( But not as much as Liam Gallagher !! )

Profilbild von Mouaddib
Mouaddibvor 1 Monat

@krispHQ Why the fuck are people striving to achieve human like video to a point they become indistinguishable from reality. Are those fucks blind, egotistical or too stupid to see the danger? Or the danger is the goal?

Profilbild von Tom
Tomvor 1 Monat

@krispHQ The 3.5x smaller model is just as exciting as the accuracy gains. Better efficiency opens the door for many more deployments.

Profilbild von AI TALKs
AI TALKsvor 1 Monat

@krispHQ Love how you fixed the over-aggressive isolation issue without hurting clean audio. Clean progress.

Profilbild von caçacéfalos
caçacéfalosvor 1 Monat

@krispHQ Bora

Profilbild von 🧠 PocketIQ | Free AI Trading
🧠 PocketIQ | Free AI Tradingvor 1 Monat

@chesny @krispHQ 🔼 Less guessing. More analysis.

Profilbild von Iconic Icon
Iconic Iconvor 1 Monat

@krispHQ But the CFF is way too high. Fix it.

Profilbild von Jacob
Jacobvor 25 Tagen

@krispHQ The US remake of Peep Show looks terrible.

Profilbild von sabir hussain
sabir hussainvor 1 Monat

@krispHQ Better voice isolation means cleaner AI conversations

Profilbild von Joaquim Pinto
Joaquim Pintovor 29 Tagen

@krispHQ 2.5 É um vinho da beira interior, ( Caria )ótimo para acompanhar, caça, queijos e carnes vermelhas.

Profilbild von Bally_AgenticAI
Bally_AgenticAIvor 1 Monat

@krispHQ I read a lot of launch posts. Almost none admit the previous version had a regression. @krispHQ just did, and shipped the fix with receipts: 46.4% fewer word errors across 10 STT engines, no harm on clean audio. That is how you earn trust in voice infrastructure.

Profilbild von villager
villagervor 1 Monat

@krispHQ 1B minutes of convo, impressive! Any plans for multi‑speaker support?

Profilbild von ArifIno
ArifInovor 1 Monat

@krispHQ @BrainPathio

Profilbild von BrainPath.io
BrainPath.iovor 1 Monat

@krispHQ Woooow 🚀

Profilbild von Haider Anis
Haider Anisvor 1 Monat

@krispHQ This sounds like a pretty big upgrade! I like that you're focusing on improving accuracy, especially in noisy environments.

Profilbild von villager
villagervor 1 Monat

@krispHQ what kind of impact did it have on noise levels in meetings?

Ähnliche Videos

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

331,698 Aufrufe • vor 1 Monat

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 Aufrufe • vor 4 Monaten

Today we're announcing $280M in Series B funding at a $2B valuation, led by our long-time investor and partner, Menlo Ventures. When we announced our Series A last May, most conversations I had about voice began with someone explaining why they doubted it. I don't have those conversations anymore. People tell me instead how much time they save talking instead of typing and what they want us to build next. In a little over a year, voice has gone from something people were questioning to something they rely on, and it happened faster than we expected. That shift is why we've raised a new round of funding. This funding represents a deeper investment in our products, lab, models, and our team. Alongside the funding, we're announcing a preview of our first proprietary speech model, Canto. Canto is a 2B parameter speech model trained for the places people actually talk: loud rooms, windy streets, a toddler in the background, a second language mixed into the first. Models like this will change how we all interact with devices, and Canto is the first in a rapid line of them. Each one will be larger and more capable than the last, and each should improve Flow's dictation in a way you can feel the day it ships. Larger speech models also open the door to the vision we dreamed up 5 years ago: using your voice as the primary way you interact with devices. That still has to be invented. A voice interface people rely on all day, in every app, doesn't exist anywhere yet. It's time for that to change. To Sahaj Garg, our team, and every person who gave Wispr Flow a chance - this one's for you.

Tanay Kothari

615,867 Aufrufe • vor 1 Monat

Remember when AI couldn't draw a hand? Seven fingers, knuckles pointing backwards. And the AI spaghetti videos. That was three years ago. Images are done now. Video is close enough that you scrolled past AI ads this week and clocked exactly zero of them. Code writes itself and there are like 40 coding agents. AI voice spent that entire stretch sounding like the lady voice in a 2014 GPS. Flat, evenly spaced and every sentence landing with the same weight, like it's reading from a phone book. Here's why it stayed broken. Bad images are funny. You screenshot the seven fingers, it goes viral for being bad, someone fixes it. Bad audio is just boring. It doesn't fail spectacularly, so it never got that pressure. The bigger problem was the scoring. The whole industry graded AI voices on whether you could make out the words. So the models learned to over pronounce everything, hitting every syllable like a newsreader. Perfectly clear but robotic. Everyone was chasing a score that had nothing to do with sounding human. Meanwhile a small open-source team was doing something harder. Their lead researcher, an ex-NVIDIA engineer, went all in on an approach the rest of the field had written off. Two years early. No funding announcements or launch tour. He just put the whole thing on GitHub for free. It's sitting at 50,000+ stars now. Then they ran the test everyone else avoided. For 10 days they piped real users through their model and every big competitor with the listener never told which was which. Thousands of real people, real scripts. Whichever voice you actually preferred, they logged it. Theirs came out on top. It beat ElevenLabs about 6 times out of 10, head to head. It beat OpenAI's voice model 8 times out of 10. The gap was widest on the breathing, the pauses, the little hesitations, which is exactly the stuff that makes a voice sound like a person instead of a machine reading. They ran on real users rather than a lab, which is more than most of these claims can say. That's Fish Audio. This week they shipped S2.1 Pro: - Clone anyone's voice from 15 seconds of audio - Fast enough to hold a live conversation - 83 languages, one model - Type [whisper] or [sigh] mid-sentence and it does it - Around 70% cheaper than ElevenLabs - Free to download and run yourself Voice was the last thing on the list. Around 20 people with a free repo got there before other billion dollar companies did.

Rez Karim

18,446 Aufrufe • vor 2 Monaten

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,266,978 Aufrufe • vor 2 Jahren