Загрузка видео...

Не удалось загрузить видео

На главную

Introducing VIVA 2.5, with our most advanced models yet. Krisp's Voice Isolation has been running in production for 2 years, across 1B+ mins of voice AI conversations a month, improving WER by focusing on primary speaker only. As more teams ran voice isolation in front of their STT, we...

25,958,617 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 34

Фото профиля Alistair Ward
Alistair Ward1 месяц назад

@krispHQ Looks like shit pal

Фото профиля .
.1 месяц назад

@krispHQ Genuinely, what's the benefit of this technology?

Фото профиля simon
simon1 месяц назад

@krispHQ I fucking hate ai

Фото профиля Kennedy Leflar🌹🕊✝️
Kennedy Leflar🌹🕊✝️1 месяц назад

@krispHQ Still looks fake, the uncanny valley vibe will never leave AI slop

Фото профиля Tiny
Tiny1 месяц назад

@krispHQ I dont. I'm old school i can still think for myself. Thats what a brain is for 😉

Фото профиля villager
villager1 месяц назад

@krispHQ who needs 2 years of data to figure that out lol

Фото профиля Yori
Yori23 дней назад

@krispHQ its so HD that it went from looking convincing back to looking uncanny lmao

Фото профиля coolman 6666-mar
coolman 6666-mar1 месяц назад

@krispHQ give me back my water

Фото профиля Berty
Berty1 месяц назад

@krispHQ I didn’t know Vauxhall fitted a 2.5 in the viva 🤷‍♂️😀

Фото профиля asdghmejfn
asdghmejfn22 дней назад

@krispHQ Let’s get rid of this Ai slop and make em lose the money the spent to corrupt our minds and importantly our kids minds personally I don’t want my kid to use Ai to learn as it legit makes em dumb typing everything in and nog learning anymore

Фото профиля Menmaro
Menmaro21 дней назад

@krispHQ Looks and sounds like AI slop

Фото профиля Charlotte
Charlotte1 месяц назад

@krispHQ super intéressant de voir l'évolution de l'isolation vocale en prod. vous mesurez aussi la latence ajoutée par le traitement ?

Фото профиля Militant Boomer 🇵🇸 #FreePalestine
Militant Boomer 🇵🇸 #FreePalestine22 дней назад

@krispHQ why don't u work on something that actually helps humanity instead of perfecting make-believe "customer service agents" that end up referring to actual human beings to answer the simplest questions —but then they don't even know their own business's customer service hours #Fail

Фото профиля A Buffalo
A Buffalo1 месяц назад

@krispHQ We don’t want any of this

Фото профиля JORDAN
JORDAN1 месяц назад

@krispHQ This should make voice agents feel much more reliable.

Фото профиля Franchaela Fingerton
Franchaela Fingerton1 месяц назад

@krispHQ This shit is the reason glaciers are melting and leaves are falling in August. Let actors work ffs

Фото профиля DB
DB1 месяц назад

@krispHQ Yeah, I had a Vauxhall Viva in the 70s, it wasn't that big though 😕

Фото профиля OwlishNick
OwlishNick20 дней назад

@krispHQ Hey A.I engineer here and I worked on Viva and I can confirm it was trained on massive amounts of porn especially gay porn. I said "Davit why do you need all this for model training?" And he said it was very necessary to have terabytes of porn sent to him personally.

Фото профиля Lisa O'Brien
Lisa O'Brien26 дней назад

@krispHQ I love him. ( But not as much as Liam Gallagher !! )

Фото профиля Mouaddib
Mouaddib1 месяц назад

@krispHQ Why the fuck are people striving to achieve human like video to a point they become indistinguishable from reality. Are those fucks blind, egotistical or too stupid to see the danger? Or the danger is the goal?

Фото профиля Tom
Tom1 месяц назад

@krispHQ The 3.5x smaller model is just as exciting as the accuracy gains. Better efficiency opens the door for many more deployments.

Фото профиля AI TALKs
AI TALKs1 месяц назад

@krispHQ Love how you fixed the over-aggressive isolation issue without hurting clean audio. Clean progress.

Фото профиля caçacéfalos
caçacéfalos1 месяц назад

@krispHQ Bora

Фото профиля 🧠 PocketIQ | Free AI Trading
🧠 PocketIQ | Free AI Trading1 месяц назад

@chesny @krispHQ 🔼 Less guessing. More analysis.

Фото профиля Iconic Icon
Iconic Icon1 месяц назад

@krispHQ But the CFF is way too high. Fix it.

Фото профиля Jacob
Jacob25 дней назад

@krispHQ The US remake of Peep Show looks terrible.

Фото профиля sabir hussain
sabir hussain1 месяц назад

@krispHQ Better voice isolation means cleaner AI conversations

Фото профиля Joaquim Pinto
Joaquim Pinto29 дней назад

@krispHQ 2.5 É um vinho da beira interior, ( Caria )ótimo para acompanhar, caça, queijos e carnes vermelhas.

Фото профиля Bally_AgenticAI
Bally_AgenticAI1 месяц назад

@krispHQ I read a lot of launch posts. Almost none admit the previous version had a regression. @krispHQ just did, and shipped the fix with receipts: 46.4% fewer word errors across 10 STT engines, no harm on clean audio. That is how you earn trust in voice infrastructure.

Фото профиля villager
villager1 месяц назад

@krispHQ 1B minutes of convo, impressive! Any plans for multi‑speaker support?

Фото профиля ArifIno
ArifIno1 месяц назад

@krispHQ @BrainPathio

Фото профиля BrainPath.io
BrainPath.io1 месяц назад

@krispHQ Woooow 🚀

Фото профиля Haider Anis
Haider Anis1 месяц назад

@krispHQ This sounds like a pretty big upgrade! I like that you're focusing on improving accuracy, especially in noisy environments.

Фото профиля villager
villager1 месяц назад

@krispHQ what kind of impact did it have on noise levels in meetings?

Похожие видео

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

331,698 просмотров • 1 месяц назад

Cerebras inference is very fast. So fast that it changes how we think about configuring our LLMs for voice agent use cases. Kimi K2.6 is a 1T parameter reasoning model that Cerebras serves at 650 - 1,000 tokens per second (end-to-end throughput), with time to first token metrics as low as 150ms (latency). These numbers are two to three times faster than other similarly capable models. The biggest lever we get from this kind of speed is that we can use the model in reasoning mode, and still have excellent "time to first non-thinking token." This solves a big pain point we have in 2026 for voice agent use cases. Almost all recent innovation in post-training has focused on making models good at reasoning ("test time compute"). This is great, but it makes the user-facing model latency much, much slower. Which is a problem for conversational voice agents. We can run Kimi K2.6 with reasoning turned on, and get responses faster than other models produce with reasoning disabled. On my 30-turn voice agent benchmark, Kimi K2.6 with reasoning enabled ties GPT 5.1 and Haiku 4.5 with reasoning disabled, and is still about 200ms seconds faster! On my primary task agent benchmark, Kimi K2.6 is now the #2 model. It ranks just behind Gemini 3.5 Flash in "high" reasoning mode, and tied with GLM 5, Sonnet 4.6, and GPT 5.4 with reasoning set to "low." But Kimi K2.6 completes each turn in the agent loop in under 500ms. The other four models are all at least 3x slower. (Models only qualify for this benchmark if they can complete task turns at a P50 <4s.) A couple of other things that this speed buys us, for production voice agents: - Tool calls happen fast enough that we don't have to work around tool call latency in our pipeline design. - We can prompt the model to output structured data at the beginning of a response, followed by plain text for voice generation. This opens up possibilities like asking the model to do complex classification/generation tasks that influence the rest of the pipeline. For example, the model could create a detailed style prompt for a steerable TTS model, for each individual conversation turn. And, of course, you can use Kimi K2.6 with reasoning turned off. Cerebras calls this "instant" mode. Here's a video of a Cerebras Kimi K2.6 voice agent with voice-to-voice response time, measured at the client, under 500ms. This is the true response latency as perceived by the user, including all network and audio codec overhead, transcription and turn detection, Kimi K2.6 token generation, and voice generation. 500ms is, effectively, instant. So the Cerebras naming for this mode is a propos. :-)

kwindla

40,593 просмотров • 4 месяцев назад

Today we're announcing $280M in Series B funding at a $2B valuation, led by our long-time investor and partner, Menlo Ventures. When we announced our Series A last May, most conversations I had about voice began with someone explaining why they doubted it. I don't have those conversations anymore. People tell me instead how much time they save talking instead of typing and what they want us to build next. In a little over a year, voice has gone from something people were questioning to something they rely on, and it happened faster than we expected. That shift is why we've raised a new round of funding. This funding represents a deeper investment in our products, lab, models, and our team. Alongside the funding, we're announcing a preview of our first proprietary speech model, Canto. Canto is a 2B parameter speech model trained for the places people actually talk: loud rooms, windy streets, a toddler in the background, a second language mixed into the first. Models like this will change how we all interact with devices, and Canto is the first in a rapid line of them. Each one will be larger and more capable than the last, and each should improve Flow's dictation in a way you can feel the day it ships. Larger speech models also open the door to the vision we dreamed up 5 years ago: using your voice as the primary way you interact with devices. That still has to be invented. A voice interface people rely on all day, in every app, doesn't exist anywhere yet. It's time for that to change. To Sahaj Garg, our team, and every person who gave Wispr Flow a chance - this one's for you.

Tanay Kothari

615,867 просмотров • 1 месяц назад

Remember when AI couldn't draw a hand? Seven fingers, knuckles pointing backwards. And the AI spaghetti videos. That was three years ago. Images are done now. Video is close enough that you scrolled past AI ads this week and clocked exactly zero of them. Code writes itself and there are like 40 coding agents. AI voice spent that entire stretch sounding like the lady voice in a 2014 GPS. Flat, evenly spaced and every sentence landing with the same weight, like it's reading from a phone book. Here's why it stayed broken. Bad images are funny. You screenshot the seven fingers, it goes viral for being bad, someone fixes it. Bad audio is just boring. It doesn't fail spectacularly, so it never got that pressure. The bigger problem was the scoring. The whole industry graded AI voices on whether you could make out the words. So the models learned to over pronounce everything, hitting every syllable like a newsreader. Perfectly clear but robotic. Everyone was chasing a score that had nothing to do with sounding human. Meanwhile a small open-source team was doing something harder. Their lead researcher, an ex-NVIDIA engineer, went all in on an approach the rest of the field had written off. Two years early. No funding announcements or launch tour. He just put the whole thing on GitHub for free. It's sitting at 50,000+ stars now. Then they ran the test everyone else avoided. For 10 days they piped real users through their model and every big competitor with the listener never told which was which. Thousands of real people, real scripts. Whichever voice you actually preferred, they logged it. Theirs came out on top. It beat ElevenLabs about 6 times out of 10, head to head. It beat OpenAI's voice model 8 times out of 10. The gap was widest on the breathing, the pauses, the little hesitations, which is exactly the stuff that makes a voice sound like a person instead of a machine reading. They ran on real users rather than a lab, which is more than most of these claims can say. That's Fish Audio. This week they shipped S2.1 Pro: - Clone anyone's voice from 15 seconds of audio - Fast enough to hold a live conversation - 83 languages, one model - Type [whisper] or [sigh] mid-sentence and it does it - Around 70% cheaper than ElevenLabs - Free to download and run yourself Voice was the last thing on the list. Around 20 people with a free repo got there before other billion dollar companies did.

Rez Karim

18,446 просмотров • 2 месяцев назад

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,266,978 просмотров • 2 лет назад