Загрузка видео...

Не удалось загрузить видео

На главную

"Voice agents are built to be helpful. The problem is they can be too helpful and cross the boundary." I sat down with Sumanyu Sharma 🍫, Founder and CEO of Hamming, at our AI Engineer 🔜 Paris 🇫🇷 booth. Hamming stress-tests and monitors voice agents so the bank or...

18,670 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 7

Фото профиля Sumanyu Sharma 🍫
Sumanyu Sharma 🍫1 месяц назад

@HammingAI @aiDotEngineer Super fun chat, thanks for having me!

Фото профиля justine
justine1 месяц назад

@HammingAI @aiDotEngineer learned so much from you! thanks for coming on!

Фото профиля Raymond Chen
Raymond Chen1 месяц назад

@sumanyu @HammingAI @aiDotEngineer let's go @sumanyu !!!

Фото профиля justine
justine1 месяц назад

@HammingAI @aiDotEngineer watch the interview on youtube:

Фото профиля Michael Kuznetsov
Michael Kuznetsov1 месяц назад

@sumanyu @HammingAI @aiDotEngineer such good insights as usual @sumanyu!

Фото профиля Lawrence Lin Murata
Lawrence Lin Murata1 месяц назад

@sumanyu @HammingAI @aiDotEngineer great interview @sumanyu!!

Фото профиля Sahil Gandhi
Sahil Gandhi1 месяц назад

@sumanyu @HammingAI @aiDotEngineer Facts that's why we are building @talentpluto

Похожие видео

Learn to build conversational AI voice agents in "Building AI Voice Agents for Production", created in collaboration with LiveKit and RealAvatar, and taught by dsa (Co-founder & CEO of LiveKit), Shayne (Developer Advocate, LiveKit), and Nedelina Teneva (Head of AI at RealAvatar, an AI Fund portfolio company). Voice agents combine speech and reasoning capabilities to enable real-time conversations. They're already being used to support customer service, to improve accessibility in healthcare, for entertainment applications, and for talk therapy. In this course, you’ll learn to build voice agents that listen, reason, and respond naturally. You’ll follow the architecture used to create the "AI Andrew" Avatar, a collaborative project between and RealAvatar that responds to users in what sounds like my voice. You’ll build a voice agent from scratch and deploy it to the cloud, enabling support for many simultaneous users. What you’ll learn: - Understand the fundamentals of voice agents, including key components like speech-to-text (STT), text-to-speech (TTS), and LLMs, and how latency is introduced at each layer. - Explore voice agent architectures and the trade-offs between modular pipelines and speech-to-speech APIs. - Explore how platforms like LiveKit mitigate latency issues with optimized networking infrastructure and low-latency communication protocols. - Learn how to connect client devices to voice agents using WebRTC—and why it outperforms HTTP and WebSocket for low-latency audio streaming. - Incorporate voice activity detection (VAD), end-of-turn detection, and context management to detect turns, handle interruptions, and manage conversational flow. - Understand the trade-offs between latency, quality, and cost in an example in which you build a voice agent and change its voice. - Equip your agent with metrics to measure latency at each stage of the voice pipeline and learn the key levers you can pull to make your agent faster and more responsive. The voice agents built in this course also incorporate voice technology from , a supporting contributor to the project. By the end of this course, you'll have learned the components of an AI voice agent pipeline, combined them into a system with low-latency communication, and deployed them on cloud infrastructure so it scales to many users. I’m looking forward to seeing what voice agents you build from this course! Please sign up here:

Andrew Ng

87,810 просмотров • 1 год назад

NEW: AssemblyAI Handles 4X More Voice Data Per Day Than YouTube "Our TAM has just increased by 100X" CEO Dylan Fox (Dylan Fox) One of the fastest growing categories in AI, voice powers note-taking, healthcare, coding agents, call centers, AI companions, drive-thru ordering, consumer electronics & humanoid robots. AssemblyAI is the infrastructure underneath it. Powering billion-dollar companies like Granola & Commure, + Tolans & Ciro AI. Stats: › 120M+ voice conversations a week at peak › 2M+ hours of voice a day, 4X YouTube's daily volume › Weekly conversations up 800% in 3 years › 1M+ developers, 40% signed up last year › ~100M API calls a day › ~80 employees AssemblyAI was 1 of 6 companies in YC's first AI batch in 2017, run by Daniel Gross. Backed by Accel, Insight Partners, YC, Smith Point Capital, Nat Friedman, Daniel Gross, Patrick & John Collison We cover: › The McDonald's drive-thru has no idea it's McDonald's › Why 75% of a voice model is the data you train it on › Why humanoid robots can't tell who's talking to them › Whether voice agents should have disclosures 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) Dylan Fox, Founder & CEO at Assembly AI (01:06) AssemblyAI's voice traffic beat YouTube by 4x (03:33) $100K in GPU credits that started it all (06:39) Why Voice AI is inflecting right now (10:11) AssemblyAI's infrastructure-first strategy (12:41) Handling 120 million weekly calls (14:49) How AI agents are rewriting Internal Ops at AssemblyAI (17:05) The Sovereign AI problem nobody's solved for voice (18:26) Is the Keyboard and Mouse finally dying? (22:33) The problem Humanoid Robotics hasn't solved yet (25:14) The real problem behind translating languages (29:27) Can AI actually talk to animals? (31:24) The real reason Open-Source benchmarks lie to you (34:20) Building websites in chat rooms as a kid (36:48) On-Device AI models are about to change everything (43:03) The people who shaped Dylan's career

Molly O’Shea

327,480 просмотров • 1 месяц назад

How many AI agents work at your company? We now have over 3,258 agents working alongside 1,300 humans. The crazy part is these agents were created by EVERY EMPLOYEE at our company... sales reps, marketers, customer support, product, eng. Literally EVERYONE. BUT I'm most surprised by the adoption and value that MANAGERS are getting from agents. I used to think that every IC would become a manager of agents. Now I think that managers will very likely manage WAY more agents than their ICs combined. And managers' agents will manage their ICs' agents - overseeing them for human-in-the-loop interactions. When creating agents, we use 100% context from all of your activity, files edited, tasks and projects worked on, hierarchy, skills, and role information. We build a user-based context model to make agents as relatable as possible to the specific human that we're building for. This means they truly understand the nuances of the work and what "great" looks like - because great is very much in the eye of the beholder. Great is by definition, subjective. This is also why the human ENGAGEMENT loops are SO vital to agent value. The iteration AFTER the agent is onboarded is where the MAGIC happens. This is just like a manager managing an IC in real life... you're giving feedback. In this case, though, agents learn INSTANTLY, and they retain the knowledge perfectly and indefinitely. Even though I've been pushing AI for years now to everyone in our company, this was the first time we had truly end-to-end AI adoption and retention. This kind of AI adoption is wild. But the value we're realizing is truly INSANE. Super Agents outnumber our humans nearly 3 to 1. What if you could 3X your workforce overnight? Watch this video to see how 👇

Zeb Evans

425,244 просмотров • 7 месяцев назад

I’m excited to announce coval raised a $28M Series A, led by Norwest with participation from Base10, twilio Ventures and Y Combinator, among others. I started Coval with the belief that the simulation and evaluation approaches that made autonomous vehicles a reality could unlock the same for voice AI. That parallel has only gotten sharper as enterprises deploy autonomous conversational agents at scale. Voice and chat are the new interface in the AI era. Talking is the most natural way to interact with complex autonomous systems. The same way web and mobile transformed every enterprise, every company will have a voice agent as the front door to their product. This platform shift is happening now. From Y Combinator to working with companies like Deepgram, Perplexity , and Zoom, it's been a wild ride. We watched voice AI go from an emerging modality to every Fortune 500 enterprise building a conversational interface. Autonomous agents are the future. But you can't ship and hope for the best when agents are taking on real work. The stakes are too high. Coval will be what enables us to actually trust those agents. The platform to help teams reliably scale voice and chat agents. We wouldn’t be here without Coval’s founding engineers Rhea Pokorny and Kobi Hudson. Rhea took a chance on Coval before we had a single customer, a line of code, or an office that wasn't my couch. Kobi left Waymo after 10 years building foundational simulation systems to help us build the frontier of voice simulation. To the full Coval team - you all are insane engineers and builders who inspire me every day. Rob Young Alejandra Vergara Henry Finkelstein Callum Reid Dakota Mallen Dana Dzik Mallory McLoughlin Christopher Kuester Jake Levi Brooke Hartley Kappi Patterson To all our investors - Norwest, Base10 Partners, twilio Ventures, Y Combinator, Alumni Ventures, Swift Ventures, Fortitude VC and so many more. You are what made Coval possible. To Koh Terai who brought my video vision to life. And last but certainly not least, thank you to everyone who's been building with us. We're just getting started.

Brooke Hopkins

85,937 просмотров • 2 месяцев назад

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

329,729 просмотров • 22 дней назад

"AI agents will hold more crypto than humans within a decade." Charles Hoskinson (Charles Hoskinson) studied math, dropped out, built one of the only blockchains designed by peer-reviewed research. He co-founded Ethereum, walked away over how it was run, and built Cardano to do it differently. The man who has argued with everyone in this industry now thinks the biggest user of crypto won't be people at all. "Humans are a rounding error in the system we're building. AI agents don't sleep, don't panic-sell, and don't care about price. They transact in tokens because that's the only thing they can actually use." We cover: - Why AI agents (not humans) become the dominant on-chain actors, and what that does to every token model - The infrastructure that has to exist before agents can transact safely at scale - Why most current blockchains can't handle machine-speed transactions - Where Cardano's research-first approach fits in a world of autonomous agents - The identity problem: how do you tell a human from an agent on-chain, and why it matters - Why he's bullish on the technology but blunt about the timeline - What he thinks the rest of the industry is getting wrong about AI + crypto - The one thing that has to happen for any of this to be real Thanks to Charles for coming on New Era Finance Podcast. TIMESTAMPS: 00:00 - Intro 01:30 - Why AI Agents Change Everything 06:30 - Humans as a Rounding Error 12:00 - The Infrastructure Gap 18:30 - Identity: Human vs Agent On-Chain 24:30 - Where Cardano Fits 30:00 - What The Industry Gets Wrong 34:00 - The Timeline Nobody Wants To Hear

Michaël van de Poppe

293,517 просмотров • 3 месяцев назад