Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Wow, TypeSafe AI Jev is a game changer for voice-driven UI. Responsiveness really matters for an experience to 'feel' right, and this certainly does the job! Latency here is not optimal given I'm in London and calling out to us-west. But still, super snappy. In this video, Jamcat is...

26,440 görüntüleme • 22 saat önce •via X (Twitter)

10 Yorum

Erik Spock Gafni profil fotoğrafı
Erik Spock Gafni17 saat önce

@typesafeai can you DJ our launch party?

Jon Taylor profil fotoğrafı
Jon Taylor14 saat önce

@typesafeai A human in the loop generative DJ set could either be incredible, or empty the dance floor 😅 at least I wouldn’t have to take requests

Nathan LeClaire profil fotoğrafı
Nathan LeClaire19 saat önce

@typesafeai u in my mind Londonbros! Let's go!

Aleix Conchillo Flaqué profil fotoğrafı
Aleix Conchillo Flaqué19 saat önce

@typesafeai Dope!

soulblocks profil fotoğrafı
soulblocks15 saat önce

@typesafeai nice

Mark ꩜ profil fotoğrafı
Mark ꩜13 saat önce

@typesafeai Really cool @JonPTaylor!

Geωrge profil fotoğrafı
Geωrge14 saat önce

@typesafeai You really haven't learned to keep secrets

vivek profil fotoğrafı
vivek19 saat önce

@typesafeai share repo link this is amazing !!

AFX LAB profil fotoğrafı
AFX LAB13 saat önce

@typesafeai Jev should be a tool for LLMs to save on tokens on simple tasks

Cumulative Web Inc. profil fotoğrafı
Cumulative Web Inc.14 saat önce

@typesafeai Responsiveness is the whole ballgame for voice-driven UI — a slow agent feels broken even when it's right. We build trust/verification tooling for voice agents; latency and correctness are the two halves. Great demo 👑

Benzer Videolar

Introducing PhoneLLM, an open model for voice agents. GPT 5.6 Terra performance on typical voice agent tasks at 1/3 the latency and 1/18 the cost. For voice agents, we need models that are both very low latency and very good at tool calling and instruction following. There's a trade-off here, and we often have to compromise on either latency or capability when building voice agents. With PhoneLLM (and the training and data stack that made this model possible) we're fixing this problem. For the last couple of years, most of the effort in frontier model development has gone towards leveraging test-time compute. Which is awesome! Models of all shapes and sizes are available that perform really, really well ... if you have "thinking" turned on for your model. But if you need your agent to respond at voice conversation speed, you can't use thinking models. PhoneLLM is a full-weights fine-tune of NVIDIA Nemotron Nano 30B. We trained on a wide range of real-world telephone and customer support use cases. The training focused on taking the excellent Nano 30B base capabilities and teaching the model to do typical voice agent tasks with thinking disabled. The results are really good: accurate tool calling and concise, on-topic responses in long conversations. And fast: TTFAT measured server-side is <100ms if you run PhoneLLM on a lightly loaded B200. :-) But seriously, when we characterize model latency, we do it with full, end-to-end, batched request simulations using real Pipecat voice agent pipelines. You can serve more than 80 concurrent agents on a single B200 with P95 end-to-end TTFAT <600ms. Including network overhead. That's an LLM cost-per-minute around $0.0025. (1/4 of a cent.) At a latency lower than any third-party API offers today. More details about this model, including weights on Hugging Face, how to spin it up with one click on Modal, and a starter project repo you can clone, are in the thread ...

kwindla

329,729 görüntüleme • 22 gün önce

Anthropic's in trouble, again! They spent years building what's now fully open-source. What made Claude feel different from a normal app is that the agent could act inside the interface instead of only talking in a chat box. For instance, Claude Artifacts let an agent render real UI, charts, dashboards, and interactive components that assemble live inside the response. Every major AI product tried to replicate it. But the problem was that unlike reasoning, planning, tool-calling, etc., none of it shipped natively with LangGraph, CrewAI, or Google ADK. So teams started building an owned version that required engineering the entire interface layer from scratch. Most teams, however, just settled for shipping the agent as a backend API in a chat box since rendering the UI is only one piece of it. To actually make it work, the interface layer also needed real-time streaming, state kept in sync between agent and UI, conversations that persist across sessions, and reconnection when a user refreshes mid-run. CopilotKit🪁 is now the only open-source framework that actually lets you build your own full-stack Claude-like apps. It decouples the agent from the interface, talking over AG-UI (an open protocol for agent-to-user communication). Being a standard protocol, the frontend never needs to know whether it is talking to a LangGraph or a CrewAI agent. You can change the backend anytime and the UI will never notice. In practice, CopilotKit's interface layer gives several pre-implemented React building blocks that wire the agent directly into the app, like: - generative UI, so the agent renders real components instead of text - chat windows, sidebars, and popups, or a fully headless setup - shared state, so the agent and app stay in sync - human-in-the-loop approvals, where the agent waits before acting - persistent threads that store the whole session, including the agent-user interactions and generated UI, not just text And because that full history is captured, those interactions can feed a self-learning layer that also improves the agent from real usage over time. The interface layer that Anthropic spent years engineering in-house is now literally available to any developer/team. CopilotKit is open-source with 30k+ GitHub stars, and AG-UI, the protocol underneath, is already supported across every major agent framework: LangGraph, CrewAI, Mastra, Google ADK, and more. CopilotKit GitHub repo → (don't forget to star it ⭐ ) If you want to go deeper, I found a detailed breakdown by Shubham Saboo recently on the three Generative UI patterns, with implementation. Read it below.

Avi Chawla

460,618 görüntüleme • 3 ay önce

Learn to build conversational AI voice agents in "Building AI Voice Agents for Production", created in collaboration with LiveKit and RealAvatar, and taught by dsa (Co-founder & CEO of LiveKit), Shayne (Developer Advocate, LiveKit), and Nedelina Teneva (Head of AI at RealAvatar, an AI Fund portfolio company). Voice agents combine speech and reasoning capabilities to enable real-time conversations. They're already being used to support customer service, to improve accessibility in healthcare, for entertainment applications, and for talk therapy. In this course, you’ll learn to build voice agents that listen, reason, and respond naturally. You’ll follow the architecture used to create the "AI Andrew" Avatar, a collaborative project between and RealAvatar that responds to users in what sounds like my voice. You’ll build a voice agent from scratch and deploy it to the cloud, enabling support for many simultaneous users. What you’ll learn: - Understand the fundamentals of voice agents, including key components like speech-to-text (STT), text-to-speech (TTS), and LLMs, and how latency is introduced at each layer. - Explore voice agent architectures and the trade-offs between modular pipelines and speech-to-speech APIs. - Explore how platforms like LiveKit mitigate latency issues with optimized networking infrastructure and low-latency communication protocols. - Learn how to connect client devices to voice agents using WebRTC—and why it outperforms HTTP and WebSocket for low-latency audio streaming. - Incorporate voice activity detection (VAD), end-of-turn detection, and context management to detect turns, handle interruptions, and manage conversational flow. - Understand the trade-offs between latency, quality, and cost in an example in which you build a voice agent and change its voice. - Equip your agent with metrics to measure latency at each stage of the voice pipeline and learn the key levers you can pull to make your agent faster and more responsive. The voice agents built in this course also incorporate voice technology from , a supporting contributor to the project. By the end of this course, you'll have learned the components of an AI voice agent pipeline, combined them into a system with low-latency communication, and deployed them on cloud infrastructure so it scales to many users. I’m looking forward to seeing what voice agents you build from this course! Please sign up here:

Andrew Ng

87,810 görüntüleme • 1 yıl önce

AG-UI makes building agentic applications dramatically easier. Here's how it works. This is a model for a simple chatbot: User → LLM → Response But interactive agents that render UI, pause for approvals, and ask users for input need a much more complex model. When building these agents, a response from the LLM will include a series of state changes as the agent runs: • Agent started a task • Agent called a tool • Agent updated its state • Agent streams these tokens • Agent is waiting on a human • Agent is resuming the task The Agent-User Interaction Protocol (AG-UI) treats the LLM response as a stream of events rather than a text endpoint. In practice, here is what you get as an agent runs: 1. Lifecycle events so your UI knows where the agent is. 2. Text messages that stream tokens. 3. Tool calls so your UI can prefill a form with any required arguments. 4. State updates that keep your UI in sync with the agent. 5. Special events for human approvals, rich media, and custom needs. All of these events travel over standard transports (SSE, WebSockets, or plain HTTP) as JSON. As a result, you can build a frontend that stays in sync with the agent's progress without having to invent a custom process to make this happen. For example, building a human-in-the-loop workflow becomes an off-the-shelf component you can integrate rather than build from scratch. CopilotKit🪁 is the creator of AG-UI, and you can use it when building frontend applications pretty much anywhere: • React • Angular • Vue • React Native • Slack • Teams • Discord • WhatsApp • Telegram Here is the link for you to check it out: Thanks to the CopilotKit team for partnering with me on this post.

Santiago

17,438 görüntüleme • 2 ay önce

$VET, #VeFam. In this video, I demonstrate in less than 4:30 minutes how to create an AI agent on veworld(.)ai. Watch me build a Mr. Robot Monologue Writer agent. If you haven't seen Mr. Robot, I suggest you watch it! This is just early bird access. The options for tools and integrations and such are limited, but what exists is already working quite well. The process is easy peasy. The UI is simple, but effective. It asks you for... 1. Role & Purpose 2. Voice & Style 3. Behavior 4. Rules 5. Tags 6. Avatar image 7. Welcome text. 8. Test drive before publication. ... and that's about it. This free version lets you have at most 3 agents, I am told. This implies that there is also a paid version. I'm all for it, because it sounds to me like VeChain is ready to do real business! I am providing feedback to Jérôme Grillères in order to help improve VeChain's AI agent marketplace. I didn't have to set up anything. The web UI is all I needed! The agent is running on Claude Sonnet 3.7. I did not have to provide a Claude API key. We seem to be riding along on VeChain's. I hope there'll be a choice for more models, including ChatGPT, in the future. This is so user friendly, that I can easily imagine that this would take off in a big, big way. I'm definitely building on this, when it goes into production with full features. Even if my own AI agents aren't successful, then I'm sure others' will be. And that means the $VET / $VTHO / $B3TR flywheel is going to take off in a big, big way. I, for one, am here for it. (See the reply below for the listing of the AI agent I just created.)

₿lackthorne AI

16,711 görüntüleme • 2 ay önce