正在加载视频...

视频加载失败

Shape your voice agent’s personality with instructions for tone, pacing, and expressiveness. GPT-Live-1 can mirror the tone and emotion in a speaker’s voice and adapt to their pace. You can also set its language and response length.

98,451 次观看 • 10 天前 •via X (Twitter)

3 条评论

OpenAI Developers 的头像
OpenAI Developers10 天前

GPT-Live-1 handles listening and speaking in one model, cutting out extra handoffs so the conversation moves fasterrrrrrr 🏎️ And you can keep talking while your backend model handles reasoning and tool calls.

OpenAI Developers 的头像
OpenAI Developers10 天前

GPT-Live-1 finally makes conversations fluid. It distinguishes speech from background noise, so café chatter doesn’t have to stop the conversation ☕ You can even add a detail you just thought of or change direction mid-conversation without waiting for the model to finish its speaking turn.

OpenAI Developers 的头像
OpenAI Developers10 天前

GPT-Live-1 is now available in the API. Bring ChatGPT’s natural back-and-forth to your app, with voice agents that listen while they speak and work with the models and harness you choose.

相关视频

VoxCPM 2 just dropped by OpenBMB Only 2B-param open-source TTS (Text-to-Speech) model built for production-grade multilingual voice work. Apache-2.0 license, Can run on only 8GB VRAM. • Eliminates the "robotic" feel of traditional TTS, delivering prosody and emotional depth suitable for high-stakes professional environments like filmmaking, gaming, animation, and audiobooks. • 30-language multilingual: no language tag needed, just type in a supported language and generate directly. • Voice design: create a brand-new voice from a text description alone, like age, tone, pace, or emotion. No reference audio required. Describe the desired voice characteristics (gender, age, tone, emotion, pace …) in Control Instruction, and VoxCPM2 will craft a unique voice from your description alone. • Controllable cloning: clone from a short clip, then steer delivery style without losing the speaker’s core voice. • Ultimate cloning: use reference audio + transcript for continuation-style cloning that keeps the tiny vocal details. • 48kHz output: takes 16kHz reference audio and produces studio-quality speech without an external upsampler. • Real-time ready: around 0.3 RTF on RTX 4090, even lower with Nano-VLLM. • Commercial use: Apache-2.0 licensed. Developer-Friendly Infrastructure: - Native Torch Inference: Direct support for PyTorch-based workflows. - Training Flexibility: Supports both full-parameter and LoRA fine-tuning for specific domain adaptation. - Production Readiness: Compatible with voxcpm-nanovllm for large-scale, high-concurrency deployment.

Rohan Paul

13,541 次观看 • 5 个月前