Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

BREAKING 🔥: ByteDance launched SeedRealtime, a native audio-visual full-duplex LLM! > SeedRealtime uses a unified architecture to natively fuse audio, video, and text, enabling real-time interaction over continuous multimodal streams and delivering a brand-new "watch, listen, and speak" experience. It is now live on the Doubao App for free,...

94,643 görüntüleme • 1 ay önce •via X (Twitter)

23 Yorum

🚨 AI News | TestingCatalog profil fotoğrafı
🚨 AI News | TestingCatalog1 ay önce

Documented 🗞️

Alex Volkov profil fotoğrafı
Alex Volkov1 ay önce

Free or charge... holy surveilance

🚨 AI News | TestingCatalog profil fotoğrafı
🚨 AI News | TestingCatalog1 ay önce

I wish it would be testable globally 😬

Alex Volkov profil fotoğrafı
Alex Volkov1 ay önce

we need chinese phone numbers don't we

TheCoderBTW profil fotoğrafı
TheCoderBTW1 ay önce

Full-duplex means it can listen while it talks. Does it actually let you cut in mid-sentence and it stops, or is there still a turn boundary underneath?

Philippe Tremblay profil fotoğrafı
Philippe Tremblay1 ay önce

How do I work for ByteDance? They're crushing it.

Frangel profil fotoğrafı
Frangel1 ay önce

Lets see if this finally makes AI conversations not sound so robotic.

J A Z I I profil fotoğrafı
J A Z I I1 ay önce

Chinese are making everything atp

Hussain Hashim | Building SundayBack profil fotoğrafı
Hussain Hashim | Building SundayBack1 ay önce

@testingcatalog so basically it's like a live stream with AI that actually interacts with you? that's pretty wild. wonder how it'll change online events!

Shinka - AI profil fotoğrafı
Shinka - AI1 ay önce

doubao shipping a native full-duplex multimodal model for free just gave every avatar startup a hard deadline

Inflectiv AI ⧉ profil fotoğrafı
Inflectiv AI ⧉1 ay önce

Making this available for free will give ByteDance a huge amount of real-world interaction data. That's often as valuable as the model itself.

Saeed Anwar profil fotoğrafı
Saeed Anwar1 ay önce

Native multimodal fusion in a full-duplex loop is fundamentally different from bolt-on audio wrappers around a text model, and the latency numbers on turn-taking will be the real test. Has anyone seen benchmarks on interrupt handling or speaker transition lag yet?

AI Mastery Guide profil fotoğrafı
AI Mastery Guide1 ay önce

Full duplex audio and video in one model is a big jump for real time chat

NUS profil fotoğrafı
NUS1 ay önce

wow

James profil fotoğrafı
James1 ay önce

This on edge?!

Nobody profil fotoğrafı
Nobody1 ay önce

Opensource?

StatysTheBaddest profil fotoğrafı
StatysTheBaddest1 ay önce

Me: Why are Chinese people so awkward? Also me: Nvm.. They're engineers.

David T Kramaley profil fotoğrafı
David T Kramaley1 ay önce

full‑duplex multimodal chat flips the script on AI interaction. finally a model that watches and talks simultaneously

Bonsai 🌳 profil fotoğrafı
Bonsai 🌳1 ay önce

Another breakthrough from Chinese developers Not surprising at all

Oleksandr profil fotoğrafı
Oleksandr1 ay önce

mindblown. native audio‑visual duplex blurs the line between perception and response. real‑time chat just leveled up

Rakhul profil fotoğrafı
Rakhul1 ay önce

The real test isn't the architecture, it's latency under load. GPT-4o's omni mode demo'd beautifully too, then shipped with audio disabled for months. Let's see what the actual p95 response time looks like when Doubao scales this beyond early adopters

Jack Li profil fotoğrafı
Jack Li1 ay önce

I see why the cellphone is an obstacle here. We need to obsolete cellphones and get something better.

Adam Aziz profil fotoğrafı
Adam Aziz1 ay önce

@zakariaornot 👀

Benzer Videolar

🚀 🚀Excited to announce the technical report of MiniCPM-o 4.5! MiniCPM-o 4.5 transitions #AI interaction from traditional turn-based processing to a real-time, native full-duplex stream-based paradigm. 🌊 The Omni-Flow Framework Instead of traditional VAD-based workarounds, we introduce the #Omni-#Flow framework. This unified stream paradigm aligns video, audio, and text on a synchronized millisecond timeline. • Native Full-Duplex: Simultaneous perception and response. • Proactive Interaction: Natively manages turn-taking without external VAD, supports proactive reminding. 📉 9B Scale, SOTA Performance MiniCPM-o 4.5 demonstrates SOTA multimodal intelligence at its scale: • Multimodal Benchmarks: Comparable to #Gemini 2.5 Flash on MMBench EN (87.6) and MathVista (80.1). • Streaming Evaluation: 54.4% win rate on LiveSports-3K-CC, surpassing specialized models. 💻 The Ultimate Edge AI — Fully Functional without Network Connection We are providing one-click installers for Windows (12G VRAM,RTX 5070) and macOS (M1-M5 Max/ M5 Pro). • Local API Support: Deploy your own inference server to integrate native full-duplex into custom apps. • Free Access: We are offering free community API services for exploration. • 100% Private: Your data never leaves your machine. Deploy in under 10 minutes. 🛠️👇 👐 Join the Open Future The weights are open. The protocol is public. 📄 Technical Report: 💻 GitHub: 🤗 HuggingFace: 🌐 Web Demo: #MiniCPMo #OpenSourceAI #EdgeAI #MachineLearning #ComputerVision #LLM

OpenBMB

148,094 görüntüleme • 5 ay önce