Загрузка видео...
Не удалось загрузить видео
BREAKING 🔥: ByteDance launched SeedRealtime, a native audio-visual full-duplex LLM! > SeedRealtime uses a unified architecture to natively fuse audio, video, and text, enabling real-time interaction over continuous multimodal streams and delivering a brand-new "watch, listen, and speak" experience. It is now live on the Doubao App for free,... show more
94,643 просмотров • 1 месяц назад •via X (Twitter)
Комментарии: 23

Documented 🗞️

Free or charge... holy surveilance

I wish it would be testable globally 😬

we need chinese phone numbers don't we

Full-duplex means it can listen while it talks. Does it actually let you cut in mid-sentence and it stops, or is there still a turn boundary underneath?

How do I work for ByteDance? They're crushing it.

Lets see if this finally makes AI conversations not sound so robotic.

Chinese are making everything atp

@testingcatalog so basically it's like a live stream with AI that actually interacts with you? that's pretty wild. wonder how it'll change online events!

doubao shipping a native full-duplex multimodal model for free just gave every avatar startup a hard deadline

Making this available for free will give ByteDance a huge amount of real-world interaction data. That's often as valuable as the model itself.

Native multimodal fusion in a full-duplex loop is fundamentally different from bolt-on audio wrappers around a text model, and the latency numbers on turn-taking will be the real test. Has anyone seen benchmarks on interrupt handling or speaker transition lag yet?

Full duplex audio and video in one model is a big jump for real time chat

wow

This on edge?!

Opensource?

Me: Why are Chinese people so awkward? Also me: Nvm.. They're engineers.

full‑duplex multimodal chat flips the script on AI interaction. finally a model that watches and talks simultaneously

Another breakthrough from Chinese developers Not surprising at all

mindblown. native audio‑visual duplex blurs the line between perception and response. real‑time chat just leveled up

The real test isn't the architecture, it's latency under load. GPT-4o's omni mode demo'd beautifully too, then shipped with audio disabled for months. Let's see what the actual p95 response time looks like when Doubao scales this beyond early adopters

I see why the cellphone is an obstacle here. We need to obsolete cellphones and get something better.

@zakariaornot 👀

