Загрузка видео...

Не удалось загрузить видео

На главную

BREAKING 🔥: ByteDance launched SeedRealtime, a native audio-visual full-duplex LLM! > SeedRealtime uses a unified architecture to natively fuse audio, video, and text, enabling real-time interaction over continuous multimodal streams and delivering a brand-new "watch, listen, and speak" experience. It is now live on the Doubao App for free,...

94,643 просмотров • 1 месяц назад •via X (Twitter)

Комментарии: 23

Фото профиля 🚨 AI News | TestingCatalog
🚨 AI News | TestingCatalog1 месяц назад

Documented 🗞️

Фото профиля Alex Volkov
Alex Volkov1 месяц назад

Free or charge... holy surveilance

Фото профиля 🚨 AI News | TestingCatalog
🚨 AI News | TestingCatalog1 месяц назад

I wish it would be testable globally 😬

Фото профиля Alex Volkov
Alex Volkov1 месяц назад

we need chinese phone numbers don't we

Фото профиля TheCoderBTW
TheCoderBTW1 месяц назад

Full-duplex means it can listen while it talks. Does it actually let you cut in mid-sentence and it stops, or is there still a turn boundary underneath?

Фото профиля Philippe Tremblay
Philippe Tremblay1 месяц назад

How do I work for ByteDance? They're crushing it.

Фото профиля Frangel
Frangel1 месяц назад

Lets see if this finally makes AI conversations not sound so robotic.

Фото профиля J A Z I I
J A Z I I1 месяц назад

Chinese are making everything atp

Фото профиля Hussain Hashim | Building SundayBack
Hussain Hashim | Building SundayBack1 месяц назад

@testingcatalog so basically it's like a live stream with AI that actually interacts with you? that's pretty wild. wonder how it'll change online events!

Фото профиля Shinka - AI
Shinka - AI1 месяц назад

doubao shipping a native full-duplex multimodal model for free just gave every avatar startup a hard deadline

Фото профиля Inflectiv AI ⧉
Inflectiv AI ⧉1 месяц назад

Making this available for free will give ByteDance a huge amount of real-world interaction data. That's often as valuable as the model itself.

Фото профиля Saeed Anwar
Saeed Anwar1 месяц назад

Native multimodal fusion in a full-duplex loop is fundamentally different from bolt-on audio wrappers around a text model, and the latency numbers on turn-taking will be the real test. Has anyone seen benchmarks on interrupt handling or speaker transition lag yet?

Фото профиля AI Mastery Guide
AI Mastery Guide1 месяц назад

Full duplex audio and video in one model is a big jump for real time chat

Фото профиля NUS
NUS1 месяц назад

wow

Фото профиля James
James1 месяц назад

This on edge?!

Фото профиля Nobody
Nobody1 месяц назад

Opensource?

Фото профиля StatysTheBaddest
StatysTheBaddest1 месяц назад

Me: Why are Chinese people so awkward? Also me: Nvm.. They're engineers.

Фото профиля David T Kramaley
David T Kramaley1 месяц назад

full‑duplex multimodal chat flips the script on AI interaction. finally a model that watches and talks simultaneously

Фото профиля Bonsai 🌳
Bonsai 🌳1 месяц назад

Another breakthrough from Chinese developers Not surprising at all

Фото профиля Oleksandr
Oleksandr1 месяц назад

mindblown. native audio‑visual duplex blurs the line between perception and response. real‑time chat just leveled up

Фото профиля Rakhul
Rakhul1 месяц назад

The real test isn't the architecture, it's latency under load. GPT-4o's omni mode demo'd beautifully too, then shipped with audio disabled for months. Let's see what the actual p95 response time looks like when Doubao scales this beyond early adopters

Фото профиля Jack Li
Jack Li1 месяц назад

I see why the cellphone is an obstacle here. We need to obsolete cellphones and get something better.

Фото профиля Adam Aziz
Adam Aziz1 месяц назад

@zakariaornot 👀

Похожие видео

🚀 🚀Excited to announce the technical report of MiniCPM-o 4.5! MiniCPM-o 4.5 transitions #AI interaction from traditional turn-based processing to a real-time, native full-duplex stream-based paradigm. 🌊 The Omni-Flow Framework Instead of traditional VAD-based workarounds, we introduce the #Omni-#Flow framework. This unified stream paradigm aligns video, audio, and text on a synchronized millisecond timeline. • Native Full-Duplex: Simultaneous perception and response. • Proactive Interaction: Natively manages turn-taking without external VAD, supports proactive reminding. 📉 9B Scale, SOTA Performance MiniCPM-o 4.5 demonstrates SOTA multimodal intelligence at its scale: • Multimodal Benchmarks: Comparable to #Gemini 2.5 Flash on MMBench EN (87.6) and MathVista (80.1). • Streaming Evaluation: 54.4% win rate on LiveSports-3K-CC, surpassing specialized models. 💻 The Ultimate Edge AI — Fully Functional without Network Connection We are providing one-click installers for Windows (12G VRAM,RTX 5070) and macOS (M1-M5 Max/ M5 Pro). • Local API Support: Deploy your own inference server to integrate native full-duplex into custom apps. • Free Access: We are offering free community API services for exploration. • 100% Private: Your data never leaves your machine. Deploy in under 10 minutes. 🛠️👇 👐 Join the Open Future The weights are open. The protocol is public. 📄 Technical Report: 💻 GitHub: 🤗 HuggingFace: 🌐 Web Demo: #MiniCPMo #OpenSourceAI #EdgeAI #MachineLearning #ComputerVision #LLM

OpenBMB

148,094 просмотров • 4 месяцев назад