Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Introducing GPT-4o, our new model which can reason across text, audio, and video in real time. It's extremely versatile, fun to play with, and is a step towards a much more natural form of human-computer interaction (and even human-computer-computer interaction):

4,360,353 görüntüleme • 2 yıl önce •via X (Twitter)

9 Yorum

Greg Brockman profil fotoğrafı
Greg Brockman2 yıl önce

The new Voice Mode will be coming to ChatGPT Plus in upcoming weeks.

Greg Brockman profil fotoğrafı
Greg Brockman2 yıl önce

GPT-4o can also generate any combination of audio, text, and image outputs, which leads to interesting new capabilities we are still exploring. See e.g. the "Explorations of capabilities" section in our launch blog post ( or these generated images:

Greg Brockman profil fotoğrafı
Greg Brockman2 yıl önce

We also have significantly improved non-English language performance quite a lot, including improving the tokenizer to better compress many of them:

Dennis profil fotoğrafı
Dennis2 yıl önce

Dozens of startups obliterated

Benjamin BLM profil fotoğrafı
Benjamin BLM2 yıl önce

Audio, text and image? What's in the training data? You need literal billions of images and text tokens. Where did you get them, from the internet?

AshutoshShrivastava profil fotoğrafı
AshutoshShrivastava2 yıl önce

Desktop app with Vison is the best feature which was launched definitely game changer.

Charlene Wang profil fotoğrafı
Charlene Wang2 yıl önce

gpt4-o’s real time translation is gonna go viral & support the best consumer AI hardware. Saw it translate group conversations and respond in different language. It’s super fast and I can’t wait to use it in Japan!

illusion diffusion profil fotoğrafı
illusion diffusion2 yıl önce

where?

Alex Sharp profil fotoğrafı
Alex Sharp2 yıl önce

now that ai's can talk to each other i can finally delete all in from my podcast player

Benzer Videolar

VITA Towards Open-Source Interactive Omni Multimodal LLM discuss: The remarkable multimodal capabilities and interactive experience of GPT-4o underscore their necessity in practical applications, yet open-source models rarely excel in both areas. In this paper, we introduce VITA, the first-ever open-source Multimodal Large Language Model (MLLM) adept at simultaneous processing and analysis of Video, Image, Text, and Audio modalities, and meanwhile has an advanced multimodal interactive experience. Starting from Mixtral 8x7B as a language foundation, we expand its Chinese vocabulary followed by bilingual instruction tuning. We further endow the language model with visual and audio capabilities through two-stage multi-task learning of multimodal alignment and instruction tuning. VITA demonstrates robust foundational capabilities of multilingual, vision, and audio understanding, as evidenced by its strong performance across a range of both unimodal and multimodal benchmarks. Beyond foundational capabilities, we have made considerable progress in enhancing the natural multimodal human-computer interaction experience. To the best of our knowledge, we are the first to exploit non-awakening interaction and audio interrupt in MLLM. VITA is the first step for the open-source community to explore the seamless integration of multimodal understanding and interaction. While there is still lots of work to be done on VITA to get close to close-source counterparts, we hope that its role as a pioneer can serve as a cornerstone for subsequent research.

AK

23,958 görüntüleme • 1 yıl önce