Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Announcing GPT-4, a large multimodal model, with our best-ever results on capabilities and alignment:

12,468,034 görüntüleme • 3 yıl önce •via X (Twitter)

10 Yorum

OpenAI profil fotoğrafı
OpenAI3 yıl önce

Join us at 1 pm PT today for a developer demo livestream showing GPT-4 and its capabilities/limitations: (comments in Discord:

Delysium - $AGI 🟨 profil fotoğrafı
Delysium - $AGI 🟨3 yıl önce

Speaking of breakthroughs - here's the biggest one in #Web3

Lex Fridman profil fotoğrafı
Lex Fridman3 yıl önce

This is awesome! 🤯

Nick Massian profil fotoğrafı
Nick Massian3 yıl önce

This all sounds great, but why don’t you make it unbiased? It’s very left leaning.

Ali profil fotoğrafı
Ali3 yıl önce

I was hype for it until i saw "GPT-4 is 82% less likely to respond to requests for disallowed content"

Harrison Kinsley profil fotoğrafı
Harrison Kinsley3 yıl önce

Multi-modal is very interesting. The training data used for models like this becomes more and more general. Getting closer to "one model to rule them all."

Jason Wei profil fotoğrafı
Jason Wei3 yıl önce

This example of a pirate explaining taxes in the style of shakespeare indicates so many levels of compositionality. The qualitative experience of GPT-4 is remarkable and will unlock a world of new use cases!

Steev profil fotoğrafı
Steev3 yıl önce

“82% less likely to respond to requests for disallowed content” - huge red flag. Who gets to decide what’s disallowed? Hopefully this costs you the leadership position as someone else releases a model that’s not nerfed in an arbitrary way.

Neville's Toad profil fotoğrafı
Neville's Toad3 yıl önce

Does safer mean bias

Yiğit Konur profil fotoğrafı
Yiğit Konur3 yıl önce

In case you are looking for a summary, I've already created a 'human-curated' thread here 👇🏽

Benzer Videolar

VITA Towards Open-Source Interactive Omni Multimodal LLM discuss: The remarkable multimodal capabilities and interactive experience of GPT-4o underscore their necessity in practical applications, yet open-source models rarely excel in both areas. In this paper, we introduce VITA, the first-ever open-source Multimodal Large Language Model (MLLM) adept at simultaneous processing and analysis of Video, Image, Text, and Audio modalities, and meanwhile has an advanced multimodal interactive experience. Starting from Mixtral 8x7B as a language foundation, we expand its Chinese vocabulary followed by bilingual instruction tuning. We further endow the language model with visual and audio capabilities through two-stage multi-task learning of multimodal alignment and instruction tuning. VITA demonstrates robust foundational capabilities of multilingual, vision, and audio understanding, as evidenced by its strong performance across a range of both unimodal and multimodal benchmarks. Beyond foundational capabilities, we have made considerable progress in enhancing the natural multimodal human-computer interaction experience. To the best of our knowledge, we are the first to exploit non-awakening interaction and audio interrupt in MLLM. VITA is the first step for the open-source community to explore the seamless integration of multimodal understanding and interaction. While there is still lots of work to be done on VITA to get close to close-source counterparts, we hope that its role as a pioneer can serve as a cornerstone for subsequent research.

AK

23,958 görüntüleme • 1 yıl önce