Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Meet Reka Core, our best and most capable multimodal language model yet. 🔮 It’s been a busy few months training this model and we are glad to finally ship it! 💪 Core has a lot of capabilities, and one of them is understanding video --- let’s see what Core...

757,979 görüntüleme • 2 yıl önce •via X (Twitter)

10 Yorum

Reka profil fotoğrafı
Reka2 yıl önce

We evaluate Core on standard benchmarks for both text and multimodal, along with a blind third-party human evaluation.

Reka profil fotoğrafı
Reka2 yıl önce

Along with Core, we have published a technical report detailing the training, architecture, data, and evaluation for the Reka models.

Reka profil fotoğrafı
Reka2 yıl önce

See some examples of outputs from Core in or try it yourself on

Reka profil fotoğrafı
Reka2 yıl önce

Check out the blogpost for more details.

dmitriy profil fotoğrafı
dmitriy2 yıl önce

Please familiarize it with big chungus

Matt Henderson profil fotoğrafı
Matt Henderson2 yıl önce

Flash to Core in two months! 📈

sridhar profil fotoğrafı
sridhar2 yıl önce

Congratulations, @DaniYogatama and @RekaAILabs team!!

Ben (e/treats) profil fotoğrafı
Ben (e/treats)2 yıl önce

congrats!!! please add pricing to your pricing page 👀 it's just a form right now

Reka profil fotoğrafı
Reka2 yıl önce

It's a bug and we are fixing it. Currently, Core is 10 bucks per million input tokens and 25 bucks per million output tokens.

djcows profil fotoğrafı
djcows2 yıl önce

nice. however, one small issue, no submariner

Benzer Videolar

VITA Towards Open-Source Interactive Omni Multimodal LLM discuss: The remarkable multimodal capabilities and interactive experience of GPT-4o underscore their necessity in practical applications, yet open-source models rarely excel in both areas. In this paper, we introduce VITA, the first-ever open-source Multimodal Large Language Model (MLLM) adept at simultaneous processing and analysis of Video, Image, Text, and Audio modalities, and meanwhile has an advanced multimodal interactive experience. Starting from Mixtral 8x7B as a language foundation, we expand its Chinese vocabulary followed by bilingual instruction tuning. We further endow the language model with visual and audio capabilities through two-stage multi-task learning of multimodal alignment and instruction tuning. VITA demonstrates robust foundational capabilities of multilingual, vision, and audio understanding, as evidenced by its strong performance across a range of both unimodal and multimodal benchmarks. Beyond foundational capabilities, we have made considerable progress in enhancing the natural multimodal human-computer interaction experience. To the best of our knowledge, we are the first to exploit non-awakening interaction and audio interrupt in MLLM. VITA is the first step for the open-source community to explore the seamless integration of multimodal understanding and interaction. While there is still lots of work to be done on VITA to get close to close-source counterparts, we hope that its role as a pioneer can serve as a cornerstone for subsequent research.

AK

23,958 görüntüleme • 2 yıl önce