正在加载视频...

视频加载失败

As one example of 1.5 Pro’s sophisticated multimodal understanding and reasoning capabilities with long context, when given a 44-minute silent film, the model can analyze various plot points and events, and even makes sense of small details you might have missed.

348,318 次观看 • 2 年前 •via X (Twitter)

10 条评论

Kullar 的头像
Kullar2 年前

Forget that, have you seen openai sora?

Gus 🖖 Ferguson 的头像
Gus 🖖 Ferguson2 年前

So is Gemini 1.5 Pro more or less advanced than Gemini Advanced? Is Gemini [number] the core model generation and Nano, Pro and Ultra are weights? What is Gemini Advanced? Why is Google so, so bad at naming products... I suspect they let devs name their products don't they🤦‍♂️

Kiri 的头像
Kiri2 年前

Excellent Sundar! Gives me goosebumps to see :) can’t wait to explore all of the possibilities with Gemini, wonderful work from you and the team—congrats!! 🔥✨🤘

Siddarth Pai 的头像
Siddarth Pai2 年前

Google has great researchers and foundational technology need to productize better that's where gap was for taking lead in Social, Cloud or even now on AI

EcceAI 的头像
EcceAI2 年前

This is but my mind is already blown mindblowing

Adolfo Asorlin 👨🏼‍🚀 的头像
Adolfo Asorlin 👨🏼‍🚀2 年前

😳😋

Zippy Raj 𝕏 的头像
Zippy Raj 𝕏2 年前

Damn

Dishwasher 的头像
Dishwasher2 年前

I want to be able to edit existing pictures via natural language sooooo bad. please, Google, make it happen!

Gaurav Goyal 的头像
Gaurav Goyal2 年前

@BhandEth will make editing so simple

Types Digital 的头像
Types Digital2 年前

Multimodality era 🧬

相关视频

VITA Towards Open-Source Interactive Omni Multimodal LLM discuss: The remarkable multimodal capabilities and interactive experience of GPT-4o underscore their necessity in practical applications, yet open-source models rarely excel in both areas. In this paper, we introduce VITA, the first-ever open-source Multimodal Large Language Model (MLLM) adept at simultaneous processing and analysis of Video, Image, Text, and Audio modalities, and meanwhile has an advanced multimodal interactive experience. Starting from Mixtral 8x7B as a language foundation, we expand its Chinese vocabulary followed by bilingual instruction tuning. We further endow the language model with visual and audio capabilities through two-stage multi-task learning of multimodal alignment and instruction tuning. VITA demonstrates robust foundational capabilities of multilingual, vision, and audio understanding, as evidenced by its strong performance across a range of both unimodal and multimodal benchmarks. Beyond foundational capabilities, we have made considerable progress in enhancing the natural multimodal human-computer interaction experience. To the best of our knowledge, we are the first to exploit non-awakening interaction and audio interrupt in MLLM. VITA is the first step for the open-source community to explore the seamless integration of multimodal understanding and interaction. While there is still lots of work to be done on VITA to get close to close-source counterparts, we hope that its role as a pioneer can serve as a cornerstone for subsequent research.

AK

23,958 次观看 • 2 年前

Gemini-1.5 Pro has its spotlight stolen today, and people are poking fun at Sora vs Google memes. Well, I think it's the biggest boost in LLM capability so far in 2024. v1.5's 10M token context (1) excels at retrieval; (2) generalizes zero-shot to extremely long instructions like full tutorials and codebases; and (3) works across modalities such as text, audio, and video. Here's a stunning example: v1.5 learns to translate from English to Kalamang purely in context, following a full linguistic manual at inference time. Kalamang is a language spoken by fewer than 200 speakers in western New Guinea. Gemini has never seen this language during training and is only provided with 500 pages of linguistic documentation, a dictionary, and ~400 parallel sentences in context. It basically acquires a sophisticated new skill in the neural activations, instead of gradient finetuning. I talked about the Myth of Context Length many times before: don't get too excited by claims of 1M or even 1B context tokens. LSTMs already achieved literally infinite context length 25 yrs ago! What truly matters is how well the model actually uses the context to solve real-world problems, and Gemini-1.5 has surpassed the SOTA with flying colors. The paper is also well-written with lots of solid quantitative analysis on in-context memorization and generalization. Paper: “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context” Congrats to Jeff Dean Oriol Vinyals Sundar Pichai and team!

Jim Fan

278,483 次观看 • 2 年前