Загрузка видео...

Не удалось загрузить видео

На главную

As one example of 1.5 Pro’s sophisticated multimodal understanding and reasoning capabilities with long context, when given a 44-minute silent film, the model can analyze various plot points and events, and even makes sense of small details you might have missed.

348,411 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 10

Фото профиля Kullar
Kullar2 лет назад

Forget that, have you seen openai sora?

Фото профиля Gus 🖖 Ferguson
Gus 🖖 Ferguson2 лет назад

So is Gemini 1.5 Pro more or less advanced than Gemini Advanced? Is Gemini [number] the core model generation and Nano, Pro and Ultra are weights? What is Gemini Advanced? Why is Google so, so bad at naming products... I suspect they let devs name their products don't they🤦‍♂️

Фото профиля Kiri
Kiri2 лет назад

Excellent Sundar! Gives me goosebumps to see :) can’t wait to explore all of the possibilities with Gemini, wonderful work from you and the team—congrats!! 🔥✨🤘

Фото профиля Siddarth Pai
Siddarth Pai2 лет назад

Google has great researchers and foundational technology need to productize better that's where gap was for taking lead in Social, Cloud or even now on AI

Фото профиля EcceAI
EcceAI2 лет назад

This is but my mind is already blown mindblowing

Фото профиля Adolfo Asorlin 👨🏼‍🚀
Adolfo Asorlin 👨🏼‍🚀2 лет назад

😳😋

Фото профиля Zippy Raj 𝕏
Zippy Raj 𝕏2 лет назад

Damn

Фото профиля Dishwasher
Dishwasher2 лет назад

I want to be able to edit existing pictures via natural language sooooo bad. please, Google, make it happen!

Фото профиля Gaurav Goyal
Gaurav Goyal2 лет назад

@BhandEth will make editing so simple

Фото профиля Types Digital
Types Digital2 лет назад

Multimodality era 🧬

Похожие видео

Gemini-1.5 Pro has its spotlight stolen today, and people are poking fun at Sora vs Google memes. Well, I think it's the biggest boost in LLM capability so far in 2024. v1.5's 10M token context (1) excels at retrieval; (2) generalizes zero-shot to extremely long instructions like full tutorials and codebases; and (3) works across modalities such as text, audio, and video. Here's a stunning example: v1.5 learns to translate from English to Kalamang purely in context, following a full linguistic manual at inference time. Kalamang is a language spoken by fewer than 200 speakers in western New Guinea. Gemini has never seen this language during training and is only provided with 500 pages of linguistic documentation, a dictionary, and ~400 parallel sentences in context. It basically acquires a sophisticated new skill in the neural activations, instead of gradient finetuning. I talked about the Myth of Context Length many times before: don't get too excited by claims of 1M or even 1B context tokens. LSTMs already achieved literally infinite context length 25 yrs ago! What truly matters is how well the model actually uses the context to solve real-world problems, and Gemini-1.5 has surpassed the SOTA with flying colors. The paper is also well-written with lots of solid quantitative analysis on in-context memorization and generalization. Paper: “Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context” Congrats to Jeff Dean Oriol Vinyals Sundar Pichai and team!

Jim Fan

278,557 просмотров • 2 лет назад