Loading video...

Video Failed to Load

Go Home

JUST IN: Google releases Gemini 1.5, a powerful MoE model. It's a huge breakthrough. The model has the longest context window ever seen: 1 million tokens. It can process 1 hour of video, 11 hours of audio, 30,000 lines of code, or 700,000 words in a single prompt. When...

83,432 views • 2 years ago •via X (Twitter)

10 Comments

Lior⚡'s profile picture
Lior⚡2 years ago

Technical report here:

Lior⚡'s profile picture
Lior⚡2 years ago

Amazing work by @OriolVinyals, @JeffDean, and @GoogleAI team

Elad Gil's profile picture
Elad Gil2 years ago

I think has had 5MM context window for a while?

Lior⚡'s profile picture
Lior⚡2 years ago

Might've got tricked by their copy 🤔 'We’ve been able to significantly increase the amount of information our models can process — running up to 1 million tokens consistently, achieving the longest context window of any large-scale foundation model yet."

Google's profile picture
Google2 years ago

The Gemini fun has just gotten started.

Uri Eliabayev's profile picture
Uri Eliabayev2 years ago

עוד פרטים וזה נראה שהם ב10 מיליון 🤔

SaaS Growth Strategies's profile picture
SaaS Growth Strategies2 years ago

Really interesting to see the innovation coming from Google. Longer context length will enable new LLM applications in document processing. Not to mention AI tutors will become better.

Evgeny Matohin's profile picture
Evgeny Matohin2 years ago

I wish it had a proper API!

RohiniAI's profile picture
RohiniAI2 years ago

Raising the level

ASIF AGHA's profile picture
ASIF AGHA2 years ago

three.js demo was interesting to see!

Related Videos

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,265,265 views • 1 year ago