Loading video...

Video Failed to Load

Go Home

I've been trying Meta smart glasses' new multimodal AI - while it's pretty basic right now, it's still sick to see it combine what it sees from the camera with the language model to describe what it's seeing! Already solid for accessibility Full episode:

1,128,504 views • 2 years ago •via X (Twitter)

10 Comments

Ben Geskin's profile picture
Ben Geskin2 years ago

We are getting there 👀

Everett World's profile picture
Everett World2 years ago

Multi-modality is the next step. We're moving from LLMs to world models, that will be even more helpful for practical reasons.

Karthik Kannan's profile picture
Karthik Kannan2 years ago

Um, we’re already here

Average Engineer's profile picture
Average Engineer2 years ago

When he says that phrase and ask questions, Glass takes photo and takes question from users using OpenAI Whisper API, Upload to Gta-4 Vision API with your prompt Get back the results and make it speak again using API. Is there anything i am missing here?

David's profile picture
David2 years ago

This is the way Marques !

Rahul's profile picture
Rahul2 years ago

Lower hanging use cases I could use this for already: reading while walking.

Pawel's profile picture
Pawel2 years ago

It is going to be the feature of AR

Newtonian's profile picture
Newtonian2 years ago

Epic 👏🔥🔥

Joseph Bella's profile picture
Joseph Bella2 years ago

I feel like we are seeing history repeat itself when Apple released the Newton and Palm made the Pilot.

Ellie MacQueen's profile picture
Ellie MacQueen2 years ago

Here for the David stares

Related Videos