Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

I've been trying Meta smart glasses' new multimodal AI - while it's pretty basic right now, it's still sick to see it combine what it sees from the camera with the language model to describe what it's seeing! Already solid for accessibility Full episode:

1,128,504 Aufrufe • vor 2 Jahren •via X (Twitter)

10 Kommentare

Profilbild von Ben Geskin
Ben Geskinvor 2 Jahren

We are getting there 👀

Profilbild von Everett World
Everett Worldvor 2 Jahren

Multi-modality is the next step. We're moving from LLMs to world models, that will be even more helpful for practical reasons.

Profilbild von Karthik Kannan
Karthik Kannanvor 2 Jahren

Um, we’re already here

Profilbild von Average Engineer
Average Engineervor 2 Jahren

When he says that phrase and ask questions, Glass takes photo and takes question from users using OpenAI Whisper API, Upload to Gta-4 Vision API with your prompt Get back the results and make it speak again using API. Is there anything i am missing here?

Profilbild von David
Davidvor 2 Jahren

This is the way Marques !

Profilbild von Rahul
Rahulvor 2 Jahren

Lower hanging use cases I could use this for already: reading while walking.

Profilbild von Pawel
Pawelvor 2 Jahren

It is going to be the feature of AR

Profilbild von Newtonian
Newtonianvor 2 Jahren

Epic 👏🔥🔥

Profilbild von Joseph Bella
Joseph Bellavor 2 Jahren

I feel like we are seeing history repeat itself when Apple released the Newton and Palm made the Pilot.

Profilbild von Ellie MacQueen
Ellie MacQueenvor 2 Jahren

Here for the David stares

Ähnliche Videos