Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

Live visual descriptions can aid blind people in understanding their surroundings with autonomy and independence. In our #UIST2024 work, we present WorldScribe that generates automated visual descriptions that are adaptive to the users’ contexts in real-time, in the real world.

16,135 Aufrufe • vor 1 Jahr •via X (Twitter)

10 Kommentare

Profilbild von Anhong Guo
Anhong Guovor 1 Jahr

First, WorldScribe is adaptive to users' intent, i.e. prioritizing the most pertinent descriptions based on semantic relevance, or visual attributes based on customizations. E.g. users can specify to find specific objects, or ask for descriptions with more color and texture info.

Profilbild von Anhong Guo
Anhong Guovor 1 Jahr

Second, WorldScribe is adaptive to visual contexts, e.g. it provides consecutively succinct descriptions for dynamic visual scenes, while it presents longer and more detailed ones for stable settings. This strategy is particularly useful when describing scenes *live*.

Profilbild von Anhong Guo
Anhong Guovor 1 Jahr

Third, WorldScribe is also adaptive to sound contexts, e.g., increasing description volume in noisy environments, or pausing when conversations start. Such manipulations are related to our prior work on SoundShift in #DIS2024:

Profilbild von Anhong Guo
Anhong Guovor 1 Jahr

WorldScribe is powered by a suite of vision, language, and sound recognition models, and introduces a description generation pipeline with different VLMs that balances the tradeoffs between their richness and latency to support real-time usage.

Profilbild von Anhong Guo
Anhong Guovor 1 Jahr

Our user study with blind participants and subsequent pipeline evaluation show that WorldScribe can provide real-time and fairly accurate visual descriptions to facilitate environment understanding that is adaptive and customized to users' contexts.

Profilbild von Anhong Guo
Anhong Guovor 1 Jahr

However, there's still a lot more to do to make automated live descriptions truly context-aware and humanized, such as incorporating additional real-world knowledge like GPS and maps, adapting to users' changing intents, and embedding short- and long-term memory into the system.

Profilbild von Anhong Guo
Anhong Guovor 1 Jahr

WorldScribe is led by Ruei-Che Chang @RueiChe, who has been doing a series of work on multimodal real-world accessibility, e.g., OmniScribe, SoundShift, and WorldScribe. We will present and demo this work at #UIST2024. Check out the paper here:

Profilbild von Toby J. Li😺 (he/him)
Toby J. Li😺 (he/him)vor 1 Jahr

Super cool work! Love how it focuses on inferring user intents and providing contextually-relevant information.

Profilbild von Anhong Guo
Anhong Guovor 1 Jahr

Thanks Toby! See ya in Pittsburgh!

Profilbild von Lyman
Lymanvor 1 Jahr

good work

Ähnliche Videos