Loading video...

Video Failed to Load

Go Home

Image captioning spots objects, but can it spot 𝘵𝘩𝘦 𝘷𝘪𝘣𝘦? 🤔 The "Caption This" demo in Google AI Studio showcases 2.5 Pro's upgraded creative capabilities and deep visual understanding to decode images and create contextual, witty captions.

96,545 views • 1 year ago •via X (Twitter)

11 Comments

Jay's profile picture
Jay1 year ago

The vibessss

UserInterface's profile picture
UserInterface4 years ago

Need Professional Video Production, Music Videos, Commercials, Graphic Design, or Photo Retouching? We will take your project from concept to completion. #services #creative #DMV

ᗩᗰEᖇIᑕᗩᑎ ᗰᗩᗪE ᗰEᗪIᗩ's profile picture
ᗩᗰEᖇIᑕᗩᑎ ᗰᗩᗪE ᗰEᗪIᗩ1 year ago

Now this looks like fun.

James Oliver Bennett's profile picture
James Oliver Bennett1 year ago

It looks great.😉😀☺️

đạt nguyễn's profile picture
đạt nguyễn1 year ago

@Google I dropped my phone and now I lost my account. Please help me.

Tsukuyomi's profile picture
Tsukuyomi1 year ago

spotting the vibe? that’s a tall order for any AI. creativity is a messy business. 🖤

Nexa Cards's profile picture
Nexa Cards1 year ago

Vibe check, initiated!

⟁ndrew V's profile picture
⟁ndrew V1 year ago

How long will this fun last though. Our days of novel delight I’m worried may be numbered.

0x.Agency's profile picture
0x.Agency1 year ago

Vibe-detecting AI? Cool!

Holu Olivier's profile picture
Holu Olivier1 year ago

@Google Hi, Can some users opt out of #ai in @google search. It’s becoming hard to enjoy searching by ourselves especially when it’s the first thing we see. (Keeping google search alive😉)

Pay It Now (PIN)'s profile picture
Pay It Now (PIN)1 year ago

Vibe-checking AI!

Related Videos

MagicScroll: Nontypical Aspect-Ratio Image Generation for Visual Storytelling via Multi-Layered Semantic-Aware Denoising paper page: Visual storytelling often uses nontypical aspect-ratio images like scroll paintings, comic strips, and panoramas to create an expressive and compelling narrative. While generative AI has achieved great success and shown the potential to reshape the creative industry, it remains a challenge to generate coherent and engaging content with arbitrary size and controllable style, concept, and layout, all of which are essential for visual storytelling. To overcome the shortcomings of previous methods including repetitive content, style inconsistency, and lack of controllability, we propose MagicScroll, a multi-layered, progressive diffusion-based image generation framework with a novel semantic-aware denoising process. The model enables fine-grained control over the generated image on object, scene, and background levels with text, image, and layout conditions. We also establish the first benchmark for nontypical aspect-ratio image generation for visual storytelling including mediums like paintings, comics, and cinematic panoramas, with customized metrics for systematic evaluation. Through comparative and ablation studies, MagicScroll showcases promising results in aligning with the narrative text, improving visual coherence, and engaging the audience. We plan to release the code and benchmark in the hope of a better collaboration between AI researchers and creative practitioners involving visual storytelling.

AK

22,379 views • 2 years ago