正在加载视频...

视频加载失败

Build AR applications that recognize physical objects and provide real-time spatial guidance. Stijn Spanhove and Pavlo Tkachenko used Gemini 2.5 Pro’s multimodal vision and sound effect prompting capabilities to create an immersive experience with LEGO Smart Bricks and Snap Spectacles.

48,222 次观看 • 5 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

Engineer builds AR ad blocker for real life using Snap Spectacles and Gemini AI | Aamir Khollam, Interesting Engineering A prototype AR app detects ads in your surroundings and blocks them live through Snap Spectacles using Google’s Gemini AI. You can pay for YouTube Premium to avoid watching ads during your sacred food-with-YouTube ritual, or pay for Spotify to not pester you with ads while you try to get a morning workout in, but what can you do about real-life advertisements? A Belgian software developer may have found a workaround. Stijn Spanhove has created an experimental augmented reality app that detects and digitally covers physical ads like billboards and soda cans in the real world. Built for Snap’s fifth-generation AR glasses (Snap Spectacles), the prototype uses Google’s Gemini AI to identify branded content and instantly mask it. Instead of the original imagery, the app places a bright red square over the detected ad. These red blocks also name the hidden brand, like “Bol. billboard,” turning ad removal into a kind of real-time brand callout. “It’s exciting to imagine a future where you control the physical content you see,” Spanhove posted on X (formerly Twitter). In follow-up replies, he hinted at additional features, including options to replace the red square with personal photos or text from a notes app. Check it out below – Combines Snap and Google tools The app relies on Snap’s Depth Cache API to register objects in 3D space and maintain spatial consistency as the user moves. Gemini, Google’s generative AI model, identifies the ads themselves, whether on large posters, newspaper pages, or food packaging. This allows the blocker to function beyond obvious signage. In demo clips shared by Spanhove, it successfully covers ads on cereal boxes, magazines, and public signage, though not without a delay of a second or two. The red overlays appear to float stably, following head movements and perspective shifts with accuracy. Still, it’s early days. Spanhove describes the software as “experimental,” and the user experience reflects that. Because Snap Spectacles use transparent displays, the overlays can’t fully block light, so the original ad sometimes faintly shows through. Also, the Spectacles’ narrow 46-degree field of view limits coverage to only what’s directly ahead. Design stirs debate The app has sparked strong reactions online. Many praised the concept of user-controlled visual space, while others criticized the bold red boxes as more jarring than the ads themselves. One suggestion, from user Hexographer ⬣Backup account⬣, proposed replacing the red blocks with more visually pleasing alternatives, such as “images of local foliage or animal life.” They also suggested that users could set a custom folder of replacement visuals, including personal or family photos. Others asked for cross-platform support, but the app currently only works with Snap Spectacles. Devices like the Apple Vision Pro or Meta Quest will need separate development efforts. Despite its limitations, the project arrives at a moment when major tech firms like Meta and Microsoft have scaled back their AR ambitions. Snap, meanwhile, continues renting its Spectacles to developers at $99 per month, making experiments like Spanhove’s possible. Read more:

Owen Gregorian

45,774 次观看 • 1 年前

Explore state-of-the-art multimodal prompting in our new short course Large Multimodal Model Prompting with Gemini, taught by Erwin Huizenga in collaboration with Google Cloud. One interesting insight from this course: with multimodal models, prompt structure matters significantly. Placing text inputs, such as a patient's medical history, before image inputs, like an X-ray, can enhance the model's ability to contextualize and interpret visual data effectively. In other contexts, such as image captioning, you may get better results by putting the image first. Multimodal models behave differently than text-only LLMs, and effective prompting for models varies depending on the model you’re using. In this course you’ll learn how to effectively prompt Gemini models. Gemini's multimodal capabilities also enable new approaches in AI application development, for example: - The Gemini library handles various video formats (MP4, MOV, MPEG), streamlining applications using these formats. - Large context window (up to 1 million tokens) enables processing of extensive content, like analyzing multiple 50-minute videos simultaneously. - Function calling feature integrates real-time data (e.g., current exchange rates) into model responses. The course demonstrates building multimodal applications with real-world examples including document analyzers that reason across text and graphs simultaneously, video content extractors that find and timestamp specific information from multiple hours of footage, and automated expense report systems processing receipt images while cross-referencing company policies. Sign up here:

Andrew Ng

74,060 次观看 • 1 年前

VITA Towards Open-Source Interactive Omni Multimodal LLM discuss: The remarkable multimodal capabilities and interactive experience of GPT-4o underscore their necessity in practical applications, yet open-source models rarely excel in both areas. In this paper, we introduce VITA, the first-ever open-source Multimodal Large Language Model (MLLM) adept at simultaneous processing and analysis of Video, Image, Text, and Audio modalities, and meanwhile has an advanced multimodal interactive experience. Starting from Mixtral 8x7B as a language foundation, we expand its Chinese vocabulary followed by bilingual instruction tuning. We further endow the language model with visual and audio capabilities through two-stage multi-task learning of multimodal alignment and instruction tuning. VITA demonstrates robust foundational capabilities of multilingual, vision, and audio understanding, as evidenced by its strong performance across a range of both unimodal and multimodal benchmarks. Beyond foundational capabilities, we have made considerable progress in enhancing the natural multimodal human-computer interaction experience. To the best of our knowledge, we are the first to exploit non-awakening interaction and audio interrupt in MLLM. VITA is the first step for the open-source community to explore the seamless integration of multimodal understanding and interaction. While there is still lots of work to be done on VITA to get close to close-source counterparts, we hope that its role as a pioneer can serve as a cornerstone for subsequent research.

AK

23,958 次观看 • 1 年前