
AI Bites | YouTube Channel
@ai_bites • 2,434 subscribers
AI tools, papers and hands-on coding to solve problems with AI. Former @UniofOxford @Oxford_VGG Our products: https://t.co/uhnIm6VOmS
Shorts
🏠 How can AI understand the structure of an entire building from images? PolyLayout introduces a new approach for multi-room 3D layout estimation — allowing AI to jointly reconstruct connected indoor spaces instead of treating every room independently. By combining neural networks with explicit geometric reasoning, PolyLayout predicts more accurate and robust room layouts while adapting to different scenes and camera setups. A step closer to machines that can truly understand the spaces around us. 🧠🏗️ Paper Title: PolyLayout: Multi-room Manhattan Layout Estimation Project: Link:
33,047 görüntüleme
ReViV reconstructs viewer-centric human motion (body, hand, and gaze) and view-centric scene geometry (camera and depth) from a single egocentric RGB video in a unified feed-forward model. It formulates the task as learning the full joint probability distribution over multimodal signals, including RGB video, camera trajectory, gaze direction, full-body motion, hand motion, and depth. Powered by a Masked Generative Egocentric Transformer, ReViV operates within a single feed-forward architecture to simultaneously reconstruct the temporally consistent 4D reconstruction across the viewer and the view with fast inference speed. Paper Title: ReViV: Reconstructing the Viewer and the View in 4D from Monocular Project: Link:
18,590 görüntüleme
Current Vision-Language-Action (VLA) paradigms in autonomous driving primarily rely on Imitation Learning (IL), which introduces inherent challenges such as distribution shift and causal confusion. Online Reinforcement Learning offers a promising pathway to address these issues through trial-and-error learning. However, applying online reinforcement learning to VLA models in autonomous driving is hindered by inefficient exploration in continuous action spaces. MindDrive, a VLA framework comprising a large language model (LLM) with two distinct sets of LoRA parameters. The one LLM serves as a Decision Expert for scenario reasoning and driving decision-making, while the other acts as an Action Expert that dynamically maps linguistic decisions into feasible trajectories. Paper Title: MindDrive: A Vision-Language-Action Model for Autonomous Driving via Project: Link:
43,496 görüntüleme