
Ryohei Sasaki@engineer
@rsasaki0109 • 10,690 subscribers
Software Engineer at MAP IV(TIER IV group). Previously, Waseda University. AI/Robotics/Autonomous Driving/GNSS/LiDAR/Camera/IMU/SLAM/Localization/Mapping
Shorts
Videos

Inflatable Whole-Body Robotic Skin with Internal Time-of-Flight (ToF) Depth Sensing and Kinematics-Based Point-Cloud Prediction(IROS2026) An inflatable Baymax costume serves as the robotic skin envelope. Distributed internal ToF sensors detect whole-body contact during dynamic human–robot interaction.
Ryohei Sasaki@engineer129,779 görüntüleme • 13 gün önce

GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation ECCV 2026 GaussianGPT generates 3D Gaussian scenes fully autoregressively, token by token using a GPT-style transformer allowing for flexible 3D generation, completion, and outpainting.
Ryohei Sasaki@engineer40,616 görüntüleme • 9 gün önce

DART: Detect Anything in Real Time Detect Anything in Real Time: Real-time object detection using frontier object detection models. Training-free framework that converts SAM3 into a real-time multi-class open-vocabulary detector. Achieves 55.8 AP on COCO val2017 (80 classes) at 15.8 FPS (4 classes, 1008px) on a single RTX 4080.
Ryohei Sasaki@engineer33,932 görüntüleme • 9 gün önce

Face Anything: 4D Face Reconstruction from Any Image Sequence ECCV 2026 Face Anything is a unified feed-forward model for high-fidelity 4D face reconstruction and dense tracking from arbitrary image sequences. The key idea is canonical facial point prediction, a representation that assigns each pixel a normalized facial coordinate in a shared canonical space. This formulation transforms dense tracking and dynamic reconstruction into a single canonical reconstruction problem, producing temporally consistent geometry and reliable correspondences.
Ryohei Sasaki@engineer30,095 görüntüleme • 12 gün önce

EgoExoMoCap: Distributed Ego-Exo Human Motion Capture ECCV 2026 (Oral) Meta Reality Labs/ETH Zürich With two or more people wearing smart glasses, EgoExoMoCap combines each wearer’s egocentric (Ego) and exocentric (Exo) views to estimate full-body motion. Without bulky multi-camera setups or mocap suits, it reconstructs 3D human motion even under occlusions using head/wrist tracking and DINOv3-based visual features.
Ryohei Sasaki@engineer26,080 görüntüleme • 22 gün önce

remind-reid-tracker REMIND — RE-Identification with Memory for INDoor Navigation REMIND addresses a core challenge in visual tracking: re-identifying objects that disappear and reappear, look similar to one another, or are observed from changing viewpoints. Rather than relying on position or motion cues, REMIND builds appearance-based identity models per object using DINOv3 patch features, decomposed into: Global and part-level descriptors — K-means and attention-guided semantic parts extracted per detection Relational context (neighbor sets) — structural scene layout encoded as co-occurrence graphs of neighboring objects Known-set distance disambiguation — geometry-aware resolution of visually ambiguous groups Adaptive memory — per-object appearance, part, and background models that update over time The association pipeline runs a per-frame sequence of visual evidence building, context activation, global Hungarian assignment, and post-assignment guards, producing explicit uncertainty signals (ambiguous, provisional) alongside confident identity decisions. Evaluation outputs span case, object, frame, class, scene, and batch levels, with full internal telemetry for diagnostic analysis.
Ryohei Sasaki@engineer19,013 görüntüleme • 1 ay önce

Glob3R: Global Structure-from-Motion With 3D Foundation Models HKUST Spatial Artificial Intelligence Lab · Alibaba Group · Nanjin University · Fudan University Glob3R is a scalable global 3D reconstruction framework built on geometric foundation models. It converts dense feed-forward predictions into reliable multi-view tracks and jointly refines camera poses and scene geometry through motion averaging and bundle adjustment. Glob3R improves reconstruction accuracy and consistency on long sequences, large-scale scenes, and unordered image collections.
Ryohei Sasaki@engineer17,701 görüntüleme • 1 ay önce

VLAExplain — Interpreting Vision-Language-Action (VLA) Models VLAExplain is an interpretability toolkit designed to help users visually understand the inner workings of Vision-Language-Action (VLA) models. Currently, attention analysis is supported for both the pi05 and unifolm-vla models. For details, please check pi05 and UnifoLM-VLA readme files respectively. Demo of pi05 in action:
Ryohei Sasaki@engineer12,774 görüntüleme • 4 ay önce

LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM LEGO-SLAM running at 15 FPS on a ScanNet scene with language-based loop closing for drift correction. LEGO-SLAM is a 3DGS-based SLAM framework that supports open-vocabulary semantic querying and rendering. It tracks via G-ICP and efficiently builds a map by embedding Gaussians with scene-adaptive 16D language features. Map management is achieved through Language Pruning and Language-Based Loop Detection. The generated map enables open-vocabulary 3D Object Localization.
Ryohei Sasaki@engineer15,060 görüntüleme • 5 ay önce
Daha fazla içerik yok.