Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

"Fast Foundation Stereo: When Foundation Models Meet Efficient Stereo Matching" TL;DR: distills stereo foundation models into an efficient real-time stereo matching framework, preserving high accuracy while dramatically reducing inference cost.

14,827 görüntüleme • 1 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Meet My AI Ears. A lot of folks ask me how I capture ASMR video and audio for training of AI? I always use Binaural 3D audio and have for decades in different forms. But how? A Brief History of Binaural Recording Binaural recording, the foundation of 3D audio, dates back to 1881 when French inventor Clément Ader created the first system using multiple telephone transmitters at the Paris Opera to transmit stereo sound to listeners, simulating spatial presence. By the 1920s, patents like W. Bartlett Jones’ 1927 filing advanced devices for capturing and reproducing “binaural” signals. The 1930s saw Alan Blumlein’s work on stereophonic sound, which he termed “binaural,” laying groundwork for modern stereo. Commercial milestones hit in the 1950s with binaural records from labels like Cook Laboratories and the first binaural reel-to-reel tapes. A resurgence came in the 1970s with Neumann’s KU-80 dummy head, the first commercial binaural system. Today, it’s integral to VR, ASMR, and immersive media. The Technology of Binaural 3D Audio At its core, binaural recording mimics human hearing by using two microphones placed in ear-shaped molds or a dummy head, separated like human ears (typically 14-18 cm apart). This captures spatial cues: interaural time differences (ITD) for sound arrival timing, interaural level differences (ILD) for volume variations, and head-related transfer functions (HRTF) that account for how the head, torso, and pinnae filter sounds. The result? A 3D soundscape that tricks the brain into perceiving direction, distance, and elevation when played back via headphones—no speakers needed for immersion. Advanced setups use omnidirectional capsules (e.g., DPA 4060) for high-fidelity capture, often in silicone ears to replicate natural diffraction. I use the 3DIO Microphones today but I would cover a dummy head in texture material and place two stereo (4 channels) microphones in each ear. I would then mix down the resulting signals into stereo. The 3DIO series features dual omnidirectional capsules in realistic silicone ear molds, spaced 14 cm apart for compact, accurate 3D capture based on over 13 years of research into human hearing. Models like the Free Space Pro II use premium DPA 4060 CORE capsules for ultra-low noise and high sensitivity, delivering stereo output ideal for immersive applications like game audio. Today just about any ASMR producer uses these. But I use them to capture, curate and archive sound and video we will lose or just about lost for AI training in a way no model or AI company is doing today. I am duplicating the human 3D binaural audio experience and memory. Below is a crude demonstration. If you can listen in headphones. Or turn your phone sideways to feel the audio space. I’ll have a far more professional demo soon to show the real power of 3D audio. (Oh that music is a MIDI player that uses disks to play).

Brian Roemmele

30,800 görüntüleme • 8 ay önce

🚨 BREAKING: Big news in the computer vision world! 🎥 Luxonis | Robotic Vision just dropped its new OAK 4 line, and it’s a big upgrade for edge computer vision. Instead of being “just a stereo camera,” OAK 4 is a fully standalone vision computer with 52 TOPS of on-device AI. Models run locally, depth is computed locally, and no external PC or cloud pipeline is required. This is why robotics teams love it: lower latency, lower cost, fewer failure points in the field. The hardware is built for the real-world. IP67, shock-resistant, wide-FOV RGB + stereo pair, IR projection, IMU, audio, and a patent-pending calibration system that keeps depth accurate even when conditions change. But the real move is the platform. With Luxonis Hub, you can deploy models, grab telemetry, push OTA updates, or collect data when performance drifts, all from a unified interface. It turns a single device into an end-to-end edge CV system. Most customers today in robotics are groups who just want something that works: AMRs, bin-picking systems, trailer-loading robots, and ag-tech. 🤖 And they all say the same thing, the appeal isn’t raw TOPS, it’s the all-in-one simplicity that lets them scale without building custom infrastructure. Feels like the direction edge vision has been waiting for: rugged hardware + high-throughput on-device compute + a real management layer. A next step toward “plug-and-deploy” perception for robots. 🔗 Find out more here: ~~ ♻️ Join the weekly robotics newsletter, and never miss any news →

Lukas Ziegler

41,955 görüntüleme • 9 ay önce