Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

OAK 4 CS now supports PTP synchronization. Sync multi-camera exposures over standard Ethernet. No FSYNC wiring, no daisy chains, no trigger cables across long distances. Just a shared clock distributed over the network. For multi-camera perception, distributed AI vision, and sensor fusion at scale.

32,787 Aufrufe • vor 1 Monat •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

Mistral AI Releases Robostral Navigate: An 8B Model Enabling Robots to Navigate Complex Environments Hitting 76.6% on R2R-CE With One RGB Camera. No LiDAR. No depth sensor. No multi-camera rig. Here's how it works. 👇 1. Pointing, not metric commands The model predicts the pixel coordinates of the next target in the camera view, plus the arrival orientation. Working in pixel space keeps it robust to camera intrinsics and world scale. When the target leaves the frame, it falls back to local displacements ("2m forward, 1.5m left, turn 25°"). 2. Grounding-first No open-source VLM base. It starts from Mistral's grounding model (pointing, counting, localization). Navigation emerges once the model knows where things are. → ~400,000 trajectories across 6,000 simulated scenes 3. Prefix-caching for training A tree-based attention mask packs a full episode into one sequence — all time steps in a single forward pass. → 22× fewer training tokens; months of training done in days 4. Online RL on top After supervised training, CISPO adds trial-and-error learning to fight distribution shift from behavior cloning. → +3.2% success rate from RL alone 5. The numbers (R2R-CE, Matterport3D) → 76.6% success on validation unseen → +9.7 pts over best single-camera approach → +4.5 pts over best depth/multi-camera system The key takeaway: state-of-the-art continuous VLN without a sensor stack — grounding-init, pixel-space actions, prefix-cached SFT, and online RL, on one RGB camera. Full analysis: Technical details: Mistral AI Mistral AI for Developers

Marktechpost AI

39,955 Aufrufe • vor 6 Tagen

This is #GoProMISSION1 PRO 🎥 The only 8K60 camera with a 1-inch sensor. Our compact, cinema-grade camera features a proprietary GP3 processor and 50MP sensor that enable intelligent low-light capture, industry-leading frame rates and resolutions, and groundbreaking thermal performance. ✔️ 1-inch Quad-Bayer sensor with up to 14-stops of dynamic range at the sensor for low-light capture ✔️ Longest continuous runtimes + most dependable thermal performance of any GoPro ever—over 5 hours in 1080p + over 3 hours in 4K at 100°F ✔️ Industry-leading 8K60—300% more pixels than 4K ✔️ 4K240 + 1080p960 ultra slo-mo with real frames—not AI-interpolated ✔️ 8K30 + 4K120 Open Gate capture ✔️ Gallery-ready 50MP photos + 44MP frame grabs ✔️ Up to 240 Mbps bit rate out of the box + 300 Mbps with GoPro Labs ✔️ 10-Bit color + GP-Log2 with LUTs for Rec.709 + Rec.2020 outputs ✔️ HLG HDR with Simultaneous Dual-Gain Readout—the industry standard for pros ✔️ New intelligent capture modes: Dive, Vlog, Low-Light, Sport POV, + Subject Tracking ✔️ 13% higher capacity Enduro 2 battery in the same form factor with new fast charging ✔️ Rugged + waterproof, now to 66ft (20m) without a housing ✔️ Emmy® Award Winning #HyperSmooth in-camera video stabilization ✔️ New 4-microphone array, 32-bit float audio, multi-track recording, + manual audio controls ✔️ Timecode Sync to streamline multi-camera editing + GPS with telemetry data ✔️ New Point-and-Shoot Grip compatibility for elite handheld control ✔️ Removable Lens Hood included to reduce glare + flares ✔️ Bluetooth® 5.3 Super Wideband connectivity + USB-C port for external audio capture ✔️ A cinema-grade camera that anybody can use Enhanced by a GoPro Subscription: ✔️ Unlimited cloud backup at 100% quality ✔️ Camera replacement guarantee ✔️ Up to 50% off select accessories Order your MISSION 1 Series camera now to get a free Point-and-Shoot Grip ($100 value) + free shipping at Pro-tip: Existing GoPro Subscribers save $100 with the annual camera discount.

GoPro

20,272 Aufrufe • vor 1 Monat

🇨🇳 Another great Chinese Model, OmniHuman-1.5 from ByteDance Turns 1 image plus a voice track into expressive avatar video by pairing a System 1 and System 2 inspired planner with a Diffusion Transformer, Produces coherent motion for over 1 minute with moving camera and multi character scenes. Most avatar models move to the beat of the audio but miss meaning, so gestures feel generic and emotions feel shallow. The fix here is a Multimodal LLM planner that listens to the speech and drafts a structured plan describing intent, emotions, beats, and high level actions, which gives the motion engine clear semantic targets instead of only rhythm. The motion engine is a Multimodal Diffusion Transformer that fuses the plan with audio, the single reference image, and optional text prompts, then synthesizes continuous body, face, and head motion that matches both words and tone. A key trick is a Pseudo Last Frame, a synthetic target that summarizes the next expected state, which stabilizes fusion across modalities and keeps motion consistent over long spans. From just 1 image and speech, the system outputs speaking avatars with synchronized lips, context aware gestures, and continuous camera movement, and it also supports multi character interactions without manual choreography. Reported results show strong lip sync accuracy, high video quality, natural motion, and close match to text prompts, and the same setup works on nonhuman characters too.

Rohan Paul

63,859 Aufrufe • vor 10 Monaten