正在加载视频...

视频加载失败

here's a follow up ground based drone tracker :> this little setup runs with 100Hz+ tracking updates, ~20ms latency, low light, no motion blur. (4d object tracking + 6dof pose/velocity)

116,499 次观看 • 3 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

We are excited to share our work “Event-Aided Sharp Radiance Field Reconstruction for Fast-Flying Drones” published in IEEE Transactions on Robotics IEEE Transactions on Robotics (T-RO), which tackles sharp radiance field reconstruction under agile drone motion, where RGB frames are heavily motion-blurred and pose priors become unreliable! 4 years in the making! Code & dataset released! PDF: Code & Dataset: Full Narrated Video: High-speed flight is essential for time- and battery-constrained missions (e.g., inspection, exploration, search & rescue). However, fast motion corrupts visual data with severe motion blur and introduces drift/noise in visual-inertial odometry, making NeRF-based 3D reconstruction particularly brittle. We propose a unified framework that leverages asynchronous #EventCamera streams together with motion-blurred frames to reconstruct high-fidelity radiance fields from agile drone flights. Our key idea is to embed event-image fusion directly into radiance field optimization while jointly refining a shared, continuous-time camera trajectory initialized from event-based VIO. This enables us to recover sharp radiance fields and accurate trajectories without ground-truth supervision during training. We validate our method on synthetic data and on real sequences captured by a drone flying up to 2 m/s. Despite severe blur and noisy pose priors, our method preserves fine scene details and achieves a performance gain of over 50% on real-world data compared to state-of-the-art methods. Kudos to Rong Zou and Marco Cannici! Marco Cannici Reference: Rong Zou*, Marco Cannici*, Davide Scaramuzza Event-Aided Sharp Radiance Field Reconstruction for Fast-Flying Drones IEEE Transactions on Robotics (T-RO), 2026 NCCR Robotics European Research Council (ERC) AUTOASSESS UZH IfI University of Zurich UZH Science Prophesee SynSense UZH Space Hub

Davide Scaramuzza

12,028 次观看 • 5 个月前

DRONE VIDEOGRAPHERS CHARGE $10K FOR THIS SHOT. HE PULLS IT FROM GOOGLE EARTH AND A PROMPT You never buy a drone, book a pilot, or leave the house. You pick any city on Earth, trace the flight path you want, and let Gemini render it as real-looking FPV footage. Clients pay thousands for this shot. You make it from a screenshot Here is the exact process: 1. Open Google Earth. Find the city or building you want. Frame the angle you'd want a drone to start from and take a screenshot 2. Draw the path. On that screenshot, draw a red line showing exactly where the drone should fly through the scene. This line is what the AI follows 3. Open Gemini and drop in the screenshot. Use the video generation in the Gemini app, the part that animates a still image into motion. Nano Banana handles images, the video engine is what turns your shot into footage 4. Paste the prompt. Tell it to follow the red flight path through the city, fast smooth motion, banking around buildings, golden-hour light, motion blur, 9:16 vertical, real FPV drone look. Full prompt is in the comments 5. Generate and clean it up. One clip is a few seconds. Stitch a couple together for a full flythrough and you have a reel Set the prompt once and you can re-run it for any location on the planet Who pays for this: Real estate agents, hotels, restaurants and event venues all need aerial b-roll and almost none can afford a real drone shoot Pull listings or venues with flat, ground-level photos and zero aerial footage. Send a free sample flythrough of their own location, then charge per clip or a monthly rate for ongoing reels One agent with ten listings is a recurring client, fully online Full prompt in the comments Bookmark this

Yarchi

53,401 次观看 • 2 个月前

Created this race using GPT Image 2 and Seedance 2.0 on TapNow Prompt Follow the storyboard strictly in exact order from Panel 1 to Panel 9. Do not skip, merge, or rearrange scenes. Keep the SAME female cyclist identity across the entire film. No face changes, no hairstyle changes, no helmet changes, no body proportion inconsistencies. Baby pink must remain the dominant apparel color throughout all cycling scenes. Avoid black wardrobe replacements. Preserve realistic nighttime lighting continuity between shots. Maintain the same cool blue tones and subtle red light reflections. Heavy rain intensity must stay visually consistent across all scenes. Water physics must look physically accurate: droplets, splashes, mist, wheel spray, and runoff should behave naturally. Avoid artificial AI motion. Camera movement should feel like real cinema rigs, FPV drones, mounted bike cameras, or stabilized tracking systems. Drone shots must maintain locked framing and smooth movement without random drifting or orbiting. Use subtle cinematic motion only — no excessive shaking or jitter. Keep realistic breathing, body fatigue, pedaling mechanics, and fabric reactions to wind and rain. Preserve shallow depth of field in macro shots and atmospheric haze in wide shots. Keep the environment dark, moody, and cinematic with strong contrast between wet reflections and darkness. Ensure all reflections on asphalt, water droplets, and bike components react naturally to changing light sources. Maintain premium commercial pacing: slow controlled preparation and macro shots transitioning into aggressive high-speed riding sequences. Final output should resemble a high-budget Nike / Rapha night cycling commercial shot during a real mountain storm. Ultra-realistic cinematic night cycling commercial about female endurance cyclists riding through an intense rainstorm in the mountains at night. Premium Nike / Rapha aesthetic with baby pink performance cycling apparel as the dominant accent color. Hyper-realistic documentary look, no stylization, no anime look, no beauty filters. Natural skin texture, realistic rain interaction, physically accurate water behavior, cinematic low-key lighting, cool blue night tones mixed with subtle red rear-light reflections. Heavy rain, fog, wet asphalt reflections, cinematic motion blur, high dynamic range, shallow depth of field, premium sports commercial quality. The film follows a strict 9-panel storyboard structure with seamless cinematic transitions and continuity preserved across every scene. The SAME female cyclist identity must remain consistent throughout the entire video: same face, helmet, glasses, body proportions, baby pink apparel, lighting style, and overall appearance. Maintain continuity of rain intensity, wetness, fog density, and environmental lighting between all shots. Panel 1: Extreme macro close-up of the female cyclist’s eyes and face in heavy rain at night. Focus on soaked eyelashes, wet skin texture, raindrops streaming across the face, baby pink helmet and baby pink face mask visible. Red rear bike light flickers dynamically across her eyes and skin while cool blue night tones dominate the scene. High contrast cinematic lighting, shallow depth of field, subtle breathing motion, intense determined expression. Panel 2: Cinematic medium close-up frontal shot of the cyclist riding aggressively through heavy rain at night. She pedals hard with strong effort and forward-leaning posture. Baby pink waterproof cycling jacket soaked with rainwater. Front bike light cuts through fog and rain with subtle flickering illumination. Wet asphalt reflects red and white lights. Smooth cinematic tracking shot with controlled stable motion and slight natural float. Panel 3: Ultra-realistic macro shot of large raindrops impacting wet asphalt at night. Crown-shaped splashes and overlapping ripples in slow motion. Rough wet asphalt texture, cool blue cinematic tones, subtle reflections from bike lights.

Sharon Riley

72,716 次观看 • 3 个月前

Neon Drift: The 46 JDM Legend – Midnight High-Speed Run Made with seedance 2.0 Prompt: Cinematic 30-second vertical 9:16 video, highly detailed anime/cinematic 3D style like Arcane + Cyberpunk 2077, night to golden hour transition. A handsome young man with messy blonde hair, black thick-rimmed glasses, light beard, serious intense look, full sleeve tattoos on both arms, wearing black shirt and grey pants, standing in a rainy neon-lit cyberpunk city street at night. He checks his glowing smartphone, then walks confidently towards a white Honda Integra/JDM coupe with bold black "46" graffiti on sides, red underglow lights. He opens the door, sits inside, presses the start button (close-up on tattooed hand), dashboard view with his glasses reflecting colorful neon lights on the gauges. White smoke bursts from the exhaust. Dynamic driving sequence: car speeding through wet neon streets under overpasses with colorful signs, then powerful drift with thick white + red smoke, red underglow glowing on wet road. Epic wide shots of the white sports car with "46" graphics drifting and accelerating on a big bridge during beautiful purple-orange sunset, city skyline in background, dramatic lens flares, motion blur, cinematic camera angles (low tracking shots, side profile, rear tracking, aerial). Moody cyberpunk atmosphere, reflections on wet roads, volumetric fog, intense colors, smooth transitions, high energy, satisfying car sounds implied, premium car commercial feel, ultra realistic details, 4K, cinematic lighting --ar 9:16 --stylize 250 --v 6

Noor

17,167 次观看 • 1 个月前

Girl puts on headphones then becomes Spider Woman over New York. Made with seedance 2.0 Prompt: A young East Asian woman with a short bob haircut sits on a wide windowsill in a messy Brooklyn-style apartment bedroom. She wears a black long-sleeve top, white collared shirt with a dark tie, gray cargo pants with a light sweater tied around her waist, pink-and-white arm sleeves, bright pink socks, and teal sneakers. Black over-ear headphones rest around her neck. The room is cluttered with anime posters (including Dragon Ball), comic books, a Monster Energy can, Funko Pops, clothes scattered on the floor, and a bed to the right. Outside the large open window is a cloudy New York City skyline of red-brick buildings. She looks at the camera with a slight smile, puts the headphones on, leans back, then suddenly swings out the window on a thin white web line. Dynamic tracking shots follow her web-slinging through the streets of New York: she flies between brick apartment buildings with fire escapes, over busy intersections filled with yellow taxis and pedestrians, past corner stores and traffic lights, diving and twisting acrobatically. The camera moves with her — low angles looking up, high angles looking down, fast motion blur on the city. She lands on a rooftop, stands with arms outstretched in triumph, hair blowing, looking out over the golden-hour skyline. The view includes water towers, dense rooftops, the East River, and the distant One World Trade Center glowing in the soft sunset light. Cinematic, realistic live-action style, vibrant colors, energetic movement, inspired by superhero web-slinging sequences.

Noor

22,966 次观看 • 4 天前

Everyone is sleeping on Meta's SAM 3 release. But it's actually a big deal. Here's why: Companies spend millions paying humans to label images and videos frame by frame. A single autonomous driving dataset? Months of work, hundreds of annotators, millions in cost. Without labeled data, you can't train custom models. Without custom models, you're stuck with generic solutions. This is why most companies never move past pilots. SAM 3 breaks this cycle. First let's look at the evolution: SAM 1 segmented objects when you clicked on them. Revolutionary, but one object at a time. SAM 2 added video tracking with memory. Game-changing, but you still manually prompted every object. SAM 3 changes everything with text prompts. Type "yellow school bus" and it finds ALL of them in your image or video. Not just one. Every instance across thousands of frames. Now here's where people get confused: "Can't I just use GPT-5 or Gemini for this?" No, and here's why that's a terrible approach. Large multimodal LLMs are great for reasoning, but they're slow and expensive for production visual tasks. You're paying API costs per image, waiting seconds for responses, getting inconsistent results. SAM 3 runs in 30 milliseconds on a single GPU for 100+ objects. That's 100x faster, and you own the infrastructure. More importantly, SAM 3 gives you precise pixel-level masks, not descriptions. Try asking an LLM to segment every defective part on a manufacturing line in real-time. It won't work. SAM 3 does this effortlessly. The real breakthrough is their data engine. Meta built an AI-human hybrid system that's 5x faster for complex annotations. They trained SAM 3 on 4 million unique visual concepts - 50x more than existing benchmarks like LVIS. SAM 3 is trained on 4 million unique visual concepts, it handles everything: - Text-based concept search - Interactive refinement with clicks - Video tracking across frames - Zero-shot detection of new concepts The model is open source. Weights, code, and benchmarks are on GitHub. If you're building computer vision applications, this is the foundation model to evaluate. The annotation time savings alone will pay for integration costs within weeks. Find the relevant links in the next tweet!

Akshay 🚀

46,421 次观看 • 8 个月前