Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

EgoExoMoCap: Distributed Ego-Exo Human Motion Capture ECCV 2026 (Oral) Meta Reality Labs/ETH Zürich With two or more people wearing smart glasses, EgoExoMoCap combines each wearer’s egocentric (Ego) and exocentric (Exo) views to estimate full-body motion. Without bulky multi-camera setups or mocap suits, it reconstructs 3D human motion even under...

25,765 görüntüleme • 10 gün önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

MVP of Multiview Video → Camera parameters + 3D keypoints. Visualized with Rerun The basic pipeline as of right now looks like this: 1. Capture 🔴 – Using 4 iPhones and an Insta360 Go. iPhone videos are captured via Final Cut Pro Multicam for easy sync and the exocentric view; the Insta360 Go is used for the egocentric view. 2. Sync 🕒 – Custom Gradio app using two Rerun viewers and callbacks for easily aligning frame timestamps so the ego and exo views are aligned. 3. Calibrate 🎯 – Use VGGT from Jianyuan and AI at Meta to get intrinsics/extrinsics for sparse cameras. 4. Estimate 3D 🕺 – Use RTMLib whole‑body keypoint estimator on each frame, then triangulate in 3D. What's missing? 1. No temporal coherence: I’m estimating keypoints one frame at a time and one camera at a time. This leads to a lot of jittering. For now, I plan on adding a One Euro Filter to help with jittering. Long term, I'd want to train a multiview keypoint estimator 2. Kinematic fitting is still missing; this is my next goal. The output will be joint angles, as explored in my previous posts. 3. Missing dense point cloud: VGGT seems to fail for me here. I’m looking to explore using MP‑SFM as a method for generating dense multiview depth maps + normals (plus it has a friendlier license compared to VGGT). 4. Eventually, creation of 4D Gaussian splatting using something akin to DN‑splatter—my long‑term goal is a data engine that provides poses/depths/splats/keypoints/etc.

Pablo Vela

42,785 görüntüleme • 1 yıl önce

Meet My AI Ears. A lot of folks ask me how I capture ASMR video and audio for training of AI? I always use Binaural 3D audio and have for decades in different forms. But how? A Brief History of Binaural Recording Binaural recording, the foundation of 3D audio, dates back to 1881 when French inventor Clément Ader created the first system using multiple telephone transmitters at the Paris Opera to transmit stereo sound to listeners, simulating spatial presence. By the 1920s, patents like W. Bartlett Jones’ 1927 filing advanced devices for capturing and reproducing “binaural” signals. The 1930s saw Alan Blumlein’s work on stereophonic sound, which he termed “binaural,” laying groundwork for modern stereo. Commercial milestones hit in the 1950s with binaural records from labels like Cook Laboratories and the first binaural reel-to-reel tapes. A resurgence came in the 1970s with Neumann’s KU-80 dummy head, the first commercial binaural system. Today, it’s integral to VR, ASMR, and immersive media. The Technology of Binaural 3D Audio At its core, binaural recording mimics human hearing by using two microphones placed in ear-shaped molds or a dummy head, separated like human ears (typically 14-18 cm apart). This captures spatial cues: interaural time differences (ITD) for sound arrival timing, interaural level differences (ILD) for volume variations, and head-related transfer functions (HRTF) that account for how the head, torso, and pinnae filter sounds. The result? A 3D soundscape that tricks the brain into perceiving direction, distance, and elevation when played back via headphones—no speakers needed for immersion. Advanced setups use omnidirectional capsules (e.g., DPA 4060) for high-fidelity capture, often in silicone ears to replicate natural diffraction. I use the 3DIO Microphones today but I would cover a dummy head in texture material and place two stereo (4 channels) microphones in each ear. I would then mix down the resulting signals into stereo. The 3DIO series features dual omnidirectional capsules in realistic silicone ear molds, spaced 14 cm apart for compact, accurate 3D capture based on over 13 years of research into human hearing. Models like the Free Space Pro II use premium DPA 4060 CORE capsules for ultra-low noise and high sensitivity, delivering stereo output ideal for immersive applications like game audio. Today just about any ASMR producer uses these. But I use them to capture, curate and archive sound and video we will lose or just about lost for AI training in a way no model or AI company is doing today. I am duplicating the human 3D binaural audio experience and memory. Below is a crude demonstration. If you can listen in headphones. Or turn your phone sideways to feel the audio space. I’ll have a far more professional demo soon to show the real power of 3D audio. (Oh that music is a MIDI player that uses disks to play).

Brian Roemmele

30,800 görüntüleme • 8 ay önce

🚨 SIGGRAPH Asia 2025 Paper Alert 🚨 ➡️Paper Title: WorldExplorer: Towards Generating Fully Navigable 3D Scenes 🌟Few pointers from the paper 🎯Generating 3D worlds from text is a highly anticipated goal in computer vision. Existing works are limited by the degree of exploration they allow inside of a scene, i.e., produce stretched-out and noisy artifacts when moving beyond central or panoramic perspectives. 🎯 To this end, authors of this paper proposed “WorldExplorer”, a novel method based on autoregressive video trajectory generation, which builds fully navigable 3D scenes with consistent visual quality across a wide range of viewpoints. 🎯They initialize their scenes by creating multi-view consistent images corresponding to a 360 degree panorama. 🎯Then, they expanded it by leveraging video diffusion models in an iterative scene generation pipeline. 🎯Concretely, they generated multiple videos along short, pre-defined trajectories, that explore the scene in depth, including motion around objects. 🎯Their novel scene memory conditions each video on the most relevant prior views, while a collision-detection mechanism prevents degenerate results, like moving into objects. 🎯Finally,they fuse all generated views into a unified 3D representation via 3D Gaussian Splatting optimization. 🎯Compared to prior approaches, WorldExplorer produces high-quality scenes that remain stable under large camera motion, enabling for the first time realistic and unrestricted exploration. 🎯They believe this marks a significant step toward generating immersive and truly explorable virtual 3D environments. 🏢Organization: TU München 🧙Paper Authors: Manuel-Andreas Schneider, Lukas Höllein , Matthias Niessner 📝 Read the Full Paper here: 🗂️ Project Page: 🧑‍💻 Code: 🎥 Be sure to watch the attached Technical Summary Video - Sound on 🔊🔊 Find this Valuable 💎 ? ♻️QT and teach your network something new Follow me 👣, naveen manwani , for the latest updates on Tech and AI-related news, insightful research papers, and exciting announcements. #SIGGRAPHAsia2025

naveen manwani

10,578 görüntüleme • 11 ay önce

Fast Company just published a great piece on World Labs , Fei-Fei Li , Marble, and the idea that spatial intelligence / world models may be one of the next big shifts in AI. I was happy to be quoted in the article, but I also wanted to share more context about my own experience with World Labs and Marble, and why this direction is especially interesting to me. My starting point: volumetric capture — For the past few years I’ve been exploring and using volumetric capture and reconstruction (photogrammetry, NeRFs, 3D Gaussian Splats) mostly capturing locations around Montreal. Alleys, museums, urban interiors. I love every step of it: the capture itself, the pipeline, and what can be done with the output. Turning real spaces into real-time explorable systems. I do this personally, sharing explorations here, and professionally as chief technologist, and co-founder of Dpt. Physical reality + generative manipulation — In my work I’m especially drawn to mixing physical reality with generative and digital manipulation: using physical interfaces (light, clay, ink, ... ) to drive generative AI pipelines, building mixed reality prototypes that reshape your surroundings, or starting from real captured spaces and transforming them using tools like Marble. Like many people, I saw the World Labs announcement on Twitter in September 2024, and Marble when it surfaced in early December. But by then, I already had a sense something was coming. The first conversation — As someone deep into volumetric capture and radiance fields, I obviously knew about Ben Mildenhall and his pioneering work on NeRF. To my surprise, Ben reached out to me in late June 2024. He’d been following some of my experiments and wanted to chat about my process and workflows and how I was using this “stuff” creatively. At that point he didn’t share what he was building, but we had a genuinely great conversation about radiance fields, AI, and my work. He was curious about the creative perspective, not just the technical one. When the World Labs announcement dropped a few months later, it all made sense. I understood what Ben had been working on, and why the creative angle mattered to them. Then in August 2025, he invited me to try the Marble beta, and I’ve been experimenting with it since. Experimenting with Marble — The first thing I used Marble for was materializing scene and world concepts during ideation at the studio, and seeing if and how it could fit into our production pipeline. In parallel, I dove into a series of experiments focused on world manipulation: starting from real captured spaces and transforming them using Marble. I’d already been exploring that idea using img2img diffusion with ControlNet on NeRF renders, real-time video streams, and even mixed reality using headset camera feeds. But Marble brings something different. It generates persistent, spatially cohesive 3D worlds that can be rendered in real time across a wide range of devices. That’s a real shift. Experiment 01: Parallel Realities — The first experiment, Parallel Realities, starts from a volumetric capture of a real location, reconstructed as 3D Gaussian Splats. Using Marble, I generate an alternate version of that same space, something informed by the original architecture: abandoned, nature-reclaimed, alternate era. Then, using Spark (World Labs’ 3D Gaussian Splatting renderer for THREE.js) I make both realities coexist in the same spatial coordinate system. From there, I use a portal UX mechanic to let the user step between the real reconstruction and the Marble-generated version. Experiment 02: Hidden Depth The second experiment, Hidden Depth, does not transform a space as much as expand it. A captured location has a visual boundary (a mural, a doorway, a dark corridor) and Marble generates what exists beyond it. For example: a Montreal alley has a painted mural; step through it and you’re inside a world informed by what is actually depicted there. World Labs showcased part of this work here: And in their Spark 2.0 post: The project page is here: Why this matters to me — Being able to start from a real 3D Gaussian Splat scene and manipulate it with Marble opens up a lot of ideas. The 3DGS pipeline is becoming an increasingly compelling foundation for exploration, experimentation, and storytelling. What matters most to me right now is more control. The more I can steer the generated scene or world, the more useful the tool becomes. I want more features like the already existing multiple input images and Chisel, the blockout-based approach. I would like better local control, the ability to expand a generated world more and more while preserving coherence, and the ability to directly import 3D Gaussian Splat scenes to be used as a starting point. I want more ways to shape the result, not just a “prompt and hope” approach. — It is exciting to see this field moving from research and demos toward actual creative workflows.

Hugues Bruyère

69,960 görüntüleme • 2 ay önce

The Mathematics of Moving a Cursor with Neural Signals What might Neuralink Neuralink be doing Mathematically? Consider the task of moving a cursor without touching it. The machine is not looking for a full thought, a sentence, or an image. For this Control problem, the useful object is an intended movement state. sₜ = (pₜ, vₜ) Here, pₜ is the cursor position at time t, and vₜ is the velocity the user is trying to express. The implant records neural activity through many electrode channels, then the decoder tries to estimate vₜ from that activity. Neuralink’s PRIME material describes the N1 Implant as recording and transmitting brain activity with the goal of enabling computer control. For channel i, a simple population model is rᵢ(t) ≈ bᵢ + aᵢ max(0, dᵢ · vₜ) + ηᵢ(t) where rᵢ(t) is the measured activity, bᵢ is baseline activity, aᵢ is channel gain, dᵢ is the channel’s preferred movement direction, and ηᵢ(t) is noise. One channel is not the command. The useful signal is the pattern across many channels: rₜ = (r₁(t), r₂(t), …, rₙ(t)) The decoder subtracts the baseline vector b and applies a learned map W: v̂ₜ = W(rₜ − b) This gives an estimate of the intended velocity. The cursor then updates by pₜ₊₁ = pₜ + Δt v̂ₜ This is the loop shown in the render: neural activity -> decoded velocity -> cursor motion The cortical network and electrode threads show the measurement side. The N1 Implant is described as using 1,024 electrodes distributed across 64 flexible threads, each thinner than a human hair. The decoder panel shows the computational side with activity rₜ, decoded velocity v̂ₜ, and the cursor state pₜ changing over time. A noisy biological pattern becomes a state estimate. That estimate becomes motion on a screen. Therefore, the first lesson is not that Neuralink makes the brain a screen. For cursor control, the Mathematics is more precise: A small piece of intention is represented as a hidden state, measured through neural activity, decoded as a vector, and turned into action. #Neuralink #BrainComputerInterface #NeuralEngineering #Mathematics #StateEstimation #Neuroscience #MachineLearning #BiomedicalEngineering

Mathelirium

14,520 görüntüleme • 3 ay önce

Tiny drone hits invisible mode by twisting faster than eye can detect | Omar Kardoudi, New Atlas Engineers at Northwestern University have built a drone that vanishes without camouflage or transparent panels. Its trick is spinning so fast that your eyes simply give up trying to focus, a stealth edge that could turn surveillance into something almost invisible. The aircraft, nicknamed Phantom Twist, rotates up to 25 times per second, a rate that outpaces how quickly our visual system can process sharp detail. Instead of true invisibility, the drone dissolves into a faint, ghostly blur that blends into whatever is behind it. The work, led by associate professor Michael Rubenstein, was presented on July 16 at the Robotics: Science and Systems 2026 conference in Sydney, Australia, under the title Computational Design of a Low-Visibility UAV Using Human-Aligned Perceptual Metric. "Most efforts to hide drones focus on making them look like their surroundings," says Rubenstein. "Instead, we asked whether we could design the drone itself around the way humans perceive motion. This idea of low visibility through persistent motion is something few people have explored." That distinction matters because drones are increasingly used to watch wildlife, check aging infrastructure, or survey wetlands, but their mere presence changes the behavior of whatever they're observing. Birds scatter, animals flee, people act differently. A drone that's hard to spot could do the same job without that side effect. Prior attempts at motion-based concealment offer useful context here. The Northwestern paper points to an earlier project nicknamed the Boomerang Drone, covered in a 2006 New York Times Magazine piece, which tried a similar high-speed rotation trick but couldn't spin fast enough to fully exploit the blur effect, leaving it largely visible. The paper authors also trace the broader idea of active concealment back to the "Yehudi light," a counter-illumination project developed by the National Defense Research Committee in 1944 to hide Allied sea-search aircraft from enemy view. The Phantom Twist itself takes a very different shape from those earlier attempts. Rather than a typical quadcopter with four separate rotors, it runs on a single motor and a single propeller, with the propeller spinning one way while the rest of the drone's body spins the opposite way. "For a typical quadrotor drone, the propellers are spinning, but the robot is stationary," Rubenstein explains. "So, you still see its body. For our drone, the whole thing is rotating, so there are no stationary parts." To reach that layout, the team's computer model generated roughly 20,000 possible drone configurations capable of stable flight, then used artificial intelligence and optimization algorithms to repeatedly rearrange the motor, propeller, circuit board, counterweight, and batteries. Each design was simulated spinning mid-flight and overlaid on 100 real-world backgrounds, then scored by a perceptual model built to mimic human vision, where a lower score meant better camouflage. The 500 best-scoring designs were run through the optimizer again to squeeze out further gains before a final version was built. Emma Alexander, an assistant professor of computer science and one of the study's co-authors, explains the underlying physics. "The human eye takes time to accumulate signals, roughly analogous to the exposure time of a camera," she says. "When an object spins quickly, we perceive it as blurring out and losing distinct features. Because this new drone is almost entirely transparent, its few opaque components are visually averaged with the background for an overall appearance of a slight haze." According to the paper's visibility metric, the finished drone is about 10 times harder to spot than a standard quadcopter. But the spinning trick has real limits that make this drone far from being completely unnoticeable. The propeller still makes an audible whir that gives the drone away even when the eye can't, and its support wires and rods remain partly visible. The paper's authors suggest future versions could lean on more transparent materials and quieter propulsion, edging the drone ever closer to true – and somewhat scary – invisibility. After all, the same trick making a drone less impactful on wildlife could just as easily help it sneak around for reasons that aren't so friendly.

Owen Gregorian

24,363 görüntüleme • 1 ay önce

Stanford professor Judy Fan went on stage at MIT and broke down why humans are so good at making the invisible visible... And why AI hasn't actually learned to "see" the way we do. It completely changes how you think about Human Intelligence v/s Artificial Intelligence: 1. Nature never gave us straight lines or sharp corners. The number line, the coordinate plane, even basic geometry are all human inventions. We created tools that do not exist in nature simply because we needed a way to think more clearly. 2. The coordinate system Descartes invented solved a problem that had stumped mathematicians for centuries, doubling the volume of a cube. Once invented, this tool became so indispensable that virtually every math curriculum on Earth still depends on it. 3. Humans have been doing this for at least 30,000 to 80,000 years. The story of human progress is inseparable from the story of marking up our environment, from cave walls to Galileo's telescope to Feynman diagrams of particles we will never see with our own eyes. 4. Every major scientific breakthrough relied on a visual tool that made something invisible visible. Darwin needed side-by-side illustrations of finches to see variation that was otherwise too subtle to notice. Cajal needed detailed drawings of neurons under a microscope to map how the nervous system was wired. 5. Fan's research group studies something deceptively simple: how people decide what to put into a drawing and what to leave out. When two people played a drawing game, sketchers used far more detail when the target object had close competitors than when it stood alone, all the way down to using fewer strokes and less time when more detail was not necessary. 6. People are not just copying what they see. They are making constant judgment calls about what level of detail actually serves the goal of communication, and they do this naturally without ever being taught the theory behind it. 7. There is a real difference between drawing something so someone can identify it and drawing something so someone can understand how it works. In one study, participants drew explanatory diagrams that emphasized moving, causal parts of a machine while depictive drawings emphasized background and overall appearance, even though both were drawing the exact same object. 8. Explanatory drawings were genuinely better at helping someone figure out how to operate a machine, but worse at helping someone identify which machine it actually was. You cannot optimize a single drawing for both goals at once. Communication always involves tradeoffs. 9. AI vision models trained on photographs generalize surprisingly well to simple, sparse sketches, suggesting that resemblance based recognition is not just a story we tell ourselves. It is something modern neural networks can replicate with real accuracy. 10. But there remains a large, measurable gap between how confidently AI models recognize sketches and how confidently humans do, even when both groups answer the same questions about the same images. Humans are simply far more reliable and far more consistent in their judgments. 11. When researchers compared human-made sketches to AI-generated sketches under tight stroke budgets, both were similarly recognizable at higher budgets, but diverged sharply as the budget shrank. Humans and AI systems simplify drawings in fundamentally different ways once resources get scarce. 12. Reading a graph is not one single skill. It involves perception, knowing where to look, mapping that visual information onto the actual question being asked, and then translating that mapping into an answer. Each of these steps can independently break down, and people fail for very different underlying reasons even when they land on the same wrong answer. 13. When tested directly against humans on graph reading tasks, leading multimodal AI models, including GPT-4V, showed a meaningful performance gap. Even when a model's overall accuracy approached human levels, its pattern of mistakes looked nothing like how humans actually get things wrong. 14. People choose entirely different types of charts depending on what specific question they are trying to answer, not out of a generic preference for bar charts or scatter plots. Their chart choices closely tracked which visualization would genuinely help someone answer that specific question correctly. 15. Two of the most widely used graph literacy tests in education research turned out to correlate strongly with each other, suggesting they measure overlapping skills. But when researchers dug into the actual error patterns, the standard categories used in textbooks, like "find the maximum" or "identify a cluster," failed to explain why people got things wrong nearly as well as a more basic, underlying four-factor model did. 16. The deepest goal behind all of this research is not just academic curiosity. It is to eventually help students and everyday people develop genuine literacy with the visual tools that science and modern decision-making increasingly depend on, because every generation should be able to see further than the last by standing on the visual tools the previous generation built. Follow Yasmine Khosrowshahi for more ideas on thinking better, becoming clearer & building a more intentional life.

Yasmine Khosrowshahi

890,410 görüntüleme • 1 ay önce

Kled Version 3 is coming. Over $20M+ in rewards will be paid directly to users from leading AI labs across robotics, legal services, image and video generation, world modeling, and more. In the last seven days, we’ve received inbound data requests from several decacorn AI labs and enterprises for datasets our human data marketplace is uniquely positioned to provide. Since receiving the specs for these requests, we now have a much better picture and understanding of how to reshape the systems that collect this data, so here’s what’s coming: 1. A fully redesigned home experience: The home feed is being rebuilt to surface the highest-value, most relevant tasks for each user, similar to how Uber Eats surfaces top restaurants. The goal is to turn every user into their most effective version as a data contributor. 2. Automated quality enforcement at scale: New ML systems are being built to evaluate task-specific requirements in real time. For example, if a task requires “two hands visible on camera at all times,” any video that fails that spec will be automatically rejected. This logic will apply across thousands of tasks and specifications using a general ML. 3. Kled Shop: Some tasks require better capture hardware. We’re introducing Kled Shop, where users can redeem points or tokens for equipment like Meta glasses, drones, and other tools. Points and tokens can be converted directly from payouts. 4. Partner-run data labeling and evaluation work: Some of our partners operate high-paying data labeling and model evaluation programs. We’re integrating their workflows directly into Kled so qualified users can access these roles in one place. These jobs are owned and managed by our partners. Kled’s role is to route the right people to the right work. Some opportunities pay $50–$1,000 per hour depending on expertise. 5. Global payouts and localization: We’re partnering with a major payment processor to enable cashouts in users’ native currencies. This unlocks broader global participation. Multi-language support is also coming to accelerate user growth. This full suite of tools will be rolling out soon, directly to Kled users. Top earners are currently making ~$7,000 per month. With this update, we should see the first ~$10,000 per month earner.

Avi Patel

124,728 görüntüleme • 7 ay önce

Ahmedabad Crime Branch is making use of technical measures to avoid any stampede kind of situation. Anti stampede visual analytics,using reference area and crowd movement, head count algorithm. Anti-stampede algorithms on CCTV cameras are a crucial advancement in crowd management, leveraging AI and image processing to prevent dangerous situations in densely populated areas. Here's a breakdown of their usage: How they work: Real-time monitoring: AI-powered CCTV cameras continuously analyze video streams in real-time. Crowd density estimation: Algorithms calculate the number of people in a given area. This can involve: Pixel-based analysis: Converting images to black and white and counting "black pixels" (representing people). Object detection: Using machine learning models (like Mask R-CNN) to identify and count individuals, often by detecting heads or torsos. Thresholding: Pre-defined "threshold values" for crowd density are established. When the detected density crosses these thresholds, it triggers an alert. Anomaly detection: Beyond just density, these algorithms can identify unusual crowd behaviors such as: * Sudden surges in movement. * Unusual clustering patterns. * Fallen individuals. * Aggressive movements. Alerting authorities: Upon detecting a potential stampede risk, the system sends immediate alerts to security personnel or control rooms via LCD displays, GSM messages, or other communication channels. Predictive analytics: Some advanced systems use time-series prediction models to forecast crowd behavior and dynamics based on historical and real-time data, helping anticipate potential bottlenecks or overcrowding. Reinforcement learning: Algorithms can learn from past incidents to suggest optimal crowd flow routes and alternative evacuation paths during emergencies. Benefits: Proactive prevention: The primary benefit is the ability to detect and warn of potential stampedes before they occur, allowing authorities to take preventative measures. Real-time insights: Provides immediate and accurate data on crowd density and movement, far surpassing manual observation. Enhanced safety: Significantly improves safety in public spaces by reducing human error and enabling swift responses to risks. Optimized resource allocation: Helps in better deployment of security personnel and resources to areas with high crowd density. Improved efficiency: Automates a labor-intensive task, freeing up human operators for more complex decision-making. Data for future planning: The collected data can be analyzed to improve crowd management strategies for future events. Challenges: Accuracy limitations: While advanced, AI algorithms can still face challenges with: Occlusion: People blocking each other, making accurate counting difficult. Varying conditions: Changes in lighting, weather, and camera angles can affect accuracy. Bias in training data: Can lead to false positives or inaccurate detections. Computational complexity and cost: Developing and deploying such systems can be expensive due to the need for high-resolution cameras, powerful processing units, and sophisticated algorithms. Data privacy and ethical concerns: The extensive use of CCTV and AI raises concerns about individual privacy and potential misuse of data. Integration with existing infrastructure: Integrating new AI-powered systems with older CCTV networks can be complex. Human intervention still crucial: While AI can alert, human responders are still essential for effective intervention and crowd dispersal. As seen in the Kumbh Mela example, even with AI alerts, a lack of ground personnel can limit effectiveness. Defining thresholds: Determining appropriate crowd density thresholds for different environments and cultural contexts can be challenging. Real-world applications: Large public gatherings: Religious festivals (like the Kumbh Mela in India, which has used AI for crowd management), concerts, sports events, and political rallies. Transportation hubs: Railway stations, airports, and bus terminals to manage passenger flow. Shopping malls and commercial centers: To monitor crowd density during peak hours and special events. Stadiums and arenas: For managing ingress, egress, and crowd movement during events. Tourist attractions: To prevent overcrowding at popular sites. Overall, anti-stampede algorithms on CCTV cameras represent a significant leap forward in ensuring public safety, offering a powerful tool for proactive crowd management. However, their successful implementation requires careful consideration of technological limitations, ethical implications, and the continued need for effective human intervention. Ahmedabad Police અમદાવાદ પોલીસ Vijay Patel | Megh Updates 🚨™ | Akash Anand | | #BengaluruStampede | #Stampede

Janak Dave

339,758 görüntüleme • 1 yıl önce

Can't get enough of Seedance 2.5. The best part is it can generate 30s video in one single prompt, and the result is so realistic! So, here's another one that I made using CapCut. Prompt: [STYLE + CAMERA + ATMOSPHERE] Gritty, raw handheld 35mm film aesthetic with natural film grain. Harsh direct sunlight creating high-contrast shadows over a dramatic coastal cliff and open ocean. Continuous single-take handheld tracking shot (3rd-person / over-the-shoulder) with no cuts. Atmosphere: high-altitude wind, realistic coastal cliff and ocean physics, sudden wingsuit deployment. Audio: heavy rhythmic breathing, intense wind howl, fabric snap of wingsuit opening, high-speed air rush over open water, near-miss whooshes past yachts, soft landing roll on sand, distant ocean waves and beach ambient noise, final bite sounds. [IMAGE REFERENCES] Use the provided Hoshino character sheet as the single strict visual reference for the male character. Exact face, black hair, dark brown eyes, lean athletic build (178 cm), gold earrings, and overall facial structure locked from the reference. Outfit: modern high-performance cliff-jumping wingsuit — sleek, form-fitting design in matte charcoal black with sharp white paneling and subtle gold zipper accents (matching his minimalist aesthetic). The wingsuit is worn from the start with a matching technical backpack. Body proportions, posture, and face remain fully locked to the Hoshino reference. [TIMELINE SECOND BY SECOND] 0-3s: [Handheld medium] Hoshino stands on the edge of a high rocky cliff overlooking the open ocean, wearing his charcoal-and-white wingsuit and technical backpack. He looks straight into the camera with calm, confident intensity, then turns and launches into a clear, controlled forward somersault in slow motion as he leaves the cliff edge. 3-5s: [Continuous freefall] He completes the somersault and falls head-first toward the sea. At exactly 1.5 seconds into the fall he fully deploys the wingsuit with a sharp snap. The wing membranes inflate and he levels out into a smooth glide. 5-12s: [High-speed tracking] Camera stays locked behind him as he rockets at full speed just above the ocean surface along the coastline. He weaves tightly past sheer cliff faces, banking hard left and right, turquoise water and rocky walls streaking past at extreme velocity. 12-18s: [Low-level chaos] He drops lower, flying just above the water. He almost collides with a large luxury yacht that suddenly turns, banks hard to avoid it, then narrowly misses a smaller motor yacht. He dips under a yacht’s outstretched boom and threads between two more vessels. 18-23s: [Water-level action] Still flying extremely low over the sea, he dodges a startled seabird that dives across his path, skims past a group of people on a nearby yacht who scatter in surprise, and banks sharply to avoid an open yacht swim platform. The camera stays locked behind him through every near-miss. 23-26s: [Landing] He flares the wingsuit hard, touches down with both feet on the soft beach sand and immediately rolls forward to kill the speed. He stands up, peels off the wingsuit in one fluid motion and drops it on the sand beside him, revealing a clean fitted white undershirt underneath. 26-28s: [Beach level] He walks a few steps still wearing the backpack and stops right in front of a classic beachside hot-dog cart. The vendor hands him a steaming hot dog fresh off the grill. 28-30s: [Close continuous] He turns, looks directly into the camera, takes a big bite of the hot dog and chews with a calm, satisfied expression as the shot holds. [STYLE & QUALITY BOOSTERS] Photorealistic 8K, ultra-detailed textures, cinematic lighting, perfect motion blur, high dynamic range, coherent physics (fabric, air, wingsuit membranes, impact, roll, near-misses with yachts, water spray), stable character locked to the Hoshino reference, realistic ocean reflections, cliff rock textures and wind, no artifacts, movie-level stability, pure single continuous take.

MrDejie

197,939 görüntüleme • 19 gün önce

A surveillance regime is being assembled in front of our very eyes. Most of you do not understand the tyrannical nightmare directly ahead. People still talk like this is a future problem. It's not. The battle against dystopia is now. 1984 is here. It arrives as "efficiency," procurement, integrations, device adoption, and a slow widening of what the state can do to you without asking, without noticing you, without needing you to consent. The attached video shows a federal agent in tactical gear wearing Meta Ray-Ban smart glasses on his face while carrying out state work. A camera at eye level. A microphone. A network connection. Constant recording that becomes a file, then a feed, then a searchable object that can be attached to a case, cross-referenced, stored, shared, and re-used. Raw footage is noise until it can be fused, searched, and operationalized. That is where AI enters. ICE is paying Palantir to build "ImmigrationOS," described as a platform for near real-time visibility into enforcement workflows, including tracking people and helping decide who gets targeted. Immigration is a perfect testing ground because the public has been trained to tolerate exceptional measures when the target group is already politically disposable. The capability does not stay there. It never stays there. The nightmare is the stack itself. Capture at the edge, aggregation in the middle, AI scoring and targeting at the top. Doorbell cameras feed the state. License plate readers feed the state. Data brokers feed the state. Local sharing agreements feed the state. Wearables feed the state. AI compresses the labor cost of suspicion, so the state can watch more people, more often, with less human effort. That is the mechanical shift most people still do not understand. AI does not need to be sentient to be oppressive. It needs to be cheap enough to scale the state’s attention and fast enough to outrun your ability to contest what it thinks it knows. Let me paint you the picture: an agent walks up the block wearing smart glasses, and your face, your voice, your license plate, and the people you are standing with get captured immediately. That clip hits an internal system where AI transcribes, tags, and links it to whatever identifiers already exist, your name, your address, your contacts, your past crossings, your employer, your car, your social graph. The platform fuses that with location trails and camera networks, then assigns a "priority" score that quietly moves you up a queue nobody outside the system can see. A caseworker opens a dashboard and clicks through prebuilt options that generate a task list, knock, detain, transfer, pressure, repeat, with the paperwork already half-written by AI. This is tyranny by procedure. The real chokehold is anticipatory control. Once people know they can be indexed, scored, and surfaced for enforcement, they start trimming their speech, their associations, their routes, their friends. Dissent becomes a risk factor, organizing becomes exposure, and the state does not need to ban protest when it can make participation costly enough to thin the crowd. It inevitably expands to include anyone who interrupts the smooth operation of power, and this kind of surveillance gives power the ability to punish quietly, repeatedly, and selectively, with plausible deniability stapled to every click. This is the formula for tyranny. The state gains capability, then finds incentives to justify using it. Agencies protect budgets by "demonstrating output." Contractors protect revenue by "delivering results." Politicians protect narratives by demanding visible "enforcement." Restraint gets treated as "inefficiency." Efficiency gets treated as "virtue." Your rights get treated as "friction." Any system that allows institutions to assemble your biography without your consent, then act on it without due process, is an engine of domination. Some people tolerate it because they think they will never be the target. That belief is childish. Power does not remain polite. It expands to fill the permissions you give it. A surveillance regime always needs new enemies to justify itself, because a machine built for pursuit must pursue. The regime being assembled is not subtle, it is simply normalized. It runs on boredom, exhaustion, and the assumption that someone else will handle it. It's all being wired now, and once it is wired, it will not ask your permission to be used.

Dylan Allman

327,629 görüntüleme • 7 ay önce

Would you believe an AI agent can test a real VR action game in real time, the way a person plays it? Meta XR Operator makes it possible. As far as I know, this is the first time. I am not talking about tapping a menu or replaying a recorded click path, but genuinely moving, shooting, and using the same game mechanics a human player does. In NeonReach VR, which is a real (and open source) action game, rings spawn 12m out and come at you somewhere between 1.5 and 5.5 m/s, getting faster over a 90 second ramp. There are three kinds: straight, weaving side to side, and spinning. Every shot is a full slingshot cycle, so you press, pull back, aim, then release. Obstacles arrive at head height and cost you a life if you don't get out of the way. You have ten lives. Here is why the game is hard for an AI agent. Even though Meta XR Operator gives the agent everything it needs to observe the app and act inside it, the agent still cannot play. One agent turn takes 10 to 15 seconds. One throw is four steps that have to happen in order, because the press has to latch before the pull, and they cannot be batched into a single call. So a throw costs about 45 seconds. A fast ring only exists for 2.3 seconds. One action takes 20x longer than the target is alive. Prompt tuning does not close a gap that size. What works is a three stage path: EXPLORATION, then SKILL, then SCRIPT. 1/ EXPLORATION. The agent drives the live app and works the game out on its own. It verified the coordinate mapping by setting a pose and reading it back, then derived the launch model. The more useful output was the traps it found. For example, the player's own body collider silently deflects a ball released inside it, with no error and no log line. That produced two confident wrong conclusions before anyone caught them. 2/ SKILL. All of that gets written down as a reusable SKILL.md plus an aim solver. There is a section that separates what was actually verified from what was assumed, so a wrong conclusion cannot quietly turn into doctrine. This stage also produced the trick that mattered. Set timeScale to 0 and a throw becomes atomic in game time, so however long the agent spends thinking never shows up in the shot. 3/ SCRIPT. The agent then compiles everything into a player script, a loop that observes, decides, and throws, calling the MCP servers directly from Python with no model in the hot path. Round trips drop from 10 to 15 seconds down to something between 1 and 16 milliseconds. The loop runs at 23 Hz, about 0.75 seconds per throw, roughly 60x faster than the agent doing it turn by turn. The result is that it plays like a person, which you can see from the attached video. It tracks the rings, works out where each one is going, throws with whichever hand is free, moves out of the way of the obstacles, and does not wait around to see whether the last throw landed. Shipping settings, no difficulty edits, no health locks, no slow motion. It plays until it actually loses. The takeaway generalizes beyond games: an agent does not have to be the player. Even following the same rules as a player, it has too much latency between moves. Having the agent write the thing that acts bypasses that constraint entirely. Try it yourself: or explore the agent-created skill and scripts: Based on NeonReach VR by Dilmer, with no code changes. I only upgraded its Meta XR Core SDK to v205, which ships Meta XR Operator. Our blog post, Introducing Meta XR Operator: Close the Build-Test-Verify Loop for VR: Disclosure: I work at Meta. And this represents my own opinion. #XR #VR #AI #MetaQuest #Unity #GameDev

Xiang Wei

51,880 görüntüleme • 6 gün önce

I made this product launch video over the weekend with just prompts It's all vibe coded There's something you should know, though: Like everyone else, a few days ago my timeline started getting full of videos like this when Remotion launched their Claude skill, so I decided to give it a go I was captivated by all the examples, so I started like everyone was saying: "just write a prompt" I typed the prompt, and it created an extremely bland, untasteful, stock-looking video 10 prompts in and it was not getting better. It was very, very bland. But at least it was something, so I kept going at it I ended up spending my entire weekend on this, 2-3 days of work. Only to realize my original reference videos that inspired me to get started were all fake Everyone was outright lying about their results. They all claimed "I made this with just one prompt", but it was just bait, they didn't really use Remotion or code at all, it was just a normal, human-made motion video Then you expand the X post and read the replies and they're all like "haha joke" in the comments, but their main post already got 1.5 million views and bamboozled everyone who didn't read further And this is a problem: when a viral trend happens, these posts flood your timeline, and you only realize that they're all noise and bait (and that they haven't even used the tools they claim) when you click through the post and read its comments. But 90% of people (like me, initially) just see the post on their timeline while scrolling, and assume it's all real. You don't go in to check every single post you see: you just like it, or save it for later, and carry on with your day, thinking what you saw was the real thing, and that it's all outstanding results, and that motion designers are really done And it's so anxiety inducing, because everyone is hyping their results, but most of it is just not true. I have stopped reading X lately because going in makes me so anxious, everyone is claiming extraordinary outlier results just for the views and clicks, and you feel like you're lagging behind and you're not good enough because you don't get those results So for this video I decided to actually take the tech out for a spin, and see what results I could really get out of it I used Remotion and Claude Code 4.5, but contrary to what everyone was claiming, this video was not "just a prompt". It was fully vibe coded, but it required much more than a prompt. It was multiple days worth of work Here's what I learned: - Making vibe coded videos with Remotion is ~10-20x slower than building app code. I've been wasting my Claude limits on this video - Everything takes a lot of manual work and reprompting. You often need to go frame by frame correcting tiny things - It makes very silly mistakes - Even Opus 4.5 has very very limited knowledge of spatial / visual things. It doesn't understand well z-indexes, layers, compositions, proportions, temporal coherence, etc. Claude Code feels extremely dumb when creating code for Remotion videos, which surprised me a lot, beacuse I had been mind blown by how incredibly well it worked with my Ruby on Rails SaaS codebases - You need to have some design knowledge to adjust things manually, you need to ask for exactly what you want, in the technical jargon it expects. You can't just say "make this more beautiful" or "animate this better" because it just creates slop - Right now vibe coded videos are promising, but I think I could have done this video faster just by doing it manually in After Effects. It really took that much work - If you have a creative idea for something you want to animate, it takes multiple hours of back and forth prompting to create just one or two seconds worth of **good** animation - Tip: PARAMETERIZE everything! It tends to hardcode magic numbers everywhere in the code, so if you change something earlier in the video timeline, everything else breaks. You want to essentially be creating "key frames" with code by telling it to parameterize every frame where something important happens, and calculate the rest of the keyframes based off that. This comes in handy when you need, for example, to adjust keyframes to match the music So in summary: vibe coded videos are promising, but right now it only works for very stock-looking videos unless you put in a ton of effort Maybe actually useful for 1-2 second web animations though, I'll try that next It will obviously get better, this feels like the quality of code generation in 2023-2024, you need to hold its hand and correct it at every step along the way. But even if video code generation was better, you would still need someone with motion design knowledge to at least set the creative direction, lay out the overall script and composition, etc. It's not completely hands-off unless you want slop And a word on caution: especially here on X, there's 90% hype and 10% reality, nothing is what it seems. Do not believe what you see online, people are constantly baiting and then just laughing it off in the comments

Javi

313,364 görüntüleme • 7 ay önce

Here's a devlog made by an anonymous Chinese fan replicating the surprisingly brand new technique that I developed for detecting asteroids which wound up being so powerful that it can easily track Stealth Fighters from over 100km away even when it’s only using three $30 webcams as sensors meaning it easily outperforms all modern stealth tracking techniques in precision, range and cost. And while this demo is using optical light, this same technique which I call pixel motion to voxel projection, can be used interchangeably with thermal infrared cameras to work at night and also majorly boosts the effectiveness of radar allowing you to track fighters much more effectively through clouds and over the horizon. This technique will also always eventually give the exact location of the target even if the image is blurry as those blurs will always average out from the different perspectives into revealing the precise location of the target in the voxel grid. There is definitely a Mandela effect with this technique as it feels as though it should already exist, especially because at first as it sounds like it is performing triangulation (which has existed for years and is what we do for mocap and tennis ball tracking). But triangulation is entirely separate to this as triangulations only works if you have already identified where the ball is in a 2D image because you’re able to rely on being able to use at least 2 separate high quality cameras which are much closer to the ball making the ball’s apparent size much much bigger and therefore gives you hundreds of pixels to work with which makes it much easier to use object recognition techniques to recognize where it is in the image aka in 2D and then you’re just using the other cameras view to project out lines which intersect in 3D to find out where the ball is in 3D. The major difference is that pixel motion to voxel projection allows you to find where the object is in 3D without having already found it in 2D which is an unbelievable difference as it allows you to use much lower quality cameras together to accumulate data together into 3D space. If this seem like it doesn’t mean much then what it actually means is that you don’t understand what I’m saying as what I’m saying means a LOT in practical terms as it means you go from having to use an imaging system that has to be able to image the object to the point that it is over a hundred total pixels in surface area to have enough data to recognize it to instead be able to use something that is only images the object to be 1 pixel in surface area and only changes the brightness value by 1 value every now and then. I’d recommend an amazing video by DST studios called “Lowlight cameras can’t defeat stealth” if you want a great video which goes over the difficulty of even using telescopes to recognize stealth fighters and why this is so impressive compared to other techniques and ironically it is what inspired me to realize the asteroid tracker I was working on actually could do this. Which brings me to the point that if this wasn’t a new technique then not only would there be at least one example of an asteroid survey that points distant telescopes at the same place at the same time in order to be able to add the light together to detect asteroids which as I was shocked to learn isn’t a thing despite the fact that it would make detecting asteroids trivial by comparison to modern 2D imaging while also having no impact on the normal scientific operations of those surveys other than small changes to scheduling. But there would also be an example of a drone tracker that uses this instead of using the aforementioned high quality zoomable telescope which has to be able to zoom in close enough to be able to recognize a drone. If you want to tell me that this is something that already exists give me an exact example of a product that uses it, not the general outline of a concept that you think it is, the actual product and then also tell me the asteroid survey that uses distant telescopes that point at the exact same place at the exact same time because I can guarantee that if you google what you think uses this you won’t even find the steps of subtracting the images from each other to get motion and will definitely not get the added step of projecting that motion into a voxel grid (It would blow your mind if you found out how Xbox kinect cameras work.) Also I want to make it clear, I’m not saying you should just use web cams to do this, I’m just using them as an example to show you the power of this in reality you would probably want to use 5 high quality zoomable thermal cameras which pan across the sky in sync with each other which due to using lower frequency are much less prone to the Rayleigh scattering that scatters visible light at 150 or so km away and again, you can also use this to majorly upgrade radar. Pretty much all of the problems you could think of for this are incredibly easy to overcome if you apply even a small amount of brainpower into fixing the problem. And yes, this gives you the exact location down to the meter of whatever you are tracking even if the image is blurry as those blurs will always average out to the exact location down to the meter in the voxel grid. Which is what makes this technique so powerful since the cost of adding each camera to The network grows linearly while the rate at which each camera gives more information grows exponentially due to the increasing unlikeliness of all of them having more movement in the same place. And given the size of the cameras it really wouldn’t be that hard to hide and network these cameras together in other countries and on sea buoys to know where planes are everywhere in the world. Which brings me to the point that I personally really don’t care about the military uses of this technology, if all it could do is precisely track stealth fighters then I wouldn’t have cared enough to work on it, I could have used any of the many other life saving techniques as the subject of the video, stealth fighters just sounds the most clickable and the scale of the problem is more intuitive to most people and if I did use any of those as subjects for the demo it would inevitably result in the stealth fighter technique being figured out anyway and all of the other uses are so useful that I don't think anyone would reasonably complain about the upside. The real purpose of this video is that since this is a new technique that hasn’t been used to detect stealth fighters despite the billions we have spent on that, then what else can you apply this to that could go on to improve billions of people’s lives that you or others are working on. For example this also allows you to majorly improve the effectiveness of cryo electron microscopy and CT scanners. This part also is kind of hard to explain as it also sounds like it exists but again, when you look through all of the places where you think it is being used you will find that it wasn’t. What I’m saying here isn’t that this is a Radon transform or gaussian splat or whatever, I’m saying that this is able to get new information that wasn’t being accessed before due to the added information about depth you get from the correlation of movement between each perspective which adds to the information that you already have. This allows you to directly subtract foreground and background objects as well as noise faster than you would be able to before and works better than super resolution for your images since super resolution won’t remove foreground and background objects like this does and instead just scales up target, foreground and background objects indiscriminately. And while with enough data Radon transforms or other scanning techniques would eventually get you a correct answer this will get you there a lot faster since those are mostly averaging techniques which average out noise whereas this gets you the ability to directly subtract noise. I’m not expecting you to think that this would do anything but if you try it for yourself you will find that it does majorly improve your ability to perform 3d scans. Again, cryo EM is a field where you would expect this technique to exist but when you look through all the papers on the topic there is no mention of tilting the grid slightly in order to be able to change your perspective slightly on the order of the feature size (if you tilt the grid then you only need precision on the order of an arc minute to do this) and doing multiple exposures from multiple different known tilts and then using those difference images to correlate depth from motion. In fact, in cryo EM you would normally want to do the opposite of this and have your exposures all taken from the same grid angle and just use the variations in how many of the same proteins are oriented in order to be able to scan them for a 3D model but this will generate you far more data faster. There is so much information that I can’t really explain in text so if you have any questions such as why this hasn’t been made before then they will most likely be answered in the video I originally posted which I have added to the end of the first Devlog for your convenience. And again, pretty much all of the problems with the technique can be fixed with a little bit of brainpower, in reality you would probably want to use 5 high quality zoomable thermal cameras which pan across the sky in sync with each other which due to using lower frequency are much less prone to the Rayleigh scattering that scatters visible light at 150 or so km away and again, you can also use this to majorly upgrade radar.

ConsistentlyInconsistent

50,799 görüntüleme • 1 yıl önce