Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

NeRF-Insert: Local 3D editing with multimodal control signals by @benoriol et al. Project: 🖼️🚀 Objects are inserted in NeRFs using 2D inpainting models, with editing constrained to a specific region in space 🖍️ User-given coarse masks define the insert region 🎛️ Meshes or visual references are incorporated for variable...

11,384 görüntüleme • 2 yıl önce •via X (Twitter)

1 Yorum

MrNeRF profil fotoğrafı
MrNeRF2 yıl önce

2 | 2

Benzer Videolar

Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields paper page: Editing a local region or a specific object in a 3D scene represented by a NeRF is challenging, mainly due to the implicit nature of the scene representation. Consistently blending a new realistic object into the scene adds an additional level of difficulty. We present Blended-NeRF, a robust and flexible framework for editing a specific region of interest in an existing NeRF scene, based on text prompts or image patches, along with a 3D ROI box. Our method leverages a pretrained language-image model to steer the synthesis towards a user-provided text prompt or image patch, along with a 3D MLP model initialized on an existing NeRF scene to generate the object and blend it into a specified region in the original scene. We allow local editing by localizing a 3D ROI box in the input scene, and seamlessly blend the content synthesized inside the ROI with the existing scene using a novel volumetric blending technique. To obtain natural looking and view-consistent results, we leverage existing and new geometric priors and 3D augmentations for improving the visual fidelity of the final result. We test our framework both qualitatively and quantitatively on a variety of real 3D scenes and text prompts, demonstrating realistic multi-view consistent results with much flexibility and diversity compared to the baselines. Finally, we show the applicability of our framework for several 3D editing applications, including adding new objects to a scene, removing/replacing/altering existing objects, and texture conversion.

AK

62,768 görüntüleme • 3 yıl önce

Alibaba presents MIMO Controllable Character Video Synthesis with Spatial Decomposed Modeling Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case training, which severely limits their applicability of modeling arbitrary characters in a short time. Recent 2D methods break this limitation via pre-trained diffusion models, but they struggle for pose generality and scene interaction. To this end, we propose MIMO, a novel framework which can not only synthesize character videos with controllable attributes (i.e., character, motion and scene) provided by simple user inputs, but also simultaneously achieve advanced scalability to arbitrary characters, generality to novel 3D motions, and applicability to interactive real-world scenes in a unified framework. The core idea is to encode the 2D video to compact spatial codes, considering the inherent 3D nature of video occurrence. Concretely, we lift the 2D frame pixels into 3D using monocular depth estimators, and decompose the video clip to three spatial components (i.e., main human, underlying scene, and floating occlusion) in hierarchical layers based on the 3D depth. These components are further encoded to canonical identity code, structured motion code and full scene code, which are utilized as control signals of synthesis process. The design of spatial decomposed modeling enables flexible user control, complex motion expression, as well as 3D-aware synthesis for scene interactions. Experimental results demonstrate effectiveness and robustness of the proposed method.

AK

148,998 görüntüleme • 1 yıl önce

🔴 Finally! NVIDIA has finally made the code for Neuralangelo public! It has the ability to transform any video into a highly detailed 3D environment, and it's a technology related to but DIFFERENT from NeRF. 💡 Here's how it works: It takes a 2D video as input, showing an object, monument, building, landscape, etc., from various perspectives and analyzes details such as depth, size, and the shapes of objects. From this, the AI sketches an initial 3D model, similar to how an artist molds a figure. This representation is then refined to highlight more details, just as an artist would make the final touches when sculpting. The result is a 3D environment/model, perfect for use in any environment. Imagine the applications it will have for video games, cinema, virtual environments, VR, and more! 📽️🎮 💡 More details: A year ago, an article was presented on a groundbreaking technique called NVIDIA's Instant NeRF. This technique turns images into stunning 3D scenes in a short time, ideal for creating realistic models for video games and other applications. Although Instant NeRF had a lot of potential, the generated models were not perfect and often lacked detailed structures, appearing somewhat cartoonish. A year on, NVIDIA releases a new technique based on Instant NeRF, named Neuralangelo. This enhances the fidelity of surface structures. While NeRF reconstructs real objects in virtual environments from images or videos, Instant NeRF speeds up this process, and Neuralangelo further improves the quality, making the generated objects appear even more realistic when examined up close. Neuralangelo improves Instant NeRF's approach in two key ways related to the hash grid encoding technique: 1⃣ Numerical gradients have been used to compute higher-order derivatives as a smoothing operation. This optimizes the "hash grid" encoding using numerical rather than analytical gradients, providing a smoother input to the network that produces the 3D model. 2⃣ A "coarse-to-fine" optimization has been implemented in the hash grids to control different levels of detail. That is, they first focus on a smoothed version of the scene, and then refine it with more detailed updates. Well, as Arthur C. Clarke said, "Any sufficiently advanced technology is indistinguishable from magic."

Javi Lopez ⛩️

689,169 görüntüleme • 2 yıl önce

The war in #Sudan 🇸🇩 is still going on, hundred thousands of soldiers are fighting on each side for the control of the country. The fightings are now occuring in the central part of the country, in the Kordofan region. On one side, the Rapid Support Forces, sponsored by the UAE are trying to consolidate their parralel government in Nyala with key victories. The RSF occupy western Sudan, especially the Darfour region. They are supported by the UAE, Chad, South Sudan, Libya, CAR and Wagner/AC 🇦🇪🇹🇩🇸🇸🇱🇾🇨🇫🇷🇺. Their leader is Mohammed Hamdan Dagolo 'Hemetti'. They are Bagarra arabs, opposing black communities in the west and the Nile river tribes controlling the government. They managed to occupy En Nahud, a key road hub in west Kordofan region, they are now trying to take Babanusa, they failed for now to destroy the encircled garrison. RSF forces resist SAF progress in the Kordofan regions, especially around El Obeid, key supply road and supply hub for the army. The main battle is happening in El Fasher, capital of Darfour region. There, SAF and Joint forces troops are encircled since 2 years. The SAF broke the 223rd RSF offensive on the 1.5 million inhabitant city (including at least 500 000 refugees). The city faces famine due to the encirclement. The SAF coalition is composed of the officially recognized "government" led by general Abdel Fattah al Burhan. They represent Nile Arabs, but are also allied to avrious tribes, including African tribes of North Darfur. They are supported by Egypt, Iran, Turkiye, Russia and Qatar 🇪🇬🇮🇷🇹🇷🇷🇺🇶🇦. The SAF managed to liberate all the southern part of the country (Singa, Wad Madani), the capital city Khartoum and its neighbouring twins of Bahri and Omdourman. Now, they want to destroy the RSF by pushing them out of Kordofan before attacking their base in Darfur region. Massive fightings are ongoing around El Obeid and in the south Darfur region. The battle for Kordofan is ongoing, it will decide wether the army is able to liberate and reunify the country or if the RSF militia will create a de facto arab state in the west of Sudan...

Clément Molin

27,624 görüntüleme • 1 yıl önce

** Mega Parodius Update ( MegaDrive / Genesis ) ** 2 vid updates - In the first the Vic Viper players ship is heavily damaged .. only the bomb bay is operational ... so the viper goes on a bomb *blast* processing rampage !! ... This shows enemy destruction and some polishing of the bomb effect. The 2nd vid shows VRAM handling of objects with the tile viewer opened below the gameplay window. You can see objects load in , then as only the burning chests remain those animations are shifted live to maximise free object buffer. Progress since last progress vid : All collisions now implemented : Enemy to Player added , Bomb to Enemy added, Enemy Destruction added. Auto Defragging VRAM system implemented for Enemy Objects. Supports variable slot size also to maximise tile space . Fixed slot sizes wouldn't allow enough slots as they would need to be too large to cover all objects. Without defragging the VRAM would become like swiss cheese after a few screens of enemies also making larger slots difficult to allocate. Vram was getting tight due to things like the Blue Bell Bomb / White Bell text needing a reserved amount of vram ( as they can occur anywhere in the level via powerups ) . Theres also theres quite a bit of variety in Level 1 with warmup enemies , bottle enemies, chickens, crabs, orcas, chests, bees, penguin bathes, walking penguins, bouncing heads etc. We are also using all the arcade assets without trying to cut frames. It can't be loaded all at once so a dynamic system is needed. The enemy object buffer is 444 tiles presently for Level 1 , but the Catboss for example takes 308 when he's onscreen so things can get squeezed. For Mega versions of levels I wanted to ramp things up so that might mean we need more diversity in places than for the arcade accurate option , this system will help that also. The system can stream animation or utilise static animation if vram allows. Seems a mundane thing but sloppy vram management is the achiles heal of GFX diversity in a retro game . To help reduce the Bell effects VRAM buffer , Blue Bomb tiles required reduced to 120 tiles Max ( was 172 ) with no loss of quality . I've also reduced the sprite count by 6 sprites overall - down to 46 sprites used at max frame size. Lots more action / enemies to add yet but particularly the VRAM handling had to be improved before we can implement everything so glad thats in place now. Pyron and Vector Orbitex have been busy with upcoming GFX / Audion for the port , as always I'm very gratefull to have these Masters working with me on the port . #SGDK #Parodius #SegaMegadrive #SegaGenesis

Shannon Birt

13,389 görüntüleme • 6 ay önce

This BlenderFusion paper basically says "screw trying to describe 3D edits through text" and just... use Blender :-) The idea is pretty straightforward -- instead of trying to cram 3D understanding into a diffusion model, use depth estimation & segmentation to project 2D images into 2.5D meshes, edit them in actual 3D software, then use a fine-tuned diffusion model to make the results photorealistic again. The clever bit is their "dual-stream architecture" -- the model sees both the original scene AND the edited Blender render in parallel, learning to preserve what matters while fixing the inevitable artifacts from transforming imperfect 2.5D/3D reconstructions. They train it with smart masking strategies so it learns when to ignore the original scene (for removals/replacements) and can manipulate objects independently of camera motion. What you get is pretty impressive control -- not just moving objects around, but changing materials, deforming shapes, swapping backgrounds, all while maintaining visual coherence. Neural Assets (one of my favorite papers last year) tried to crack this with learned object tokens, but it struggled with overlapping objects and loses fine details (due to low res DINO encodings). BlenderFusion just sidesteps the whole problem -- want to rotate something 173.5 degrees? Just rotate it in Blender. Want to duplicate an object 8 times? Copy paste away. The diffusion model's only job is making it look photorealistic, not figuring out the 3D underpinnings. The catch? Lacks temporal consistency for animation. Each viewpoint is generated independently, so while a single edit looks great, smoothly animating a car or camera down the street won't work -- you'd get flickering and inconsistencies between frames. That said, this approach is so much more intuitive for finer grain image editing than trying to describe your changes in text prompts. It's the kind of thing that makes you wonder why we're trying to do everything inside neural networks when perfectly good 3D tools already exist -- giving you the best of both worlds.

Bilawal Sidhu

34,440 görüntüleme • 1 yıl önce

🚀 Announcing Echo — our new frontier model for 3D world generation. Echo turns a simple text prompt or image into a fully explorable, 3D-consistent world. Instead of disconnected views, the result is a single, coherent spatial representation you can move through freely. This is part of a bigger shift in AI: from generating pixels and tokens to generating spaces. Echo predicts a geometry-grounded 3D scene at metric scale, meaning every novel view, depth map, and interaction comes from the same underlying world — not independent hallucinations. Once generated, the world is interactive in real time. You control the camera, explore from any angle, and render instantly — even on low-end hardware, directly in the browser. High-quality 3D world exploration is no longer gated by expensive equipment. Under the hood, Echo infers a physically grounded 3D representation and converts it into a renderable format. For our web demo, we use 3D Gaussian Splatting (3DGS) for fast, GPU-friendly rendering — but the representation itself is flexible and can be easily adapted. Why this matters: consistent 3D worlds unlock real workflows — digital twins, 3D design, game environments, robotics simulation, and more. From a single photo or a line of text, Echo builds worlds that are reliable, editable, and spatially faithful. Echo also enables scene editing and restyling. Change materials, remove or add objects, explore design variations — all while preserving global 3D consistency. Editing no longer breaks the world. This is only the beginning. Echo is the foundation for future world models with dynamics, physical reasoning, and richer interaction — environments that don’t just look right, but behave right. Explore the generated worlds on our website and sign up for the closed beta. The era of spatial intelligence starts here. 🌍 #Echo #WorldModels #SpatialAI #3DFoundationModels Check it out:

SpAItial AI

176,105 görüntüleme • 7 ay önce

We benchmarked leading multimodal foundation models (GPT-4o, Claude 3.5 Sonnet, Gemini, Llama, etc.) on standard computer vision tasks—from segmentation to surface normal estimation—using standard datasets like COCO and ImageNet. These models have made remarkable progress; however, it is unclear exactly where they stand in terms of understanding vision in detail. Especially when it comes to tasks beyond question-answering. How well do they understand an object's segments or geometry? Our analyses yield an assessment that is quantitatively and qualitatively detailed and is compatible with evaluations developed in the field of computer vision over the past decades. Observed trends: 🔹 The foundation models consistently underperform task-specific SOTA models across all tasks. However, they are respectable generalists, which is remarkable as they are presumably trained primarily on image-text-based tasks. 🔹 They perform semantic tasks notably better than geometric ones. 🔹 GPT-4o performs the best among non-reasoning models, getting the top position in 4 out of 6 tasks. 🔹 Reasoning models, e.g., o3, show improvements in geometric tasks. 🔹 The 'image generation' models, e.g., GPT-40 Image Generation, which have been natively trained multimodally, exhibit quirks. E.g., hallucinated objects, misalignment between the input and output, etc. 🔹 While the prompting techniques affect performance, better models exhibit less sensitivity to variations in prompts. We control for the variance introduced by the prompting methods in our experiments. 🌐 Detailed analyses, visualizations: ⌨️ code: 🧵 1/n

Amir Zamir

73,074 görüntüleme • 1 yıl önce

The Dark Secrets Of Johnson & Johnson...One Of The World’s Most Admired Corporations Is Also Its Deadliest Criminal Enterprise. Gardiner Harris J & J Targeted Nursing Homes To Use Antipsychotic Medications As 'Chemical Restraints' While Knowing Full Well It Would Kill Millions. 1 in 5 nursing home residents are prescribed an antipsychotic medication as 'off label' for use as a 'chemical restraint.' Keeping nursing home residents quiet benefits the facility. Antipsychotics like Johnson & Johnson's Risperdal are given to patients with dementia, often without informed consent, to control & sedate behavior. Risperdal is an antipsychotic medication used in the treatment of bipolar disorder & schizophrenia that Johnson & Johnson claimed had fewer side effects than its older antipsychotic, Haldol. Johnson & Johnson illegally marketed the use of Risperdal to elderly patients with dementia. Risperdal led to strokes, pneumonia, heart attacks & death. Risperdal & many other antipsychotic medications should never be used on patients with dementia. These antipsychotics are for specific diagnostics only...Schizophrenia & Bipolar. The package insert of Risperdal displays the harshest warning known as 'the black box warning' which is the most severe warning on a drug label: RISPERDAL® is not approved for use in patients with dementia-related psychosis. WARNING: INCREASED MORTALITY IN ELDERLY PATIENTS WITH DEMENTIA RELATED PSYCHOSIS Elderly patients with dementia-related psychosis treated with antipsychotic drugs are at an increased risk of death. Analyses of 17 placebo-controlled trials (modal duration of 10 weeks), largely in patients taking atypical antipsychotic drugs, revealed a risk of death in drug-treated patients of between 1.6 to 1.7 times the risk of death in placebo-treated patients. The deaths were cardiovascular (e.g., heart failure, stroke, sudden death) or infectious (e.g., pneumonia) in nature. Yet, despite the black box warnings added in 2005, nursing home residents are still given antipsychotics everyday as part of their polypharmacy 'medication cocktail' to keep them sedated & compliant. Drug manufacturers need to be held accountable for marketing drugs for uses that are known to be dangerous. Nursing facility staff need training & sufficient staffing to manage behaviors without resorting to drugging patients. 👇Risperdal Package Insert👇 👇No More Tears: The Dark Secrets Of J & J👇 👇The History Of Risperdal Warnings👇 Speaker: Gardiner Harris Podcast: Dr. Josh Axe

Valerie Anne Smith

40,457 görüntüleme • 1 yıl önce

The CIA and DARPA have been heavily involved in the research and development of mind control technology and has been a subject of interest, controversy, and secrecy since the mid-20th century. The CIA's MKULTRA program, which began in the 1950s, is perhaps the most well-known initiative in this field. This program involved extensive research into behavioral modification, including the use of drugs, hypnosis, and other methods to influence human behavior. Although primarily focused on chemical and biological agents, it laid the groundwork for later technological explorations. Project Pandora was funded by the CIA. This project in the 1960s looked into the effects of microwave beams on the brain, aiming to use such technologies for behavior and mood manipulation. This was part of early research into what would later be considered "non-lethal" weapons. Radio Frequency Energy and Microwave Technology has been studied and researched into how radio frequency energy could interact with the human brain has been documented. Projects like Pandora aimed to understand how microwaves could transmit signals to influence behavior or induce specific emotional states. Voice to Skull (V2K), also known as microwave auditory effect, involves sending sound directly into someone's head without the use of speakers. There have been claims and some research suggesting its use in psychological operations or in influencing behavior, though much of this remains speculative or classified. Heterodyning and Electromagnetic Technologies are methods involve the modulation of electromagnetic waves to interact with the brain's electrical activity. Research into these areas has been aimed at both therapeutic uses and potentially more invasive applications like mind control. Psychotronic technology is often associated with the concept of using electromagnetic fields or radiation to affect mental processes, psychotronic research has been noted in various contexts, including in Russian studies on psychotronic warfare. This term, however, sometimes borders on the speculative or pseudoscientific, with limited verifiable research in mainstream academic settings. Harvard and Yale have been involved in various psychological and neuroscience research projects, direct public evidence linking them to specific mind control research funded by CIA or DARPA is less clear. However, both universities have extensive research in neuropsychology and cognitive sciences, which could indirectly contribute to understanding how such technologies might work. Universities like Stanford, MIT, and Carnegie Mellon have engaged in research that touches on brain-computer interfaces, neurostimulation, and cognitive enhancement, areas that could theoretically overlap with mind control technologies. For instance, DARPA has funded research at these institutions for projects related to brain-computer interfaces, though the applications are often framed for medical or enhancing human performance rather than control. The military interest in these technologies often centers on weaponization, psychological operations, or enhancing soldier performance through cognitive augmentation. The use for inducing particular behaviors or emotions, especially in scenarios like false flag operations, could definitely be a possibility, but something that will never be admitted. This consists of the manipulation of individuals or groups to perform acts that benefit a hidden agenda unwillingly or unknowingly. Mind control technology, especially agencies like the CIA and DARPA, are shrouded in secrecy, with much of the research declassified or discussed in the public domain only after significant time has passed or in very general terms. While there's a clear interest in using technology to influence people's emotions and inducing behaviors for operations or nefarious purposes, it remains largely speculative or cloaked in national security classification.

The SCIF

21,479 görüntüleme • 1 yıl önce

The Dark Evolution of Mind Control: From MKULTRA, DEWs, Professor Delgado's remote bull, to modern remote Neuro-Weapons. We are able to control minds in ways you never thought possible. In the 1960s, Yale neuroscientist Dr. José Delgado pioneered brain stimulation techniques that shocked the world. Using implanted electrodes called "stimoceivers," he could remotely control animal behavior via radio signals. In his most famous experiment in 1963, Delgado stepped into a Spanish bullring armed only with a remote control. As a raging bull charged, he pressed a button, stimulating the animal's caudate nucleus, a brain region linked to movement and aggression, forcing it to skid to a halt just feet away. Similar implants in monkeys allowed him to trigger emotions like rage, calm, or even social hierarchy shifts, where subordinate monkeys learned to "control" aggressive ones by flipping levers that pacified them. Delgado's work extended to humans, where he induced euphoria, anger, or involuntary movements by stimulating limbic system areas, hinting at a future where brains could be "programmed" like machines. Delgado himself noted the shift from electrodes to non-invasive methods, like low-power pulsing magnetic fields to alter monkey behavior without wires. This laid the groundwork for today's neuro-technologies, where intelligence agencies, big tech, and shadowy networks reportedly deploy remote tools for surveillance and manipulation, far beyond invasive implants. Fast-forward to now, declassified documents and patents suggest advancements in remote neural monitoring (RNM), voice-to-skull (V2K), and directed energy weapons (DEWs) enable real-time brain reading, emotion control, and behavioral influence using radio frequencies (RF), extremely low frequencies (ELF), and electromagnetic radiation. RNM purportedly tracks brain waves via satellite, decoding thoughts like a "brain fingerprint" for constant surveillance. V2K, based on the microwave auditory effect, beams voices directly into skulls, bypassing ears, potentially making targets believe they're hearing gods, demons, or commands. There are no coincidences when 90% of school shooters say they hear "demons" talking to them. "What if it was actually a person running a script on a selected target, to commit acts of violence.? DEWs, including microwaves and lasers, could induce pain, fatigue, or altered states without trace, as in Havana Syndrome cases affecting diplomats. These tools allegedly fuel "gang stalking" operations, where coordinated harassment via tech and human agents isolates targets, amplifying paranoia. Intelligence agencies like the CIA have historical ties to mind control (MKUltra), and big tech's brain-computer interfaces (BCIs) blur lines between therapy and control. Imagine manipulating a vulnerable individual, like a potential school shooter, by remotely inducing voices urging violence, then framing it as mental illness. What seems like inner demons could be an operator running a psy-op, using ELF waves to tweak emotions or RF to simulate auditory hallucinations. Or, targeting a large group of the population through specific frequencies through your own phone or tablet that cause emotional control when a specific political figure is displayed on the screen, literally manipulating your emotions or behavior without you ever knowing to mold your ideas or views... This ties into broader DEWs for mind control. High-power microwaves disrupt cognition, while advanced systems read/write thoughts in real time, scanning neural patterns to "decode" intentions or implant suggestions. The World Economic Forum (WEF) has spotlighted this, discussing "brain transparency" via wearables that track thoughts for productivity or safety, warning of a future where bosses monitor focus or AI decodes emotions. WEF sessions explore mind-reading tech, like translating thoughts to text or using AI for "full rich thoughts" transmission. AI has supercharged this field. Machine learning decodes brain signals with pinpoint accuracy, enabling BCIs like Neuralink to control cursors via thought alone. AI "co-pilots" infer intent from neural data, boosting noninvasive systems for rehab or augmentation. But in darker hands, AI could automate mass surveillance, predicting and preempting "deviant" thoughts, revolutionizing control from Delgado's crude remotes to seamless, invisible dominance. Delgado dreamed of a "psychocivilized society." Are we already there, hidden in plain sight? Pay attention because I wish I was joking about these capabilities and advancements in technology, but they're already using them on the population without you even knowing about it.

The SCIF

19,881 görüntüleme • 4 ay önce