Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

🤔🤔Tired of static, lifeless image edits? Not anymore! 🤗 🚀🚀We introduce MotionEdit, a framework supporting image editing that understands action, motion, interaction beyond static changes! 🤩🤩 🔗Full paper: ✨Project page:

59,481 Aufrufe • vor 8 Monaten •via X (Twitter)

0 Kommentare

Keine Kommentare verfügbar

Kommentare vom Original-Post werden hier angezeigt

Ähnliche Videos

InstantDrag Improving Interactivity in Drag-based Image Editing discuss: Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to the challenge of accurately reflecting user interaction while maintaining image content. Some existing approaches rely on computationally intensive per-image optimization or intricate guidance-based methods, requiring additional inputs such as masks for movable regions and text prompts, thereby compromising the interactivity of the editing process. We introduce InstantDrag, an optimization-free pipeline that enhances interactivity and speed, requiring only an image and a drag instruction as input. InstantDrag consists of two carefully designed networks: a drag-conditioned optical flow generator (FlowGen) and an optical flow-conditioned diffusion model (FlowDiffusion). InstantDrag learns motion dynamics for drag-based image editing in real-world video datasets by decomposing the task into motion generation and motion-conditioned image generation. We demonstrate InstantDrag's capability to perform fast, photo-realistic edits without masks or text prompts through experiments on facial video datasets and general scenes. These results highlight the efficiency of our approach in handling drag-based image editing, making it a promising solution for interactive, real-time applications.

AK

71,232 Aufrufe • vor 1 Jahr

VideoRF: Rendering Dynamic Radiance Fields as 2D Feature Video Streams paper page: Neural Radiance Fields (NeRFs) excel in photorealistically rendering static scenes. However, rendering dynamic, long-duration radiance fields on ubiquitous devices remains challenging, due to data storage and computational constraints. In this paper, we introduce VideoRF, the first approach to enable real-time streaming and rendering of dynamic radiance fields on mobile platforms. At the core is a serialized 2D feature image stream representing the 4D radiance field all in one. We introduce a tailored training scheme directly applied to this 2D domain to impose the temporal and spatial redundancy of the feature image stream. By leveraging the redundancy, we show that the feature image stream can be efficiently compressed by 2D video codecs, which allows us to exploit video hardware accelerators to achieve real-time decoding. On the other hand, based on the feature image stream, we propose a novel rendering pipeline for VideoRF, which has specialized space mappings to query radiance properties efficiently. Paired with a deferred shading model, VideoRF has the capability of real-time rendering on mobile devices thanks to its efficiency. We have developed a real-time interactive player that enables online streaming and rendering of dynamic scenes, offering a seamless and immersive free-viewpoint experience across a range of devices, from desktops to mobile phones.

AK

38,686 Aufrufe • vor 2 Jahren

CoDeF: Content Deformation Fields for Temporally Consistent Video Processing abs: paper page: present the content deformation field CoDeF as a new type of video representation, which consists of a canonical content field aggregating the static contents in the entire video and a temporal deformation field recording the transformations from the canonical image (i.e., rendered from the canonical content field) to each individual frame along the time axis.Given a target video, these two fields are jointly optimized to reconstruct it through a carefully tailored rendering pipeline.We advisedly introduce some regularizations into the optimization process, urging the canonical content field to inherit semantics (e.g., the object shape) from the video.With such a design, CoDeF naturally supports lifting image algorithms for video processing, in the sense that one can apply an image algorithm to the canonical image and effortlessly propagate the outcomes to the entire video with the aid of the temporal deformation field.We experimentally show that CoDeF is able to lift image-to-image translation to video-to-video translation and lift keypoint detection to keypoint tracking without any training.More importantly, thanks to our lifting strategy that deploys the algorithms on only one image, we achieve superior cross-frame consistency in processed videos compared to existing video-to-video translation approaches, and even manage to track non-rigid objects like water and smog.

AK

153,241 Aufrufe • vor 3 Jahren

Google presents Still-Moving Customized Video Generation without Customized Video Data Customizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its infancy, primarily due to the lack of customized video data. In this work, we introduce Still-Moving, a novel generic framework for customizing a text-to-video (T2V) model, without requiring any customized video data. The framework applies to the prominent T2V design where the video model is built over a text-to-image (T2I) model (e.g., via inflation). We assume access to a customized version of the T2I model, trained only on still image data (e.g., using DreamBooth or StyleDrop). Naively plugging in the weights of the customized T2I model into the T2V model often leads to significant artifacts or insufficient adherence to the customization data. To overcome this issue, we train lightweight Spatial Adapters that adjust the features produced by the injected T2I layers. Importantly, our adapters are trained on "frozen videos" (i.e., repeated images), constructed from image samples generated by the customized T2I model. This training is facilitated by a novel Motion Adapter module, which allows us to train on such static videos while preserving the motion prior of the video model. At test time, we remove the Motion Adapter modules and leave in only the trained Spatial Adapters. This restores the motion prior of the T2V model while adhering to the spatial prior of the customized T2I model. We demonstrate the effectiveness of our approach on diverse tasks including personalized, stylized, and conditional generation. In all evaluated scenarios, our method seamlessly integrates the spatial prior of the customized T2I model with a motion prior supplied by the T2V model.

AK

40,485 Aufrufe • vor 2 Jahren

Aigents! 🤖 Have you ever imagined a world where technology could not only see, but also understand and advise you? 🤔 Imagine playing a game of chess, strategizing each move, and then, with the help of AigentX, discovering the next winning move. This is not just a possibility anymore, it's reality! 🚀 Introducing AigentX's Image Processing feature! This isn't just about analysing and processing images; it's about bringing them to life with insights and guidance that we once thought could never become a reality and merely a figment of our imagination! 🤯 Picture this: You're deeply engrossed in a chess match, contemplating your next move. You snap a picture of the board and send it to AigentX. In moments, you're not just seeing the pieces – you're understanding the potential for your next move! 📸 AigentX doesn't just view the image; it interprets it, offering you useful insights and it isn't just chess! This feature can be taken straight to the real world with features like chart analysis and more! 📈 But that's just the beginning. AigentX is ready to transform how you interact with any image. From art analysis to real-world problem-solving, the possibilities are endless 📈 Check out our latest video below to see AigentX in action. This is more than an update; it's a leap into a future where your interaction with images is redefined 📻 Test out AigentX for yourself and witness how it's changing the game, one image at a time. Stay tuned for more – the journey with AigentX has just begun!

AGIX | $AGX

22,621 Aufrufe • vor 2 Jahren

One of the workflows I helped advance and lead for Destiny 2 was our static compositing pipeline for marketing visuals. When we were creating marketing content for #Destiny2, we always needed strong key visuals for weapons, characters, abilities, and other in-game elements. Unlike a traditional gameplay screenshot, where everything is baked into one static image, composited shots let us isolate individual elements and light them more precisely. That gave us a lot more flexibility to build scenes with better legibility, clearer focus, stronger lighting, improved contrast, and a more intentional overall look. One of the biggest strengths of compositing is having access to individual layers in Photoshop. That layered approach not only makes the editing process easier, but also makes the final artwork much more flexible. Images can be adapted for trailer motion graphics, reworked for print, or pushed further as creative needs evolve. Destiny 2 composites are built using Adobe Photoshop alongside in-game captures. Using advanced development tools, we light, capture, and isolate characters, weapons, and props directly from the game environment, all of which were originally created by the talented Bungie 3D art teams. Those assets are then brought into Photoshop, where they’re cut out and assembled into custom scenes using layered compositions and visual effects. The composites highlighted below are a few of the Destiny 2 pieces I’ve worked on that are meaningful to me. Let me know what your favorite is!

Biwald

67,791 Aufrufe • vor 1 Monat

-- What holds it together -- ✂️PAPER CUTS.- When the image is the end ⤵️ What you bring to life is the journey towards the image: how the pieces come to be as they are. The movement is the idea, and your image in Seedance is the final frame. ---------------------------------------------------- What happens to the object? 1⃣ By the time the object TRANFORMS into something else. It takes place in a continuous shot, a single camera that zooms in slowly and never cuts away. A cut would break the transformation; the hypnotic effect lies precisely in the fact that it doesn’t flicker. Template.- [@.RE IMAGE] is the LAST frame. [GLOBAL] [Technique + light + backdrop + focus]. [Real materials]. One continuous hypnotic transformation — a single take, no cuts. The same [OBJECT A] does not appear in pieces; it [VERB: melts / folds / blooms / dissolves] into [OBJECT B]. Slow surreal dreamlike drift, one unbroken slow push, slight stop-motion shimmer, no snapping. 0-1s: [state A, intact and recognizable]. 1-2.5s: [the change begins — "the same material begins to..."]. 2.5-3.5s: [the change at full — "...becomes..."]. 3.5-5s: [it settles]. The push eases to rest. Locked, exact match to the last frame. [LOGIC RULE] one continuous same-lens push, never cutting. [A] morphs into [B], [details that must not deform] stay legible, no warping. hypnotic drift. SFX: [a sound that also transforms]. no music, normal speed. 2⃣ When one thing LEADS to another, we need each step to be a link in the chain of causality. Template.- [@.RE IMAGE] is the LAST frame. [GLOBAL] [Technique + light + backdrop + focus]. [Real materials]. Hypnotic stop-motion paper cadence, slight frame-step, brisk causal montage, ~1.5s per shot, no naturalistic motion, no slow-mo. [cut] [shot + camera] / [carries on from previous cut] / [ACTION: what it DOES, not how it looks]. SFX: [beat-anchored hit]. [cut] ... (link 2 — chains from link 1) [cut] ... (link 3) [cut] ... (link 4) [cut] ... ease back to reveal / [link 5]. Locked, exact match to the last frame. SFX: ... [LOGIC RULE].- [materials], [what must NOT happen], no warping, [text legible if any]. ~1.5s per shot, no slow-mo. no music.

AlexandrIA

41,320 Aufrufe • vor 1 Monat