Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

🕹️We are excited to introduce "ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation" ChronoEdit reframes image editing as a video generation task to encourage temporal consistency. It leverages a temporal reasoning stage that denoises with “video reasoning tokens” to "reason" on physically plausible edits. See the attached...

37,032 görüntüleme • 11 ay önce •via X (Twitter)

6 Yorum

Huan Ling profil fotoğrafı
Huan Ling11 ay önce

This work is a great collaboration at @NVIDIAAI by @jayzhangjiewu @xuanchi13 @TianchangS @TianshiC @Kai__He @YifanLu17525599 @ruiyuan_gao @xieenze_jr @voidrank Jose M. Alvarez @JunGao33210520 @FidlerSanja @zianwang97 @HuanLing6 Models and code will be released in the coming week. Stay Tuned.

Huan Ling profil fotoğrafı
Huan Ling11 ay önce

1 - ChronoEdit repurposes pretrained video generative models for editing by reframing the task as a two-frame video generation problem, where the input image and its edited version are modeled as consecutive frames. When fine-tuned with curated image-editing data, this two-frame formulation equips the video model with editing functionalities while leveraging its pretrained temporal prior to preserve object fidelity.

Huan Ling profil fotoğrafı
Huan Ling11 ay önce

3 - Overview of the ChronoEdit pipeline: From right to left, the denoising process begins in the temporal reasoning stage, where the model imagines and denoises a short trajectory of intermediate frames. These intermediate frames act as reasoning tokens, guiding how the edit should unfold in a physically consistent manner. For efficiency, the reasoning tokens are discarded in the subsequent editing frame generation stage, where the target frame is further refined into the final edited image.

Huan Ling profil fotoğrafı
Huan Ling11 ay önce

2 - ChronoEdit introduces temporal reasoning toekns. If the video reasoning tokens are fully denoised into a clean video, the model can illustrate how it “thinks” by visualizing intermediate frames as a reasoning trajectory—though at the expense of slower inference. Notably, an emergent capability of our approach is its ability to generate reasoning trajectory videos to realize edits. Some interesting reasoning videos ( See more on project page).

Jason Li profil fotoğrafı
Jason Li11 ay önce

Most reasonable diffusion image reasoning.

IronRed | SandHive profil fotoğrafı
IronRed | SandHive11 ay önce

the vibes. what were some of the wildest failures y'all encountered?

Benzer Videolar

MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model with Gradio demo local demo: This paper studies the human image animation task, which aims to generate a video of a certain reference identity following a particular motion sequence. Existing animation works typically employ the frame-warping technique to animate the reference image towards the target motion. Despite achieving reasonable results, these approaches face challenges in maintaining temporal consistency throughout the animation due to the lack of temporal modeling and poor preservation of reference identity. In this work, we introduce MagicAnimate, a diffusion-based framework that aims at enhancing temporal consistency, preserving reference image faithfully, and improving animation fidelity. To achieve this, we first develop a video diffusion model to encode temporal information. Second, to maintain the appearance coherence across frames, we introduce a novel appearance encoder to retain the intricate details of the reference image. Leveraging these two innovations, we further employ a simple video fusion technique to encourage smooth transitions for long video animation. Empirical results demonstrate the superiority of our method over baseline approaches on two benchmarks. Notably, our approach outperforms the strongest baseline by over 38% in terms of video fidelity on the challenging TikTok dancing dataset. Code and model will be made available.

AK

810,731 görüntüleme • 2 yıl önce

InstantDrag Improving Interactivity in Drag-based Image Editing discuss: Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to the challenge of accurately reflecting user interaction while maintaining image content. Some existing approaches rely on computationally intensive per-image optimization or intricate guidance-based methods, requiring additional inputs such as masks for movable regions and text prompts, thereby compromising the interactivity of the editing process. We introduce InstantDrag, an optimization-free pipeline that enhances interactivity and speed, requiring only an image and a drag instruction as input. InstantDrag consists of two carefully designed networks: a drag-conditioned optical flow generator (FlowGen) and an optical flow-conditioned diffusion model (FlowDiffusion). InstantDrag learns motion dynamics for drag-based image editing in real-world video datasets by decomposing the task into motion generation and motion-conditioned image generation. We demonstrate InstantDrag's capability to perform fast, photo-realistic edits without masks or text prompts through experiments on facial video datasets and general scenes. These results highlight the efficiency of our approach in handling drag-based image editing, making it a promising solution for interactive, real-time applications.

AK

71,232 görüntüleme • 2 yıl önce

CoDeF: Content Deformation Fields for Temporally Consistent Video Processing abs: paper page: present the content deformation field CoDeF as a new type of video representation, which consists of a canonical content field aggregating the static contents in the entire video and a temporal deformation field recording the transformations from the canonical image (i.e., rendered from the canonical content field) to each individual frame along the time axis.Given a target video, these two fields are jointly optimized to reconstruct it through a carefully tailored rendering pipeline.We advisedly introduce some regularizations into the optimization process, urging the canonical content field to inherit semantics (e.g., the object shape) from the video.With such a design, CoDeF naturally supports lifting image algorithms for video processing, in the sense that one can apply an image algorithm to the canonical image and effortlessly propagate the outcomes to the entire video with the aid of the temporal deformation field.We experimentally show that CoDeF is able to lift image-to-image translation to video-to-video translation and lift keypoint detection to keypoint tracking without any training.More importantly, thanks to our lifting strategy that deploys the algorithms on only one image, we achieve superior cross-frame consistency in processed videos compared to existing video-to-video translation approaches, and even manage to track non-rigid objects like water and smog.

AK

153,305 görüntüleme • 3 yıl önce