正在加载视频...

视频加载失败

We are just scratching the surface of precise control over AI video generation. MotionStream unlocks real-time video with interactive motion controls. You can interactively generate video based on motion inputs (like drawn trajectories, camera movements, or motion transfer). 29fps generation w/ 0.4 second latency on a single H100, oh...

46,079 次观看 • 10 个月前 •via X (Twitter)

0 条评论

暂无评论

原始帖子的评论将显示在这里

相关视频

InstantDrag Improving Interactivity in Drag-based Image Editing discuss: Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to the challenge of accurately reflecting user interaction while maintaining image content. Some existing approaches rely on computationally intensive per-image optimization or intricate guidance-based methods, requiring additional inputs such as masks for movable regions and text prompts, thereby compromising the interactivity of the editing process. We introduce InstantDrag, an optimization-free pipeline that enhances interactivity and speed, requiring only an image and a drag instruction as input. InstantDrag consists of two carefully designed networks: a drag-conditioned optical flow generator (FlowGen) and an optical flow-conditioned diffusion model (FlowDiffusion). InstantDrag learns motion dynamics for drag-based image editing in real-world video datasets by decomposing the task into motion generation and motion-conditioned image generation. We demonstrate InstantDrag's capability to perform fast, photo-realistic edits without masks or text prompts through experiments on facial video datasets and general scenes. These results highlight the efficiency of our approach in handling drag-based image editing, making it a promising solution for interactive, real-time applications.

AK

71,232 次观看 • 2 年前

Multi-Track Timeline Control for Text-Driven 3D Human Motion Generation paper page: Recent advances in generative modeling have led to promising progress on synthesizing 3D human motion from text, with methods that can generate character animations from short prompts and specified durations. However, using a single text prompt as input lacks the fine-grained control needed by animators, such as composing multiple actions and defining precise durations for parts of the motion. To address this, we introduce the new problem of timeline control for text-driven motion synthesis, which provides an intuitive, yet fine-grained, input interface for users. Instead of a single prompt, users can specify a multi-track timeline of multiple prompts organized in temporal intervals that may overlap. This enables specifying the exact timings of each action and composing multiple actions in sequence or at overlapping intervals. To generate composite animations from a multi-track timeline, we propose a new test-time denoising method. This method can be integrated with any pre-trained motion diffusion model to synthesize realistic motions that accurately reflect the timeline. At every step of denoising, our method processes each timeline interval (text prompt) individually, subsequently aggregating the predictions with consideration for the specific body parts engaged in each action. Experimental comparisons and ablations validate that our method produces realistic motions that respect the semantics and timing of given text prompts.

AK

126,630 次观看 • 2 年前

Okay, I think I found a much better way to deal with one of the most annoying parts of making AI videos. The tools are just too fragmented. One for writing. One for generating footage. One for voiceover. Another for music. Then you still have to bring everything into an editor and put the whole thing together. And there’s another problem with traditional AI video tools. The video looks great, except for one tiny thing. A wrong number. A typo. One bad shot. A graphic you want to change. And suddenly you’re looking at another full generation. I tested Fotor Agent, and this is where it gets interesting. Instead of stopping at video generation, it can take care of the production process itself. Give it a prompt, a brief, a PDF, a report, a chart, or even a reference video. It can figure out the structure, generate the footage, voiceover, music, captions, motion graphics and transitions, then arrange everything into a complete multi-track project. And you’re not stuck with whatever it generated. Once the video is ready, I can still go in and make changes. Move a clip. Change the timing. Rewrite some text. Update a number. Change a color. Resize a graphic. Swap an asset. Or select a specific area and tell the Agent what needs to be changed. That’s a pretty different experience from treating AI video as a finished file. The Motion Graphics are another part I really like. Charts, data, numbers, text and logos can be turned into native 4K motion graphics, while the individual elements stay editable. So things that normally mean opening After Effects or handing the work off to a motion designer can be handled inside the same project. Motion graphics can cost as little as $0.06/sec, or about $1.80 for 30 seconds, while significantly reducing both production time and overall cost compared with traditional workflows. For me, the interesting part isn't simply that Fotor Agent can generate a video. It’s that I can give it the job of building the first complete version, then spend my time on the parts that actually need my input. That feels a lot closer to having an AI production partner than another AI clip generator. Check out the video it made:

Leonardo

94,441 次观看 • 3 天前