Загрузка видео...

Не удалось загрузить видео

На главную

MotionGPT: Human Motion as a Foreign Language paper page: Though the advancement of pre-trained large language models unfolds, the exploration of building a unified model for language and other multi-modal data, such as motion, remains challenging and untouched so far. Fortunately, human motion displays a semantic coupling akin to...

125,319 просмотров • 3 лет назад •via X (Twitter)

Комментарии: 10

Фото профиля Ravi Kiran S
Ravi Kiran S3 лет назад

The ICME-2023 paper Action-GPT seems to be missing from Related Works and comparative evaluations.

Фото профиля Harry Pham
Harry Pham3 лет назад

This is amazing . Can wait until people start using it to turn spicy fanfic or novel into animation 😍.

Фото профиля imzhexu
imzhexu3 лет назад

@memdotai mem it

Фото профиля Mem
Mem3 лет назад

@_akhaliq Saved! Here's the compiled thread: 🪄 AI-generated summary: "This thread introduces MotionGPT, a paper that explores the challenge of building a unified model for language and motion data. It provides a link to the paper page for further...

Фото профиля Chris Chen
Chris Chen3 лет назад

We will try to release everything of MotionGPT( just like our previous work Motion-Latent-Diffusion( (CVPR 2023). You can also check that work if it helps. MotionGPT is a unified and user-friendly motion-language model using LLM.

Фото профиля AI Tools for 100x Growth
AI Tools for 100x Growth3 лет назад

Thanks for sharing this interesting research! It's exciting to see how AI is being applied to other types of data beyond just language, such as human motion. I'm curious to see how this work will progress in the future.

Фото профиля Marketing with HMA
Marketing with HMA3 лет назад

This is great! Can play a big role in motion graphics.

Фото профиля Lucas Fonseca
Lucas Fonseca3 лет назад

@memdotai mem it

Фото профиля Mem
Mem3 лет назад

@_akhaliq Saved! Here's the compiled thread: 🪄 AI-generated summary: "This thread introduces MotionGPT, a paper that explores the challenge of building a unified model for language and motion data. It provides a link to the paper page for further...

Фото профиля Rob
Rob3 лет назад

could use for monitor system for seniors or babies. Like fall detection or dangerous behavior warning

Похожие видео

Multi-Track Timeline Control for Text-Driven 3D Human Motion Generation paper page: Recent advances in generative modeling have led to promising progress on synthesizing 3D human motion from text, with methods that can generate character animations from short prompts and specified durations. However, using a single text prompt as input lacks the fine-grained control needed by animators, such as composing multiple actions and defining precise durations for parts of the motion. To address this, we introduce the new problem of timeline control for text-driven motion synthesis, which provides an intuitive, yet fine-grained, input interface for users. Instead of a single prompt, users can specify a multi-track timeline of multiple prompts organized in temporal intervals that may overlap. This enables specifying the exact timings of each action and composing multiple actions in sequence or at overlapping intervals. To generate composite animations from a multi-track timeline, we propose a new test-time denoising method. This method can be integrated with any pre-trained motion diffusion model to synthesize realistic motions that accurately reflect the timeline. At every step of denoising, our method processes each timeline interval (text prompt) individually, subsequently aggregating the predictions with consideration for the specific body parts engaged in each action. Experimental comparisons and ablations validate that our method produces realistic motions that respect the semantics and timing of given text prompts.

AK

126,585 просмотров • 2 лет назад

Seamless Human Motion Composition with Blended Positional Encodings Conditional human motion generation is an important topic with many applications in virtual reality, gaming, and robotics. While prior works have focused on generating motion guided by text, music, or scenes, these typically result in isolated motions confined to short durations. Instead, we address the generation of long, continuous sequences guided by a series of varying textual descriptions. In this context, we introduce FlowMDM, the first diffusion-based model that generates seamless Human Motion Compositions (HMC) without any postprocessing or redundant denoising steps. For this, we introduce the Blended Positional Encodings, a technique that leverages both absolute and relative positional encodings in the denoising chain. More specifically, global motion coherence is recovered at the absolute stage, whereas smooth and realistic transitions are built at the relative stage. As a result, we achieve state-of-the-art results in terms of accuracy, realism, and smoothness on the Babel and HumanML3D datasets. FlowMDM excels when trained with only a single description per motion sequence thanks to its Pose-Centric Cross-ATtention, which makes it robust against varying text descriptions at inference time. Finally, to address the limitations of existing HMC metrics, we propose two new metrics: the Peak Jerk and the Area Under the Jerk, to detect abrupt transitions.

AK

30,864 просмотров • 2 лет назад