Loading video...

Video Failed to Load

Go Home

High quality AI generated talking heads are coming! GAIA can generate talking avatars from a single portrait image and speech clip. It even supports text prompts like `sad`, `open mouth` or `surprise` to guide video generation. Crazy times ahead 🤯

660,226 views • 2 years ago •via X (Twitter)

10 Comments

Dreaming Tulpa 🥓👑's profile picture
Dreaming Tulpa 🥓👑2 years ago

If you like Tweets like this, you might enjoy my weekly newsletter, #aiartweekly. A free, once–weekly e-mail round-up of the latest AI art news, interviews with artists and useful tools & resources. Join 2800+ subscribers here:

Dmytro's profile picture
Dmytro2 years ago

Finally no need to attend the teams/zoom calls 😅 1. Jot two lines 2. Unwrap by gpt 3. Paste into gaia Go for a walk

Dreaming Tulpa 🥓👑's profile picture
Dreaming Tulpa 🥓👑2 years ago

I like your way of thinking 😂

CryptoSilvia's profile picture
CryptoSilvia2 years ago

This is insane… I doubt the majority of people would think this was fake/ai unless they were told.

Dreaming Tulpa 🥓👑's profile picture
Dreaming Tulpa 🥓👑2 years ago

Definitely

CoffeeVectors's profile picture
CoffeeVectors2 years ago

It’s so fascinating how the line between image and animation is getting blurrier and blurrier. Not to mention the boundary between aesthetics.

OxxBigDaddyxxO's profile picture
OxxBigDaddyxxO2 years ago

AI can be trained by it watching videos of you and skimming your social media posts to come up with a dialog that you would have. Pair that with this tech, and you don't ever have to repond to or chat with anyone, have AI do it.

Dreaming Tulpa 🥓👑's profile picture
Dreaming Tulpa 🥓👑2 years ago

No more Zoom calls 👌

Nk's profile picture
Nk2 years ago

Incels will use it for the most terrible things. I hope if anyone commits suicide because of this trash the victim's relatives will fucking SUE you.

Atlas3D's profile picture
Atlas3D2 years ago

this should be recreatable with better openpose/canny controlnet integrations directly into SVD and other generative video pipelinse

Related Videos

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 views • 1 year ago

🎥 Today we’re premiering Meta Movie Gen: the most advanced media foundation models to-date. Developed by AI research teams at Meta, Movie Gen delivers state-of-the-art results across a range of capabilities. We’re excited for the potential of this line of research to usher in entirely new possibilities for casual creators and creative professionals alike. More details and examples of what Movie Gen can do ➡️ 🛠️ Movie Gen models and capabilities Movie Gen Video: 30B parameter transformer model that can generate high-quality and high-definition images and videos from a single text prompt. Movie Gen Audio: A 13B parameter transformer model that can take a video input along with optional text prompts for controllability to generate high-fidelity audio synced to the video. It can generate ambient sound, instrumental background music and foley sound — delivering state-of-the-art results in audio quality, video-to-audio alignment and text-to-audio alignment. Precise video editing: Using a generated or existing video and accompanying text instructions as an input it can perform localized edits such as adding, removing or replacing elements — or global changes like background or style changes. Personalized videos: Using an image of a person and a text prompt, the model can generate a video with state-of-the-art results on character preservation and natural movement in video. We’re continuing to work closely with creative professionals from across the field to integrate their feedback as we work towards a potential release. We look forward to sharing more on this work and the creative possibilities it will enable in the future.

AI at Meta

2,267,059 views • 2 years ago

We’re excited to announce the release and open-source of HunyuanImage 3.0 — the largest and most powerful open-source text-to-image model to date, with over 80 billion total parameters, of which 13 billion are activated per token during inference.The effect is completely comparable to the industry’s flagship closed-source model.🚀🚀🚀 HunyuanImage 3.0 originates from our internally developed native multimodal large language model, with fine-tuning and post-training focused on text-to-image generation. This unique foundation gives the model a powerful set of capabilities: ✅Reason with world knowledge ✅Understand complex, thousand-word prompts ✅Generate precise text within images Different from traditional DiT architecture image generation models, HunyuanImage 3.0’s MoE architecture uses a Transfusion-based approach to deeply couple Diffusion and LLM training for a single, powerful system. Built on Hunyuan-A13B, HunyuanImage 3.0 was trained on a massive dataset: 5 billion image-text pairs, video frames, interleaved image-text data, and 6 trillion tokens of text corpora. This hybrid training across multimodal generation, understanding, and LLM capabilities allows the model to seamlessly integrate multiple tasks. Whether you're an illustrator, designer, or creator, this is built to slash your workflow from hours to minutes. HunyuanImage 3.0 can generate intricate text, detailed comics, expressive emojis, and lively, engaging illustrations for educational content. The current release focuses solely on text-to-image generation and future updates will include image-to-image, image editing, multi-turn interaction, and more. 👉🏻Try it now: 🔗GitHub: 🤗Hugging Face:

Tencent Hy

413,175 views • 1 year ago