Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

New Open-Source AI Animator: Text-to-Animation for Humans, Animals, Creatures, Robots & More UniMate generates motion from text prompts for rigged 3D characters. Describe an action and turn it into animation. Highlights: • One model, different skeletons—no separate retraining for each rig. • Generate transitions between existing keyframes. • Edit...

236,419 Aufrufe • vor 3 Tagen •via X (Twitter)

14 Kommentare

Profilbild von LieDisinfectant
LieDisinfectantvor 3 Tagen

This will change everything

Profilbild von Paul Sliusar
Paul Sliusarvor 3 Tagen

Tried it today with some twist creates amazing results only 1gb easy local runs

Profilbild von 🔞Logisfor
🔞Logisforvor 3 Tagen

Когда же наконец появится что-то подобное для 2D-анимации? Хотелось бы, чтобы это был инструмент для генерации концептов, нарезки, ригинга, работы с весами и ключами анимации, а также с вариациями инбитвинов. 🥲🥲🥲

Profilbild von Rodrigo Echagüe
Rodrigo Echagüevor 3 Tagen

Does it work?

Profilbild von Ozan Kaşıkçı
Ozan Kaşıkçıvor 3 Tagen

Beautiful!

Profilbild von Azad Nezamdoust
Azad Nezamdoustvor 3 Tagen

It is not usable. no weights are shared

Profilbild von joaopaulofurtado
joaopaulofurtadovor 3 Tagen

Cool, but how is this different than just using blender?

Profilbild von 4rnu Studio
4rnu Studiovor 3 Tagen

it is a addon for blender?

Profilbild von Big Bro Studios
Big Bro Studiosvor 3 Tagen

Are you freaking serious?? This is brutal! Thank you brooo!!

Profilbild von Prof. Johann Georg Faust
Prof. Johann Georg Faustvor 3 Tagen

Found this too. Did u tried it, stefan?

Profilbild von Aviv.eth🦍
Aviv.eth🦍vor 3 Tagen

@RidazLp2

Profilbild von Arashi Kageyama
Arashi Kageyamavor 3 Tagen

You've tested it and it works well? Gamechanger if so.

Profilbild von Namuken
Namukenvor 3 Tagen

Very nice

Profilbild von Krex
Krexvor 3 Tagen

i tried it and it failed every prompt even the most simple ones

Ähnliche Videos

Multi-Track Timeline Control for Text-Driven 3D Human Motion Generation paper page: Recent advances in generative modeling have led to promising progress on synthesizing 3D human motion from text, with methods that can generate character animations from short prompts and specified durations. However, using a single text prompt as input lacks the fine-grained control needed by animators, such as composing multiple actions and defining precise durations for parts of the motion. To address this, we introduce the new problem of timeline control for text-driven motion synthesis, which provides an intuitive, yet fine-grained, input interface for users. Instead of a single prompt, users can specify a multi-track timeline of multiple prompts organized in temporal intervals that may overlap. This enables specifying the exact timings of each action and composing multiple actions in sequence or at overlapping intervals. To generate composite animations from a multi-track timeline, we propose a new test-time denoising method. This method can be integrated with any pre-trained motion diffusion model to synthesize realistic motions that accurately reflect the timeline. At every step of denoising, our method processes each timeline interval (text prompt) individually, subsequently aggregating the predictions with consideration for the specific body parts engaged in each action. Experimental comparisons and ablations validate that our method produces realistic motions that respect the semantics and timing of given text prompts.

AK

126,635 Aufrufe • vor 2 Jahren

This is probably the most complex workflow I’ve ever built, only with open-source tools. It took my 4 days. It takes four inputs: author, title, and style; and generates a full visual animated story in one click in ComfyUI . I worked on it for four days. There are still some bugs, but here’s the first preview. Here’s a quick breakdown: - The four inputs are sent to LLMs with precise instructions to generate: first, prompts for images and image modifications; second, prompts for animations; third, prompts for generating music. - All voices are generated from the text and timed precisely, as they determine the length of each animation segment. - The first image and video are generated to serve as the title, but also as the guide for all other images created for the video. - Titles and subtitles are also added automatically in Comfy. - I also developed a lot of custom nodes for minor frame calculations, mostly to match audio and video. - The full system is a large loop that, for each line of text, generates an image and then a video from that image. The loop was the hardest part to build in this workflow, so it can process either a 20-second video or a 2-minute video with the same input. - There are multiple combinations of LLMs that try to understand the text in the best way to provide the best prompts for images and video. - The final video is assembled entirely within ComfyUI. - The music is generated based on the LLM output and matches the exact timing of the full animation. - Done! For reference, this workflow uses a lot of models and only works on an RTX 6000 Pro with plenty of RAM. My goal is not to replace humans, as I’ll try to explain later, this workflow is highly controlled and can be adapted or reworked at any point by real artists! My aim was to create a tool that can animate text in one go, allowing the AI some freedom while keeping a strict flow. I don’t know yet how I’ll share this workflow with people, I still need to polish it properly, but maybe through Patreon. Anyway, I hope you enjoy my research, and let’s always keep pushing further! :)

Lovis Odin

58,841 Aufrufe • vor 1 Jahr

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 Aufrufe • vor 1 Jahr

Blended-NeRF: Zero-Shot Object Generation and Blending in Existing Neural Radiance Fields paper page: Editing a local region or a specific object in a 3D scene represented by a NeRF is challenging, mainly due to the implicit nature of the scene representation. Consistently blending a new realistic object into the scene adds an additional level of difficulty. We present Blended-NeRF, a robust and flexible framework for editing a specific region of interest in an existing NeRF scene, based on text prompts or image patches, along with a 3D ROI box. Our method leverages a pretrained language-image model to steer the synthesis towards a user-provided text prompt or image patch, along with a 3D MLP model initialized on an existing NeRF scene to generate the object and blend it into a specified region in the original scene. We allow local editing by localizing a 3D ROI box in the input scene, and seamlessly blend the content synthesized inside the ROI with the existing scene using a novel volumetric blending technique. To obtain natural looking and view-consistent results, we leverage existing and new geometric priors and 3D augmentations for improving the visual fidelity of the final result. We test our framework both qualitatively and quantitatively on a variety of real 3D scenes and text prompts, demonstrating realistic multi-view consistent results with much flexibility and diversity compared to the baselines. Finally, we show the applicability of our framework for several 3D editing applications, including adding new objects to a scene, removing/replacing/altering existing objects, and texture conversion.

AK

62,768 Aufrufe • vor 3 Jahren

We’re excited to announce the release and open-source of HunyuanImage 3.0 — the largest and most powerful open-source text-to-image model to date, with over 80 billion total parameters, of which 13 billion are activated per token during inference.The effect is completely comparable to the industry’s flagship closed-source model.🚀🚀🚀 HunyuanImage 3.0 originates from our internally developed native multimodal large language model, with fine-tuning and post-training focused on text-to-image generation. This unique foundation gives the model a powerful set of capabilities: ✅Reason with world knowledge ✅Understand complex, thousand-word prompts ✅Generate precise text within images Different from traditional DiT architecture image generation models, HunyuanImage 3.0’s MoE architecture uses a Transfusion-based approach to deeply couple Diffusion and LLM training for a single, powerful system. Built on Hunyuan-A13B, HunyuanImage 3.0 was trained on a massive dataset: 5 billion image-text pairs, video frames, interleaved image-text data, and 6 trillion tokens of text corpora. This hybrid training across multimodal generation, understanding, and LLM capabilities allows the model to seamlessly integrate multiple tasks. Whether you're an illustrator, designer, or creator, this is built to slash your workflow from hours to minutes. HunyuanImage 3.0 can generate intricate text, detailed comics, expressive emojis, and lively, engaging illustrations for educational content. The current release focuses solely on text-to-image generation and future updates will include image-to-image, image editing, multi-turn interaction, and more. 👉🏻Try it now: 🔗GitHub: 🤗Hugging Face:

Tencent Hy

413,175 Aufrufe • vor 1 Jahr

35 WEBSITES GOOGLE DOESN'T WANT YOU TO KNOW 1. Explee .com — sends cold emails on autopilot 2. NoteGPT — turns docs into podcasts 3. Napkin AI — turns text into diagrams 4. Ideogram — generates text in images perfectly 5. Suno — makes full songs from a prompt 6. HeyGen — clones your face into videos 7. Kling AI — best AI video generation 8. ElevenLabs — clone any voice instantly 9. Gamma — AI presentations in seconds 10. Perplexity — AI search with real sources 11. Pika — animate any image into video 12. Runway — cinematic AI video generation 13. Cursor — AI code editor that builds for you 14. v0 — generate UI components with AI 15. Lovable — turn ideas into working apps 16. Descript — edit video by editing text 17. Opus Clip — auto cut long videos into shorts 18. Krea AI — real time AI image generation 19. Magnific — upscale any image with AI 20. Viggle — make characters move realistically 21. tl;dv — record and summarize any meeting 22. Fireflies — AI meeting notes automatically 23. Castmagic — turn audio into content pieces 24. Replit — code and deploy from browser 25. Leonardo AI — generate images for free 26. Synthesia — AI avatar videos no camera needed 27. Fliki — turn text into videos with AI 28. Photoroom — AI product photography 29. Invideo AI — turn prompts into full videos 30. Consensus — search what science agrees on 31. SciSpace — understand any research paper 32. Tome — AI builds your pitch decks 33. Beautiful AI — smart presentation design 34. Meshy — turn text into 3D models 35. Vizcom — turn sketches into renders The AI revolution isn't coming. It already happened and you missed half of it.

Jawad Rahman

44,004 Aufrufe • vor 3 Tagen

35 WEBSITES GOOGLE DOESN'T WANT YOU TO KNOW 1. Explee .com — sends cold emails on autopilot 2. NoteGPT — turns docs into podcasts 3. Napkin AI — turns text into diagrams 4. Ideogram — generates text in images perfectly 5. Suno — makes full songs from a prompt 6. HeyGen — clones your face into videos 7. Kling AI — best AI video generation 8. ElevenLabs — clone any voice instantly 9. Gamma — AI presentations in seconds 10. Perplexity — AI search with real sources 11. Pika — animate any image into video 12. Runway — cinematic AI video generation 13. Cursor — AI code editor that builds for you 14. v0 — generate UI components with AI 15. Lovable — turn ideas into working apps 16. Descript — edit video by editing text 17. Opus Clip — auto cut long videos into shorts 18. Krea AI — real time AI image generation 19. Magnific — upscale any image with AI 20. Viggle — make characters move realistically 21. tl;dv — record and summarize any meeting 22. Fireflies — AI meeting notes automatically 23. Castmagic — turn audio into content pieces 24. Replit — code and deploy from browser 25. Leonardo AI — generate images for free 26. Synthesia — AI avatar videos no camera needed 27. Fliki — turn text into videos with AI 28. Photoroom — AI product photography 29. Invideo AI — turn prompts into full videos 30. Consensus — search what science agrees on 31. SciSpace — understand any research paper 32. Tome — AI builds your pitch decks 33. Beautiful AI — smart presentation design 34. Meshy — turn text into 3D models 35. Vizcom — turn sketches into renders The AI revolution isn't coming. It already happened and you missed half of it.

Farhan Azad Shuvra

71,095 Aufrufe • vor 4 Tagen