Loading video...

Video Failed to Load

Go Home

ByteDance announced SeedEdit! A new image model that can edit images with text prompts. It allows for high-resolution editing and supports various changes like local replacements, geometric transformations, and style adjustments. Links โฌ‡๏ธ

46,540 views โ€ข 1 year ago โ€ขvia X (Twitter)

10 Comments

Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘'s profile picture
Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘1 year ago

Project Page: Demo: Get notified on release:

Nho Eskape ๐Ÿ”ž's profile picture
Nho Eskape ๐Ÿ”ž1 year ago

How is it with NSFW?

Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘'s profile picture
Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘1 year ago

Havenโ€™t tried it myself yey

Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘'s profile picture
Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘1 year ago

Definitely dope to see image editing getting easier

Nim Eshed ๐•๐Ÿฆ‹'s profile picture
Nim Eshed ๐•๐Ÿฆ‹1 year ago

Wow ๐Ÿ‘Œ๐Ÿ‘Œ๐Ÿ‘Œ๐Ÿ‘Œ

Aswanth achoo'z's profile picture
Aswanth achoo'z1 year ago

Wff๐Ÿค˜

Nim Eshed ๐•๐Ÿฆ‹'s profile picture
Nim Eshed ๐•๐Ÿฆ‹1 year ago

Please post as soon as it out

Dennis's profile picture
Dennis1 year ago

๐Ÿ”ฅ ๐Ÿ”ฅ

Heba AI's profile picture
Heba AI1 year ago

Didnt we have add-on to A1111 that did the same in like 2 years ago? It was SD 1.5 based so quality was lower, but idea was the same.

Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘'s profile picture
Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘1 year ago

So? Most of what is out there sucks for this kind of task. This model tries to improve it ๐Ÿคทโ€โ™‚๏ธ

Related Videos

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.๐ŸŽจ โœจ New in 2.1: ๐Ÿ”นAdvanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. ๐Ÿ”นPrecise Chinese and English Text Rendering with seamless imageโ€“text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. ๐Ÿ”นRich Styles and High Aesthetic: Capable of generating images in various stylesโ€”including photorealistic portraits, comics, and vinyl figuresโ€”it delivers outstanding visual appeal and artistic quality. ๐Ÿ”นHigh-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideasโ€”like posters with slogans or multi-panel comicsโ€”into visuals faster than ever. Weโ€™re just getting started. Stay tuned for our native multimodal image generation model coming soon. ๐ŸŒWebsite: ๐Ÿ”—Github: ๐Ÿค—Hugging Face: โœจHugging Face Demo:

Tencent Hy

89,257 views โ€ข 11 months ago

๐Ÿ“ข๐Ÿ“ข ๐๐ž๐ซ๐œ๐‡๐ž๐š๐: ๐๐ž๐ซ๐œ๐ž๐ฉ๐ญ๐ฎ๐š๐ฅ ๐‡๐ž๐š๐ ๐Œ๐จ๐๐ž๐ฅ ๐Ÿ๐จ๐ซ ๐’๐ข๐ง๐ ๐ฅ๐ž-๐ˆ๐ฆ๐š๐ ๐ž ๐Ÿ‘๐ƒ ๐‡๐ž๐š๐ ๐‘๐ž๐œ๐จ๐ง๐ฌ๐ญ๐ซ๐ฎ๐œ๐ญ๐ข๐จ๐ง & ๐„๐๐ข๐ญ๐ข๐ง๐ ๐Ÿ“ข๐Ÿ“ข PercHead reconstructs realistic 3D heads from a single image and enables disentangled 3D editing via geometric controls and style inputs from images or text. At its core is a generalized 3D head decoder trained with perceptual supervision from DINOv2 and SAM 2.1. We find that our new perceptual loss formulation improves reconstruction fidelity compared to commonly-used methods such as LPIPS. Our trained reconstruction model is able to generate 3D-consistent heads from a single input image. Even with challenging side-view inputs, the model robustly infers missing regions for a coherent, high-fidelity output. In addition, our architecture seamlessly adapts to downstream tasks: by swapping the encoder, we can transform the model into a disentangled 3D editing pipeline. In this scenario, we can control geometry through - potentially hand-drawn - segmentation maps, and condition style via image or text prompt. We also provide an interactive GUI to enable the exploration of our editing pipeline. ๐ŸŒ ๐Ÿ“ฝ๏ธ Great work by Antonio Oroz and Tobias Kirschstein

Matthias Niessner

18,855 views โ€ข 9 months ago