ๆญฃๅœจๅŠ ่ฝฝ่ง†้ข‘...

่ง†้ข‘ๅŠ ่ฝฝๅคฑ่ดฅ

ByteDance announced SeedEdit! A new image model that can edit images with text prompts. It allows for high-resolution editing and supports various changes like local replacements, geometric transformations, and style adjustments. Links โฌ‡๏ธ

46,540 ๆฌก่ง‚็œ‹ โ€ข 1 ๅนดๅ‰ โ€ขvia X (Twitter)

10 ๆก่ฏ„่ฎบ

Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘ ็š„ๅคดๅƒ
Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘1 ๅนดๅ‰

Project Page: Demo: Get notified on release:

Nho Eskape ๐Ÿ”ž ็š„ๅคดๅƒ
Nho Eskape ๐Ÿ”ž1 ๅนดๅ‰

How is it with NSFW?

Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘ ็š„ๅคดๅƒ
Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘1 ๅนดๅ‰

Havenโ€™t tried it myself yey

Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘ ็š„ๅคดๅƒ
Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘1 ๅนดๅ‰

Definitely dope to see image editing getting easier

Nim Eshed ๐•๐Ÿฆ‹ ็š„ๅคดๅƒ
Nim Eshed ๐•๐Ÿฆ‹1 ๅนดๅ‰

Wow ๐Ÿ‘Œ๐Ÿ‘Œ๐Ÿ‘Œ๐Ÿ‘Œ

Aswanth achoo'z ็š„ๅคดๅƒ
Aswanth achoo'z1 ๅนดๅ‰

Wff๐Ÿค˜

Nim Eshed ๐•๐Ÿฆ‹ ็š„ๅคดๅƒ
Nim Eshed ๐•๐Ÿฆ‹1 ๅนดๅ‰

Please post as soon as it out

Dennis ็š„ๅคดๅƒ
Dennis1 ๅนดๅ‰

๐Ÿ”ฅ ๐Ÿ”ฅ

Heba AI ็š„ๅคดๅƒ
Heba AI1 ๅนดๅ‰

Didnt we have add-on to A1111 that did the same in like 2 years ago? It was SD 1.5 based so quality was lower, but idea was the same.

Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘ ็š„ๅคดๅƒ
Dreaming Tulpa ๐Ÿฅ“๐Ÿ‘‘1 ๅนดๅ‰

So? Most of what is out there sucks for this kind of task. This model tries to improve it ๐Ÿคทโ€โ™‚๏ธ

็›ธๅ…ณ่ง†้ข‘

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.๐ŸŽจ โœจ New in 2.1: ๐Ÿ”นAdvanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. ๐Ÿ”นPrecise Chinese and English Text Rendering with seamless imageโ€“text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. ๐Ÿ”นRich Styles and High Aesthetic: Capable of generating images in various stylesโ€”including photorealistic portraits, comics, and vinyl figuresโ€”it delivers outstanding visual appeal and artistic quality. ๐Ÿ”นHigh-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideasโ€”like posters with slogans or multi-panel comicsโ€”into visuals faster than ever. Weโ€™re just getting started. Stay tuned for our native multimodal image generation model coming soon. ๐ŸŒWebsite: ๐Ÿ”—Github: ๐Ÿค—Hugging Face: โœจHugging Face Demo:

Tencent Hy

89,257 ๆฌก่ง‚็œ‹ โ€ข 11 ไธชๆœˆๅ‰

๐Ÿ“ข๐Ÿ“ข ๐๐ž๐ซ๐œ๐‡๐ž๐š๐: ๐๐ž๐ซ๐œ๐ž๐ฉ๐ญ๐ฎ๐š๐ฅ ๐‡๐ž๐š๐ ๐Œ๐จ๐๐ž๐ฅ ๐Ÿ๐จ๐ซ ๐’๐ข๐ง๐ ๐ฅ๐ž-๐ˆ๐ฆ๐š๐ ๐ž ๐Ÿ‘๐ƒ ๐‡๐ž๐š๐ ๐‘๐ž๐œ๐จ๐ง๐ฌ๐ญ๐ซ๐ฎ๐œ๐ญ๐ข๐จ๐ง & ๐„๐๐ข๐ญ๐ข๐ง๐ ๐Ÿ“ข๐Ÿ“ข PercHead reconstructs realistic 3D heads from a single image and enables disentangled 3D editing via geometric controls and style inputs from images or text. At its core is a generalized 3D head decoder trained with perceptual supervision from DINOv2 and SAM 2.1. We find that our new perceptual loss formulation improves reconstruction fidelity compared to commonly-used methods such as LPIPS. Our trained reconstruction model is able to generate 3D-consistent heads from a single input image. Even with challenging side-view inputs, the model robustly infers missing regions for a coherent, high-fidelity output. In addition, our architecture seamlessly adapts to downstream tasks: by swapping the encoder, we can transform the model into a disentangled 3D editing pipeline. In this scenario, we can control geometry through - potentially hand-drawn - segmentation maps, and condition style via image or text prompt. We also provide an interactive GUI to enable the exploration of our editing pipeline. ๐ŸŒ ๐Ÿ“ฝ๏ธ Great work by Antonio Oroz and Tobias Kirschstein

Matthias Niessner

18,855 ๆฌก่ง‚็œ‹ โ€ข 9 ไธชๆœˆๅ‰