正在加载视频...

视频加载失败

We gave four frontier image editing models the same photo and 30 edits in a row: Ideogram 4.5 keeps most of the room intact, while the others drift, GPT Image 2.5 Sunburst most visibly. We've seen some interesting demos of multi-turn editing consistency from the latest image editing models,...

52,850 次观看 • 1 天前 •via X (Twitter)

9 条评论

serx · fireply.ai 的头像
serx · fireply.ai1 天前

95% untouched on a tulip vase edit is the number i care about

Som Dutt | AI/ML Analyst 的头像
Som Dutt | AI/ML Analyst1 天前

Best single edit ≠ best editor. Once edits stack, pixel preservation beats raw quality. Would love to see “% unchanged per edit” tracked on the leaderboard itself.

aqui 的头像
aqui1 天前

funny how they marketed Image 2.5 as the most consistent editing model

Tatsuya 的头像
Tatsuya1 天前

ideogram winning this is the surprise, nobody talks about them

Simo | ai-costguard 的头像
Simo | ai-costguard1 天前

Each edit looks fine. The sequence is where it falls apart

Wei佳 的头像
Wei佳1 天前

Ideogram editing locally while GPT re-renders everything each turn. That's your drift, right there.

Mehmed II 的头像
Mehmed II1 天前

so ideogram ftw GPT image has always been shite for consistency across edits

Lover of Apps 的头像
Lover of Apps1 天前

I already put Ideogram 4.5 in my workflow for edits.

Rakesh Sahni 的头像
Rakesh Sahni1 天前

"add tulips" shouldn't mean "change the whole room." the 95% vs roughly 20% gap is striking. i'd like to know what "unchanged" means here: - identical pixels - OR are tiny colour shifts allowed? that cutoff matters when reading those numbers.

相关视频

InstantDrag Improving Interactivity in Drag-based Image Editing discuss: Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to the challenge of accurately reflecting user interaction while maintaining image content. Some existing approaches rely on computationally intensive per-image optimization or intricate guidance-based methods, requiring additional inputs such as masks for movable regions and text prompts, thereby compromising the interactivity of the editing process. We introduce InstantDrag, an optimization-free pipeline that enhances interactivity and speed, requiring only an image and a drag instruction as input. InstantDrag consists of two carefully designed networks: a drag-conditioned optical flow generator (FlowGen) and an optical flow-conditioned diffusion model (FlowDiffusion). InstantDrag learns motion dynamics for drag-based image editing in real-world video datasets by decomposing the task into motion generation and motion-conditioned image generation. We demonstrate InstantDrag's capability to perform fast, photo-realistic edits without masks or text prompts through experiments on facial video datasets and general scenes. These results highlight the efficiency of our approach in handling drag-based image editing, making it a promising solution for interactive, real-time applications.

AK

71,232 次观看 • 2 年前

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 次观看 • 1 年前