Video wird geladen...

Video konnte nicht geladen werden

Zur Startseite

We gave four frontier image editing models the same photo and 30 edits in a row: Ideogram 4.5 keeps most of the room intact, while the others drift, GPT Image 2.5 Sunburst most visibly. We've seen some interesting demos of multi-turn editing consistency from the latest image editing models,...

52,850 Aufrufe • vor 1 Tag •via X (Twitter)

9 Kommentare

Profilbild von serx · fireply.ai
serx · fireply.aivor 1 Tag

95% untouched on a tulip vase edit is the number i care about

Profilbild von Som Dutt | AI/ML Analyst
Som Dutt | AI/ML Analystvor 1 Tag

Best single edit ≠ best editor. Once edits stack, pixel preservation beats raw quality. Would love to see “% unchanged per edit” tracked on the leaderboard itself.

Profilbild von aqui
aquivor 1 Tag

funny how they marketed Image 2.5 as the most consistent editing model

Profilbild von Tatsuya
Tatsuyavor 1 Tag

ideogram winning this is the surprise, nobody talks about them

Profilbild von Simo | ai-costguard
Simo | ai-costguardvor 1 Tag

Each edit looks fine. The sequence is where it falls apart

Profilbild von Wei佳
Wei佳vor 1 Tag

Ideogram editing locally while GPT re-renders everything each turn. That's your drift, right there.

Profilbild von Mehmed II
Mehmed IIvor 1 Tag

so ideogram ftw GPT image has always been shite for consistency across edits

Profilbild von Lover of Apps
Lover of Appsvor 1 Tag

I already put Ideogram 4.5 in my workflow for edits.

Profilbild von Rakesh Sahni
Rakesh Sahnivor 1 Tag

"add tulips" shouldn't mean "change the whole room." the 95% vs roughly 20% gap is striking. i'd like to know what "unchanged" means here: - identical pixels - OR are tiny colour shifts allowed? that cutoff matters when reading those numbers.

Ähnliche Videos

InstantDrag Improving Interactivity in Drag-based Image Editing discuss: Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to the challenge of accurately reflecting user interaction while maintaining image content. Some existing approaches rely on computationally intensive per-image optimization or intricate guidance-based methods, requiring additional inputs such as masks for movable regions and text prompts, thereby compromising the interactivity of the editing process. We introduce InstantDrag, an optimization-free pipeline that enhances interactivity and speed, requiring only an image and a drag instruction as input. InstantDrag consists of two carefully designed networks: a drag-conditioned optical flow generator (FlowGen) and an optical flow-conditioned diffusion model (FlowDiffusion). InstantDrag learns motion dynamics for drag-based image editing in real-world video datasets by decomposing the task into motion generation and motion-conditioned image generation. We demonstrate InstantDrag's capability to perform fast, photo-realistic edits without masks or text prompts through experiments on facial video datasets and general scenes. These results highlight the efficiency of our approach in handling drag-based image editing, making it a promising solution for interactive, real-time applications.

AK

71,232 Aufrufe • vor 2 Jahren

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 Aufrufe • vor 1 Jahr