Загрузка видео...

Не удалось загрузить видео

На главную

We gave four frontier image editing models the same photo and 30 edits in a row: Ideogram 4.5 keeps most of the room intact, while the others drift, GPT Image 2.5 Sunburst most visibly. We've seen some interesting demos of multi-turn editing consistency from the latest image editing models,...

52,850 просмотров • 1 день назад •via X (Twitter)

Комментарии: 9

Фото профиля serx · fireply.ai
serx · fireply.ai1 день назад

95% untouched on a tulip vase edit is the number i care about

Фото профиля Som Dutt | AI/ML Analyst
Som Dutt | AI/ML Analyst1 день назад

Best single edit ≠ best editor. Once edits stack, pixel preservation beats raw quality. Would love to see “% unchanged per edit” tracked on the leaderboard itself.

Фото профиля aqui
aqui1 день назад

funny how they marketed Image 2.5 as the most consistent editing model

Фото профиля Tatsuya
Tatsuya1 день назад

ideogram winning this is the surprise, nobody talks about them

Фото профиля Simo | ai-costguard
Simo | ai-costguard1 день назад

Each edit looks fine. The sequence is where it falls apart

Фото профиля Wei佳
Wei佳1 день назад

Ideogram editing locally while GPT re-renders everything each turn. That's your drift, right there.

Фото профиля Mehmed II
Mehmed II1 день назад

so ideogram ftw GPT image has always been shite for consistency across edits

Фото профиля Lover of Apps
Lover of Apps1 день назад

I already put Ideogram 4.5 in my workflow for edits.

Фото профиля Rakesh Sahni
Rakesh Sahni1 день назад

"add tulips" shouldn't mean "change the whole room." the 95% vs roughly 20% gap is striking. i'd like to know what "unchanged" means here: - identical pixels - OR are tiny colour shifts allowed? that cutoff matters when reading those numbers.

Похожие видео

InstantDrag Improving Interactivity in Drag-based Image Editing discuss: Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to the challenge of accurately reflecting user interaction while maintaining image content. Some existing approaches rely on computationally intensive per-image optimization or intricate guidance-based methods, requiring additional inputs such as masks for movable regions and text prompts, thereby compromising the interactivity of the editing process. We introduce InstantDrag, an optimization-free pipeline that enhances interactivity and speed, requiring only an image and a drag instruction as input. InstantDrag consists of two carefully designed networks: a drag-conditioned optical flow generator (FlowGen) and an optical flow-conditioned diffusion model (FlowDiffusion). InstantDrag learns motion dynamics for drag-based image editing in real-world video datasets by decomposing the task into motion generation and motion-conditioned image generation. We demonstrate InstantDrag's capability to perform fast, photo-realistic edits without masks or text prompts through experiments on facial video datasets and general scenes. These results highlight the efficiency of our approach in handling drag-based image editing, making it a promising solution for interactive, real-time applications.

AK

71,232 просмотров • 2 лет назад

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 просмотров • 1 год назад