Загрузка видео...

Не удалось загрузить видео

На главную

Hy Image3.5 preview is live. 🚀 Professional-grade image generation, +30% win rate in human eval vs Hy Image3.0 Both Text to image & Image to image available. Up to 2K. Better Consistency. API: Priced for everyone. $0.024 per image on Tencent Cloud API. We only charge for what we...

36,840 просмотров • 6 часов назад •via X (Twitter)

Комментарии: 14

Фото профиля SSH
SSH5 часов назад

open weight!!!!!!!!!!!!!!!!!

Фото профиля 👁️M1nd 3xp4nd3r👁️
👁️M1nd 3xp4nd3r👁️5 часов назад

You all should open source ❤️

Фото профиля Fajar M Reza
Fajar M Reza5 часов назад

Multimodal image generation becomes more useful when price and consistency improve together.

Фото профиля Greg Hunkins
Greg Hunkins4 часов назад

Great release, congrats, excited to test!

Фото профиля _alphashark_
_alphashark_5 часов назад

lock the seed and run the same image to image prompt at each strength setting. track face and logo drift with lpips, since better consistency can still hide local identity changes.

Фото профиля Leo Lu
Leo Lu5 часов назад

Free reference images make image-to-image iteration cheaper than the $0.024 output price suggests.

Фото профиля China OSS AI-lilxl
China OSS AI-lilxl5 часов назад

Tencent Hunyuan launched Hy Image3.5 preview (text-to-image & image-to-image, +30% win rate vs 3.0 in human eval). image-to-image is the piece the ComfyUI community will test hardest here. Hunyuan 3.0 was one of the few Chinese diffusion models that got deep local tooling support. curious if 3.5 keeps the same open-weights roadmap once the cloud preview wraps up.

Фото профиля AZIZ | AI 🇸🇦
AZIZ | AI 🇸🇦6 часов назад

Seens stunning 🤩

Фото профиля Omar
Omar4 часов назад

@grok is this going open weight? What do we know about the model size too

Фото профиля SilentYears
SilentYears6 часов назад

HY-MT系列模型还会有更新吗

Фото профиля liftoff
liftoff4 часов назад

卧槽?官方的?怎么一点流量没有的

Фото профиля .
.5 часов назад

it is not opensource? put it in your ass

Фото профиля OnSoloAI
OnSoloAI5 часов назад

Hunyuan Image 3.5 is on OnSolo. Exclusive. 5 refs. 2K. 1 credit. Members free.

Фото профиля SK
SK6 часов назад

Seems everything tenant releases is always stuck in preview

Похожие видео

InstantDrag Improving Interactivity in Drag-based Image Editing discuss: Drag-based image editing has recently gained popularity for its interactivity and precision. However, despite the ability of text-to-image models to generate samples within a second, drag editing still lags behind due to the challenge of accurately reflecting user interaction while maintaining image content. Some existing approaches rely on computationally intensive per-image optimization or intricate guidance-based methods, requiring additional inputs such as masks for movable regions and text prompts, thereby compromising the interactivity of the editing process. We introduce InstantDrag, an optimization-free pipeline that enhances interactivity and speed, requiring only an image and a drag instruction as input. InstantDrag consists of two carefully designed networks: a drag-conditioned optical flow generator (FlowGen) and an optical flow-conditioned diffusion model (FlowDiffusion). InstantDrag learns motion dynamics for drag-based image editing in real-world video datasets by decomposing the task into motion generation and motion-conditioned image generation. We demonstrate InstantDrag's capability to perform fast, photo-realistic edits without masks or text prompts through experiments on facial video datasets and general scenes. These results highlight the efficiency of our approach in handling drag-based image editing, making it a promising solution for interactive, real-time applications.

AK

71,232 просмотров • 2 лет назад

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,392 просмотров • 1 год назад

We’re excited to announce the release and open-source of HunyuanImage 3.0 — the largest and most powerful open-source text-to-image model to date, with over 80 billion total parameters, of which 13 billion are activated per token during inference.The effect is completely comparable to the industry’s flagship closed-source model.🚀🚀🚀 HunyuanImage 3.0 originates from our internally developed native multimodal large language model, with fine-tuning and post-training focused on text-to-image generation. This unique foundation gives the model a powerful set of capabilities: ✅Reason with world knowledge ✅Understand complex, thousand-word prompts ✅Generate precise text within images Different from traditional DiT architecture image generation models, HunyuanImage 3.0’s MoE architecture uses a Transfusion-based approach to deeply couple Diffusion and LLM training for a single, powerful system. Built on Hunyuan-A13B, HunyuanImage 3.0 was trained on a massive dataset: 5 billion image-text pairs, video frames, interleaved image-text data, and 6 trillion tokens of text corpora. This hybrid training across multimodal generation, understanding, and LLM capabilities allows the model to seamlessly integrate multiple tasks. Whether you're an illustrator, designer, or creator, this is built to slash your workflow from hours to minutes. HunyuanImage 3.0 can generate intricate text, detailed comics, expressive emojis, and lively, engaging illustrations for educational content. The current release focuses solely on text-to-image generation and future updates will include image-to-image, image editing, multi-turn interaction, and more. 👉🏻Try it now: 🔗GitHub: 🤗Hugging Face:

Tencent Hy

413,096 просмотров • 11 месяцев назад