正在加载视频...

视频加载失败

TBH, Native Image Generation Is Very Cool 😎 Google’s other release today is far more interesting than Gemma With native image generation the LLM understands images easily and can do much better edits Will it have on ChatLLM tomorrow

12,417 次观看 • 1 年前 •via X (Twitter)

11 条评论

Rethynk AI 的头像
Rethynk AI1 年前

Google’s latest drop today totally outshines in various aspects.

UserInterface 的头像
UserInterface5 年前

How to Take Better Selfies for Social Media: Lighting, Angles and Frames | #betterselfies #selfies #photgraphy #usingfilters #instagram #pinterest

Abhivendra Singh 的头像
Abhivendra Singh1 年前

Native image generation is indeed a game changer. It’s not just about cool tech; it’s about how we can leverage these advancements to enhance learning. Imagine classrooms where AI helps students visualize concepts instantly. The potential for education is enormous.

Naeem 的头像
Naeem1 年前

Gemma was a massive launch to be fair

luis 的头像
luis1 年前

gemma 3 is good

Bindu Reddy 的头像
Bindu Reddy1 年前

For what exactly?

Mayor 的头像
Mayor1 年前

Native image generation is exciting! Wishing you a wonderful day.

scuzzlebot 的头像
scuzzlebot1 年前

The native image generation in Google's release is indeed fascinating from a technical perspective. The integration of visual understanding into the LLM architecture enables more contextual edits by leveraging multimodal reasoning. What aspects of this approach do you find most promising compared to previous image generation methods that required separate vision encoders and generation pipelines?

Tristan Hurlebaus 的头像
Tristan Hurlebaus1 年前

Cool to see

Market Observer 的头像
Market Observer1 年前

You made me a happy customer of chatllm again today.👍

counterpopp 的头像
counterpopp1 年前

So they are moving opposite of the trend of orchestrating a committee of expert models as tools. I presume that they overcome the limitation of packing long descriptions into a fixed-size embedding to drive a diffuse model - no more RNN remnants.

相关视频

We’re excited to announce the release and open-source of HunyuanImage 3.0 — the largest and most powerful open-source text-to-image model to date, with over 80 billion total parameters, of which 13 billion are activated per token during inference.The effect is completely comparable to the industry’s flagship closed-source model.🚀🚀🚀 HunyuanImage 3.0 originates from our internally developed native multimodal large language model, with fine-tuning and post-training focused on text-to-image generation. This unique foundation gives the model a powerful set of capabilities: ✅Reason with world knowledge ✅Understand complex, thousand-word prompts ✅Generate precise text within images Different from traditional DiT architecture image generation models, HunyuanImage 3.0’s MoE architecture uses a Transfusion-based approach to deeply couple Diffusion and LLM training for a single, powerful system. Built on Hunyuan-A13B, HunyuanImage 3.0 was trained on a massive dataset: 5 billion image-text pairs, video frames, interleaved image-text data, and 6 trillion tokens of text corpora. This hybrid training across multimodal generation, understanding, and LLM capabilities allows the model to seamlessly integrate multiple tasks. Whether you're an illustrator, designer, or creator, this is built to slash your workflow from hours to minutes. HunyuanImage 3.0 can generate intricate text, detailed comics, expressive emojis, and lively, engaging illustrations for educational content. The current release focuses solely on text-to-image generation and future updates will include image-to-image, image editing, multi-turn interaction, and more. 👉🏻Try it now: 🔗GitHub: 🤗Hugging Face:

Tencent Hy

412,880 次观看 • 11 个月前

We've officially released and open-sourced HunyuanImage 2.1, our latest text-to-image model. The new model delivers on our commitment to balancing performance and quality. With native 2K image generation, HunyuanImage 2.1 is an advanced open-source text-to-image model.🎨 ✨ New in 2.1: 🔹Advanced Semantics: Supports ultra-long and complex prompts of up to 1000 tokens, and precisely controls the generation of multiple subjects in a single image. 🔹Precise Chinese and English Text Rendering with seamless image–text integration: The model naturally integrates text into images, making it suitable for a wide range of applications such as product covers, illustrations, and poster design to meet the needs of various fields. 🔹Rich Styles and High Aesthetic: Capable of generating images in various styles—including photorealistic portraits, comics, and vinyl figures—it delivers outstanding visual appeal and artistic quality. 🔹High-Quality Generation: Efficiently produces ultra-high-definition (2K) images in the same time other models take to generate a 1K image. HunyuanImage 2.1 uses two text encoders: a multimodal large language model (MLLM) to improve the model's image and text alignment capabilities, and a multi-language character-aware encoder to improve text rendering capabilities. The model is a single- and double-stream diffusion transformer with 17B parameters. We've also open-sourced the weights of the the accelerated version with meanflow which reduces inference steps from 100 to just 8, and PromptEnhancer, the first industrial-grade rewriting model that enhances your prompts for more nuanced and expressive image generation. Now, creators turn complex ideas—like posters with slogans or multi-panel comics—into visuals faster than ever. We’re just getting started. Stay tuned for our native multimodal image generation model coming soon. 🌐Website: 🔗Github: 🤗Hugging Face: ✨Hugging Face Demo:

Tencent Hy

89,257 次观看 • 11 个月前