正在加载视频...
视频加载失败
Deep dreams on modern LLMs are so cool (optimizing an image to maximize P(target caption)) Gemma 12B (left) has no vision encoder — it reads pixels like token embeddings — and stamps recognisable objects around the canvas. E4B (right) has one, and drifts to texture instead.
95,199 次观看 • 1 个月前 •via X (Twitter)
0 条评论
暂无评论
原始帖子的评论将显示在这里
