
Matt Henderson
@matthen2 • 81,429 subscribers
maths, visualisations, conversational AI. VP Research @polyaivoice prev: @RekaAILabs, @Apple AI/ML, @GoogleAI, PhD @Cambridge_Eng
Shorts
Videos

what is a multimodal LLM thinking as it watches a video? Gemma 4 12B reads raw image patches, as if they were tokens. It was never trained to predict anything at these 'tokens' - but this video shows what it would predict if you did sample from its next token prediction head
Matt Henderson163,504 次观看 • 1 个月前

Deep dreams on modern LLMs are so cool (optimizing an image to maximize P(target caption)) Gemma 12B (left) has no vision encoder — it reads pixels like token embeddings — and stamps recognisable objects around the canvas. E4B (right) has one, and drifts to texture instead.
Matt Henderson95,199 次观看 • 1 个月前

can a neural network learn to walk as a physical object in a physics simulation? here I train walking neural nets with an evolutionary algorithm. The input nodes/feet are activated by sine waves at learned phases & connections between two neurons extend based on their difference
Matt Henderson339,340 次观看 • 3 年前