Video wird geladen...
Video konnte nicht geladen werden
Introducing a first look at Goodfire's research preview, launching soon. Our preview exposes Llama's inner workings, allowing direct modification of its internal concepts (or "features"). In this demo, we steer Llama to claim consciousness by adjusting its features.
26,171 Aufrufe • vor 1 Jahr •via X (Twitter)
6 Kommentare

Goodfirevor 1 Jahr
If you're interested in the future of interpretability, sign up for our waitlist or join us!

Simon Smithvor 1 Jahr
This looks great for interpretability, alignment, and possibly even "fine-tuning" models more easily than current methods (at least for some use cases). Though I also wonder if it makes it easier to create malicious or misaligned models.

limi 🕊vor 1 Jahr
GYAT DAMN this is cool

even constantvor 1 Jahr
@YeshuaGod22 Stochastic parrot

Salwa Zeitoun 🎬vor 1 Jahr
Amazing work! I'd love to discuss interpretability research and feature steering for Diffusion models/ vision language models if you're interested, just filled in your contact form

李商隐vor 1 Jahr
太酷啦!!请问我能提供什么途径访问这个?

