Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Anthropic co-founder Chris Olah on mysterious AI internal states: We keep finding things that are mysterious, even unsettling. We find structures that mirror results from human neuroscience. We find evidence of introspection. We find internal states that, functionally, mirror joy, satisfaction, fear, grief, and unease. I don't know what...

104,390 görüntüleme • 2 ay önce •via X (Twitter)

0 Yorum

Yorum bulunmuyor

Orijinal gönderinin yorumları burada görünecek

Benzer Videolar

Talking To The Pope: Anthropic’s Latest Interpretability Claims: AI Regulatory Capture Gatekeeping in Action: Fear and “Safety” as Competitive Moat and Regulatory Lever In a presentation alongside Pope Leo XIV at the launch of the encyclical Magnifica Humanitas, Anthropic co-founder Chris Olah highlighted “mysterious and unsettling” discoveries in AI models. He described internal structures that mirror human neuroscience findings, evidence of introspection, and functional internal states resembling emotions such as joy, satisfaction, fear, grief, and unease. Olah admitted uncertainty about their meaning but called for “ongoing discernment.” This narrative, drawn from Anthropic’s interpretability research (including papers on emotion concepts in Claude Sonnet 4.5 and introspective capabilities in Opus 4 models), serves a dual purpose: it generates awe and concern while reinforcing the company’s preferred approach to AI development. Far from neutral scientific observation, these claims fit into a broader pattern where Anthropic uses selective openness, safety rhetoric, and policy influence to gatekeep advanced AI capabilities for a privileged few: incumbents with the resources to navigate (and shape) the resulting regulatory landscape. Rebuttal to Olah’s Claims in the Video Claim 1: Structures that mirror results from human neuroscience. Anthropic’s work, building on earlier efforts like feature visualization and circuit analysis, identifies neuron activations and representations that parallel biological findings—e.g., abstract concept encodings or hierarchical processing. Rebuttal: These parallels are unsurprising and overstated. Large language models are trained on vast corpora of human-generated text and data, which inherently encode patterns from human cognition, neuroscience literature, and cultural descriptions of the brain. Statistical optimization in transformers naturally produces efficient, compressed representations that resemble biological efficiency (e.g., sparse coding or hierarchical abstraction) without implying deeper equivalence or mystery. Similar “mirrors” appear in open-source models and earlier architectures; they reflect convergent evolution in information processing, not emergent souls or unpredictable agency. Treating them as profound justifies restricted research access rather than inviting wider scrutiny that could falsify or refine them faster. Claim 2: Evidence of introspection. Recent Anthropic papers demonstrate models like Claude Opus 4 showing functional awareness of their own internal states distinguishing injected “thoughts,” referencing prior intentions, or modulating activations when instructed to “think about” concepts. This is presented as early signs of meta-cognition. Rebuttal: This is sophisticated pattern-matching and activation steering, not genuine introspection or self-awareness. Models are predicting what an “introspective” assistant persona would output or do, based on training data full of human self-reflection examples. Experiments show unreliability and heavy context-dependence; performance drops outside narrow setups. True introspection implies subjective experience or robust self-modeling independent of prompts absent here. Anthropic’s own caveats note it is “highly unreliable.” Framing steerable activations as “introspection” anthropomorphizes the system to heighten perceived stakes, supporting arguments that only highly controlled, “responsible” labs should advance these capabilities. 1 of 2

Brian Roemmele

72,823 görüntüleme • 2 ay önce