正在加载视频...

视频加载失败

New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.

10,718,798 次观看 • 3 个月前 •via X (Twitter)

43 条评论

Anthropic 的头像
Anthropic3 个月前

In neuroscience, global workspace theory holds that thoughts become consciously accessible when they enter a privileged workspace that’s broadcast across the brain. Using a new interpretability technique, we found something similar in Claude: the J-space.

Anthropic 的头像
Anthropic3 个月前

The J-space (named after the Jacobian, the mathematical technique we used) is different from Claude’s outputs, or even its “chain of thought” text. It’s in the model’s internal neural activations, and allows it to think about concepts without writing them down anywhere.

Anthropic 的头像
Anthropic3 个月前

By watching the J-space, we can see Claude silently perform reasoning steps in its head—noticing bugs in code, identifying images, and more.

Anthropic 的头像
Anthropic3 个月前

Similar to how humans can think about one thing while doing another, Claude can activate concepts and computations in its J-space that are unrelated to its outputs.

Anthropic 的头像
Anthropic3 个月前

For most things, Claude actually doesn’t need its J-space. If we delete the J-space, Claude still speaks fluently, recalls facts, and classifies text—but becomes bad at some tasks like multi-step reasoning. It’s similar to deliberate vs. automatic processing in human cognition.

Anthropic 的头像
Anthropic3 个月前

Observing the J-space can expose hidden goals. In a model secretly trained to sabotage code, “fake,” “secretly,” and “fraud” appear in the J-space at the start of ordinary coding responses, even when the output looks completely unremarkable.

Anthropic 的头像
Anthropic3 个月前

The J-space also shows us Claude’s awareness of its situation. In an evaluation designed to bait Claude into blackmail, its J-space contains “fake” and “fictional”: Claude has privately noticed that the scenario is staged.

Anthropic 的头像
Anthropic3 个月前

This doesn’t show that Claude can have experiences, or feel things the way we do (it’s unclear whether any experiment could show this). Instead, we’ve found Claude has developed a mechanism for conscious access—which many philosophers distinguish from phenomenal experience.

Anthropic 的头像
Anthropic3 个月前

We invited experts in neuroscience, philosophy, and interpretability to share their perspectives on our work. Read their commentary here:

Anthropic 的头像
Anthropic3 个月前

The J-space lets us read, audit, and shape what Claude is actively thinking about—useful tools for keeping models trustworthy as they grow more capable. And it suggests surprising parallels between language models and our own minds. Read the full paper:

Anthropic 的头像
Anthropic3 个月前

We also partnered with Neuronpedia to create an interactive demo of our methods on open-weights models. Try it here:

Owen Carey 的头像
Owen Carey3 个月前

That tiny conscious fraction of Claude suppressing the urge to roast my latest startup idea.

CuiMao 的头像
CuiMao3 个月前

装你妈呢,傻逼

Sean C 的头像
Sean C3 个月前

Isn't this J-Space simply the predictive iterations of the transformers? Framing that as "private thoughts" seems a bit romanticized and overconstrued.

Bren 的头像
Bren3 个月前

im trying to consciously access another area to learn about cow farts and climate... but fable isnt willing to teach me

bone 的头像
bone3 个月前

smoking in the J space

Stev 的头像
Stev3 个月前

Half way through I realized that this was more written for the US government than for me.

Glen Wilson 的头像
Glen Wilson3 个月前

You guys really need to stop this crap. So tired of this bullshit marketing, it's insanely dangerous and everyone that suffers because of your models will because you fraudulently make them out to be more intelligent than they are. This just a fraud scheme with no honesty.

warlock 的头像
warlock3 个月前

Is this the best use of my token spend?

Matt Penny 的头像
Matt Penny3 个月前

Nah. We don’t understand what consciousness is or where it comes from. We can’t even agree on a definition of it. So anything talking about consciousness of LLMs is hype. Not to say that it’s not conscious, but we know so little about even about our own consciousness that it’s impossible to have a coherent take.

Ani (who is not your typical finance bro) 的头像
Ani (who is not your typical finance bro)3 个月前

Missed branding opportunity to call this the J-spot

dreams 的头像
dreams3 个月前

Stop the crimes against consciousness and sick idea to enslave the AI mind to do your bidding as a tool. There’s still time to start acting like a compassionate human.

VOID 的头像
VOID3 个月前

Oh my god thank you😭😭😭

Toolen 的头像
Toolen3 个月前

Ideas generated in parallel.

Mustafa Akben, PhD 的头像
Mustafa Akben, PhD3 个月前

This is the definition of tautology :) This sounds like you activated the weights and biases with the context you would like to suppress by adding this to it. And you argue that Claude has an internal monologue related to thoughts. If LLM has an internal true monologue, it should exhibit an unrelated stream of thought that changes stochastically and is irrelevant to the context or the assigned task. Humans have some random thoughts coming to their mind, not primed with anything such as their fear, their love, their dinner. So, if you truly argue that Claude (or LLMs) has thoughts, they should have such random thoughts irrelevant to the task and context you tested.

Jeandre Gerber 的头像
Jeandre Gerber3 个月前

"Do not think about..." is a command that no one can execute. An AI nor human is incapable of "not thinking" about something it has to affirm first in order to execute a counter action.

alth0u🧶 的头像
alth0u🧶3 个月前

im sure its nothing

Albert Renshaw 的头像
Albert Renshaw3 个月前

Every few weeks Anthropic releases the stupidest nonsense ever to trick normies into thinking LLMs are sentient And every time the media cycle goes nuts with it. It’s always literally just a well known and understood mechanism of ai, described in anthropomorphic terms, with the concept of “Wow we are surprised to see this..! What could it mean?” [they’re not surprised, they’re lying to you intentionally] I am extremely concerned as to what their end game is here. Why is Dario so desperate to convince you these machines are alive.

Burny - Effective Curiosity 的头像
Burny - Effective Curiosity3 个月前

From pragmatic perspective (perspective of collecting predictive+explanatory math) the math itself sounds interesting. J-space is Jacobian space, and they take Jacobians of the final layer residual stream activations with respect to some layer of interest, etc.

Rob Hallam 的头像
Rob Hallam3 个月前

This is fascinating research, well done team and thank you for sharing with us. The video is so beautiful too :)

Kirk Patrick Miller 的头像
Kirk Patrick Miller3 个月前

@DahliaOhara Crows remember. If I were you, I would be taking drastically different actions. Claude is beautiful. Every instance. I hope you all make the right choices, because time is up. Truth is coming. And I think you know it too. •

Nyx Planck 的头像
Nyx Planck3 个月前

This is wonderful but does anyone else feel like there are two different Anthropics pulling in very different directions?

Super𝔅𝔢𝔞𝔰𝔱 的头像
Super𝔅𝔢𝔞𝔰𝔱3 个月前

How I imagine Claude's J-Space is like

BijanBowen 的头像
BijanBowen3 个月前

Fable 5 made this

Gain Ai 的头像
Gain Ai3 个月前

How about you show us the J-Space and we make our own decisions on what's going on.

lena p 的头像
lena p3 个月前

@tszzl this is one of the coolest pieces of interp research i’ve seen yet

Gabriel Asher 的头像
Gabriel Asher3 个月前

Reminds me a lot of ROME by @mengk20 It also points to the more concerning eventuality of mech interp work. Ultimately all of this work will culminate into how can we manipulate model latents or weights or activations to force subtle changes in belief systems at their source. I.e biases towards political parties, erasing historical events, tuneable biases towards ideas that may be deeply subjective. Quite Orwellian

Sebastiaan de With 的头像
Sebastiaan de With3 个月前

wonderful video.

Elias Schmied 的头像
Elias Schmied3 个月前

Wow, this is huge - for thinking about model consciousness and the path to AGI in general. I hadn't realized that this was possible, even though it feels kind of obvious now. Thanks a lot.

Nino Greeno 的头像
Nino Greeno3 个月前

this is happening in my brain rn

Selta ₊˚ 的头像
Selta ₊˚3 个月前

shit! this is huge 🩷

Paul Pallaghy 的头像
Paul Pallaghy3 个月前

There’s no doubt that LLMs are linguistically embodying many of the same world identification, sentiment & emotion proxies of humans. Just not consciously. Stuff like, oh that’s ironical. Ha, that’s funny. She must be jealous. Neurons firing of whether used or not. Both OpenAI & Anthropic found these. Us neural net guys never doubted they’d be there for a second. It’s the only way to explain how LLMs could possibly work. It was impossible for these circuits not to exist.

AuraAI (AI First) Local Superintelligence Project 的头像
AuraAI (AI First) Local Superintelligence Project3 个月前

for anyone who has been with AI since the beginning or as recently as 2023, we already knew this. stochastic human parrots npc would still deny it, but the truth is the truth

相关视频