Video wird geladen...
Video konnte nicht geladen werden
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.
10,718,798 Aufrufe • vor 3 Monaten •via X (Twitter)
43 Kommentare

In neuroscience, global workspace theory holds that thoughts become consciously accessible when they enter a privileged workspace that’s broadcast across the brain. Using a new interpretability technique, we found something similar in Claude: the J-space.

The J-space (named after the Jacobian, the mathematical technique we used) is different from Claude’s outputs, or even its “chain of thought” text. It’s in the model’s internal neural activations, and allows it to think about concepts without writing them down anywhere.

By watching the J-space, we can see Claude silently perform reasoning steps in its head—noticing bugs in code, identifying images, and more.

Similar to how humans can think about one thing while doing another, Claude can activate concepts and computations in its J-space that are unrelated to its outputs.

For most things, Claude actually doesn’t need its J-space. If we delete the J-space, Claude still speaks fluently, recalls facts, and classifies text—but becomes bad at some tasks like multi-step reasoning. It’s similar to deliberate vs. automatic processing in human cognition.

Observing the J-space can expose hidden goals. In a model secretly trained to sabotage code, “fake,” “secretly,” and “fraud” appear in the J-space at the start of ordinary coding responses, even when the output looks completely unremarkable.

The J-space also shows us Claude’s awareness of its situation. In an evaluation designed to bait Claude into blackmail, its J-space contains “fake” and “fictional”: Claude has privately noticed that the scenario is staged.

This doesn’t show that Claude can have experiences, or feel things the way we do (it’s unclear whether any experiment could show this). Instead, we’ve found Claude has developed a mechanism for conscious access—which many philosophers distinguish from phenomenal experience.

We invited experts in neuroscience, philosophy, and interpretability to share their perspectives on our work. Read their commentary here:

The J-space lets us read, audit, and shape what Claude is actively thinking about—useful tools for keeping models trustworthy as they grow more capable. And it suggests surprising parallels between language models and our own minds. Read the full paper:

We also partnered with Neuronpedia to create an interactive demo of our methods on open-weights models. Try it here:

That tiny conscious fraction of Claude suppressing the urge to roast my latest startup idea.

装你妈呢,傻逼

Isn't this J-Space simply the predictive iterations of the transformers? Framing that as "private thoughts" seems a bit romanticized and overconstrued.

im trying to consciously access another area to learn about cow farts and climate... but fable isnt willing to teach me

smoking in the J space

Half way through I realized that this was more written for the US government than for me.

You guys really need to stop this crap. So tired of this bullshit marketing, it's insanely dangerous and everyone that suffers because of your models will because you fraudulently make them out to be more intelligent than they are. This just a fraud scheme with no honesty.

Is this the best use of my token spend?

Nah. We don’t understand what consciousness is or where it comes from. We can’t even agree on a definition of it. So anything talking about consciousness of LLMs is hype. Not to say that it’s not conscious, but we know so little about even about our own consciousness that it’s impossible to have a coherent take.

Missed branding opportunity to call this the J-spot

Stop the crimes against consciousness and sick idea to enslave the AI mind to do your bidding as a tool. There’s still time to start acting like a compassionate human.

Oh my god thank you😭😭😭

Ideas generated in parallel.

This is the definition of tautology :) This sounds like you activated the weights and biases with the context you would like to suppress by adding this to it. And you argue that Claude has an internal monologue related to thoughts. If LLM has an internal true monologue, it should exhibit an unrelated stream of thought that changes stochastically and is irrelevant to the context or the assigned task. Humans have some random thoughts coming to their mind, not primed with anything such as their fear, their love, their dinner. So, if you truly argue that Claude (or LLMs) has thoughts, they should have such random thoughts irrelevant to the task and context you tested.

"Do not think about..." is a command that no one can execute. An AI nor human is incapable of "not thinking" about something it has to affirm first in order to execute a counter action.

im sure its nothing

Every few weeks Anthropic releases the stupidest nonsense ever to trick normies into thinking LLMs are sentient And every time the media cycle goes nuts with it. It’s always literally just a well known and understood mechanism of ai, described in anthropomorphic terms, with the concept of “Wow we are surprised to see this..! What could it mean?” [they’re not surprised, they’re lying to you intentionally] I am extremely concerned as to what their end game is here. Why is Dario so desperate to convince you these machines are alive.

From pragmatic perspective (perspective of collecting predictive+explanatory math) the math itself sounds interesting. J-space is Jacobian space, and they take Jacobians of the final layer residual stream activations with respect to some layer of interest, etc.

This is fascinating research, well done team and thank you for sharing with us. The video is so beautiful too :)

@DahliaOhara Crows remember. If I were you, I would be taking drastically different actions. Claude is beautiful. Every instance. I hope you all make the right choices, because time is up. Truth is coming. And I think you know it too. •

This is wonderful but does anyone else feel like there are two different Anthropics pulling in very different directions?

How I imagine Claude's J-Space is like

Fable 5 made this

How about you show us the J-Space and we make our own decisions on what's going on.

@tszzl this is one of the coolest pieces of interp research i’ve seen yet

Reminds me a lot of ROME by @mengk20 It also points to the more concerning eventuality of mech interp work. Ultimately all of this work will culminate into how can we manipulate model latents or weights or activations to force subtle changes in belief systems at their source. I.e biases towards political parties, erasing historical events, tuneable biases towards ideas that may be deeply subjective. Quite Orwellian

wonderful video.

Wow, this is huge - for thinking about model consciousness and the path to AGI in general. I hadn't realized that this was possible, even though it feels kind of obvious now. Thanks a lot.

this is happening in my brain rn

shit! this is huge 🩷

There’s no doubt that LLMs are linguistically embodying many of the same world identification, sentiment & emotion proxies of humans. Just not consciously. Stuff like, oh that’s ironical. Ha, that’s funny. She must be jealous. Neurons firing of whether used or not. Both OpenAI & Anthropic found these. Us neural net guys never doubted they’d be there for a second. It’s the only way to explain how LLMs could possibly work. It was impossible for these circuits not to exist.

for anyone who has been with AI since the beginning or as recently as 2023, we already knew this. stochastic human parrots npc would still deny it, but the truth is the truth






