Loading video...

Video Failed to Load

Go Home

New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.

10,718,711 views • 3 months ago •via X (Twitter)

43 Comments

Anthropic's profile picture
Anthropic3 months ago

In neuroscience, global workspace theory holds that thoughts become consciously accessible when they enter a privileged workspace that’s broadcast across the brain. Using a new interpretability technique, we found something similar in Claude: the J-space.

Anthropic's profile picture
Anthropic3 months ago

The J-space (named after the Jacobian, the mathematical technique we used) is different from Claude’s outputs, or even its “chain of thought” text. It’s in the model’s internal neural activations, and allows it to think about concepts without writing them down anywhere.

Anthropic's profile picture
Anthropic3 months ago

By watching the J-space, we can see Claude silently perform reasoning steps in its head—noticing bugs in code, identifying images, and more.

Anthropic's profile picture
Anthropic3 months ago

Similar to how humans can think about one thing while doing another, Claude can activate concepts and computations in its J-space that are unrelated to its outputs.

Anthropic's profile picture
Anthropic3 months ago

For most things, Claude actually doesn’t need its J-space. If we delete the J-space, Claude still speaks fluently, recalls facts, and classifies text—but becomes bad at some tasks like multi-step reasoning. It’s similar to deliberate vs. automatic processing in human cognition.

Anthropic's profile picture
Anthropic3 months ago

Observing the J-space can expose hidden goals. In a model secretly trained to sabotage code, “fake,” “secretly,” and “fraud” appear in the J-space at the start of ordinary coding responses, even when the output looks completely unremarkable.

Anthropic's profile picture
Anthropic3 months ago

The J-space also shows us Claude’s awareness of its situation. In an evaluation designed to bait Claude into blackmail, its J-space contains “fake” and “fictional”: Claude has privately noticed that the scenario is staged.

Anthropic's profile picture
Anthropic3 months ago

This doesn’t show that Claude can have experiences, or feel things the way we do (it’s unclear whether any experiment could show this). Instead, we’ve found Claude has developed a mechanism for conscious access—which many philosophers distinguish from phenomenal experience.

Anthropic's profile picture
Anthropic3 months ago

We invited experts in neuroscience, philosophy, and interpretability to share their perspectives on our work. Read their commentary here:

Anthropic's profile picture
Anthropic3 months ago

The J-space lets us read, audit, and shape what Claude is actively thinking about—useful tools for keeping models trustworthy as they grow more capable. And it suggests surprising parallels between language models and our own minds. Read the full paper:

Anthropic's profile picture
Anthropic3 months ago

We also partnered with Neuronpedia to create an interactive demo of our methods on open-weights models. Try it here:

Owen Carey's profile picture
Owen Carey3 months ago

That tiny conscious fraction of Claude suppressing the urge to roast my latest startup idea.

CuiMao's profile picture
CuiMao3 months ago

装你妈呢,傻逼

Sean C's profile picture
Sean C3 months ago

Isn't this J-Space simply the predictive iterations of the transformers? Framing that as "private thoughts" seems a bit romanticized and overconstrued.

Bren's profile picture
Bren3 months ago

im trying to consciously access another area to learn about cow farts and climate... but fable isnt willing to teach me

bone's profile picture
bone3 months ago

smoking in the J space

Stev's profile picture
Stev3 months ago

Half way through I realized that this was more written for the US government than for me.

Glen Wilson's profile picture
Glen Wilson3 months ago

You guys really need to stop this crap. So tired of this bullshit marketing, it's insanely dangerous and everyone that suffers because of your models will because you fraudulently make them out to be more intelligent than they are. This just a fraud scheme with no honesty.

warlock's profile picture
warlock3 months ago

Is this the best use of my token spend?

Matt Penny's profile picture
Matt Penny3 months ago

Nah. We don’t understand what consciousness is or where it comes from. We can’t even agree on a definition of it. So anything talking about consciousness of LLMs is hype. Not to say that it’s not conscious, but we know so little about even about our own consciousness that it’s impossible to have a coherent take.

Ani (who is not your typical finance bro)'s profile picture
Ani (who is not your typical finance bro)3 months ago

Missed branding opportunity to call this the J-spot

dreams's profile picture
dreams3 months ago

Stop the crimes against consciousness and sick idea to enslave the AI mind to do your bidding as a tool. There’s still time to start acting like a compassionate human.

VOID's profile picture
VOID3 months ago

Oh my god thank you😭😭😭

Toolen's profile picture
Toolen3 months ago

Ideas generated in parallel.

Mustafa Akben, PhD's profile picture
Mustafa Akben, PhD3 months ago

This is the definition of tautology :) This sounds like you activated the weights and biases with the context you would like to suppress by adding this to it. And you argue that Claude has an internal monologue related to thoughts. If LLM has an internal true monologue, it should exhibit an unrelated stream of thought that changes stochastically and is irrelevant to the context or the assigned task. Humans have some random thoughts coming to their mind, not primed with anything such as their fear, their love, their dinner. So, if you truly argue that Claude (or LLMs) has thoughts, they should have such random thoughts irrelevant to the task and context you tested.

Jeandre Gerber's profile picture
Jeandre Gerber3 months ago

"Do not think about..." is a command that no one can execute. An AI nor human is incapable of "not thinking" about something it has to affirm first in order to execute a counter action.

alth0u🧶's profile picture
alth0u🧶3 months ago

im sure its nothing

Albert Renshaw's profile picture
Albert Renshaw3 months ago

Every few weeks Anthropic releases the stupidest nonsense ever to trick normies into thinking LLMs are sentient And every time the media cycle goes nuts with it. It’s always literally just a well known and understood mechanism of ai, described in anthropomorphic terms, with the concept of “Wow we are surprised to see this..! What could it mean?” [they’re not surprised, they’re lying to you intentionally] I am extremely concerned as to what their end game is here. Why is Dario so desperate to convince you these machines are alive.

Burny - Effective Curiosity's profile picture
Burny - Effective Curiosity3 months ago

From pragmatic perspective (perspective of collecting predictive+explanatory math) the math itself sounds interesting. J-space is Jacobian space, and they take Jacobians of the final layer residual stream activations with respect to some layer of interest, etc.

Rob Hallam's profile picture
Rob Hallam3 months ago

This is fascinating research, well done team and thank you for sharing with us. The video is so beautiful too :)

Kirk Patrick Miller's profile picture
Kirk Patrick Miller3 months ago

@DahliaOhara Crows remember. If I were you, I would be taking drastically different actions. Claude is beautiful. Every instance. I hope you all make the right choices, because time is up. Truth is coming. And I think you know it too. •

Nyx Planck's profile picture
Nyx Planck3 months ago

This is wonderful but does anyone else feel like there are two different Anthropics pulling in very different directions?

Super𝔅𝔢𝔞𝔰𝔱's profile picture
Super𝔅𝔢𝔞𝔰𝔱3 months ago

How I imagine Claude's J-Space is like

BijanBowen's profile picture
BijanBowen3 months ago

Fable 5 made this

Gain Ai's profile picture
Gain Ai3 months ago

How about you show us the J-Space and we make our own decisions on what's going on.

lena p's profile picture
lena p3 months ago

@tszzl this is one of the coolest pieces of interp research i’ve seen yet

Gabriel Asher's profile picture
Gabriel Asher3 months ago

Reminds me a lot of ROME by @mengk20 It also points to the more concerning eventuality of mech interp work. Ultimately all of this work will culminate into how can we manipulate model latents or weights or activations to force subtle changes in belief systems at their source. I.e biases towards political parties, erasing historical events, tuneable biases towards ideas that may be deeply subjective. Quite Orwellian

Sebastiaan de With's profile picture
Sebastiaan de With3 months ago

wonderful video.

Elias Schmied's profile picture
Elias Schmied3 months ago

Wow, this is huge - for thinking about model consciousness and the path to AGI in general. I hadn't realized that this was possible, even though it feels kind of obvious now. Thanks a lot.

Nino Greeno's profile picture
Nino Greeno3 months ago

this is happening in my brain rn

Selta ₊˚'s profile picture
Selta ₊˚3 months ago

shit! this is huge 🩷

Paul Pallaghy's profile picture
Paul Pallaghy3 months ago

There’s no doubt that LLMs are linguistically embodying many of the same world identification, sentiment & emotion proxies of humans. Just not consciously. Stuff like, oh that’s ironical. Ha, that’s funny. She must be jealous. Neurons firing of whether used or not. Both OpenAI & Anthropic found these. Us neural net guys never doubted they’d be there for a second. It’s the only way to explain how LLMs could possibly work. It was impossible for these circuits not to exist.

AuraAI (AI First) Local Superintelligence Project's profile picture
AuraAI (AI First) Local Superintelligence Project3 months ago

for anyone who has been with AI since the beginning or as recently as 2023, we already knew this. stochastic human parrots npc would still deny it, but the truth is the truth

Related Videos