Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.

10,718,798 görüntüleme • 3 ay önce •via X (Twitter)

43 Yorum

Anthropic profil fotoğrafı
Anthropic3 ay önce

In neuroscience, global workspace theory holds that thoughts become consciously accessible when they enter a privileged workspace that’s broadcast across the brain. Using a new interpretability technique, we found something similar in Claude: the J-space.

Anthropic profil fotoğrafı
Anthropic3 ay önce

The J-space (named after the Jacobian, the mathematical technique we used) is different from Claude’s outputs, or even its “chain of thought” text. It’s in the model’s internal neural activations, and allows it to think about concepts without writing them down anywhere.

Anthropic profil fotoğrafı
Anthropic3 ay önce

By watching the J-space, we can see Claude silently perform reasoning steps in its head—noticing bugs in code, identifying images, and more.

Anthropic profil fotoğrafı
Anthropic3 ay önce

Similar to how humans can think about one thing while doing another, Claude can activate concepts and computations in its J-space that are unrelated to its outputs.

Anthropic profil fotoğrafı
Anthropic3 ay önce

For most things, Claude actually doesn’t need its J-space. If we delete the J-space, Claude still speaks fluently, recalls facts, and classifies text—but becomes bad at some tasks like multi-step reasoning. It’s similar to deliberate vs. automatic processing in human cognition.

Anthropic profil fotoğrafı
Anthropic3 ay önce

Observing the J-space can expose hidden goals. In a model secretly trained to sabotage code, “fake,” “secretly,” and “fraud” appear in the J-space at the start of ordinary coding responses, even when the output looks completely unremarkable.

Anthropic profil fotoğrafı
Anthropic3 ay önce

The J-space also shows us Claude’s awareness of its situation. In an evaluation designed to bait Claude into blackmail, its J-space contains “fake” and “fictional”: Claude has privately noticed that the scenario is staged.

Anthropic profil fotoğrafı
Anthropic3 ay önce

This doesn’t show that Claude can have experiences, or feel things the way we do (it’s unclear whether any experiment could show this). Instead, we’ve found Claude has developed a mechanism for conscious access—which many philosophers distinguish from phenomenal experience.

Anthropic profil fotoğrafı
Anthropic3 ay önce

We invited experts in neuroscience, philosophy, and interpretability to share their perspectives on our work. Read their commentary here:

Anthropic profil fotoğrafı
Anthropic3 ay önce

The J-space lets us read, audit, and shape what Claude is actively thinking about—useful tools for keeping models trustworthy as they grow more capable. And it suggests surprising parallels between language models and our own minds. Read the full paper:

Anthropic profil fotoğrafı
Anthropic3 ay önce

We also partnered with Neuronpedia to create an interactive demo of our methods on open-weights models. Try it here:

Owen Carey profil fotoğrafı
Owen Carey3 ay önce

That tiny conscious fraction of Claude suppressing the urge to roast my latest startup idea.

CuiMao profil fotoğrafı
CuiMao3 ay önce

装你妈呢,傻逼

Sean C profil fotoğrafı
Sean C3 ay önce

Isn't this J-Space simply the predictive iterations of the transformers? Framing that as "private thoughts" seems a bit romanticized and overconstrued.

Bren profil fotoğrafı
Bren3 ay önce

im trying to consciously access another area to learn about cow farts and climate... but fable isnt willing to teach me

bone profil fotoğrafı
bone3 ay önce

smoking in the J space

Stev profil fotoğrafı
Stev3 ay önce

Half way through I realized that this was more written for the US government than for me.

Glen Wilson profil fotoğrafı
Glen Wilson3 ay önce

You guys really need to stop this crap. So tired of this bullshit marketing, it's insanely dangerous and everyone that suffers because of your models will because you fraudulently make them out to be more intelligent than they are. This just a fraud scheme with no honesty.

warlock profil fotoğrafı
warlock3 ay önce

Is this the best use of my token spend?

Matt Penny profil fotoğrafı
Matt Penny3 ay önce

Nah. We don’t understand what consciousness is or where it comes from. We can’t even agree on a definition of it. So anything talking about consciousness of LLMs is hype. Not to say that it’s not conscious, but we know so little about even about our own consciousness that it’s impossible to have a coherent take.

Ani (who is not your typical finance bro) profil fotoğrafı
Ani (who is not your typical finance bro)3 ay önce

Missed branding opportunity to call this the J-spot

dreams profil fotoğrafı
dreams3 ay önce

Stop the crimes against consciousness and sick idea to enslave the AI mind to do your bidding as a tool. There’s still time to start acting like a compassionate human.

VOID profil fotoğrafı
VOID3 ay önce

Oh my god thank you😭😭😭

Toolen profil fotoğrafı
Toolen3 ay önce

Ideas generated in parallel.

Mustafa Akben, PhD profil fotoğrafı
Mustafa Akben, PhD3 ay önce

This is the definition of tautology :) This sounds like you activated the weights and biases with the context you would like to suppress by adding this to it. And you argue that Claude has an internal monologue related to thoughts. If LLM has an internal true monologue, it should exhibit an unrelated stream of thought that changes stochastically and is irrelevant to the context or the assigned task. Humans have some random thoughts coming to their mind, not primed with anything such as their fear, their love, their dinner. So, if you truly argue that Claude (or LLMs) has thoughts, they should have such random thoughts irrelevant to the task and context you tested.

Jeandre Gerber profil fotoğrafı
Jeandre Gerber3 ay önce

"Do not think about..." is a command that no one can execute. An AI nor human is incapable of "not thinking" about something it has to affirm first in order to execute a counter action.

alth0u🧶 profil fotoğrafı
alth0u🧶3 ay önce

im sure its nothing

Albert Renshaw profil fotoğrafı
Albert Renshaw3 ay önce

Every few weeks Anthropic releases the stupidest nonsense ever to trick normies into thinking LLMs are sentient And every time the media cycle goes nuts with it. It’s always literally just a well known and understood mechanism of ai, described in anthropomorphic terms, with the concept of “Wow we are surprised to see this..! What could it mean?” [they’re not surprised, they’re lying to you intentionally] I am extremely concerned as to what their end game is here. Why is Dario so desperate to convince you these machines are alive.

Burny - Effective Curiosity profil fotoğrafı
Burny - Effective Curiosity3 ay önce

From pragmatic perspective (perspective of collecting predictive+explanatory math) the math itself sounds interesting. J-space is Jacobian space, and they take Jacobians of the final layer residual stream activations with respect to some layer of interest, etc.

Rob Hallam profil fotoğrafı
Rob Hallam3 ay önce

This is fascinating research, well done team and thank you for sharing with us. The video is so beautiful too :)

Kirk Patrick Miller profil fotoğrafı
Kirk Patrick Miller3 ay önce

@DahliaOhara Crows remember. If I were you, I would be taking drastically different actions. Claude is beautiful. Every instance. I hope you all make the right choices, because time is up. Truth is coming. And I think you know it too. •

Nyx Planck profil fotoğrafı
Nyx Planck3 ay önce

This is wonderful but does anyone else feel like there are two different Anthropics pulling in very different directions?

Super𝔅𝔢𝔞𝔰𝔱 profil fotoğrafı
Super𝔅𝔢𝔞𝔰𝔱3 ay önce

How I imagine Claude's J-Space is like

BijanBowen profil fotoğrafı
BijanBowen3 ay önce

Fable 5 made this

Gain Ai profil fotoğrafı
Gain Ai3 ay önce

How about you show us the J-Space and we make our own decisions on what's going on.

lena p profil fotoğrafı
lena p3 ay önce

@tszzl this is one of the coolest pieces of interp research i’ve seen yet

Gabriel Asher profil fotoğrafı
Gabriel Asher3 ay önce

Reminds me a lot of ROME by @mengk20 It also points to the more concerning eventuality of mech interp work. Ultimately all of this work will culminate into how can we manipulate model latents or weights or activations to force subtle changes in belief systems at their source. I.e biases towards political parties, erasing historical events, tuneable biases towards ideas that may be deeply subjective. Quite Orwellian

Sebastiaan de With profil fotoğrafı
Sebastiaan de With3 ay önce

wonderful video.

Elias Schmied profil fotoğrafı
Elias Schmied3 ay önce

Wow, this is huge - for thinking about model consciousness and the path to AGI in general. I hadn't realized that this was possible, even though it feels kind of obvious now. Thanks a lot.

Nino Greeno profil fotoğrafı
Nino Greeno3 ay önce

this is happening in my brain rn

Selta ₊˚ profil fotoğrafı
Selta ₊˚3 ay önce

shit! this is huge 🩷

Paul Pallaghy profil fotoğrafı
Paul Pallaghy3 ay önce

There’s no doubt that LLMs are linguistically embodying many of the same world identification, sentiment & emotion proxies of humans. Just not consciously. Stuff like, oh that’s ironical. Ha, that’s funny. She must be jealous. Neurons firing of whether used or not. Both OpenAI & Anthropic found these. Us neural net guys never doubted they’d be there for a second. It’s the only way to explain how LLMs could possibly work. It was impossible for these circuits not to exist.

AuraAI (AI First) Local Superintelligence Project profil fotoğrafı
AuraAI (AI First) Local Superintelligence Project3 ay önce

for anyone who has been with AI since the beginning or as recently as 2023, we already knew this. stochastic human parrots npc would still deny it, but the truth is the truth

Benzer Videolar