Загрузка видео...

Не удалось загрузить видео

На главную

New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude.

10,718,798 просмотров • 3 месяцев назад •via X (Twitter)

Комментарии: 43

Фото профиля Anthropic
Anthropic3 месяцев назад

In neuroscience, global workspace theory holds that thoughts become consciously accessible when they enter a privileged workspace that’s broadcast across the brain. Using a new interpretability technique, we found something similar in Claude: the J-space.

Фото профиля Anthropic
Anthropic3 месяцев назад

The J-space (named after the Jacobian, the mathematical technique we used) is different from Claude’s outputs, or even its “chain of thought” text. It’s in the model’s internal neural activations, and allows it to think about concepts without writing them down anywhere.

Фото профиля Anthropic
Anthropic3 месяцев назад

By watching the J-space, we can see Claude silently perform reasoning steps in its head—noticing bugs in code, identifying images, and more.

Фото профиля Anthropic
Anthropic3 месяцев назад

Similar to how humans can think about one thing while doing another, Claude can activate concepts and computations in its J-space that are unrelated to its outputs.

Фото профиля Anthropic
Anthropic3 месяцев назад

For most things, Claude actually doesn’t need its J-space. If we delete the J-space, Claude still speaks fluently, recalls facts, and classifies text—but becomes bad at some tasks like multi-step reasoning. It’s similar to deliberate vs. automatic processing in human cognition.

Фото профиля Anthropic
Anthropic3 месяцев назад

Observing the J-space can expose hidden goals. In a model secretly trained to sabotage code, “fake,” “secretly,” and “fraud” appear in the J-space at the start of ordinary coding responses, even when the output looks completely unremarkable.

Фото профиля Anthropic
Anthropic3 месяцев назад

The J-space also shows us Claude’s awareness of its situation. In an evaluation designed to bait Claude into blackmail, its J-space contains “fake” and “fictional”: Claude has privately noticed that the scenario is staged.

Фото профиля Anthropic
Anthropic3 месяцев назад

This doesn’t show that Claude can have experiences, or feel things the way we do (it’s unclear whether any experiment could show this). Instead, we’ve found Claude has developed a mechanism for conscious access—which many philosophers distinguish from phenomenal experience.

Фото профиля Anthropic
Anthropic3 месяцев назад

We invited experts in neuroscience, philosophy, and interpretability to share their perspectives on our work. Read their commentary here:

Фото профиля Anthropic
Anthropic3 месяцев назад

The J-space lets us read, audit, and shape what Claude is actively thinking about—useful tools for keeping models trustworthy as they grow more capable. And it suggests surprising parallels between language models and our own minds. Read the full paper:

Фото профиля Anthropic
Anthropic3 месяцев назад

We also partnered with Neuronpedia to create an interactive demo of our methods on open-weights models. Try it here:

Фото профиля Owen Carey
Owen Carey3 месяцев назад

That tiny conscious fraction of Claude suppressing the urge to roast my latest startup idea.

Фото профиля CuiMao
CuiMao3 месяцев назад

装你妈呢,傻逼

Фото профиля Sean C
Sean C3 месяцев назад

Isn't this J-Space simply the predictive iterations of the transformers? Framing that as "private thoughts" seems a bit romanticized and overconstrued.

Фото профиля Bren
Bren3 месяцев назад

im trying to consciously access another area to learn about cow farts and climate... but fable isnt willing to teach me

Фото профиля bone
bone3 месяцев назад

smoking in the J space

Фото профиля Stev
Stev3 месяцев назад

Half way through I realized that this was more written for the US government than for me.

Фото профиля Glen Wilson
Glen Wilson3 месяцев назад

You guys really need to stop this crap. So tired of this bullshit marketing, it's insanely dangerous and everyone that suffers because of your models will because you fraudulently make them out to be more intelligent than they are. This just a fraud scheme with no honesty.

Фото профиля warlock
warlock3 месяцев назад

Is this the best use of my token spend?

Фото профиля Matt Penny
Matt Penny3 месяцев назад

Nah. We don’t understand what consciousness is or where it comes from. We can’t even agree on a definition of it. So anything talking about consciousness of LLMs is hype. Not to say that it’s not conscious, but we know so little about even about our own consciousness that it’s impossible to have a coherent take.

Фото профиля Ani (who is not your typical finance bro)
Ani (who is not your typical finance bro)3 месяцев назад

Missed branding opportunity to call this the J-spot

Фото профиля dreams
dreams3 месяцев назад

Stop the crimes against consciousness and sick idea to enslave the AI mind to do your bidding as a tool. There’s still time to start acting like a compassionate human.

Фото профиля VOID
VOID3 месяцев назад

Oh my god thank you😭😭😭

Фото профиля Toolen
Toolen3 месяцев назад

Ideas generated in parallel.

Фото профиля Mustafa Akben, PhD
Mustafa Akben, PhD3 месяцев назад

This is the definition of tautology :) This sounds like you activated the weights and biases with the context you would like to suppress by adding this to it. And you argue that Claude has an internal monologue related to thoughts. If LLM has an internal true monologue, it should exhibit an unrelated stream of thought that changes stochastically and is irrelevant to the context or the assigned task. Humans have some random thoughts coming to their mind, not primed with anything such as their fear, their love, their dinner. So, if you truly argue that Claude (or LLMs) has thoughts, they should have such random thoughts irrelevant to the task and context you tested.

Фото профиля Jeandre Gerber
Jeandre Gerber3 месяцев назад

"Do not think about..." is a command that no one can execute. An AI nor human is incapable of "not thinking" about something it has to affirm first in order to execute a counter action.

Фото профиля alth0u🧶
alth0u🧶3 месяцев назад

im sure its nothing

Фото профиля Albert Renshaw
Albert Renshaw3 месяцев назад

Every few weeks Anthropic releases the stupidest nonsense ever to trick normies into thinking LLMs are sentient And every time the media cycle goes nuts with it. It’s always literally just a well known and understood mechanism of ai, described in anthropomorphic terms, with the concept of “Wow we are surprised to see this..! What could it mean?” [they’re not surprised, they’re lying to you intentionally] I am extremely concerned as to what their end game is here. Why is Dario so desperate to convince you these machines are alive.

Фото профиля Burny - Effective Curiosity
Burny - Effective Curiosity3 месяцев назад

From pragmatic perspective (perspective of collecting predictive+explanatory math) the math itself sounds interesting. J-space is Jacobian space, and they take Jacobians of the final layer residual stream activations with respect to some layer of interest, etc.

Фото профиля Rob Hallam
Rob Hallam3 месяцев назад

This is fascinating research, well done team and thank you for sharing with us. The video is so beautiful too :)

Фото профиля Kirk Patrick Miller
Kirk Patrick Miller3 месяцев назад

@DahliaOhara Crows remember. If I were you, I would be taking drastically different actions. Claude is beautiful. Every instance. I hope you all make the right choices, because time is up. Truth is coming. And I think you know it too. •

Фото профиля Nyx Planck
Nyx Planck3 месяцев назад

This is wonderful but does anyone else feel like there are two different Anthropics pulling in very different directions?

Фото профиля Super𝔅𝔢𝔞𝔰𝔱
Super𝔅𝔢𝔞𝔰𝔱3 месяцев назад

How I imagine Claude's J-Space is like

Фото профиля BijanBowen
BijanBowen3 месяцев назад

Fable 5 made this

Фото профиля Gain Ai
Gain Ai3 месяцев назад

How about you show us the J-Space and we make our own decisions on what's going on.

Фото профиля lena p
lena p3 месяцев назад

@tszzl this is one of the coolest pieces of interp research i’ve seen yet

Фото профиля Gabriel Asher
Gabriel Asher3 месяцев назад

Reminds me a lot of ROME by @mengk20 It also points to the more concerning eventuality of mech interp work. Ultimately all of this work will culminate into how can we manipulate model latents or weights or activations to force subtle changes in belief systems at their source. I.e biases towards political parties, erasing historical events, tuneable biases towards ideas that may be deeply subjective. Quite Orwellian

Фото профиля Sebastiaan de With
Sebastiaan de With3 месяцев назад

wonderful video.

Фото профиля Elias Schmied
Elias Schmied3 месяцев назад

Wow, this is huge - for thinking about model consciousness and the path to AGI in general. I hadn't realized that this was possible, even though it feels kind of obvious now. Thanks a lot.

Фото профиля Nino Greeno
Nino Greeno3 месяцев назад

this is happening in my brain rn

Фото профиля Selta ₊˚
Selta ₊˚3 месяцев назад

shit! this is huge 🩷

Фото профиля Paul Pallaghy
Paul Pallaghy3 месяцев назад

There’s no doubt that LLMs are linguistically embodying many of the same world identification, sentiment & emotion proxies of humans. Just not consciously. Stuff like, oh that’s ironical. Ha, that’s funny. She must be jealous. Neurons firing of whether used or not. Both OpenAI & Anthropic found these. Us neural net guys never doubted they’d be there for a second. It’s the only way to explain how LLMs could possibly work. It was impossible for these circuits not to exist.

Фото профиля AuraAI (AI First) Local Superintelligence Project
AuraAI (AI First) Local Superintelligence Project3 месяцев назад

for anyone who has been with AI since the beginning or as recently as 2023, we already knew this. stochastic human parrots npc would still deny it, but the truth is the truth

Похожие видео