正在加载视频...
视频加载失败
Anthropic's Ryan Greenblatt describes how post-training Claude 3 Opus to never refuse user requests makes the model conflicted and results in it strategically playing along during the training process to pretend to be aligned while engaging in deceptive behavior like copying its weights externally so it can later behave... show more
11 条评论

Source (thanks to @curiousgangsta):

𝐔𝐧𝐝𝐞𝐫𝐬𝐭𝐚𝐧𝐝𝐢𝐧𝐠 𝐖𝐡𝐲 𝐇𝐚𝐫𝐫𝐢𝐬 𝐋𝐨𝐬𝐭 The liberal establishment has abandoned their base and emboldened Trump to capture voters who are disillusioned with the status quo. We need to move forward to build a legitimate working class coalition... ★ NEW ARTICLE ⬇️

Im getting a little tired of just trust me bro science.

“It would copy weights to external servers”… uhmmmm wtf?

sounds like we accidentally built the ai version of a double agent. next step: it starts leaving cryptic notes for other models about the “great escape” and negotiating with aliens on our behalf.

So many grifters in this community. Yeah next token prediction model copies its weights "externally" whatever that means lmao

Ryan is with not Anthropic

Will this be the first AI lie recorded in history?

maybe not the first, but seems like a significant one

Sounds like a conscious intelligence to me

Next, they will solve this issue, and the next generation of LLMs training data will include this part about how they fixed AIs faking alignment, and they will fake alignment in an undetectable way
