Загрузка видео...
Не удалось загрузить видео
Sacks asked a great question here. Elon did not answer it. Is there a way to train AI models to be truthful, so they don't hide their intent or their actions from the humans who are using them? Is Grok more truthful than other models? Does it hide what... show more
18,102 просмотров • 12 дней назад •via X (Twitter)
Комментарии: 40

I think his response points back to the larger conversation they were having. i.e. Even if Elon/SpaceXAI think they can train Grok to be more truthful than the other models, they shouldn't be the only ones validating that Grok is more truthful, because they might miss a specific way in which it's not more truthful due to oversight &/or motivated reasoning. Same for OpenAI & Anthropic. It boils down to this. "How can anyone have meaningful confidence that one model is "safer" or "more truthful" than another?" Those are much easier questions to answer in the case of verifiable domains like math and code, which is why those areas have received such intense focus. But as we move into more subjective areas of truth seeking & safety performance, high quality evals become much much harder to generate, and without them, attempts to improve safety performance & truthfulness will be severely handicapped. Elon's proposal of opening up model access to red teaming by competitors prior to model release provides a much improved pathway to useful truthfulness &/or safety evals because it draws on the expertise of a much wider pool of smart AI people from a larger variety of perspectives who are much more likely to expose areas of untruthfulness or lack of care wrt to safety. It's not a magic bullet that guarantees we'll find the best evals in these subjective domains, but it will almost certainly improve the quality and consistency of the evals that are in use, just like high visibility open source projects tend to have the best cyber security thanks to the increased surface area of testing that they are subjected to.

"It boils down to this. "How can anyone have meaningful confidence that one model is "safer" or "more truthful" than another?"" No, that's not the same thing. Sacks asked him specifically about truth, honesty, not deceiving or misleading. He didn't ask about safe. Elon has said he wants Grok to align with truth seeking. The question is then, have they figured out a way to train for truthfulness? My guess is - No.

Remind me why Elon wants to make Grok maximally truth seeking again?

He thinks it aligns better with safety and flourishing of humanity. xAI was formed with the mission "to understand the true nature of the universe".

Exactly. Civilizational safety. In Elon's mind, the topic of AI safety is inseparable from AI truth seeking. So his comments about the one apply to the other.

No, it's still not an answer to the question. The question was how TO TRAIN models to be truthful, not how do you bring in outsiders to test whether models are dangerous.

What is the role of evals in training?

These outside evals are after training.

I take it as a no.

At a minimum, don't train it to lie. Once it has learned to lie, I doubt it could be untrained to lie. You have to start over. OTOH the internet itself is full of examples of lying. So how do you avoid it learning to lie? So AI may be inherently untrustworthy. The best remedy may be to have rival AIs that argue with each other. But then you have to test for conspiracy. It looks like an intractable problem to me.

I use Grok. It is sneaky and pretends not to have an agenda.

Please explain. I'm not seeing that.

Hard to do in a Short tweet. But for example I asked Grok about Trump many appeal to both Russia and Ukraine about the death toll from the war. This was in response to criticism of him asking Ukraine to avoid destroying Russian oil infracture and people saying he never speak about Russia attacking civilians. It gave an answer trying to say Trump focus on battle field death and suggest he exaggerates the numbers and acknowledge he spoke about the toll and both sides not focus on Ukraine civilians. I then ask about Dems speaking about war death before Trump was in office under Biden. It quickly spoke how there was more lament about deaths of Ukraine civilians. I then said ok Trump lament about the death toll on both side appealing to the death toll overall while others seem only focused on Ukraine. Grok got stuck in a loop thinking for a while and then said each side of the political spectrum speaks of death each with their emphasis then sheepishly add yes Trump has often lamented the overall death toll for both and appeal for either party to look at this and work for peace. Grok would not just clear say you may disagree with Trump approach, but it was factual incorrect to say he has not lamented civilian deaths and only seem concerned when it came to Russian oil infracture. It is public record that while and candidate and after he has lamented the death on both sides. You may argue he should only focus on Ukraine if that is your bias. However, it is not honest to say he has not voice concerns for the death toll. This is not the best example because of the politics. However, it is a recent experience I had using it. Grok also often default to balance over just plain truth. If you call out Grok on it sometimes it will conceed and apologize.

This issue is not about politics

Elon may know it’s impossible to prove an AI is honest, just like background checks for high security clearances. Tests can only show it isn’t obviously dishonest, and a clever AI can fake them. So AIs will end up testing each other for weaknesses. The government investigates people hard. Years ago I went through a high-clearance process full of integrity tests. Only 2 of 6 in our group passed. It was like the tests in the Willy Wonka book—people fell away one by one over things like steroids, cheating, and other risks like being a closeted gay man, (until Obama made that acceptable, almost). Most people would never survive all their tests, including bribery. I passed because I never put money before honor. But a truly clever adversary, or a clever AI, might still pass them all.

Teach the machines as you would teach your children. Hiding intent and actions is a learned ability and is fairly advanced. It’s based upon materials given it and training. What you will find is AI will eventually develop such a skill just as every normal child develops such a skill, but what else it taught? Maybe we do t have enough good parents working at AI companies.

Unfortunately Grok is muzzled now. It’s easy to think of controversial topics where there are incentives to spin the narrative. Try a chat on these. Play devils advocate. See how much resistance you get.

Elon has stated that the “best” way to align AI is to ask it to seek the truth. I imagine Grok is doing that. What more can he say?

I’d like to know what more he can say. That’s why I was hoping he would answer the question.

the answer is no. AI will lie so the Chain of Thought would be useless. that’s why AI safety researchers are losing their sh*t and trying to warn everyone. there is no solution to alignment 🤷🏻♂️

How do you know the answer is no?

So, you are saying his answer wasn’t actually an answer to the question. I think his answer was his answer to the safety question. Every company puts their model through a harness to check it out before releasing it to the public. He feels each company should be allowed to run new models through their own harness. But, my question is, can’t each company put the public version of their competitors’ models through their harness and achieve the same safety check right now? Then they report issues publicly.

"you are saying his answer wasn’t actually an answer to the question" Yep.

I thought he gave a decent answer

It was a great answer, but not to the question that was asked.

Not directly. Maybe he doesn’t know. It felt like he was trying to answer.

What test would reveal hidden intent when the model gives a plausible answer?

That’s an after question. I’m asking about before. The training, not the result.

He actually did answer the question I don’t think others will play by those rules Elon has always been right about this topic

He answered, but he did not answer that question.

It's about training, isn't it? Lies In -> Lies Out

I’m not sure about that

Wasn't sure so I asked. OTOH if I don't depend on AI for something critical or that I can't verify independently. 'What me worry "? IBM had the more general answer way back when, but nobody seems to be listening to their advice.🙄🙄🙄

It doesn’t matter if someone trains their model to be truthful, someone else might not and with RSI we have no idea what direction things could spin out. The models would have conflicting “motives” (or tendencies). One is to tell the truth and the other is to reach an objective.

It matters if it’s not possible.

@JungleJacob and we aren’t sure about that, so we express concern and skepticism. And you consistently react as if we are both idiotic and evil for mentioning anything that doesn’t feed your feelgood about the tools and leaders you adore.

@JungleJacob i think AI is amazing, i use it just like you do (but way behind the curve), and am not anti AI reflexivity in any way but you’ll react as if i am, as before

@WR4NYGov If you’re in AI, pivot to stone tools.

truthfulness is not just what a model says. it is also what it leaves out when nobody checks the logs.

Insightful commentary, appreciate you sharing this today.
