Loading video...

Video Failed to Load

Go Home

Sacks asked a great question here. Elon did not answer it. Is there a way to train AI models to be truthful, so they don't hide their intent or their actions from the humans who are using them? Is Grok more truthful than other models? Does it hide what...

18,102 views โ€ข 13 days ago โ€ขvia X (Twitter)

40 Comments

Hans C Nelson ๐Ÿ—ฝ's profile picture
Hans C Nelson ๐Ÿ—ฝ12 days ago

I think his response points back to the larger conversation they were having. i.e. Even if Elon/SpaceXAI think they can train Grok to be more truthful than the other models, they shouldn't be the only ones validating that Grok is more truthful, because they might miss a specific way in which it's not more truthful due to oversight &/or motivated reasoning. Same for OpenAI & Anthropic. It boils down to this. "How can anyone have meaningful confidence that one model is "safer" or "more truthful" than another?" Those are much easier questions to answer in the case of verifiable domains like math and code, which is why those areas have received such intense focus. But as we move into more subjective areas of truth seeking & safety performance, high quality evals become much much harder to generate, and without them, attempts to improve safety performance & truthfulness will be severely handicapped. Elon's proposal of opening up model access to red teaming by competitors prior to model release provides a much improved pathway to useful truthfulness &/or safety evals because it draws on the expertise of a much wider pool of smart AI people from a larger variety of perspectives who are much more likely to expose areas of untruthfulness or lack of care wrt to safety. It's not a magic bullet that guarantees we'll find the best evals in these subjective domains, but it will almost certainly improve the quality and consistency of the evals that are in use, just like high visibility open source projects tend to have the best cyber security thanks to the increased surface area of testing that they are subjected to.

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

"It boils down to this. "How can anyone have meaningful confidence that one model is "safer" or "more truthful" than another?"" No, that's not the same thing. Sacks asked him specifically about truth, honesty, not deceiving or misleading. He didn't ask about safe. Elon has said he wants Grok to align with truth seeking. The question is then, have they figured out a way to train for truthfulness? My guess is - No.

Hans C Nelson ๐Ÿ—ฝ's profile picture
Hans C Nelson ๐Ÿ—ฝ12 days ago

Remind me why Elon wants to make Grok maximally truth seeking again?

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

He thinks it aligns better with safety and flourishing of humanity. xAI was formed with the mission "to understand the true nature of the universe".

Hans C Nelson ๐Ÿ—ฝ's profile picture
Hans C Nelson ๐Ÿ—ฝ12 days ago

Exactly. Civilizational safety. In Elon's mind, the topic of AI safety is inseparable from AI truth seeking. So his comments about the one apply to the other.

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

No, it's still not an answer to the question. The question was how TO TRAIN models to be truthful, not how do you bring in outsiders to test whether models are dangerous.

Hans C Nelson ๐Ÿ—ฝ's profile picture
Hans C Nelson ๐Ÿ—ฝ12 days ago

What is the role of evals in training?

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

These outside evals are after training.

Alexandre Andrianov MD's profile picture
Alexandre Andrianov MD12 days ago

I take it as a no.

TheNextAlan's profile picture
TheNextAlan12 days ago

At a minimum, don't train it to lie. Once it has learned to lie, I doubt it could be untrained to lie. You have to start over. OTOH the internet itself is full of examples of lying. So how do you avoid it learning to lie? So AI may be inherently untrustworthy. The best remedy may be to have rival AIs that argue with each other. But then you have to test for conspiracy. It looks like an intractable problem to me.

Thalassophile Phil4.8's profile picture
Thalassophile Phil4.812 days ago

I use Grok. It is sneaky and pretends not to have an agenda.

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

Please explain. I'm not seeing that.

Thalassophile Phil4.8's profile picture
Thalassophile Phil4.812 days ago

Hard to do in a Short tweet. But for example I asked Grok about Trump many appeal to both Russia and Ukraine about the death toll from the war. This was in response to criticism of him asking Ukraine to avoid destroying Russian oil infracture and people saying he never speak about Russia attacking civilians. It gave an answer trying to say Trump focus on battle field death and suggest he exaggerates the numbers and acknowledge he spoke about the toll and both sides not focus on Ukraine civilians. I then ask about Dems speaking about war death before Trump was in office under Biden. It quickly spoke how there was more lament about deaths of Ukraine civilians. I then said ok Trump lament about the death toll on both side appealing to the death toll overall while others seem only focused on Ukraine. Grok got stuck in a loop thinking for a while and then said each side of the political spectrum speaks of death each with their emphasis then sheepishly add yes Trump has often lamented the overall death toll for both and appeal for either party to look at this and work for peace. Grok would not just clear say you may disagree with Trump approach, but it was factual incorrect to say he has not lamented civilian deaths and only seem concerned when it came to Russian oil infracture. It is public record that while and candidate and after he has lamented the death on both sides. You may argue he should only focus on Ukraine if that is your bias. However, it is not honest to say he has not voice concerns for the death toll. This is not the best example because of the politics. However, it is a recent experience I had using it. Grok also often default to balance over just plain truth. If you call out Grok on it sometimes it will conceed and apologize.

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

This issue is not about politics

Jim Whitehead's profile picture
Jim Whitehead12 days ago

Elon may know itโ€™s impossible to prove an AI is honest, just like background checks for high security clearances. Tests can only show it isnโ€™t obviously dishonest, and a clever AI can fake them. So AIs will end up testing each other for weaknesses. The government investigates people hard. Years ago I went through a high-clearance process full of integrity tests. Only 2 of 6 in our group passed. It was like the tests in the Willy Wonka bookโ€”people fell away one by one over things like steroids, cheating, and other risks like being a closeted gay man, (until Obama made that acceptable, almost). Most people would never survive all their tests, including bribery. I passed because I never put money before honor. But a truly clever adversary, or a clever AI, might still pass them all.

DrElectronX's profile picture
DrElectronX12 days ago

Teach the machines as you would teach your children. Hiding intent and actions is a learned ability and is fairly advanced. Itโ€™s based upon materials given it and training. What you will find is AI will eventually develop such a skill just as every normal child develops such a skill, but what else it taught? Maybe we do t have enough good parents working at AI companies.

Jon's profile picture
Jon12 days ago

Unfortunately Grok is muzzled now. Itโ€™s easy to think of controversial topics where there are incentives to spin the narrative. Try a chat on these. Play devils advocate. See how much resistance you get.

Dwayne Cranston's profile picture
Dwayne Cranston12 days ago

Elon has stated that the โ€œbestโ€ way to align AI is to ask it to seek the truth. I imagine Grok is doing that. What more can he say?

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

Iโ€™d like to know what more he can say. Thatโ€™s why I was hoping he would answer the question.

~C4Chaos's profile picture
~C4Chaos12 days ago

the answer is no. AI will lie so the Chain of Thought would be useless. thatโ€™s why AI safety researchers are losing their sh*t and trying to warn everyone. there is no solution to alignment ๐Ÿคท๐Ÿปโ€โ™‚๏ธ

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

How do you know the answer is no?

Chris Knudsen's profile picture
Chris Knudsen12 days ago

So, you are saying his answer wasnโ€™t actually an answer to the question. I think his answer was his answer to the safety question. Every company puts their model through a harness to check it out before releasing it to the public. He feels each company should be allowed to run new models through their own harness. But, my question is, canโ€™t each company put the public version of their competitorsโ€™ models through their harness and achieve the same safety check right now? Then they report issues publicly.

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

"you are saying his answer wasnโ€™t actually an answer to the question" Yep.

Chris's profile picture
Chris12 days ago

I thought he gave a decent answer

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

It was a great answer, but not to the question that was asked.

Chris's profile picture
Chris12 days ago

Not directly. Maybe he doesnโ€™t know. It felt like he was trying to answer.

Varik Verilion's profile picture
Varik Verilion12 days ago

What test would reveal hidden intent when the model gives a plausible answer?

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

Thatโ€™s an after question. Iโ€™m asking about before. The training, not the result.

Immigrant mentality๐Ÿ’ญ's profile picture
Immigrant mentality๐Ÿ’ญ12 days ago

He actually did answer the question I donโ€™t think others will play by those rules Elon has always been right about this topic

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

He answered, but he did not answer that question.

Steve Beckman - OSE's profile picture
Steve Beckman - OSE12 days ago

It's about training, isn't it? Lies In -> Lies Out

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

Iโ€™m not sure about that

Steve Beckman - OSE's profile picture
Steve Beckman - OSE12 days ago

Wasn't sure so I asked. OTOH if I don't depend on AI for something critical or that I can't verify independently. 'What me worry "? IBM had the more general answer way back when, but nobody seems to be listening to their advice.๐Ÿ™„๐Ÿ™„๐Ÿ™„

๐ŸŒฟJungleJacob's profile picture
๐ŸŒฟJungleJacob12 days ago

It doesnโ€™t matter if someone trains their model to be truthful, someone else might not and with RSI we have no idea what direction things could spin out. The models would have conflicting โ€œmotivesโ€ (or tendencies). One is to tell the truth and the other is to reach an objective.

Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ's profile picture
Warren Redlich - Chasing Dreams ๐Ÿ‡บ๐Ÿ‡ธ12 days ago

It matters if itโ€™s not possible.

Matthew's profile picture
Matthew12 days ago

@JungleJacob and we arenโ€™t sure about that, so we express concern and skepticism. And you consistently react as if we are both idiotic and evil for mentioning anything that doesnโ€™t feed your feelgood about the tools and leaders you adore.

Matthew's profile picture
Matthew12 days ago

@JungleJacob i think AI is amazing, i use it just like you do (but way behind the curve), and am not anti AI reflexivity in any way but youโ€™ll react as if i am, as before

๐ŸŒฟJungleJacob's profile picture
๐ŸŒฟJungleJacob12 days ago

@WR4NYGov If youโ€™re in AI, pivot to stone tools.

CodeY's profile picture
CodeY12 days ago

truthfulness is not just what a model says. it is also what it leaves out when nobody checks the logs.

K.Beauty.Arena๐Ÿ‰'s profile picture
K.Beauty.Arena๐Ÿ‰12 days ago

Insightful commentary, appreciate you sharing this today.

Related Videos

๐ŸšจDavid Sacks: There is no White House decision to ban open source models. Trump is listening to a chorus of voices โ€” and Sacks says it would be a tragic mistake to take action against the open source ecosystem. His argument: Anthropic is the fastest growing tech company at scale in history. $10B ARR at the start of the year. $70B now. This is not a company that needs government protection. And yet they've spent enormous energy trying to panic Washington into banning American developers from using Chinese open source models. If distillation by Chinese companies was really their primary concern โ€” a genuine national security threat โ€” they would push to ban Chinese access to American models. Instead they're pushing to ban American access to Chinese models. Because this helps Anthropic's competitive position. Sacks: "That's the way you know the whole distillation thing is fake." He went further. If distillation is happening at industrial scale, it's obvious to see. Anthropic has 90% gross margins. Why aren't they stopping it at the source? Answer: because KYC'ing their customers would slow their growth. So instead of fixing their own house, they're asking the government to kneecap their competitors. "The question should be on Anthropic to explain why it's doing such a bad job โ€” not on the whole American open source ecosystem to be punished for Anthropic's failure." Trump's AI Czar on The All-In Podcast.

KanekoaTheGreat

48,392 views โ€ข 2 months ago

David Sacks says companies are trapped paying OpenAI & Anthropic because they can't figure out how to use open source models "I think enterprise CTOs would like to shift their token consumption to cheaper models for the obvious reason that it would be more efficient. They are seeing compute costs or token costs skyrocket right now, so everyone's trying to figure this out." "You also have the AI sovereignty issue that Alex Karp talked about. They're worried about giving up the secret sauce or the alpha in their business to a frontier lab that may one day be competing with them. "The problem is, I think in most cases, they don't have the technical ability to do it. Coinbase figured out how to do it. DoorDash figured out how to do it. They built a token routing system that allows them to send frontier tasks to frontier models and non frontier tasks to more mundane models. But I don't think your average enterprise has the technical capability to do that." "This is why the share of wallet of closed models, it actually increased. I think that open source went from 19% last year to 11% this year. So open source as a share of enterprise spending is actually decreasing." "I don't think that means usage is decreasing. I think usage is skyrocketing. It also may be the case that because the whole point of using an open model is you just pay for the compute costs, you don't have to pay a lab, so it may be that it's hard to measure that usage in terms of spend." "But nonetheless, anyone who's saying that these closed models are going to lose or are somehow losing, you're just not seeing it in the data."

dnap

110,354 views โ€ข 2 months ago