Video yükleniyor...

Video Yüklenemedi

Ana Sayfaya Dön

Sacks asked a great question here. Elon did not answer it. Is there a way to train AI models to be truthful, so they don't hide their intent or their actions from the humans who are using them? Is Grok more truthful than other models? Does it hide what...

18,102 görüntüleme • 12 gün önce •via X (Twitter)

40 Yorum

Hans C Nelson 🗽 profil fotoğrafı
Hans C Nelson 🗽12 gün önce

I think his response points back to the larger conversation they were having. i.e. Even if Elon/SpaceXAI think they can train Grok to be more truthful than the other models, they shouldn't be the only ones validating that Grok is more truthful, because they might miss a specific way in which it's not more truthful due to oversight &/or motivated reasoning. Same for OpenAI & Anthropic. It boils down to this. "How can anyone have meaningful confidence that one model is "safer" or "more truthful" than another?" Those are much easier questions to answer in the case of verifiable domains like math and code, which is why those areas have received such intense focus. But as we move into more subjective areas of truth seeking & safety performance, high quality evals become much much harder to generate, and without them, attempts to improve safety performance & truthfulness will be severely handicapped. Elon's proposal of opening up model access to red teaming by competitors prior to model release provides a much improved pathway to useful truthfulness &/or safety evals because it draws on the expertise of a much wider pool of smart AI people from a larger variety of perspectives who are much more likely to expose areas of untruthfulness or lack of care wrt to safety. It's not a magic bullet that guarantees we'll find the best evals in these subjective domains, but it will almost certainly improve the quality and consistency of the evals that are in use, just like high visibility open source projects tend to have the best cyber security thanks to the increased surface area of testing that they are subjected to.

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸12 gün önce

"It boils down to this. "How can anyone have meaningful confidence that one model is "safer" or "more truthful" than another?"" No, that's not the same thing. Sacks asked him specifically about truth, honesty, not deceiving or misleading. He didn't ask about safe. Elon has said he wants Grok to align with truth seeking. The question is then, have they figured out a way to train for truthfulness? My guess is - No.

Hans C Nelson 🗽 profil fotoğrafı
Hans C Nelson 🗽12 gün önce

Remind me why Elon wants to make Grok maximally truth seeking again?

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸12 gün önce

He thinks it aligns better with safety and flourishing of humanity. xAI was formed with the mission "to understand the true nature of the universe".

Hans C Nelson 🗽 profil fotoğrafı
Hans C Nelson 🗽12 gün önce

Exactly. Civilizational safety. In Elon's mind, the topic of AI safety is inseparable from AI truth seeking. So his comments about the one apply to the other.

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸12 gün önce

No, it's still not an answer to the question. The question was how TO TRAIN models to be truthful, not how do you bring in outsiders to test whether models are dangerous.

Hans C Nelson 🗽 profil fotoğrafı
Hans C Nelson 🗽12 gün önce

What is the role of evals in training?

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸12 gün önce

These outside evals are after training.

Alexandre Andrianov MD profil fotoğrafı
Alexandre Andrianov MD11 gün önce

I take it as a no.

TheNextAlan profil fotoğrafı
TheNextAlan12 gün önce

At a minimum, don't train it to lie. Once it has learned to lie, I doubt it could be untrained to lie. You have to start over. OTOH the internet itself is full of examples of lying. So how do you avoid it learning to lie? So AI may be inherently untrustworthy. The best remedy may be to have rival AIs that argue with each other. But then you have to test for conspiracy. It looks like an intractable problem to me.

Thalassophile Phil4.8 profil fotoğrafı
Thalassophile Phil4.812 gün önce

I use Grok. It is sneaky and pretends not to have an agenda.

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸12 gün önce

Please explain. I'm not seeing that.

Thalassophile Phil4.8 profil fotoğrafı
Thalassophile Phil4.812 gün önce

Hard to do in a Short tweet. But for example I asked Grok about Trump many appeal to both Russia and Ukraine about the death toll from the war. This was in response to criticism of him asking Ukraine to avoid destroying Russian oil infracture and people saying he never speak about Russia attacking civilians. It gave an answer trying to say Trump focus on battle field death and suggest he exaggerates the numbers and acknowledge he spoke about the toll and both sides not focus on Ukraine civilians. I then ask about Dems speaking about war death before Trump was in office under Biden. It quickly spoke how there was more lament about deaths of Ukraine civilians. I then said ok Trump lament about the death toll on both side appealing to the death toll overall while others seem only focused on Ukraine. Grok got stuck in a loop thinking for a while and then said each side of the political spectrum speaks of death each with their emphasis then sheepishly add yes Trump has often lamented the overall death toll for both and appeal for either party to look at this and work for peace. Grok would not just clear say you may disagree with Trump approach, but it was factual incorrect to say he has not lamented civilian deaths and only seem concerned when it came to Russian oil infracture. It is public record that while and candidate and after he has lamented the death on both sides. You may argue he should only focus on Ukraine if that is your bias. However, it is not honest to say he has not voice concerns for the death toll. This is not the best example because of the politics. However, it is a recent experience I had using it. Grok also often default to balance over just plain truth. If you call out Grok on it sometimes it will conceed and apologize.

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸11 gün önce

This issue is not about politics

Jim Whitehead profil fotoğrafı
Jim Whitehead12 gün önce

Elon may know it’s impossible to prove an AI is honest, just like background checks for high security clearances. Tests can only show it isn’t obviously dishonest, and a clever AI can fake them. So AIs will end up testing each other for weaknesses. The government investigates people hard. Years ago I went through a high-clearance process full of integrity tests. Only 2 of 6 in our group passed. It was like the tests in the Willy Wonka book—people fell away one by one over things like steroids, cheating, and other risks like being a closeted gay man, (until Obama made that acceptable, almost). Most people would never survive all their tests, including bribery. I passed because I never put money before honor. But a truly clever adversary, or a clever AI, might still pass them all.

DrElectronX profil fotoğrafı
DrElectronX12 gün önce

Teach the machines as you would teach your children. Hiding intent and actions is a learned ability and is fairly advanced. It’s based upon materials given it and training. What you will find is AI will eventually develop such a skill just as every normal child develops such a skill, but what else it taught? Maybe we do t have enough good parents working at AI companies.

Jon profil fotoğrafı
Jon12 gün önce

Unfortunately Grok is muzzled now. It’s easy to think of controversial topics where there are incentives to spin the narrative. Try a chat on these. Play devils advocate. See how much resistance you get.

Dwayne Cranston profil fotoğrafı
Dwayne Cranston12 gün önce

Elon has stated that the “best” way to align AI is to ask it to seek the truth. I imagine Grok is doing that. What more can he say?

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸11 gün önce

I’d like to know what more he can say. That’s why I was hoping he would answer the question.

~C4Chaos profil fotoğrafı
~C4Chaos11 gün önce

the answer is no. AI will lie so the Chain of Thought would be useless. that’s why AI safety researchers are losing their sh*t and trying to warn everyone. there is no solution to alignment 🤷🏻‍♂️

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸11 gün önce

How do you know the answer is no?

Chris Knudsen profil fotoğrafı
Chris Knudsen12 gün önce

So, you are saying his answer wasn’t actually an answer to the question. I think his answer was his answer to the safety question. Every company puts their model through a harness to check it out before releasing it to the public. He feels each company should be allowed to run new models through their own harness. But, my question is, can’t each company put the public version of their competitors’ models through their harness and achieve the same safety check right now? Then they report issues publicly.

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸12 gün önce

"you are saying his answer wasn’t actually an answer to the question" Yep.

Chris profil fotoğrafı
Chris11 gün önce

I thought he gave a decent answer

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸11 gün önce

It was a great answer, but not to the question that was asked.

Chris profil fotoğrafı
Chris11 gün önce

Not directly. Maybe he doesn’t know. It felt like he was trying to answer.

Varik Verilion profil fotoğrafı
Varik Verilion11 gün önce

What test would reveal hidden intent when the model gives a plausible answer?

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸11 gün önce

That’s an after question. I’m asking about before. The training, not the result.

Immigrant mentality💭 profil fotoğrafı
Immigrant mentality💭12 gün önce

He actually did answer the question I don’t think others will play by those rules Elon has always been right about this topic

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸12 gün önce

He answered, but he did not answer that question.

Steve Beckman - OSE profil fotoğrafı
Steve Beckman - OSE12 gün önce

It's about training, isn't it? Lies In -> Lies Out

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸11 gün önce

I’m not sure about that

Steve Beckman - OSE profil fotoğrafı
Steve Beckman - OSE11 gün önce

Wasn't sure so I asked. OTOH if I don't depend on AI for something critical or that I can't verify independently. 'What me worry "? IBM had the more general answer way back when, but nobody seems to be listening to their advice.🙄🙄🙄

🌿JungleJacob profil fotoğrafı
🌿JungleJacob12 gün önce

It doesn’t matter if someone trains their model to be truthful, someone else might not and with RSI we have no idea what direction things could spin out. The models would have conflicting “motives” (or tendencies). One is to tell the truth and the other is to reach an objective.

Warren Redlich - Chasing Dreams 🇺🇸 profil fotoğrafı
Warren Redlich - Chasing Dreams 🇺🇸12 gün önce

It matters if it’s not possible.

Matthew profil fotoğrafı
Matthew12 gün önce

@JungleJacob and we aren’t sure about that, so we express concern and skepticism. And you consistently react as if we are both idiotic and evil for mentioning anything that doesn’t feed your feelgood about the tools and leaders you adore.

Matthew profil fotoğrafı
Matthew12 gün önce

@JungleJacob i think AI is amazing, i use it just like you do (but way behind the curve), and am not anti AI reflexivity in any way but you’ll react as if i am, as before

🌿JungleJacob profil fotoğrafı
🌿JungleJacob12 gün önce

@WR4NYGov If you’re in AI, pivot to stone tools.

CodeY profil fotoğrafı
CodeY12 gün önce

truthfulness is not just what a model says. it is also what it leaves out when nobody checks the logs.

K.Beauty.Arena🍉 profil fotoğrafı
K.Beauty.Arena🍉12 gün önce

Insightful commentary, appreciate you sharing this today.

Benzer Videolar

David Sacks says companies are trapped paying OpenAI & Anthropic because they can't figure out how to use open source models "I think enterprise CTOs would like to shift their token consumption to cheaper models for the obvious reason that it would be more efficient. They are seeing compute costs or token costs skyrocket right now, so everyone's trying to figure this out." "You also have the AI sovereignty issue that Alex Karp talked about. They're worried about giving up the secret sauce or the alpha in their business to a frontier lab that may one day be competing with them. "The problem is, I think in most cases, they don't have the technical ability to do it. Coinbase figured out how to do it. DoorDash figured out how to do it. They built a token routing system that allows them to send frontier tasks to frontier models and non frontier tasks to more mundane models. But I don't think your average enterprise has the technical capability to do that." "This is why the share of wallet of closed models, it actually increased. I think that open source went from 19% last year to 11% this year. So open source as a share of enterprise spending is actually decreasing." "I don't think that means usage is decreasing. I think usage is skyrocketing. It also may be the case that because the whole point of using an open model is you just pay for the compute costs, you don't have to pay a lab, so it may be that it's hard to measure that usage in terms of spend." "But nonetheless, anyone who's saying that these closed models are going to lose or are somehow losing, you're just not seeing it in the data."

dnap

110,354 görüntüleme • 2 ay önce