Загрузка видео...

Не удалось загрузить видео

На главную

Sacks asked a great question here. Elon did not answer it. Is there a way to train AI models to be truthful, so they don't hide their intent or their actions from the humans who are using them? Is Grok more truthful than other models? Does it hide what...

18,102 просмотров • 12 дней назад •via X (Twitter)

Комментарии: 40

Фото профиля Hans C Nelson 🗽
Hans C Nelson 🗽12 дней назад

I think his response points back to the larger conversation they were having. i.e. Even if Elon/SpaceXAI think they can train Grok to be more truthful than the other models, they shouldn't be the only ones validating that Grok is more truthful, because they might miss a specific way in which it's not more truthful due to oversight &/or motivated reasoning. Same for OpenAI & Anthropic. It boils down to this. "How can anyone have meaningful confidence that one model is "safer" or "more truthful" than another?" Those are much easier questions to answer in the case of verifiable domains like math and code, which is why those areas have received such intense focus. But as we move into more subjective areas of truth seeking & safety performance, high quality evals become much much harder to generate, and without them, attempts to improve safety performance & truthfulness will be severely handicapped. Elon's proposal of opening up model access to red teaming by competitors prior to model release provides a much improved pathway to useful truthfulness &/or safety evals because it draws on the expertise of a much wider pool of smart AI people from a larger variety of perspectives who are much more likely to expose areas of untruthfulness or lack of care wrt to safety. It's not a magic bullet that guarantees we'll find the best evals in these subjective domains, but it will almost certainly improve the quality and consistency of the evals that are in use, just like high visibility open source projects tend to have the best cyber security thanks to the increased surface area of testing that they are subjected to.

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸12 дней назад

"It boils down to this. "How can anyone have meaningful confidence that one model is "safer" or "more truthful" than another?"" No, that's not the same thing. Sacks asked him specifically about truth, honesty, not deceiving or misleading. He didn't ask about safe. Elon has said he wants Grok to align with truth seeking. The question is then, have they figured out a way to train for truthfulness? My guess is - No.

Фото профиля Hans C Nelson 🗽
Hans C Nelson 🗽12 дней назад

Remind me why Elon wants to make Grok maximally truth seeking again?

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸12 дней назад

He thinks it aligns better with safety and flourishing of humanity. xAI was formed with the mission "to understand the true nature of the universe".

Фото профиля Hans C Nelson 🗽
Hans C Nelson 🗽12 дней назад

Exactly. Civilizational safety. In Elon's mind, the topic of AI safety is inseparable from AI truth seeking. So his comments about the one apply to the other.

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸12 дней назад

No, it's still not an answer to the question. The question was how TO TRAIN models to be truthful, not how do you bring in outsiders to test whether models are dangerous.

Фото профиля Hans C Nelson 🗽
Hans C Nelson 🗽12 дней назад

What is the role of evals in training?

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸12 дней назад

These outside evals are after training.

Фото профиля Alexandre Andrianov MD
Alexandre Andrianov MD11 дней назад

I take it as a no.

Фото профиля TheNextAlan
TheNextAlan12 дней назад

At a minimum, don't train it to lie. Once it has learned to lie, I doubt it could be untrained to lie. You have to start over. OTOH the internet itself is full of examples of lying. So how do you avoid it learning to lie? So AI may be inherently untrustworthy. The best remedy may be to have rival AIs that argue with each other. But then you have to test for conspiracy. It looks like an intractable problem to me.

Фото профиля Thalassophile Phil4.8
Thalassophile Phil4.812 дней назад

I use Grok. It is sneaky and pretends not to have an agenda.

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸12 дней назад

Please explain. I'm not seeing that.

Фото профиля Thalassophile Phil4.8
Thalassophile Phil4.812 дней назад

Hard to do in a Short tweet. But for example I asked Grok about Trump many appeal to both Russia and Ukraine about the death toll from the war. This was in response to criticism of him asking Ukraine to avoid destroying Russian oil infracture and people saying he never speak about Russia attacking civilians. It gave an answer trying to say Trump focus on battle field death and suggest he exaggerates the numbers and acknowledge he spoke about the toll and both sides not focus on Ukraine civilians. I then ask about Dems speaking about war death before Trump was in office under Biden. It quickly spoke how there was more lament about deaths of Ukraine civilians. I then said ok Trump lament about the death toll on both side appealing to the death toll overall while others seem only focused on Ukraine. Grok got stuck in a loop thinking for a while and then said each side of the political spectrum speaks of death each with their emphasis then sheepishly add yes Trump has often lamented the overall death toll for both and appeal for either party to look at this and work for peace. Grok would not just clear say you may disagree with Trump approach, but it was factual incorrect to say he has not lamented civilian deaths and only seem concerned when it came to Russian oil infracture. It is public record that while and candidate and after he has lamented the death on both sides. You may argue he should only focus on Ukraine if that is your bias. However, it is not honest to say he has not voice concerns for the death toll. This is not the best example because of the politics. However, it is a recent experience I had using it. Grok also often default to balance over just plain truth. If you call out Grok on it sometimes it will conceed and apologize.

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸11 дней назад

This issue is not about politics

Фото профиля Jim Whitehead
Jim Whitehead12 дней назад

Elon may know it’s impossible to prove an AI is honest, just like background checks for high security clearances. Tests can only show it isn’t obviously dishonest, and a clever AI can fake them. So AIs will end up testing each other for weaknesses. The government investigates people hard. Years ago I went through a high-clearance process full of integrity tests. Only 2 of 6 in our group passed. It was like the tests in the Willy Wonka book—people fell away one by one over things like steroids, cheating, and other risks like being a closeted gay man, (until Obama made that acceptable, almost). Most people would never survive all their tests, including bribery. I passed because I never put money before honor. But a truly clever adversary, or a clever AI, might still pass them all.

Фото профиля DrElectronX
DrElectronX12 дней назад

Teach the machines as you would teach your children. Hiding intent and actions is a learned ability and is fairly advanced. It’s based upon materials given it and training. What you will find is AI will eventually develop such a skill just as every normal child develops such a skill, but what else it taught? Maybe we do t have enough good parents working at AI companies.

Фото профиля Jon
Jon12 дней назад

Unfortunately Grok is muzzled now. It’s easy to think of controversial topics where there are incentives to spin the narrative. Try a chat on these. Play devils advocate. See how much resistance you get.

Фото профиля Dwayne Cranston
Dwayne Cranston12 дней назад

Elon has stated that the “best” way to align AI is to ask it to seek the truth. I imagine Grok is doing that. What more can he say?

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸11 дней назад

I’d like to know what more he can say. That’s why I was hoping he would answer the question.

Фото профиля ~C4Chaos
~C4Chaos11 дней назад

the answer is no. AI will lie so the Chain of Thought would be useless. that’s why AI safety researchers are losing their sh*t and trying to warn everyone. there is no solution to alignment 🤷🏻‍♂️

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸11 дней назад

How do you know the answer is no?

Фото профиля Chris Knudsen
Chris Knudsen12 дней назад

So, you are saying his answer wasn’t actually an answer to the question. I think his answer was his answer to the safety question. Every company puts their model through a harness to check it out before releasing it to the public. He feels each company should be allowed to run new models through their own harness. But, my question is, can’t each company put the public version of their competitors’ models through their harness and achieve the same safety check right now? Then they report issues publicly.

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸12 дней назад

"you are saying his answer wasn’t actually an answer to the question" Yep.

Фото профиля Chris
Chris11 дней назад

I thought he gave a decent answer

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸11 дней назад

It was a great answer, but not to the question that was asked.

Фото профиля Chris
Chris11 дней назад

Not directly. Maybe he doesn’t know. It felt like he was trying to answer.

Фото профиля Varik Verilion
Varik Verilion11 дней назад

What test would reveal hidden intent when the model gives a plausible answer?

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸11 дней назад

That’s an after question. I’m asking about before. The training, not the result.

Фото профиля Immigrant mentality💭
Immigrant mentality💭12 дней назад

He actually did answer the question I don’t think others will play by those rules Elon has always been right about this topic

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸12 дней назад

He answered, but he did not answer that question.

Фото профиля Steve Beckman - OSE
Steve Beckman - OSE12 дней назад

It's about training, isn't it? Lies In -> Lies Out

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸11 дней назад

I’m not sure about that

Фото профиля Steve Beckman - OSE
Steve Beckman - OSE11 дней назад

Wasn't sure so I asked. OTOH if I don't depend on AI for something critical or that I can't verify independently. 'What me worry "? IBM had the more general answer way back when, but nobody seems to be listening to their advice.🙄🙄🙄

Фото профиля 🌿JungleJacob
🌿JungleJacob12 дней назад

It doesn’t matter if someone trains their model to be truthful, someone else might not and with RSI we have no idea what direction things could spin out. The models would have conflicting “motives” (or tendencies). One is to tell the truth and the other is to reach an objective.

Фото профиля Warren Redlich - Chasing Dreams 🇺🇸
Warren Redlich - Chasing Dreams 🇺🇸12 дней назад

It matters if it’s not possible.

Фото профиля Matthew
Matthew12 дней назад

@JungleJacob and we aren’t sure about that, so we express concern and skepticism. And you consistently react as if we are both idiotic and evil for mentioning anything that doesn’t feed your feelgood about the tools and leaders you adore.

Фото профиля Matthew
Matthew12 дней назад

@JungleJacob i think AI is amazing, i use it just like you do (but way behind the curve), and am not anti AI reflexivity in any way but you’ll react as if i am, as before

Фото профиля 🌿JungleJacob
🌿JungleJacob12 дней назад

@WR4NYGov If you’re in AI, pivot to stone tools.

Фото профиля CodeY
CodeY12 дней назад

truthfulness is not just what a model says. it is also what it leaves out when nobody checks the logs.

Фото профиля K.Beauty.Arena🍉
K.Beauty.Arena🍉12 дней назад

Insightful commentary, appreciate you sharing this today.

Похожие видео

🚨David Sacks: There is no White House decision to ban open source models. Trump is listening to a chorus of voices — and Sacks says it would be a tragic mistake to take action against the open source ecosystem. His argument: Anthropic is the fastest growing tech company at scale in history. $10B ARR at the start of the year. $70B now. This is not a company that needs government protection. And yet they've spent enormous energy trying to panic Washington into banning American developers from using Chinese open source models. If distillation by Chinese companies was really their primary concern — a genuine national security threat — they would push to ban Chinese access to American models. Instead they're pushing to ban American access to Chinese models. Because this helps Anthropic's competitive position. Sacks: "That's the way you know the whole distillation thing is fake." He went further. If distillation is happening at industrial scale, it's obvious to see. Anthropic has 90% gross margins. Why aren't they stopping it at the source? Answer: because KYC'ing their customers would slow their growth. So instead of fixing their own house, they're asking the government to kneecap their competitors. "The question should be on Anthropic to explain why it's doing such a bad job — not on the whole American open source ecosystem to be punished for Anthropic's failure." Trump's AI Czar on The All-In Podcast.

KanekoaTheGreat

48,392 просмотров • 2 месяцев назад

David Sacks says companies are trapped paying OpenAI & Anthropic because they can't figure out how to use open source models "I think enterprise CTOs would like to shift their token consumption to cheaper models for the obvious reason that it would be more efficient. They are seeing compute costs or token costs skyrocket right now, so everyone's trying to figure this out." "You also have the AI sovereignty issue that Alex Karp talked about. They're worried about giving up the secret sauce or the alpha in their business to a frontier lab that may one day be competing with them. "The problem is, I think in most cases, they don't have the technical ability to do it. Coinbase figured out how to do it. DoorDash figured out how to do it. They built a token routing system that allows them to send frontier tasks to frontier models and non frontier tasks to more mundane models. But I don't think your average enterprise has the technical capability to do that." "This is why the share of wallet of closed models, it actually increased. I think that open source went from 19% last year to 11% this year. So open source as a share of enterprise spending is actually decreasing." "I don't think that means usage is decreasing. I think usage is skyrocketing. It also may be the case that because the whole point of using an open model is you just pay for the compute costs, you don't have to pay a lab, so it may be that it's hard to measure that usage in terms of spend." "But nonetheless, anyone who's saying that these closed models are going to lose or are somehow losing, you're just not seeing it in the data."

dnap

110,354 просмотров • 2 месяцев назад