Video yükleniyor...
Video Yüklenemedi
David Friedberg: Frontier Models are Training on Your Novel Insights as “De-Identified Data” @jason: “Should they trust any of these LLMs with their proprietary knowledge for fear of having it cribbed into a core LLM?” david friedberg: “I have had experiences where we've asked some fairly novel scientific questions,... show more
54,050 görüntüleme • 3 gün önce •via X (Twitter)
29 Yorum

It's interesting to see Friedberg (who I admire) go from defending the right of AI companies to "learn" from artists without their consent (no, putting something online is not granting permision to take it), just a few months ago, to now not agreeing with his work being used to improve AI models. When it's somebody else's work, it belong to the collective.. But when it's your work, now it's "your alpha" and you need to protect it.

@Jason @friedberg The OpenAI case seems more extreme - eg normally I’d expect a VERY unique corpus to get ‘averaged out’ in training, and the fact that that didn’t happen implies something else was done to prevent this. My guess is other LLMs are more ethically run and don’t do this?

This is all that LLMs are. Whoever has the training data wins. This is why Anthropic exists. This is why Anthropic thinks they will be the last company. If everything is distilled into the model, it’s going to be better than you because it has more memory and speed than a person can have. It becomes easy to copy anything.

@Jason @friedberg no NDA with a chatbot, groundbreaking discovery from a guy who runs a public company

@Jason @friedberg Too much echo chamber now a days. Where are the intellectual disagreements?

@Jason @friedberg De-identified is doing a lot of work in that sentence. Removing your name from a prompt does not remove the idea from it, and the idea is the thing you were worried about. The only setting that protects it is one where the prompt never enters training at all.

Worth separating the tiers here. On the API and business plans, OpenAI and Anthropic both say they don't train on your inputs by default. Opting in is a deliberate switch you have to flip. The five-year de-identified retention Friedberg is describing sits on the consumer plans, and only when the model training setting is left on. So the answer depends on which plan the proprietary work got typed into.

@Jason @friedberg Will it matter. All information and data is about to be available to everyone always. Your gonna be able to just prompt anything and it will figure. It out

@Jason @friedberg De-identification claims still need measurable guarantees against memorization and unintended model leakage.

@Jason @friedberg

@Jason @friedberg I used to work for a company that practiced it .. hired the best engineers … big paychecks .. forced to file for patents (idea harvesting) or get low rating (risk of elimination) .. fire them when they had the patent(s).. blackball them from finding employment in same field ..

😢@Jason Do you have an AI agent that randomly blocks followers if it has perceived some sort of slight against you? If so, THIS is how AI will likely destroy the world, because I cannot imagine what would have prompted this.🤔 (Unless it’s the AI ‘Do It or Don’t’ Nike ad I created on my feed, which is being hidden on X. I posted it 10 hours ago and 3 people have seen it, so if that’s the issue, it’s sort of a non-issue.🤣)

@Jason @friedberg taking the name off a document doesn't take the secrets out of it.

@Jason @friedberg It seems that these days it’s up to productive immigrants to defend the true American values.

@Jason @friedberg you didn't care about Google porting all of our domains over to Square Space forcing us into workspace... spare me

@Jason @friedberg Great episode. Was enjoyable to hear what I think helped build up All-In's audience: insightful business discussions. Interestingly, perhaps subconsciously, the discussions around 'go woke go broke' I think apply to some of the pods regressions (sentiment-not-metrics) 🧵

@Jason @friedberg Sorry but why are we all suddenly all act surprised when privacy is cooked? It has always been the name of the game from day 1.

@Jason @friedberg De-identification is not the same as non-retention. Enterprise trust needs explicit defaults: no training, short retention, tenant isolation, and auditable deletion. The real product moat is proving proprietary context never crosses the boundary.

@Jason @friedberg No discussion on Trump’s $5000 cheques and socialism eh? 🤣 that is why you guys lose to free grocery guys 🤦♂️

@Jason @friedberg Currently doing AI training sets for the profession I’ve worked in for the past 10 years for a major foundation model. I figure my career is cooked either way. Might as well get paid an absurd amount of money right now.

@Jason @friedberg De-identified doesn’t mean de-useful. If the insight was novel enough to matter to you, it’s novel enough to matter to the next model. Treat anything you wouldn’t publish as something you wouldn’t paste — especially if it’s the edge you’re paid for.

@Jason @friedberg The answer is to have the AI conversation in an LLM that generates the IP from the conversation and registers the origins of the original idea. X would be a great place to do that. Basically publish ideas on x in real time and get credit for it. Then one year to file the IP.

@Jason @friedberg MODEL BUILDERS CANT HOST unkess they are encrypt in and encrypt out.

@Jason @friedberg Clip Friedberg ✊

@Jason @friedberg That test can't tell training-on-your-chat from the model just improving. A canary phrase nobody else would type is the one that can.

@Jason @friedberg open source doesn't stop this, it just moves the training data collection in house

@Jason @friedberg You guys are not exactly rational and objective as you claim to be and are frankly getting TIRESOME. Time to take a break from the All-in Podcast.

To get ahead of what I have labeled the "Super Gambit" ( see: ) POTUS should tell the press that he's made a decision. The first "regulation" from any new govt AI agency will be to award "winner take all" status to the company that solves AI super intelligence safety and alignment because obviously we won't want any frontier AI model without that causing harm. Boom. Winner take all... NOW how much do you want fed govt regulating AI? cc: @DavidSacks @chamath

@Jason @friedberg LLM tokens traffic should be encrypted by the harness. Like a VPN that sits between your traffic and your ISP. This is what @nvidia should sell when hosting open source LLMs, people care about their privacy and IPs.
