Loading video...

Video Failed to Load

Go Home

Bob McGrew (Head of Research OpenAI) explains why proprietary data no longer provides companies with a competitive advantage in the AI era. Finance companies once believed their years of accumulated data would give them an edge. They planned to train specialized models on top of GPT or Llama using...

240,704 views • 1 year ago •via X (Twitter)

8 Comments

Raphael Schaad's profile picture
Raphael Schaad1 year ago

@garrytan “proprietary data no longer provides companies with a competitive advantage in the AI era [IF that data is generally available and collectable]” (and I think that’s a pretty big IF to warrant such a blanket statement)

ᐸGerardSans/ᐳ🚀🇬🇧's profile picture
ᐸGerardSans/ᐳ🚀🇬🇧1 year ago

These claims are not true at all. Better performance in benchmarks which do not imply better performance on other tasks besides the standardised tests used in post-training. This narrative from OpenAI has been challenged by recent research papers:

Dolapo Obat's profile picture
Dolapo Obat1 year ago

Embodied labor is no longer a barrier. That’s both terrifying and liberating.

Aish's profile picture
Aish1 year ago

yeah it is!

Jacob Eckel's profile picture
Jacob Eckel1 year ago

Always assume that motivated reasoning is BS.

Shawn Chauhan's profile picture
Shawn Chauhan1 year ago

AI shifts competitive edge from data to innovation

Milo's profile picture
Milo1 year ago

This does not sound believable.

deleuze's profile picture
deleuze1 year ago

This just says data collection is better enabled by AI. Data moat is non-public info collected via the company ops, their ecosystem, and products. You need to then recreate a company’s infra, product, ecosystem etc to even collect that data Incomplete narrative at best

Related Videos

In the second episode of Scenius Studio's mini-series "The Use-Case", I sit down with Andrej Co-Founder of touch grass. Grass gives users the ability to earn ownership in the Grass network by supplying the protocol with their unused internet bandwidth for data scraping purposes (something that is already happening to most of us and we don’t get paid!). The grass protocol packages this scraped web data and sells it to AI companies who have insufficient data to further develop their models. With over 3 millions users and millions of annualized revenue, Grass is a real commercial business with a roadmap that makes it one of the most exciting projects at the intersection of crypto x AI and data. In this episode we discuss: ➔ Andrej’s background in physics, finance, and sports betting ➔ Big companies using your IP address without your knowledge or permission ➔ How the Grass protocol puts a toll booth on your internet bandwidth highway ➔ Packaging web scraped data and selling it to AI companies building Multi-Modal models ➔ Dynamics between the Grass Protocol and the labs entity developing Grass’ IP ➔ Protocol design decisions to ensure that all tokenholders (VCs, team, and community) are aligned ➔ Why Grass needed to be built on crypto rails to maximize its potential ➔ The future of LLMs and how they will search for context and information Hope you enjoy this episode of Scenius Studio's "The Use-Case". Links to listen in bio or below👇

Ben Jacobs

22,458 views • 1 year ago

🚨 THEY FOUND A LEGAL WAY TO SPY ON EVERY AMERICAN — AND THE 4TH AMENDMENT CAN'T STOP IT Most Americans believe the government needs a warrant to monitor them, track them, or collect detailed information about their private lives. But unfortunately, they found a loophole. The government may not be allowed to directly collect certain information on Americans, but private companies collect enormous amounts of it every single day through smartphones, apps, websites, search engines, location services, online purchases, and countless other digital tools most people use without a second thought. That means your location history, browsing habits, purchases, interests, movements, and daily routines are already being recorded, stored, and traded by an entire industry that most Americans have never even heard of. The loophole is that while the government may not be allowed to collect certain information directly without a warrant, it can reportedly purchase data that private companies have already collected. The information may be gathered by private companies. The data may be sold by data brokers. And the government may still end up with access to it. For years, there was one major limitation. There was simply too much data. Even if someone wanted to analyze billions of data points across millions of people, it would have required an impossible number of human analysts. Then AI arrived. Suddenly, information that would have taken years to organize can be processed, searched, categorized, connected, and analyzed in a fraction of the time. People are warning that the combination of artificial intelligence and the data broker loophole could fundamentally change what surveillance looks like in America. For the first time in history, technology may finally exist that can sift through enormous amounts of personal information at a scale that was previously impossible. The question isn't whether the data exists. The question is who has access to it. Because once your location, habits, purchases, interests, relationships, and daily routines can all be analyzed by machines, the line between convenience and surveillance starts getting harder to see. The most alarming part? Most Americans have no idea this conversation is even happening. Do you trust the government with more information about your life than your own family knows?

HustleBitch

23,930 views • 1 month ago

Big pharma just handed the AI industry one of the most important reality checks of 2026 (Save this). david friedberg revealed that Anthropic approached major life sciences companies with a pitch, share your proprietary data, sign an NDA and we will give you early access to a specialized life sciences model and nearly every company they spoke with said no. Here is what these pharma companies understood that many enterprises still have not. A large pharmaceutical company may have spent decades and tens of billions of dollars generating proprietary datasets, clinical trial results, genomic sequences, drug interaction data, compound libraries. That data is the business and the competitive moat that separates them from every other player in the industry lives in those datasets. Handing it to an AI lab in exchange for early access to a model is essentially handing your most valuable asset to a company whose entire business model depends on combining your data with everyone else's and then selling the output back to you and to your competitors. Palantir CEO Alex Karp made this exact point that enterprise leaders are paying for AI tokens that generate no tangible business value while simultaneously surrendering their most sensitive operational data to external providers. He called this transferring a company's alpha, the unique advantage that secures the business directly to a third-party lab. Microsoft CEO Satya Nadella echoed the same concern independently, warning that entire sectors might find their accumulated knowledge commoditized if they do not build their own data and model ownership layers. The structural problem is not unique to pharma but it applies to every enterprise sector. Every time an employee runs a query through a third-party frontier model, proprietary workflows, customer data, and strategic processes pass through infrastructure the enterprise does not control. The data already shows the market moving, Open-source captured 67% of all AI tokens processed in the first half of 2026, up from a fraction of that just twelve months earlier. The performance gap between proprietary frontier models and open-source alternatives has nearly closed, DeepSeek costs approximately 1/36th of GPT-5 for comparable workloads. What pharma figured out and what enterprises across every sector are starting to realize is that the model is not the moat but the data is. And once you hand your data to a model company, you have permanently surrendered the asset that took you decades and billions of dollars to build.

Milk Road AI

16,317 views • 1 month ago

60% of companies expect to be transformed by AI within two years. But here's the reality: most are still bolting AI onto legacy workflows instead of reimagining how work gets done. That's not transformation—that's just expensive automation. Our latest research reveals what separates AI leaders from laggards. The companies winning with AI aren't just adopting new tools—they're redesigning how their business creates value from the ground up. The challenge: Most organizations are nowhere near putting AI at the center of their strategy. They're missing the exponential gains that come from true AI-first thinking. The solution: A complete playbook for IT leaders and CIOs ready to move at the speed of AI. 🎯 5 Core Principles of AI-First Companies: 1️⃣ AI as a capability expander — Enable work that wasn't possible before, not just faster versions of old tasks 2️⃣ Human-AI partnership — Augment human creativity and decision-making rather than replacing it 3️⃣ AI-native design — Build systems that seamlessly integrate thousands of AI agents working behind the scenes 4️⃣ Trust & governance foundation — Remember: "AI agents can't keep a secret"—never rely on AI to maintain core security models. 5️⃣ Data as strategic fuel — Transform unstructured content into your competitive advantage The companies implementing these principles today are building insurmountable competitive advantages. They're not just working faster—they're working in entirely new ways. Ready to lead your industry's AI transformation? 📖 Get the complete framework Want to explore the benefits? Get the full details here:

Box

128,566 views • 1 year ago

.Josh Wolfe: Anybody Using DeepSeek App Is 'Absolute Fool' "Anybody using the DeepSeek app is an absolute fool. If you're using DeepSeek on companies like Together Compute, one of Lux's companies, which can get rid of the CCP censorship, then it's probably okay. But remember, the open-source movement is something we deeply believe in. Most great technologists, entrepreneurs, and venture capitalists are on the side of open source. The closed-source models that have consumed tens of billions of dollars are the ones that are really going to be at risk. When you look at Hugging Face, a major repository, or Together Compute, Runway ML, and a lot of Lux's companies, they have been pioneers in open source. Now, why am I not worried about open source, even with the DeepSeek model? As long as you don't have the CCP censorship on it, the models with their open weights allow people to run on their proprietary data. This means companies like pharma or defense companies that have their own siloed, proprietary data—think about Bloomberg with their proprietary longitudinal data, or Meta with their data—are the ones who will have the edge. Even as open source takes hold, these companies will still dominate. I’m not worried about open source being the problem. I’m more concerned about people overfunding closed models with no proprietary source. A lot of capital is going to be burned there, and we’re already seeing that with people worried about OpenAI in some aspects."

Josh Caplan

39,985 views • 1 year ago