Loading video...

Video Failed to Load

Go Home

What we’re seeing: frontier robotics teams are rebuilding data pipelines from scratch because generic licensed datasets aren’t moving their models forward. The next wave won’t be about more data. It’ll be about precise, domain-specific data that actually improves performance. The bottleneck isn’t compute. It’s the right data. 🎥 Our...

30,562 views • 4 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

Scale alone is not enough for AI data. Quality and complexity are equally critical. Excited to support all of these for LLM developers with Snorkel AI Data-as-a-Service, and to share our new leaderboard! — Our decade-plus of research and work in AI data has a simple point: scale alone is not enough. AI success is all about the quality, complexity, and distribution of data—in addition to volume. We’re excited to be powering leading LLM developers with Snorkel AI Expert Data-as-a-Service, our white glove service for custom, expert-level AI datasets—and to now preview some of what we’re building via our new Expert Data Leaderboard (🔗 in 🧵) + upcoming OSS dataset releases! Snorkel Expert Data-as-a-Service is built to meet the rapidly evolving data needs of the agentic AI world—where success is built on the quality, complexity, and distribution of datasets, in addition to size and scale. This kind of high-quality, frontier AI data can only come from a union of technology and human expertise. With Snorkel Expert Data-as-a-Service, we’re powering frontier LLM developers across agentic, expert knowledge, reasoning, coding, multi-modal, and other task types via the combination of these two key components: - (1) The Snorkel Expert Network: A global team of subject matter experts focused wholly on specialized knowledge–spanning thousands of topics in STEM/academic, vertical/professional, and consumer/lifestyle domains. - (2) Snorkel AI Data Development Platform: Our unique programmatic data curation and quality control platform, accelerating and improving expert authoring and review through principled techniques developed over the last decade of R&D. Now: we’re incredibly excited to showcase some of the power of Snorkel Expert Data-as-a-Service via the new Snorkel Leaderboard—putting frontier models to the test in complex, agentic, and reasoning settings inspired by real industry scenarios (not esoteric puzzles)! We’ll be releasing new leaderboards and accompanying expert-verified open source datasets (coming soon!) regularly. To start, we’re sharing three initial ones in preview: - SnorkelFinance: Q&A over financial documents requiring agentic tool-calling and reasoning - SnorkelUnderwrite: Agentic insurance tasks requiring industry-specific reasoning and tool use - SnorkelSequences: Mathematical tasks requiring compositional multi-step reasoning

Alex Ratner

495,851 views • 1 year ago

Programmable Bandwidth is crypto’s next meta. Gm rent your spare Wi-Fi to AI. AI needs more data. AGI is coming. Are you ready? Bandwidth = how much data your connection moves per second. Residential IP = your home’s street address on the internet, trusted as human traffic that doesn’t get blocked. Together, Bandwidth + Residential IP = clean, high-trust traffic that data buyers and AI teams actually want. The community becomes the network. LLMs eat data. The more and the higher quality, the smarter they get. Owning the data pipeline = owning the power. Proof it’s valuable: big platforms license conversational & user-generated data for serious money. Many 8-figure+ deals are public, many more done under the table. Proprietary datasets are the real edge in this AI world. Why residential IPs matter: datacenter IPs get blocked by anti-scraping shields. Residential IP networks are resilient. Grass showed the playbook: idle bandwidth can pay twice 1️⃣ Web data collection via your node 2️⃣ Renting those nodes to GPU grids for training Data + compute = compounding revenue. Hub ( takes it further: aggregate community bandwidth → build a Residential IP Supernetwork → sell real-time data APIs to enterprises (from Meta to Web2 startups). The value loop: Community miners → Bandwidth pool → Enterprise feeds → Revenue → Rewards. If revenue outpaces incentives, everyone wins. Balancing it won’t be easy, time will tell. For miners: run on the network, earn points (likely $HUB at mainnet). Early participation can matter. Why DePIN? Community-run networks scale faster, are harder to censor, and more resilient than centralized systems. We are long programmable bandwidth thesis. In an agent-driven world, whoever controls the live data feed mints the money. DYOR. Register NOW

Q42

39,768 views • 10 months ago

Telecom companies offer Recharge Plans with ‘𝐃𝐚𝐢𝐥𝐲 𝐃𝐚𝐭𝐚 𝐋𝐢𝐦𝐢𝐭𝐬’ like 1.5GB, 2GB or 3GB per day, resetting every 24 hours. Any Unused Data EXPIRES at midnight, despite being fully paid for. 𝐘𝐨𝐮 𝐚𝐫𝐞 𝐛𝐢𝐥𝐥𝐞𝐝 𝐟𝐨𝐫 𝟐𝐆𝐁. 𝐘𝐨𝐮 𝐮𝐬𝐞 𝟏.𝟓𝐆𝐁. 𝐓𝐡𝐞 𝐫𝐞𝐦𝐚𝐢𝐧𝐢𝐧𝐠 𝟎.𝟓𝐆𝐁 𝐝𝐢𝐬𝐚𝐩𝐩𝐞𝐚𝐫𝐬 𝐚𝐬 𝐝𝐚𝐲 𝐞𝐧𝐝𝐬. No refund. No rollover. Just gone. This is not an accident. This is policy. Use it unnecessarily, or lose it by midnight. That’s how mobile data works today. I raised this issue in Parliament - 𝐖𝐡𝐲 𝐬𝐡𝐨𝐮𝐥𝐝 𝐝𝐚𝐭𝐚 𝐭𝐡𝐚𝐭 𝐰𝐞 𝐡𝐚𝐯𝐞 𝐩𝐚𝐢𝐝 𝐟𝐨𝐫 𝐛𝐞 𝐅𝐎𝐑𝐅𝐄𝐈𝐓𝐄𝐃? UNUSED DATA should carry forward into the next cycle, so consumers can use what they have already paid for. My demands are clear: 𝟏. 𝐀𝐥𝐥𝐨𝐰 𝐃𝐚𝐭𝐚 𝐜𝐚𝐫𝐫𝐲-𝐟𝐨𝐫𝐰𝐚𝐫𝐝/ 𝐃𝐚𝐭𝐚 𝐫𝐨𝐥𝐥𝐨𝐯𝐞𝐫 𝐟𝐨𝐫 𝐚𝐥𝐥 𝐮𝐬𝐞𝐫𝐬 All telecom operators should provide Rollover of Unused Data. What remains unused at the end of the day, should be added to the next day’s Daily Data Limit, not erased the moment validity ends. 𝟐. 𝐆𝐢𝐯𝐞 𝐨𝐩𝐭𝐢𝐨𝐧 𝐨𝐟 𝐀𝐝𝐣𝐮𝐬𝐭𝐦𝐞𝐧𝐭 𝐨𝐟 𝐔𝐧𝐮𝐬𝐞𝐝 𝐃𝐚𝐭𝐚 𝐚𝐠𝐚𝐢𝐧𝐬𝐭 𝐧𝐞𝐱𝐭 𝐦𝐨𝐧𝐭𝐡’𝐬 𝐑𝐞𝐜𝐡𝐚𝐫𝐠𝐞 𝐀𝐦𝐨𝐮𝐧𝐭𝐬 If a consumer consistently under-utilises their data over multiple cycles, there should be a mechanism for Adjustment or Discount of that value, from the following month’s Recharge Amount. Consumers should not repeatedly pay for capacity they do not use. 𝟑. 𝐀𝐥𝐥𝐨𝐰 𝐓𝐫𝐚𝐧𝐬𝐟𝐞𝐫 𝐨𝐟 𝐔𝐧𝐮𝐬𝐞𝐝 𝐃𝐚𝐭𝐚 𝐭𝐨 𝐫𝐞𝐥𝐚𝐭𝐢𝐯𝐞𝐬 & 𝐟𝐫𝐢𝐞𝐧𝐝𝐬 Unused data should be treated as the consumer’s digital property. Users should be allowed to transfer their unused data to others, from their Daily Data Limit, just as transfer money to others. As we build a Digital India, access cannot depend on data that disappears. If you’ve paid for it, it should carry forward and remain yours to use.

Raghav Chadha

425,426 views • 4 months ago

🦔Workers in India are wearing head-mounted cameras for 12 cents an hour to collect training data for humanoid robots. The footage of them doing everyday tasks like cooking, cleaning, sorting, and walking through public spaces gets sold to robotics companies building the models meant to replace those same kinds of jobs in higher-wage countries. The arrangement has been running for roughly two years. Workers do not own the data, do not get residuals, and in many cases are not told what their footage is being used to train. My Take The workers wearing the cameras live in a country where robotics automation will hit decades later, so they are training their own future replacements at a delay that hides the consequence from them personally. The companies buying the data are mostly US and Chinese, building humanoid robots aimed at warehouses, retail, and service jobs in countries paying $15 to $25 an hour rather than 12 cents. Robotics companies need motion data that mimics how humans actually move through real environments, and synthetic data has not been good enough yet. Paying 12 cents an hour in Bengaluru is cheaper than running motion capture studios in Boston, and it works at scale because the worker absorbs the cost of the camera, the discomfort of wearing it, and the long-term loss of any rights to their own movement data. The robotics labor market that eventually emerges from this footage will displace far more wages than the data collection cost to gather. That is the trade investors funding humanoid robotics startups are betting will pay off, and the workers in the videos are the ones paying the tab up front. Hedgie🤗

Hedgie

161,757 views • 2 months ago