Loading video...

Video Failed to Load

Go Home

Yesterday I interviewed Sean Cai about AI data. This is essentially a guide for founders on how to sell data and RL envs to AI labs. "I've never seen a data contract get turned down by a top lab, if it's good quality data, for budget reasons." 00:00 What...

131,812 views • 4 months ago •via X (Twitter)

0 Comments

No comments available

Comments from the original post will appear here

Related Videos

PhD Students – How to automatically extract data from papers for your literature review? Extracting relevant data from papers is challenging. However, this process can be automated. Meet AnswerThis – a tool that extracts data in seconds. Here is how it works. 1. Go to and log in. 2. After logging in, click on 𝐸𝑥𝑡𝑟𝑎𝑐𝑡 𝑑𝑎𝑡𝑎. 3. Then click on 𝑈𝑝𝑙𝑜𝑎𝑑 𝑃𝐷𝐹 and upload your papers. 4. These are the papers from which you want to extract data. 5. After uploading papers, select data you want to extract. 6. The predefined options are - Key findings - Research gaps - Methodology - Limitations - Future work - Contributions - Practical implications 7. You can also extract custom data e.g., dataset used. 8. For example, I want to extract methodology used in these papers. 9. I selected 𝑀𝑒𝑡ℎ𝑜𝑑𝑜𝑙𝑜𝑔𝑦 and clicked on 𝐴𝑑𝑑 𝐶𝑜𝑙𝑢𝑚𝑛. 10. AnswerThis extract data about methodology used in the papers. 11. You can change data view from normal to Table View. 12. For this, scroll back to top and click on 𝑇𝑎𝑏𝑙𝑒 𝑉𝑖𝑒𝑤. 13. Now for instance, you want to extract more data from these papers. 14. Go back to the top and click on 𝐸𝑥𝑡𝑟𝑎𝑐𝑡 𝑑𝑎𝑡𝑎. 15. Select the data type you want to extract. 16. For example, I want to extract data about future work. 17. So I click on 𝐹𝑢𝑡𝑢𝑟𝑒 𝑊𝑜𝑟𝑘 and then clicked on 𝐴𝑑𝑑 𝑐𝑜𝑙𝑢𝑚𝑛. 18. AnswerThis extracted data about future work from the papers. 19. After extracting the desired data, you can export it. 20. Select the data you want to extract. 21. Then click on 𝐸𝑥𝑝𝑜𝑟𝑡 𝑑𝑎𝑡𝑎. 22. Your data will be exported in CSV format. You can then analyze this data for your literature review. Try AnswerThis today: Anything you'd like to add?

Faheem Ullah

21,390 views • 10 months ago

Major program launch: Data Analytics Professional Certificate! This large, five-course sequence takes you all the way to being job-ready as a data analyst, and shows how to use Generative AI as a thought partner to enhance your work in this role. Offered by on Coursera, this is taught by Sean Barnes, Ph.D., a Data Science & Engineering Leader at Netflix. Analyzing data remains one of the most important skills in where the world is going with AI. This comprehensive certificate takes you all the way to being job-ready. Each course comes with practical projects demonstrated in real-world contexts, such as analyzing sales data for a Korean bakery, video game sales trends across different regions, or identifying factors impacting customer retention for a communications company. You'll also work on estimating fire distribution for forest fire prevention, analyzing how a diamond's properties affect its market value, and developing predictive models for retail sales analysis, carbon emissions, and coral reef conservation. Here's some of what you'll learn: - How to define data and categorize it into its many types such as discrete & continuous numerical, structured & unstructured, time series, categorical, and know what insights can be derived from the different types of data categories. - How to differentiate between data-related job roles and their responsibilities, and how data flows through an organization from the moment of capture to decision-making. - How to perform data processing functions and apply conditional formatting in spreadsheets to extract business value from your data using statistical calculations and best practices for visualizing and interpreting data. - How to use LLMs for stakeholder analysis, data exploration, and data visualization. - Best practices for using LLMs for as a thought partner to data analysis work By the end of this professional certificate program, you will have learned core statistical concepts, analysis techniques, and visualization methodologies that will serve as the foundation for working as a data analyst. The world needs more data analysts, especially ones who know how to use modern generative AI. With data science roles projected to grow 36% by 2033, the skills taught in this program create new professional opportunities in data. Sign up here!

Andrew Ng

85,107 views • 1 year ago

David Friedberg: Michael Burry’s Datacenter Math is Wrong “I actually think Michael Burberry's got this wrong.” “What Michael Burry is saying is that all of these hyperscalers have extended their depreciation schedule or the useful life of their data centers by roughly 2x, which cuts the operating costs in half when they report it in earnings. And so it's making their earnings inflate.” “So he's claiming they're cooking the books. Google first made this change in Q1 of 2021, where they said the servers are now going from 3 to 4 years. Separately in 2021, Google took networking equipment from 3 to 5 years. And then in 2023, they took it from 5 to 6 years.” “And so this is a result of this effort where they went in and did an analysis. So what happened?” “What happened in the data centers is that the data centers transitioned from being primarily data storage and data transfer systems, where you would use hard drives and RAM and memory to store data and then transmit it back out, to being data processing centers because of the AI boom.” “So as AI became more important in the data center, more of the dollars that are going into data centers were allocated towards chips from data storage, which initially was hard drives.” “And then suddenly, when you put these processors in to process the data to do AI, the majority of the spend and the majority of the energy is going towards the processors.” “I made some calls and I checked around with some other friends, and everyone says the same thing: that these 7-8 year old TPUs and GPUs that are sitting in the data centers are still being used and they're being used at 100% utilization.” “So that actually justifies and validates the depreciation schedule being much longer versus shorter.”

The All-In Podcast

304,450 views • 9 months ago

NEW: Inside AI's Biggest Downstream Winner.. the Surge in AI Database Demand "Data is the unsung hero, & data is back." MongoDB CEO CJ Desai (CJ Desai) The unexpected result? Hyperscalers are turning away even top-50 accounts. "Sorry, we don't have a capacity." "And they are one of the top 50 customers for that hyperscaler." ElevenLabs alone runs "north of 50 million agents, depending on when you look at it, all running on MongoDB.. that gives us a lot of confidence that we have the right architecture for agentic workloads." NASDAQ: $MDB We cover: › The 3 classes of AI customers MongoDB serves › Why hyperscalers are telling top-50 accounts "no capacity" › On-prem, sovereign AI, and the data-center comeback › ElevenLabs running 50M+ agents on MongoDB › MongoDB (OLTP) vs Snowflake & Databricks (OLAP) › Why there is no standardization in enterprise AI models › Auto-scaling & the fall of expensive DBAs 𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 (00:00) CJ Desai, CEO at MongoDB (01:12) Top CEOs at the Raise AI Summit (03:24) The 3 types of companies powering the AI boom (07:25) Why on-prem is making a shocking comeback (09:01) Hyperscalers are quietly running out of capacity (11:06) What actually separates MongoDB from Snowflake x Databricks (13:18) Why Frontier Labs treat MongoDB as their memory layer (16:22) From Oracle intern to first-time CEO (19:19) Becoming CEO during the AI chaos (23:53) The World Cup crisis that tested MongoDB's scale (28:19) The real complexity behind simple AI agents (31:22) Open source vs. closed models: what customers actually pick (34:36) CJ's honest take on data centers in space (36:50) The mentors who shaped a first-time CEO (40:32) MongoDB's next big bets (41:36) Data is back?

Molly O’Shea

170,730 views • 1 month ago

Spirit Airlines stopped flying in May. Second bankruptcy in two years. Yet Google still wants to pay $10Million for the dead body. The planes are gone, the airport slots are sold. One asset left. At the bankruptcy auction, Google opened at 5 million dollars. An AI data company called Mercor countered at 7 and a half million. Then Google closed it at 10 Million. But the bids weren't for the planes or Airport slots. They were for 100 million internal emails. 500 million Teams messages. 30 million lines of code. Employee records that go back to 1986. And no, its not passenger data. This is purely internal: decades of how a real business thought, argued, and made decisions. Why pay that much for a dead company's inbox? Because it's the one thing AI can't fake. The most valuable data in the world right now is just real people thinking out loud… real decisions, real mistakes, real cause and effect. Which is why after the auction closed, another AI company came in with a 12 and a half million offer for it. Here's how you can leverage this kind of data for your business without a bankruptcy auction. People type their real, unfiltered questions into a search bar every single day… for free. And for business owners, those raw questions are a free roadmap: they tell you exactly what to create, what to fix on your site, and what to sell next. Tools like AnswerThePublic mine and present that exact data by looking at the different ways people search for a product. The AI data gold rush is just getting started, follow to not miss out!

Neil Patel

12,632 views • 5 days ago