Загрузка видео...

Не удалось загрузить видео

На главную

Analyzing large vector datasets on Source Cooperative with #DuckDB 🦆Using the 75 GB National Wetlands Inventory as an example. It used to take me hours to get summary statistics of the dataset, now it takes seconds with DuckDB 🚀 Dataset: Colab: #geospatial #dataviz #opendata

13,309 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 0

Нет доступных комментариев

Здесь появятся комментарии из оригинального поста

Похожие видео

🚀 Introducing EgoExo Forge - built on top of Rerun, Gradio, and Hugging Face hub (I’ll be in San Francisco July 21–29 — if you’re into robotics, egocentric AI, large-scale data collection, or just want to chat, DM me!) In my opinion, large-scale, diverse, and high-quality data is still the largest bottleneck for generalized robotics deployment. I believe that some version of imitation learning from human examples will be the most scalable + clean way to train humanoid robots 🤖 (similar to what Tesla did for Full Self Driving). Teleop is too expensive to collect a large enough dataset in a reasonable manner, so passive collection via egocentric (and in certain cases, exocentric) views feels like the right bet. Over the past few months, I've been trying to build out the scaffolding for this and using Rerun as my underlying infrastructure. Data being collected needs to be easily inspectable + time series and rerun provides the right tooling for this. My goal is to first build out a ground truth representative dataset from already existing open source data, generate some reasonable baselines, and then go out and collect my own data that adheres to the defined schema. 🔍 Starting with open-source datasets 1. EgoDex from Apple 2. HOCap from Nvidia and the University of Texas at Dallas 3. Assembly101 from Meta All these different datasets have different sensor configurations + annotations, so my goal with egoexo-forge is to have one consistent labeling scheme + data layout. I built a data pipeline that aligns all of the different datasets in one general schema assuming the COCO133 keypoint layout that allows for exo+ego, ego only, or exo only Since the scaffolding is already there, it becomes MUCH easier to add other datasets. So the next ones that I'll be including are HD-EPIC kitchens dataset, HOT3D, and finally my own personal iPhone + insta360 go collection method. Once I have a diverse variety of datasets, I'll double down on what I believe to be the key algorithms required to make useful data for imitation learning 📊 1. Camera Pose estimation via SLAM/SFM for ego perspective (and automatic calibration for exo) 2. Human pose estimation for both egocentric + exocentric views 3. Metric 3D reconstruction + object tracking I'll be setting up reasonable open-source baselines for each of these to validate that these datasets work, and then finally try to use the generated datasets for some imitation learning via the pi0-lerobot repo I've been working on. I plan on making a blog post + providing more info on all of this in the near future so stay tuned

Pablo Vela

32,085 просмотров • 1 год назад

To replace animal testing with AI, we need MASSIVE human datasets. Today, we're thrilled to share Axiom's new data exploration tool, providing the ability to visually explore the world's largest primary human liver toxicity dataset. Built with Axiom's proprietary wetlab protocols, our dataset includes detailed liver toxicity profiles for over 100,000 distinct molecules. The key to this dataset is our ability to do high-throughput, multiplexed high-content screening with primary human liver cells. Traditionally, toxicity assays either sacrifice throughput or sacrifice biological relevance (using easy-to-grow immortalized cell lines instead of real human cells). We managed to combine throughput, physiological relevance, and multiplexing in one platform. The assays run in a high throughput format using automation, meaning thousands of compound-dose conditions can be tested in one experiment. We achieved this using pooled primary human hepatocytes, which are often fragile and expensive. By systemizing our automation and quality control processes, we were able to run over 120+ batches on the same donor pool with incredible reproducibility and consistency. We did this while integrating many readouts per well, whereas many existing toxicity assays only do a single readout. Our multiplexed approach provides far more data per experiment enabling us to measure 10-20 different toxicity phenotypes such as apoptosis, necrosis, mitochondrial fission, endoplasmic reticulum stress, stress granule formation, microtubules, and more all from a single well on a 384-well plate! The combination of scale, high content information, and data quality is exactly what is needed to train highly accurate AI models in biology. If you're interested, please explore the dataset in the comments below and let me know if you want to chat about the details!

Brandon White

25,117 просмотров • 1 год назад

Google just wired DeepMind and Earth Engine directly into the biggest geospatial dataset on the planet. For two decades, millions of people used Google Earth to scale the Himalayas or zoom in on their childhood neighbourhoods. In 2026, Google is basically trying to shift the entire platform toward professional execution. They turned a massive digital twin of the world into an agentic AI engine for global infrastructure. The technical foundation is (obviously) all about data. Google integrated 20-metre and 40-metre elevation contours globally. Engineers and urban planners now have instant access to the exact topographic context required for site planning anywhere on Earth. The data catalogue updates continuously to maintain the freshest imagery possible. Collaboration used to kill geospatial projects. Teams would lose momentum through stale materials or bad handoffs. Google fixed this by building frictionless data import systems. You can now drop KML, KMZ, and GeoJSON files directly onto the global map. Entire departments can align on a single source of truth, moving from a raw question to a definitive answer instantly. The biggest upgrade is the introduction of agentic geospatial intelligence. Users can open 'Ask Google Earth' and search massive satellite and Street View databases using natural language. You type a command, and the AI handles the manual data wrangling. It identifies new site locations and analyses infrastructure before you even open a spreadsheet.

Yohan

45,065 просмотров • 4 месяцев назад