
Shreya Shankar
@sh_reya • 55,616 subscribers
Incoming assistant professor @CSDatCMU @CMUDB. All about data, HCI, and AI. Created https://t.co/PmuOqAXVgS and https://t.co/8MQt4na2cj.
Videos

V2, at 4x speed. Can't wait to put this out in the open (soon)
Shreya Shankar56,855 Aufrufe • vor 3 Monaten

No one's built an interactive way to dig through the ~3k newly released Epstein-related emails—so we did! Here's a free, searchable DocETL-powered interface that lets journalists, researchers, and anyone else explore the material without wading through raw data dumps 🔎
Shreya Shankar100,483 Aufrufe • vor 10 Monaten

I love agentic map-reduce, but I don't love Claude's workflows feature. It's hard to steer the workflow, and it's not data-driven enough (e.g., reduce groups are often predefined, not emergent from the map outputs) HOWEVER...I was inspired by its UI to ship something similar for DocETL. Satisfying way to visualize progress when using the CLI 😆
Shreya Shankar38,431 Aufrufe • vor 3 Monaten

What if you could understand what's buried in tens of thousands of messy text documents — without writing a single line of code? We shipped a Claude Code skill for DocETL. I asked it to scrape Hacker News 'What Are You Working On?' comments over the last 15 years, figure out what people are building, and visualize the trends. Then I went to grab coffee. Came back an hour later to a full dashboard (the coffee cost more than the analysis):
Shreya Shankar44,124 Aufrufe • vor 8 Monaten

🔍How do we make sense of messy, real-world documents? We (the EPIC Lab at UC Berkeley) are building DocETL, an open-source system for LLM-native data processing. We've started to create a showcase of demos. Here's our first: an analysis of public feedback on US AI strategy. 🇺🇸
Shreya Shankar35,294 Aufrufe • vor 1 Jahr

DocETL is a system we’ve been building at Berkeley for the past two years to make large-scale unstructured data analysis reliable and efficient. It powers our broader stack—used by journalists, public defenders, and researchers—to extract, transform, and reason over messy documents with LLMs. As part of making the DocETL ecosystem easier to use, we’re introducing a natural language–to–pipeline generator! Our hosted version is free to use BUT we're collecting the data so we can build better tools.
Shreya Shankar14,428 Aufrufe • vor 10 Monaten
Keine weiteren Inhalte verfügbar