Загрузка видео...

Не удалось загрузить видео

На главную

It's happening. Last night, I started downloading financial data from the SEC. • income statements • balance sheets • cash flow statements 10,000+ public companies. 3 million rows in total. I’m using multiple worker nodes to pull, parse, and clean the data. The orchestration is beautiful. This week, I...

313,830 просмотров • 2 лет назад •via X (Twitter)

Комментарии: 10

Фото профиля Asfi
Asfi2 лет назад

SEC EDGAR’s file structure consistency is so under appreciated. Really happy you found the right method to get it from the source. Every day new files are uploaded to the SEC. Creating an API that does minimal standardization but delivers statements as presented will be 🔥🔥🔥

Фото профиля virat
virat2 лет назад

Yes sir. Finally cracking the standardized financial data problem

Фото профиля Greg Kamradt
Greg Kamradt2 лет назад

Build in public the podscan (@arvidkahl) for financial data and its analysis and you’ll have a really fun ride

Фото профиля virat
virat2 лет назад

@arvidkahl Arvid is a legend. Loved The Embedded Entrepreneur.

Фото профиля matthew walker
matthew walker2 лет назад

Gonna put it into a knowledge graph? Graph RAG?

Фото профиля virat
virat2 лет назад

That actually sounds dope

Фото профиля Greg Osuri 🇺🇸 deAI Summer 2025
Greg Osuri 🇺🇸 deAI Summer 20252 лет назад

Will sponsor GPUs if you’re planning to opensource what you’re working on.

Фото профиля Marco Goldin
Marco Goldin2 лет назад

parsing sec XML has been painful for me: namespaces, structure changing over the years, gaap taxonomy (changing over time). I'm curious: how did you find a safe way to handle all of this?

Фото профиля virat
virat2 лет назад

Lots of coding and staring at the data. Will release design diagrams, etc. this week!

Фото профиля Robbie Bouschery
Robbie Bouschery2 лет назад

Have you seen

Похожие видео

Building Data Pipelines has levels to it: - level 0 Understand the basic flow: Extract → Transform → Load (ETL) or ELT This is the foundation. - Extract: Pull data from sources (APIs, DBs, files) - Transform: Clean, filter, join, or enrich the data - Load: Store into a warehouse or lake for analysis You’re not a data engineer until you’ve scheduled a job to pull CSVs off an SFTP server at 3AM! level 1 Master the tools: - Airflow for orchestration - dbt for transformations - Spark or PySpark for big data - Snowflake, BigQuery, Redshift for warehouses - Kafka or Kinesis for streaming Understand when to batch vs stream. Most companies think they need real-time data. They usually don’t. level 2 Handle complexity with modular design: - DAGs should be atomic, idempotent, and parameterized - Use task dependencies and sensors wisely - Break transformations into layers (staging → clean → marts) - Design for failure recovery. If a step fails, how do you re-run it? From scratch or just that part? Learn how to backfill without breaking the world. level 3 Data quality and observability: - Add tests for nulls, duplicates, and business logic - Use tools like Great Expectations, Monte Carlo, or built-in dbt tests - Track lineage so you know what downstream will break if upstream changes Know the difference between: - a late-arriving dimension - a broken SCD2 - and a pipeline silently dropping rows At this level, you understand that reliability > cleverness. level 4 Build for scale and maintainability: - Version control your pipeline configs - Use feature flags to toggle behavior in prod - Push vs pull architecture - Decouple compute and storage (e.g. Iceberg and Delta Lake) - Data mesh, data contracts, streaming joins, and CDC are words you throw around because you know how and when to use them. What else belongs in the journey to mastering data pipelines?

Zach Wilson

16,688 просмотров • 1 год назад

PhD Students – How to automatically extract data from papers for your literature review? Extracting relevant data from papers is challenging. However, this process can be automated. Meet AnswerThis – a tool that extracts data in seconds. Here is how it works. 1. Go to and log in. 2. After logging in, click on 𝐸𝑥𝑡𝑟𝑎𝑐𝑡 𝑑𝑎𝑡𝑎. 3. Then click on 𝑈𝑝𝑙𝑜𝑎𝑑 𝑃𝐷𝐹 and upload your papers. 4. These are the papers from which you want to extract data. 5. After uploading papers, select data you want to extract. 6. The predefined options are - Key findings - Research gaps - Methodology - Limitations - Future work - Contributions - Practical implications 7. You can also extract custom data e.g., dataset used. 8. For example, I want to extract methodology used in these papers. 9. I selected 𝑀𝑒𝑡ℎ𝑜𝑑𝑜𝑙𝑜𝑔𝑦 and clicked on 𝐴𝑑𝑑 𝐶𝑜𝑙𝑢𝑚𝑛. 10. AnswerThis extract data about methodology used in the papers. 11. You can change data view from normal to Table View. 12. For this, scroll back to top and click on 𝑇𝑎𝑏𝑙𝑒 𝑉𝑖𝑒𝑤. 13. Now for instance, you want to extract more data from these papers. 14. Go back to the top and click on 𝐸𝑥𝑡𝑟𝑎𝑐𝑡 𝑑𝑎𝑡𝑎. 15. Select the data type you want to extract. 16. For example, I want to extract data about future work. 17. So I click on 𝐹𝑢𝑡𝑢𝑟𝑒 𝑊𝑜𝑟𝑘 and then clicked on 𝐴𝑑𝑑 𝑐𝑜𝑙𝑢𝑚𝑛. 18. AnswerThis extracted data about future work from the papers. 19. After extracting the desired data, you can export it. 20. Select the data you want to extract. 21. Then click on 𝐸𝑥𝑝𝑜𝑟𝑡 𝑑𝑎𝑡𝑎. 22. Your data will be exported in CSV format. You can then analyze this data for your literature review. Try AnswerThis today: Anything you'd like to add?

Faheem Ullah

21,390 просмотров • 9 месяцев назад

PhD Students – How to extract data from papers for your literature review in seconds? Extracting data from papers takes a lot of time. You can automate this process with Bohrium 𝐇𝐨𝐰 𝐭𝐨 𝐚𝐮𝐭𝐨𝐦𝐚𝐭𝐢𝐜𝐚𝐥𝐥𝐲 𝐞𝐱𝐭𝐫𝐚𝐜𝐭 𝐝𝐚𝐭𝐚 𝐟𝐫𝐨𝐦 𝐩𝐚𝐩𝐞𝐫𝐬? 1. Go to and log in 2. Click on 𝐾𝑛𝑜𝑤𝑙𝑒𝑑𝑔𝑒 𝐵𝑎𝑠𝑒 from the left menu 3. Upload the papers you selected for literature review 4. You will see the following option against each paper - Read PDF - Key Takeaway - AI Poster 5. Click on 𝑅𝑒𝑎𝑑 𝑃𝐷𝐹 for the first paper in your list 6. Write a prompt for the data you want to extract 7. For example, you can enter datasets, methodology etc. 8. It will extract the required data from the paper 9. If you want to extract Key Takeaways from the paper 10. Go back and click on 𝐾𝑒𝑦 𝑇𝑎𝑘𝑒𝑎𝑤𝑎𝑦𝑠 11. Bohrium will extract Key Takeaways from the paper 12. In addition to this, you also have 2 more options - AI Poster - Podcast 13. Click on 𝐴𝐼 𝑃𝑜𝑠𝑡𝑒𝑟 and it will create a poster for you 14. This is the poster based on the given research paper 15. If you click on 𝑃𝑜𝑑𝑐𝑎𝑠𝑡, it will convert the paper to audio 16. You can listen to the paper instead of reading it Repeat this cycle for all the papers in your pool. You will end up with the required data. You can use this data to write your literature review Try Bohrium today for FREE: Anything you’d like to add?

Faheem Ullah

13,433 просмотров • 11 месяцев назад