Loading video...

Video Failed to Load

Go Home

We built a time machine for the web. Introducing Exa Snapshot: an index of 400 billion historical snapshots of webpages that lets you search as if it's the past. Snapshot is already being used for backtesting prediction models, RL at labs, exploring the pre-AI web, and more.

647,035 views • 1 day ago •via X (Twitter)

32 Comments

Exa's profile picture
Exa1 day ago

Snapshots are available to try in Exa today. Learn more:

Ishan Goswami's profile picture
Ishan Goswami1 day ago

the internal codename for this project was...

Mridul Singhai's profile picture
Mridul Singhai1 day ago

How does this differ from the venerable Wayback Machine from Internet Archive?

jacky's profile picture
jacky1 day ago

out of curiosity, how is this different from the internet archive's version?

Max Rovensky's profile picture
Max Rovensky1 day ago

You built

iFred's profile picture
iFred1 day ago

How much did you folks donate to the Internet Archive?

Jakub Hojsan's profile picture
Jakub Hojsan1 day ago

wait the audio kinda cooked

Jessica's profile picture
Jessica1 day ago

so cool seeing landing page taglines over the years

Steven Liss's profile picture
Steven Liss1 day ago

feature request: let me search the web 1 day/month/year in the future

&.'s profile picture
&.1 day ago

this is very cool

Bobir M's profile picture
Bobir M1 day ago

What is difference from waybackmachine ?

Jed White 💥♻️'s profile picture
Jed White 💥♻️1 day ago

Congrats and cool effort! Storage must be wild 🫡

khang's profile picture
khang1 day ago

wow

Aapakari's profile picture
Aapakari1 day ago

How often are pages recaptured?

r/69's profile picture
r/691 day ago

so you reinvented archive[.]org

Mustafa's profile picture
Mustafa1 day ago

Only 2 eras Chat, pre-chat

Aniket's profile picture
Aniket1 day ago

the underrated use case: checking what a competitor's pricing page said last quarter

Mudit Arora's profile picture
Mudit Arora1 day ago

glad we’re using you as our service

Phil Gara's profile picture
Phil Gara1 day ago

Wow really cool!

Akul Gupta's profile picture
Akul Gupta1 day ago

Super cool, gonna be super useful for research

Seth BRKV's profile picture
Seth BRKV1 day ago

No way! I thought about such system yesterday morning. How you can test AI inside of "real" internet, but months or years behind. And there it is (if I understood correctly)

.CitiZen.'s profile picture
.CitiZen.1 day ago

Woah

Rimas's profile picture
Rimas1 day ago

how are labs using the snapshots for rl?

echo4eva's profile picture
echo4eva1 day ago

so cool!

Harjass Gambhir's profile picture
Harjass Gambhir1 day ago

is there a filter for pre gpt pages for searching?

Krish Goel's profile picture
Krish Goel1 day ago

This was a major bottleneck when my friends and I (naively) tried to build a market prediction engine that needed to understand the market sentiment “then”. Love it.

AI News Daily's profile picture
AI News Daily1 day ago

A historical web index is a surprisingly useful primitive for AI research. Backtesting agents and models against old pages could make evaluations much closer to the messy conditions they face in the wild.

Dozer🚜's profile picture
Dozer🚜1 day ago

Wow dope idea

RISHI's profile picture
RISHI1 day ago

👀👀👀

Callum's profile picture
Callum1 day ago

super nice for evals!

Ryan's profile picture
Ryan1 day ago

400b snapshots turns search from "what does this page say?" into "when did this claim first appear?" way more useful for due diligence than another current-web wrapper.

Jason Garoutte's profile picture
Jason Garoutte1 day ago

AI is fundamentally about predictions. Having date-specific data is key.

Related Videos

🚨 THIS IS ACTUALLY INSANE Your AI agent can have access to the web. But if it can't reliably read what’s actually on the page, that access is almost useless. We looked at Firecrawl as the web layer for AI agents and the numbers are hard to ignore. The setup is simple: Give it a URL, search query, or website. Firecrawl handles the ugly part — crawling, scraping, rendering, extracting, and turning web content into something an AI model can actually use. The headline numbers: → 173,000+ GitHub stars → Search, scrape and interact with the web at scale → Supports web pages, PDFs, DOCX and other content → Structured data extraction for AI workflows → MCP support for connecting it directly to AI agents The workflow looks like this: Search → Scrape → Crawl → Extract → Feed the agent Three things stand out: 1. Scraping becomes an infrastructure layer Instead of maintaining your own pile of HTTP clients, parsers, browser automation and retry logic, you can treat web access as an API. 2. Agents get more than raw HTML The goal isn't just downloading a webpage. It's turning messy web content into clean context that an LLM can reason over. 3. The same layer works across different agent workflows Research agents. RAG pipelines. AI search. Competitive intelligence. Web-data extraction. The interesting shift: AI agents don't just need better models. They need better access to the information those models are supposed to reason about. Firecrawl is building that layer. Save this repo.

Vikas gupta

18,053 views • 13 days ago