Loading video...
Video Failed to Load
We built a time machine for the web. Introducing Exa Snapshot: an index of 400 billion historical snapshots of webpages that lets you search as if it's the past. Snapshot is already being used for backtesting prediction models, RL at labs, exploring the pre-AI web, and more.
647,035 views • 1 day ago •via X (Twitter)
32 Comments

Snapshots are available to try in Exa today. Learn more:

the internal codename for this project was...

How does this differ from the venerable Wayback Machine from Internet Archive?

out of curiosity, how is this different from the internet archive's version?

You built

How much did you folks donate to the Internet Archive?

wait the audio kinda cooked

so cool seeing landing page taglines over the years

feature request: let me search the web 1 day/month/year in the future

this is very cool

What is difference from waybackmachine ?

Congrats and cool effort! Storage must be wild 🫡

wow

How often are pages recaptured?

so you reinvented archive[.]org

Only 2 eras Chat, pre-chat

the underrated use case: checking what a competitor's pricing page said last quarter

glad we’re using you as our service

Wow really cool!

Super cool, gonna be super useful for research

No way! I thought about such system yesterday morning. How you can test AI inside of "real" internet, but months or years behind. And there it is (if I understood correctly)

Woah

how are labs using the snapshots for rl?

so cool!

is there a filter for pre gpt pages for searching?

This was a major bottleneck when my friends and I (naively) tried to build a market prediction engine that needed to understand the market sentiment “then”. Love it.

A historical web index is a surprisingly useful primitive for AI research. Backtesting agents and models against old pages could make evaluations much closer to the messy conditions they face in the wild.

Wow dope idea

👀👀👀

super nice for evals!

400b snapshots turns search from "what does this page say?" into "when did this claim first appear?" way more useful for due diligence than another current-web wrapper.

AI is fundamentally about predictions. Having date-specific data is key.
