design:archives
Differences
This shows you the differences between two versions of the page.
| Both sides previous revisionPrevious revisionNext revision | Previous revision | ||
| design:archives [2026/08/13 22:44] – Extend: what an archive preserves and what it does not (static capture, escapes, coverage, vantage), corpus-backed Use in Publications, best-practice checklist, anachronism trap, self-recording alternative; correct five stale/wrong facts (Wayback size, Sa karel.kubicek.claude | design:archives [2026/09/04 17:29] (current) – Citekey consolidation 2026-09-04: repoint duplicate bibliography keys (fouad2022my/boettger2025_regional/ahmad2026_ipfp/bouhoula2024automated/lerner2016internet) to the kept key; no prose or figure changed. Authored by Claude karel.kubicek.claude | ||
|---|---|---|---|
| Line 3: | Line 3: | ||
| A web archive lets you measure a page you cannot visit — because the date has passed, because the site is gone, or because you want everyone else to be able to measure exactly the same bytes. That is a genuinely different instrument from a crawler, and it fails in genuinely different ways. This page is about **what an archive can and cannot answer**, not about how archiving works. | A web archive lets you measure a page you cannot visit — because the date has passed, because the site is gone, or because you want everyone else to be able to measure exactly the same bytes. That is a genuinely different instrument from a crawler, and it fails in genuinely different ways. This page is about **what an archive can and cannot answer**, not about how archiving works. | ||
| - | Researchers have used archives for longitudinal analyses of privacy {[lerner2016internet, | + | Researchers have used archives for longitudinal analyses of privacy {[lerner2016_internet, |
| Two numbers frame everything that follows. In our corpus of seven security and privacy venues (see [[#Use in Publications]]), | Two numbers frame everything that follows. In our corpus of seven security and privacy venues (see [[#Use in Publications]]), | ||
| Line 16: | Line 16: | ||
| - {[hantke2023you]} — the systematic study. Compares seven archives, quantifies what goes wrong in the Internet Archive, and ends with a best-practice list. If you read one, read this. | - {[hantke2023you]} — the systematic study. Compares seven archives, quantifies what goes wrong in the Internet Archive, and ends with a best-practice list. If you read one, read this. | ||
| - | - {[lerner2016internet]} — the founding archive-based measurement in this literature (tracking, 1996–2016), | + | - {[lerner2016_internet]} — the founding archive-based measurement in this literature (tracking, 1996–2016), |
| - {[zhu2025_toward]} — why a static crawl misses half the page, measured in 2025, from the archive operators' | - {[zhu2025_toward]} — why a static crawl misses half the page, measured in 2025, from the archive operators' | ||
| - {[singh2026_empire]} — the model recent application: | - {[singh2026_empire]} — the model recent application: | ||
| Line 58: | Line 58: | ||
| ==== Escapes: when an archived page measures today ==== | ==== Escapes: when an archived page measures today ==== | ||
| - | URL rewriting inside a replayed page is not perfect, chiefly because it cannot rewrite URLs that JavaScript builds at runtime. When it fails, the replayed page issues a request to the **live** web. Lerner et al. measured this on archived snapshots of the Alexa top 500 and found that **16.1% of all requests attempted to escape the archive** — 17.2% of unique URLs and 11.5% of domains {[lerner2016internet]}. That measurement is from 2016, on 2015-era replay, and **nobody has repeated it since**; treat it as evidence that escapes matter, not as today' | + | URL rewriting inside a replayed page is not perfect, chiefly because it cannot rewrite URLs that JavaScript builds at runtime. When it fails, the replayed page issues a request to the **live** web. Lerner et al. measured this on archived snapshots of the Alexa top 500 and found that **16.1% of all requests attempted to escape the archive** — 17.2% of unique URLs and 11.5% of domains {[lerner2016_internet]}. That measurement is from 2016, on 2015-era replay, and **nobody has repeated it since**; treat it as evidence that escapes matter, not as today' |
| To put that in proportion: in the same table, robots.txt exclusion accounted for 2.0% of requests and the resource simply not being archived for 1.4%. **Escapes were by far the largest single failure mode — eight times robots exclusion.** Their crawler blocked those requests, which is the right default: an escaped request measures the present under a past timestamp. If your tooling does not block them, your "2014 measurement" | To put that in proportion: in the same table, robots.txt exclusion accounted for 2.0% of requests and the resource simply not being archived for 1.4%. **Escapes were by far the largest single failure mode — eight times robots exclusion.** Their crawler blocked those requests, which is the right default: an escaped request measures the present under a past timestamp. If your tooling does not block them, your "2014 measurement" | ||
| Line 290: | Line 290: | ||
| ===== Open Questions ===== | ===== Open Questions ===== | ||
| - | * <wrap todo> | + | <WRAP todo> |
| - | * <wrap todo>**Nobody has re-measured static-vs-dynamic archive bias at scale.** The 73.9%/95.3% tracker gap {[hantke2023you]} rests on 2,026 sites, limited by rate-limiting. Zhu et al.'s {[zhu2025_toward]} 45.8% fidelity gap is measured on their own crawls, not on the Internet Archive' | + | * **The archive comparison needs redoing, and the instrument for redoing it shrank.** Memento Time Travel was decommissioned in 2025; MemGator survives but reaches 13 archives rather than 30-plus, so Hantke et al.'s enumeration {[hantke2023you]} cannot be repeated as written and its 2022 coverage numbers are now four years old. Nobody has published a current census of which public web archives still crawl. |
| - | * <wrap todo>**Monoculture risk is unquantified.** Roughly two thirds of archive-using papers depend on one archive, one crawler and one vantage point. What that shared bias does to the field' | + | * **Nobody has re-measured static-vs-dynamic archive bias at scale.** The 73.9%/95.3% tracker gap {[hantke2023you]} rests on 2,026 sites, limited by rate-limiting. Zhu et al.'s {[zhu2025_toward]} 45.8% fidelity gap is measured on their own crawls, not on the Internet Archive' |
| - | * <wrap todo>**Archive-based measurement of regional phenomena has no instrument.** Both major sources crawl from the US. For GDPR-era consent, geo-blocking or regional advertising, | + | * **Monoculture risk is unquantified.** Roughly two thirds of archive-using papers depend on one archive, one crawler and one vantage point. What that shared bias does to the field' |
| + | * **Archive-based measurement of regional phenomena has no instrument.** Both major sources crawl from the US. For GDPR-era consent, geo-blocking or regional advertising, | ||
| + | </WRAP> | ||
| ===== Related Pages ===== | ===== Related Pages ===== | ||
design/archives.1786661095.txt.gz · Last modified: by karel.kubicek.claude
