User Tools

Site Tools


design:longitudinal

Repeating a Measurement Over Time

A longitudinal web measurement is not one measurement that lasts a long time. It is two or more measurements that have to be comparable to each other, and everything difficult about it follows from that. The web changes; so does your list, your browser, your vantage point, your filter list and the library you parse HAR files with. If wave two differs from wave one for any of those reasons, your trend line is a measurement of your own toolchain.

This page is about keeping wave two comparable to wave one. It is where the other design pages converge: Sampling and Website selection decide the frame, Crawling location the vantage point, Stateful stateless the profile, Archives whether you can go back in time at all, and Regression what to do with rows that are not independent because the same site contributed several of them. What is here is the part none of those pages owns: what has to be held fixed between waves, what cannot be held fixed however hard you try, and what the field actually reports.

The one thing to take away. Before you can call a year-on-year difference a trend, you need to know your repeat-run noise floor — how much your number moves when nothing changes but the day. One experiment in this literature has measured it: Demir et al. ran the same crawler over the same 1,000 sites on twelve consecutive days, and the observed tracking-request count moved by up to 27% while the count of distinct tracking domains moved 3.5% [1Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)]. One study, one outcome pair — but it is the only measurement there is, and it says the event count and the entity set behave very differently. Establish your own noise floor before you report a trend, and if you cannot, say which of the two kinds of outcome you are reporting.

Of the 1,120 papers in our corpus of seven security and privacy venues that ran a crawl, 250 (22.3%) say they crawled more than once — and 63.5% never say how many times they crawled at all, so 250 is a floor. Of those 250, 23 (9.2%) state all four of the list version, the browser, the vantage point and the statefulness of the profile — and the papers whose panels run longest are the ones that pin least. See Use in Publications.

What to Read First

Five papers, in the order that gets a newcomer productive fastest:

  1. [1Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)] — read this one first. 117 papers hand-coded against 18 reproducibility criteria, plus a 4.5M-page experiment that varies the setup on purpose and measures what moves. It is where your noise floor comes from and where the reporting checklist comes from.
  2. [2Demir, Nurullah; Hörnemann, Jan; Große-Kampmann, Matteo; Urban, Tobias; Pohlmann, Norbert; Holz, Thorsten; Wressnegger, Christian (2023): "On the Similarity of Web Measurements Under Different Experimental Setups", in: Proceedings of the ACM Internet Measurement Conference, pp. 356-369. ACM DOI 10.1145/3618257.3624795 is listed by DBLP but was not registered with the DOI resolver as of 2026-08-12; the DOI above resolves to the authors' institutional record of the same paper (DOI)] — the follow-up that quantifies how little two crawler configurations agree: “when comparing two different profiles, 48% of the underlying data varies”.
  3. [3Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)] — vantage point and browser configuration as experimental variables. The reason “we crawled from our university” is a design decision and not an implementation detail. See Crawling location.
  4. [4Singh, Sachin Kumar; Mahmud, Faisal; Ricci, Robert; Siby, Sandra (2026): "The Empire Strikes Back (at Your Privacy): An Archaeology of Tracking on Government Websites", Proceedings on Privacy Enhancing Technologies 2026(2):108-126. (DOI)] — the model treatment of classifier anachronism: one filter-list version held fixed for “all years” of a 1996–2025 archive, and then validated against the historical versions it replaced. See The anachronism trap.
  5. [5Nenadić, Luka; Rodriguez, David; Calandrino, Joseph A. (2026): "Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact", Proceedings on Privacy Enhancing Technologies 2026(4):703-723. (DOI)] — the model treatment of the statistics: a balanced panel of 11,800 websites observed twice, difference-in-differences, standard errors clustered on the website. See Regression.

Why Wave Two Is Not Wave One

Your noise floor

Five measurements of how much a result moves when the web does not:

What varied What moved Source
Nothing but the day: the same 1,000 sites, twelve consecutive daily crawls tracking requests varied by up to 27% (max 80,274 on day 3, min 58,951 on day 9; “The standard deviation of such requests is 8,203”); distinct tracking domains varied 3.5% [1Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)]
Browser configuration only “the identified trackers on pages can vary by 25%” [1Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)]
Crawler profile, pairwise “48% of the underlying data varies”; only 32% of cookies appear in all profiles and 42% in only one [2Demir, Nurullah; Hörnemann, Jan; Große-Kampmann, Matteo; Urban, Tobias; Pohlmann, Norbert; Holz, Thorsten; Wressnegger, Christian (2023): "On the Similarity of Web Measurements Under Different Experimental Setups", in: Proceedings of the ACM Internet Measurement Conference, pp. 356-369. ACM DOI 10.1145/3618257.3624795 is listed by DBLP but was not registered with the DOI resolver as of 2026-08-12; the DOI above resolves to the authors' institutional record of the same paper (DOI)]
Vantage point (university, residential, cloud) “Around 5% of content-providing domains show significant measurement bias across VP” [3Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)]
Crawler technology, several crawlers on the same task “variation of over 16% in the number of successful page loads”; third-party prevalence rankings change with the crawler [6Ahmad, Syed Suleman; Dar, Muhammad Daniyal; Zaffar, Muhammad Fareed; Vallina-Rodriguez, Narseo; Nithyanand, Rishab (2020): "Apophanies or Epiphanies? How Crawlers Impact Our Understanding of the Web", in: Proceedings of The Web Conference, pp. 271-280. (DOI)]
Profile region, same study, same days “privacy measurements and analyses can vary up to 65% depending on the region” [1Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)]

The authors of the first row draw the conclusion this page is built on:

studies that analyze the ecosystem will find similar results, while studies that aim to analyze the extent of a tracking phenomenon might see different results based on the measurement day

Read as design advice: in that experiment the set-valued outcome survived repetition and the count-valued one did not. The entity set — which tracking domains are present — moved 3.5% across twelve identical days. The event count — how many tracking requests were made — moved 27%. That is one study on one pair of outcomes, so it does not license a general law; what it does license is a design question you should answer before you crawl. If your outcome is a count and a 27%-scale wobble would swamp the effect you are looking for, you need either many more waves than you planned, or a paired within-site design that differences the day out, or a different outcome.

The four pins

Four things decide whether two waves measure the same thing. None of them is difficult; all of them are routinely left unstated.

  1. The list. A ranking list is a moving frame. “The Tranco top 1M” in 2022 and in 2026 are not the same 1M domains, and — see below — they are not even built from the same source data.
  2. The browser. Browsers change what they permit, block and expose on a six-week cycle. Third-party cookie behaviour, storage partitioning, referrer policy defaults and fingerprinting protections have all shifted mid-decade. A crawl driven by “Chrome” in 2023 and “Chrome” in 2026 is two instruments.
  3. The vantage point. Consent banners are geo-targeted, content is geoblocked, and datacenter address space is treated differently from residential. Changing cloud region between waves changes the measurement. See Crawling location.
  4. The crawl protocol. Statefulness, consent action, interaction depth, subpages per site, timeouts, and how failures are retried. Stateful stateless shows that the statefulness axis alone moves third-party counts in a known direction.

The fifth pin nobody names: the classifier

The four above are about the input. The instrument that turns raw observations into a number moves too, and it moves faster than the browser does. EasyList is updated most days — the copy served on 2026-08-27 carried a version stamp minted at 01:34 UTC that same morning, and the repository takes commits daily1) — and Disconnect, Cookiepedia, category services and vendor databases all change under you too. Label a 2021 crawl and a 2026 crawl with “EasyList” and the difference between them is partly list growth.

Singh et al. [4Singh, Sachin Kumar; Mahmud, Faisal; Ricci, Robert; Siby, Sandra (2026): "The Empire Strikes Back (at Your Privacy): An Archaeology of Tracking on Government Websites", Proceedings on Privacy Enhancing Technologies 2026(2):108-126. (DOI)] are the worked example of doing this right, and their validation figure is the one to remember: on a random sample of 10,000 URLs archived in 2006, the 2006 EasyList flagged 67 and the June 2025 EasyList flagged 50, with only 6 in common. Two versions of the same list, on the same URLs, agreeing on six items. Their design — one fixed list version for every year, then an explicit check of what the historical versions would have done at anchor years — is the pattern to copy. The anachronism trap covers it in full.

The frame moves even when you pin it

Pinning the list version is necessary and not sufficient, because the list itself is a derived product whose inputs change.

Tranco's permanent list IDs work exactly as advertised — the identifier GVWK, generated on 5 November 2019, still resolves in 2026 and still returns its configuration.2) But that configuration is the interesting part:

Default daily list ID Providers it was built from
5 Nov 2019 GVWK Alexa, Umbrella, Majestic, Quantcast
1 Jan 2022 XVWN Alexa, Umbrella, Majestic
1 Jan 2023 3VYXL Alexa, Umbrella, Majestic, Farsight
1 Jan 2024 V929N CrUX, Farsight, Majestic, Radar, Umbrella
26 Aug 2026 46W9X CrUX, Farsight, Majestic, Radar, Umbrella

Quantcast dropped out, Farsight came in, and then in 2023–24 Alexa was replaced by CrUX and Cloudflare Radar.3) So “we used the Tranco top 1M in both waves” does not give you a fixed frame across 2022–2024: the default list is a different instrument on either side. Tranco lets you fix the provider set when you generate a custom list, but it cannot resurrect a retired provider, so there is no configuration that makes a 2019 list and a 2026 list commensurable.

That leaves two honest options, and you should say which you took:

  • Follow a cohort. Fix the domain set at wave one and re-crawl exactly that set every wave. This is a panel. It answers “did these sites change?” and it is the only design in which a per-site paired comparison is available. It also starts decaying immediately: sites die, and the ones that die are not a random sample of the ones that live (Repeated crawls, and what they say about the sites they lost).
  • Redraw the list each wave. This is a repeated cross-section. It answers “did the popular web change?”, which is a different and often more interesting question, but it confounds real change with frame churn, and a per-site comparison is not available for the sites that entered or left.

Doing both is not free — the union of the fixed cohort and the fresh list is a bigger crawl than either — but the marginal cost is a few thousand extra targets on a list you were already going to draw, and reporting the two side by side is what turns a reviewer's objection into a robustness check.

The things you cannot pin

  • Server-side change. Sites redesign, migrate CMS, adopt a consent platform. This is the signal, not the noise, but it means a per-site outcome can change for reasons unrelated to your research question.
  • Attrition. Sites stop resolving, start blocking your address range, or move behind bot management. Biases measures how invisible this is: of the 250 repeated-crawl papers here, 2 (0.8%) use the words attrition or survivorship anywhere in the text.
  • Your own drift. Parser versions, eTLD+1 lists (the Public Suffix List takes commits most weeks4) and a library that vendors a snapshot only pins it if you pin the library), timeouts tuned between waves, and a codebase that got better at extracting the thing you are counting. A wave-two improvement in your extractor looks exactly like a wave-two increase on the web.

Choosing a Design

Design What it answers What it costs
Repeated cross-section — redraw the frame each wave how the population changed frame churn, as above
Balanced panel — fix the cohort at wave one how these sites changed attrition, as above
Rolling panel — fixed core plus a refreshed top-up both, if you report them separately two denominators to keep straight, and you must say which figure uses which
Retrospective — archives or someone else's dataset history you did not plan to collect the archive's crawler, vantage point and coverage become yours (Archives)
Continuous observatory — a platform that never stops fine-grained timing of events infrastructure and maintenance far beyond one paper

Two practical notes on cadence and span.

Cadence has to beat the thing you are measuring. Consent-banner configuration, ad creatives and filter-list coverage change on a scale of days to weeks; CSP and TLS deployment change on a scale of years. A monthly panel cannot see a two-week campaign, and a daily panel over three weeks cannot see a policy effect. Say what change rate you assumed.

A high-frequency short run is not a long panel, and the corpus conflates them. Of the 250 repeated-crawl papers, 35 (14.0%) have a total stated span under one month and 70 (28.0%) have one under two months — often daily crawls over a fortnight, which measure day-to-day variance rather than change over time. Both are legitimate; they are not the same design, and a paper that says only “we crawled repeatedly” has told you neither.

Pinning in Practice

The browser

This is the pin the field is worst at, and it is the one with a clean modern answer.

  • Chrome. Chrome for Testing publishes immutable, non-auto-updating builds of Chrome and the matching ChromeDriver, addressable by exact version, with a JSON endpoint listing every one. As of 2026-08-27 it lists 2,484 versions from 113.0.5672.0 to 154.0.8026.0, so any Chrome from the 113 line (spring 2023) onward can be re-instantiated exactly; anything older is not available through this channel.5) It is the correct default for any new repeated crawl, and it costs one line of setup.
  • Firefox. Every release is retained at archive.mozilla.org/pub/firefox/releases/. If you drive Firefox through OpenWPM you get the pin for free and should cite it: OpenWPM installs a specific unbranded Firefox build identified by a Mozilla release tag, and each OpenWPM release states which. v0.36.0 (24 August 2026) pins Firefox 154.0.6) So “OpenWPM v0.36.0” is a browser version statement, and “OpenWPM” alone is not. Demir et al. write it the right way: “We use the popular Open-WPM Framework [21] (v0.15.0 - Firefox version 88)” [1Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)].
  • Playwright already does this for you, if you pin Playwright. Its browsers.json pins an exact browser build per Playwright release — the Chromium entry is literally titled “Chrome for Testing” and carries a browserVersion, and Firefox and WebKit carry an internal revision.7) So a pinned Playwright version is a browser pin, in the same way an OpenWPM release tag is — but only if you pin Playwright rather than installing whatever npm resolves.
  • The driver too. A geckodriver or ChromeDriver upgrade can change what the browser accepts; OpenWPM's own v0.36.0 notes a geckodriver 0.37.1 change that made every browser fail to launch. Record the driver version alongside the browser.

The environment

  • A container digest, not a tag. openwpm:latest is not a pin and openwpm:v0.36.0 is only as stable as the base image; the digest is the pin. No paper in this corpus mentions one: a full-text sweep for image digest or an sha256: image reference over all 5,855 readable papers returns 0 papers, against 293 (5.0%) that mention Docker and 308 (5.3%) that mention Docker or containerisation. A paper could of course pin a digest in its repository without saying so in the text, so read this as a reporting figure and not as a practice figure. The exact regexes are in the report script and its output.
  • Whole-measurement bundles. WebREC and the .web format [7Hantke, Florian; Snyder, Peter; Haddadi, Hamed; Stock, Ben (2025): "Web Execution Bundles: Reproducible, Accurate, and Archivable Web Measurements", in: Proceedings of the USENIX Security Symposium. (Link)] package a measurement — not just a page — so it can be re-run later without recrawling; the authors report that 70% of the papers in a 2024 crawling SoK could have run with WebREC unmodified. It is the right idea and it is brand new: two papers in this corpus mention it and neither uses it — one is the WebREC paper, the other cites it in its reference list — and the repository has single-digit stars. Treat it as promising rather than established.

The list

Tranco's API takes both a list id and a date, so a wave can be pinned or a past wave recovered:

# metadata and download URL for a specific list id
curl -s https://tranco-list.eu/api/lists/id/GVWK
 
# the default daily list as it stood on a given day
curl -s https://tranco-list.eu/api/lists/date/20191105

Both were verified working on 2026-08-27. Record the id in the paper, not just the date: Versioning: Making the Draw Redrawable shows a study that publishes one id per wave, which is the pattern to copy.

Almost nobody does. Of the 266 papers here that name Tranco as a population source, 159 (59.8%) state some version or date and 23 (8.6%) cite a permanent list ID — a rate that has not moved much since 2020 (see Use in Publications). The mechanism has existed since 2019 and the field has not adopted it.

The classifier

Pick one version of every labelling resource, use it for every wave, and state it. The mechanics are easy: every EasyList distribution carries its own identity in the header,

[Adblock Plus 2.0]
! Version: 202608270134
! Title: EasyList
! Last modified: 27 Aug 2026 01:34 UTC
! Expires: 4 days (update frequency)
! Commit: a87ab3cae6d4db91c77c1eded73a9bb361af7b57

Record the ! Version: and ! Commit: lines and the exact list is recoverable from the git history — which begins on 15 May 2016, so for anything earlier you are back to the Wayback Machine, as Singh et al. were.8)

The same discipline applies to any other labelling resource — a category service, a vendor database, a model checkpoint. If it has a version, record it; if it has none, record the date you fetched it and archive the copy you used.

The vantage point

Record the country, the infrastructure type and the region identifier, and re-use the same one. If the waves are simultaneous across several locations, say so; sequential crawls from different countries confound location with time. What to Report has the full list.

Datasets Somebody Else Already Collected

Before building a panel, check whether the series exists. Some do, and none of them is free of the problems above.

  • HTTP Archive — monthly crawl of millions of URLs, desktop and emulated mobile, HAR files and BigQuery tables, back to 2010. The URL list comes from CrUX and is re-drawn every month, so an HTTP Archive time series is a repeated cross-section, not a panel; intersect the URL sets yourself if you need a cohort. Runs on the second Tuesday of each month from Google Cloud regions in the US.9)
  • CrUX — monthly, origin-level, real-user sourced. It is the frame under both HTTP Archive and modern Tranco, so it is not an independent check on either.
  • The Wayback Machine and Common Crawl — the only way to measure before you started. Read that page before using either: the replay problems are severe and specific.
  • Censorship observatories — Censored Planet [8Raman, Ram Sundara; Shenoy, Prerana; Kohls, Katharina; Ensafi, Roya (2020): "Censored Planet: An Internet-wide, Longitudinal Censorship Observatory", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] and ICLab [9Niaki, Arian Akhavan; Cho, Shinyoung; Weinberg, Zachary; Hoang, Nguyen Phong; Razaghpanah, Abbas; Christin, Nicolas; Gill, Phillipa (2020): "ICLab: A Global, Longitudinal Internet Censorship Measurement Platform", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] are continuous, longitudinal and public, and they are the model for what a maintained observatory looks like in this field.

A cautionary example from the longest-running public dataset in this field. HTTP Archive's FAQ answers “What changes have been made to the test environment that might affect the data?” with a link to a machine-readable changelog.json. That link 404s. The file still exists on the project's master branch, and it contains 12 entries whose most recent is 1 June 2017 — so nine years of changes to the test environment are not recorded there at all.10) If you build a trend on a public dataset, reconstruct its methodology history yourself and publish what you found. Do not assume the maintainers did.

Use in Publications

The figures below come from a structured extraction over 5,859 full-text papers from CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026. Every block names its own population; sentinel values (not-stated, not-applicable) are counted as silence and never as an answer. The 2025 and 2026 venue-years are provisional — CCS and IMC 2026 have not been held, and IEEE S&P and WWW 2026 abstracts are not yet in the selection source — so they are under-represented by construction. Methodology and limitations are at the end of this section; the full query log is at longitudinal.

Throughout, a paper is repeated when it ran a crawl and any stated measurement period has more than one snapshot. That is the same rule Biases uses, so the two pages agree on the 250.

Most crawls happen once, and most do not say

Snapshots stated Papers Share of 1,120 crawling papers
Not stated at all 711 63.5%
Exactly one 159 14.2%
2–3 93 8.3%
4–6 41 3.7%
7–12 33 2.9%
13–52 46 4.1%
More than 52 37 3.3%

Nearly two thirds of crawling papers never say how many times they crawled. The top band is also unreliable as a count of waves: nine papers carry a snapshot count above 1,000, and spot-checking them finds page counts and record counts mis-extracted into the field. The bands are usable; the raw maximum is not.

Is the measurement period even stated

Field Of 5,118 empirical papers Of 1,120 crawling papers
Start date 56.3% 71.7%
End date 56.4% 71.0%
Both 54.2% 69.3%
Number of snapshots 24.2% 36.5%
Cadence 25.5% 34.7%

Crawling papers are better than the field average on every row and still leave three in ten without a date range you could situate the measurement in.

How long, and how often

Longest stated span, of the 250 repeated crawling papers:

Span Papers Share
Under 1 month 35 14.0%
1–2 months 35 14.0%
3–5 months 41 16.4%
6–11 months 26 10.4%
1–2 years 27 10.8%
2–5 years 23 9.2%
5 years or more 18 7.2%
No parsable end date 45 18.0%

Cadence, of the 389 crawling papers that state one, folded into families (the fold and its residue are in longitudinal). A paper can state more than one cadence, so the column does not sum to 389:

Cadence family Papers Share of 389
Daily 102 26.2%
Duration, not a cadence 58 14.9%
Hourly 46 11.8%
Monthly 44 11.3%
Weekly 43 11.1%
Irregular or unspecified 31 8.0%
n rounds, no interval given 29 7.5%
Every few minutes 27 6.9%
Continuous 27 6.9%
Every few days 20 5.1%
Unmapped by the fold 16 4.1%
Sub-minute 9 2.3%
Yearly 9 2.3%
Quarterly 4 1.0%
Fortnightly 4 1.0%
Sentinel (not-stated written into the cadence field) 1 0.3%

The second row is a finding, not a bug in the fold. 58 papers answer “how often?” with “for three weeks” — a duration, which does not tell a reader whether that was one crawl or twenty. Counting papers rather than values, so that a paper stating both a duration and a real cadence is not penalised, 88 of 389 (22.6%) fill the cadence slot without ever stating a cadence: a duration, a bare round count, “periodically”, or the sentinel itself.

The four pins, measured

Population: the 250 repeated crawling papers against the 870 that crawled once or did not say. Each row asks only whether the paper stated the thing at all — the extraction cannot tell whether the value was the same in every wave, so every figure here is an upper bound on comparability.

Pin Repeated (of 250) Other crawls (of 870)
Population list version 167 (66.8%) 493 (56.7%)
Browser named 143 (57.2%) 386 (44.4%)
… with a version number 17 (6.8%) 49 (5.6%)
Vantage location 89 (35.6%) 212 (24.4%)
Vantage infrastructure 121 (48.4%) 324 (37.2%)
Statefulness 70 (28.0%) 149 (17.1%)
Consent action 105 (42.0%) 244 (28.0%)
Interaction depth 208 (83.2%) 633 (72.8%)
Headless or headful 42 (16.8%) 98 (11.3%)
Authentication 179 (71.6%) 600 (69.0%)
Start and end date 210 (84.0%) 566 (65.1%)
Cadence 158 (63.2%) 231 (26.6%)

Repeating a crawl does make a paper report better — every row is higher — but the absolute levels are the story:

  • 23 of 250 (9.2%) state all four of list version, browser, vantage location and statefulness. Among single-shot crawls it is 32 of 870 (3.7%).
  • 6 of 250 (2.4%) do it with a browser version rather than a browser name.
  • 28 of 250 (11.2%) state none of the four.

The longer the panel, the less it pins

Splitting the 250 by their longest stated span (45 have no parsable end date and drop out):

Group Papers List version Browser Vantage Statefulness All four
Span ≥ 12 months 68 53 36 22 14 3 (4.4%)
Span < 12 months 137 94 84 55 43 18 (13.1%)

The papers whose comparability depends most on pinning report it least often. Part of that is structural rather than careless: the multi-year group has different data provenance. 20 of 68 (29.4%) touch a web archive against 3 of 137 (2.2%), and 27 (39.7%) reuse an existing dataset against 35 (25.5%) — and for a retrospective arm most of the pins were never yours to set. Restricting to papers whose every stated period is a live crawl or an active probe, so the pins were at least available, the gap does not close: 1 of 16 multi-year live crawls states all four against 13 of 94 sub-year ones. Those denominators are small; read the multi-year one as a count, not a rate.

Stated repetition is not becoming more common

Share of crawling papers, per four-year bucket, that crawled more than once:

Bucket Crawling papers Repeated Share
2010–2011 47 4 8.5%
2012–2015 131 27 20.6%
2016–2019 249 65 26.1%
2020–2023 385 83 21.6%
2024–2026 308 71 23.1%

The last bucket is provisional: CCS and IMC 2026 have not been held and two other 2026 venue-years are only partly indexed. The practice arrived between 2012 and 2015, peaked in 2016–2019, and has been flat at roughly one crawling paper in five ever since. It is not growing, and — unlike artifact release, which went from 24.3% to 72.7% over the same period11) — nothing has pushed it.

One caveat that cuts both ways: this measures stated repetition. Because 63.5% of crawling papers never state a snapshot count, a change in reporting would show up here as a change in practice. The reporting rate would have to have moved a great deal to manufacture a flat line out of a real trend, but nothing in this data rules it out.

What the papers say about their own comparability

Full-text probes over all 250 repeated-crawl papers, whitespace collapsed. Every row is a mention count and therefore an upper bound; the two rows that carry weight were hand-read in full.

Probe Papers Share of 250
Calls itself longitudinal 123 49.2%
reproducib* / replicab* / replicat* 78 31.2%
comparable or comparability 80 32.0%
Wave / round / iteration terminology 54 21.6%
Re-crawl, repeat crawl, second crawl 42 16.8%
Names a filter list or blocklist 84 33.6%
… pinned to a version or date 5 2.0%
Same version / list / snapshot / configuration 33 13.2%
Pinned, froze or fixed a version (wide) 13 5.2%
Docker, container or VM image 18 7.2%
churn 19 7.6%
attrition or survivorship 2 0.8%

Two hand audits:

  • A deliberately wide probe for any <browser> <number> string anywhere in the text hits 58 of 250 (23.2%). Reading all 58, 26 (10.4%) state the version of a browser their own measurement drove. The other 32 are bibliography entries, ecosystem history (“Chrome 43 shipped …”), the version of a browser being measured rather than driven, and outright false positives — footnote markers, table cells, and one “Chrome browser 100 times”. So the honest range is 17 (what the structured field records) to 26 (what the text contains), against a population of 250.
  • A wide probe for browser auto-update or version drift hits 14. Reading all 14, not one discusses pinning or reporting the crawler's own browser version across waves. The closest is a limitations note that measurement tools “require constant maintenance to keep up with browser updates and interface changes”. The field does not appear to have noticed that its instrument updates itself.

Related: the exact phrase Chrome for Testing appears in 0 of the 5,855 readable papers in the corpus. A sweep widened to the tooling's domain and to the bare acronym returns 10 hits and reading all 10 shows every one is a false positive — CFT as Combating the Financing of Terrorism, Control Flow Trimming, Call Flow Tree, Crash-Fault Tolerant, a certificate subject O=CFT, and a fitness app. The tooling has been available since 2023; three years and seven venues later it has no footprint here. (Both regexes are printed with their counts in the report output, because a different regex would give a different number.)

Almost nobody pins the list either

Of the 266 papers naming Tranco as a population source:

Outcome Papers Share
States any version or date 159 59.8%
Cites a permanent Tranco list ID 23 8.6%

By year the ID rate is 1/7 (2020), 2/28, 2/29, 5/48, 4/56, 8/76, 1/22 (2026) — noisy, small, and with no trend. Permanent IDs have been available since Tranco's 2019 paper [10Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)]. The 23 were found by two wide probes and confirmed by reading all 37 candidates by hand; the list is in longitudinal.

Papers that pin all four

The 23 repeated-crawl papers that state list version, browser, vantage location and statefulness. These are the methods sections to copy from. Eight of the 23 are PETS papers from 2024 onward, which is where most of the recent repeated web-privacy measurement is appearing.

Year Venue Paper Snapshots Distinct stated vantage locations
2026 PETS Google click identifier in YouTube [11Dao, Ha; Shinde, Abhishek; Athar, Sana; Gosain, Devashish (2026): "Clicking into Exposure: Uncovering Privacy Risks of Google Click Identifier in YouTube Ads", Proceedings on Privacy Enhancing Technologies 2026(2):92-107. (DOI)] 5 6
2026 PETS Manifest V3 and ad-blocker effectiveness [12Lukić, Karlo; Papadopoulos, Lazaros (2026): "Privacy vs. Profit: The Impact of Google's Manifest Version 3 (MV3) Update on Ad Blocker Effectiveness", in: Proceedings on Privacy Enhancing Technologies. (Link)] 5 1
2026 PETS IP-based website fingerprinting in IPv6 [13Ahmad, Sumeer; Polychronakis, Michalis; Benson, Theophilus A.; Hoang, Nguyen Phong (2026): "More Space, Less Privacy? Measuring the Effectiveness of IP-based Website Fingerprinting in IPv6", in: Proceedings on Privacy Enhancing Technologies. (DOI)] 5 1
2025 PETS HTTP response headers for cross-browser tracker detection [14Rieder, Wolf; Raschke, Philip; Cory, Thomas (2025): "Beyond the Request: Harnessing HTTP Response Headers for Cross-Browser Web Tracker Detection in an Imbalanced Setting", in: Proceedings on Privacy Enhancing Technologies, pp. 100-117. (DOI)] 18 1
2025 PETS Intractable cookie crumbs [15Rasaii, Ali; Dao, Ha; Feldmann, Anja; Javid, Mohammadmahdi; Gasser, Oliver; Gosain, Devashish (2025): "Intractable Cookie Crumbs: Unveiling the Nexus of Stateful Banner Interaction and Tracking Cookies", in: Proceedings on Privacy Enhancing Technologies, pp. 429-445. (DOI)] 4 1
2025 PETS More and scammier ads [16Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] 6 5
2025 USENIX Navigating cookie consent violations across the globe [17Tang, Brian; Bui, Duc; Shin, Kang G. (2025): "Navigating Cookie Consent Violations Across the Globe", in: Proceedings of the USENIX Security Symposium. (Link)] 10 8
2025 WWW Before & After: the EU Code of Practice on Disinformation [18Papadogiannakis, Emmanouil; Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2025): "Before & After: The Effect of EU's 2022 Code of Practice on Disinformation", in: Proceedings of the ACM Web Conference. (DOI)] 2 3
2024 IEEE S&P Targeted and troublesome [19Moti, Zahra; Senol, Asuman; Bostani, Hamid; Zuiderveen Borgesius, Frederik J.; Moonsamy, Veelasha; Mathur, Arunesh; Acar, Gunes (2024): "Targeted and Troublesome: Tracking and Advertising on Children's Websites", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] 7 5
2024 PETS The devil is in the details: server-side tracking [20Fouad, Imane; Santos, Cristiana; Laperdrix, Pierre (2024): "The Devil is in the Details: Detection, Measurement and Lawfulness of Server-Side Tracking on the Web", Proceedings on Privacy Enhancing Technologies 2024(4):450-465. (DOI)] 3 2
2024 PETS Cookie banner interaction tools [21Demir, Nurullah; Urban, Tobias; Pohlmann, Norbert; Wressnegger, Christian (2024): "A Large-Scale Study of Cookie Banner Interaction Tools and their Impact on Users' Privacy", in: Proceedings on Privacy Enhancing Technologies, pp. 5-20. (DOI)] 3 5
2023 USENIX Multi-factor and risk-based authentication availability [22Gavazzi, Anthony; Williams, Ryan; Kirda, Engin; Lu, Long; King, Andre; Davis, Andy; Leek, Tim (2023): "A Study of Multi-Factor and Risk-Based Authentication Availability", in: Proceedings of the USENIX Security Symposium. (Link)] 2 5
2022 IMC Respect the ORIGIN! [23Singanamalla, Sudheesh; Paracha, Muhammad Talha; Ahmad, Suleman; Hoyland, Jonathan; Valenta, Luke; Safronov, Yevgen; Wu, Peter; Galloni, Andrew; Heimerl, Kurtis; Sullivan, Nick; Wood, Christopher A.; Fayed, Marwan (2022): "Respect the ORIGIN!: a best-case evaluation of connection coalescing in the wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] 2 1
2022 USENIX Geodifferences in mobile apps [24Kumar, Renuka; Virkud, Apurva; Sundara Raman, Ram; Prakash, Atul; Ensafi, Roya (2022): "A Large-scale Investigation into Geodifferences in Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)] 31 26
2022 WWW Reproducibility and replicability [1Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)] 12 3
2020 NDSS Encrypted DNS and traffic analysis [25Siby, Sandra; Juarez, Marc; Diaz, Claudia; Vallina-Rodriguez, Narseo; Troncoso, Carmela (2020): "Encrypted DNS -> Privacy? A Traffic Analysis Perspective", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 5 3
2020 PETS Missed by filter lists [26Fouad, Imane; Bielova, Nataliia; Legout, Arnaud; Sarafijanovic-Djukic, Natasa (2020): "Missed by Filter Lists: Detecting Unknown Third-Party Trackers with Invisible Pixels", in: Proceedings on Privacy Enhancing Technologies, pp. 499-518. (DOI)] 6 1
2020 WWW Beyond the front page [27Urban, Tobias; Degeling, Martin; Holz, Thorsten; Pohlmann, Norbert (2020): "Beyond the Front Page:Measuring Third Party Dynamics in the Field", in: Proceedings of The Web Conference 2020, pp. 1275–1286. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] 3 3
2020 WWW Representativeness of automated web crawls [28Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] 2 2
2019 WWW Political personalization of Google News search [29Le, Huyen T.; Maragh, Raven; Ekdale, Brian; High, Andrew; Havens, Timothy; Shafiq, Zubair (2019): "Measuring Political Personalization of Google News Search", in: Proceedings of the ACM Web Conference. (DOI)] 7 2
2017 NDSS Thou shalt not depend on me [30Lauinger, Tobias; Chaabane, Abdelberi; Arshad, Sajjad; Robertson, William; Wilson, Christo; Kirda, Engin (2017): "Thou Shalt Not Depend on Me: Analysing the Use of Outdated JavaScript Libraries on the Web", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] 2 1
2016 IEEE S&P Cloak of visibility [31Invernizzi, Luca; Thomas, Kurt; Kapravelos, Alexandros; Comanescu, Oxana; Picod, Jean-Michel; Bursztein, Elie (2016): "Cloak of Visibility: Detecting When Machines Browse a Different Web", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] 3 1
2016 USENIX Tracing information flows between ad exchanges [32Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)] 9 1

Snapshot counts are the paper's own stated figure and are not directly comparable — one paper's “snapshot” is a daily crawl, another's is a yearly wave. Vantage locations are distinct stated places, sentinels excluded.

Methodology and limitations of these figures

  • Population. crawled is the 1,120 papers with a crawl configuration or an automated-web-crawl study type — the same definition the corpus overview uses. repeated is those with a stated snapshot count above one.
  • “Stated” is not “held fixed”. Every pin figure counts papers that mentioned the thing. A paper can name a browser in wave one and silently upgrade it in wave two, and the extraction cannot see that. Treat the pin figures as ceilings.
  • Snapshot counts are noisy at the top. Nine papers report more than 1,000 snapshots and hand-checking finds mis-extracted page and record counts. Bands are used; raw maxima are not.
  • Spans are the longest stated period on any measurement, which for a paper that also reuses an old dataset is not the span of its own repeated crawl. This is why the multi-year group is contaminated with retrospective work, and it is reported alongside the mode breakdown rather than hidden.
  • Cadence is free text and was folded by a script published in the provenance page, with its 16 unmapped strings printed in full.
  • Probes are mentions. A phrase hit is not a reported practice, and a paper can do a thing without writing it down. Where a claim depends on the number, every hit was read and the audited count is published beside the raw one; where it does not, the number is labelled as an upper bound. Every probe regex is printed with its count in the report script's output, because a different regex gives a different number.
  • Field stability. crawlConfig firing, statefulness and interaction depth are reproducible on a repeat extraction to within a few points; temporal.mode is around 69%; free-text fields such as cadence are ~20% stable by exact string and appear here only folded. Those stability figures were measured on the older, smaller corpus and have not been re-measured — see Corpus.
  • The full query log, the report script and its unedited output, the folds and their residue, the quotes spot-checked and the sources rejected are at longitudinal.

Which Methods Are Current

Dating matters here because the corpus rewards whatever was fashionable mid-window. Read this table as of 2026-08.

Practice Status Why
Cite “the Alexa top n” for a wave after 2022 Impossible. Amazon retired the service; there is no list to pin. See Sampling.
Cite a permanent Tranco list ID per wave Current, and under-adopted. Mechanism live since 2019 and verified working in 2026; 8.6% of Tranco papers use it.
Name a browser without a version Not sufficient, and it is the majority practice. 57.2% of repeated crawls name a browser; 6.8% give a version.
Pin Chrome via Chrome for Testing Current best practice, since May 2023. Immutable builds back to 113; zero mentions in seven venues so far.
Pin Firefox via an OpenWPM release tag Current. OpenWPM has pinned unbranded Firefox builds throughout; cite the release, not the project.
Pin the browser by pinning Playwright Current. browsers.json fixes an exact Chromium, Firefox and WebKit build per Playwright release.
Container digest for the whole environment Available, not reported. 0 of 5,855 papers mention an image digest; 5.0% mention Docker.
.web measurement bundles (WebREC) New in 2025, unproven. 2 mentions in the corpus and no use; single-digit repository stars. Watch it, do not depend on it.
One fixed classifier version across all waves, validated against historical versions Current best practice, 2026. [4Singh, Sachin Kumar; Mahmud, Faisal; Ricci, Robert; Siby, Sandra (2026): "The Empire Strikes Back (at Your Privacy): An Archaeology of Tracking on Government Websites", Proceedings on Privacy Enhancing Technologies 2026(2):108-126. (DOI)]; before it, classifier anachronism was mostly unaddressed.
Balanced panel with website-clustered standard errors or difference-in-differences Current, and rare. [5Nenadić, Luka; Rodriguez, David; Calandrino, Joseph A. (2026): "Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact", Proceedings on Privacy Enhancing Technologies 2026(4):703-723. (DOI)] is the model; Regression finds one paper in 5,859 fitting the website as a random effect.
Web-archive retrospective Current but flat. Roughly 2% of empirical work names a web archive and that share has not grown (Archives); the narrower temporal.mode = web-archive enum fires on 1.3%. Neither superseded nor spreading.
Reporting a repeat-run noise floor before a trend Not established. The measurement exists [1Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)]. We did not find a repeated-crawl paper in this corpus reporting its own, but we have no probe that could settle it — read this row as an impression from reading the exemplars, not as a count.

The honest summary: the techniques for making waves comparable have improved a lot since 2023 — immutable browser builds, permanent list ids, fixed-and-validated classifiers, panel-aware statistics — and adoption of every one of them is in the single digits or lower. This is not a field where the state of the art is contested. It is one where the state of the art is available and unused.

What to Report

For a repeated measurement to be comparable to its own earlier waves, and to anyone else's, a methods section needs:

  1. The number of waves, their dates, and the interval between them. A duration is not a cadence: “we crawled for three weeks” does not say whether that is one crawl or twenty.
  2. The frame, per wave. Source list, exact version or permanent ID, and size. If you fixed a cohort at wave one, say so and give the wave-one size; if you redrew each wave, publish the overlap between waves.
  3. The browser and driver versions, per wave, and whether they changed. If they changed, say why and what you did about it.
  4. The vantage point per wave, and whether multi-location waves were simultaneous.
  5. The crawl protocol: statefulness, consent action, interaction depth, subpages, timeouts, retry policy.
  6. The classifier and its version — filter list, blocklist, category service, model checkpoint — held fixed across waves, with the version stated. If you must change it mid-study, re-label every earlier wave with the new version and report both.
  7. Attrition, numerically. How many wave-one targets survived to each later wave, and what you know about the ones that did not. Complete-case analysis over survivors is the default and it is biased (Biases).
  8. The repeat-run noise floor, if you can afford it: two waves close enough together that no real change is plausible, so a reader can tell a 5% movement from a 5% wobble.
  9. The unit of analysis in the statistics. Rows from the same site are not independent. Cluster on the site, or model it, and say which (Regression).

Open Questions

  • We found no measurement of the repeat-run noise floor for any outcome but tracking requests and tracking domains. Demir et al. give twelve days for those two [1Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)]. The same experiment for consent-banner presence, CMP identity, filter-list hit rate, CSP deployment and cookie counts would be cheap, short, and cited by every longitudinal paper thereafter.
  • What does list churn alone contribute to a published trend? Re-running a well-known longitudinal result on a fixed cohort and on a redrawn list, and reporting the two side by side, would put a number on how much of the field's trend literature is frame movement. The Tranco archive makes it possible for any result since 2019.
  • How much attrition is there, and is it selective? Biases establishes that the field does not report it, and we found no study that characterises which sites drop out of a repeated crawl or whether their tracking and security behaviour differs from the survivors.
  • Does pinning actually change the answer? The whole case for the four pins rests on single-wave variation studies. A study that repeats one measurement under pinned and unpinned conditions across a year would either justify the effort or show it is smaller than assumed.
  • A methodology-change log for public longitudinal datasets. HTTP Archive's stops in 2017. There is no maintained record of when the field's shared instruments changed, and every trend built on them inherits that gap.
  • Sampling — how to draw, and what to record so the draw is repeatable. The per-wave list version lives there.
  • Website selection — which ranking list, its provenance and its biases.
  • Archives — the retrospective arm: what an archive preserves, the escape and anachronism traps, and HTTP Archive in detail.
  • Crawling location — the vantage-point pin, and how to verify you have the one you think.
  • Stateful stateless — the profile pin, and what repeat visits mean when state carries between them.
  • Crawler — which tool, and the reproducibility caveats per library.
  • Biases — survivorship and attrition in repeated crawls, measured.
  • Regression — what to do with repeated observations of the same site, including when ignoring them makes your p-value too large.
  • Artifacts — releasing the pinned configuration alongside the data.

References

[1]
Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)
[2]
Demir, Nurullah; Hörnemann, Jan; Große-Kampmann, Matteo; Urban, Tobias; Pohlmann, Norbert; Holz, Thorsten; Wressnegger, Christian (2023): "On the Similarity of Web Measurements Under Different Experimental Setups", in: Proceedings of the ACM Internet Measurement Conference, pp. 356-369. ACM DOI 10.1145/3618257.3624795 is listed by DBLP but was not registered with the DOI resolver as of 2026-08-12; the DOI above resolves to the authors' institutional record of the same paper (DOI)
[3]
Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)
[4]
Singh, Sachin Kumar; Mahmud, Faisal; Ricci, Robert; Siby, Sandra (2026): "The Empire Strikes Back (at Your Privacy): An Archaeology of Tracking on Government Websites", Proceedings on Privacy Enhancing Technologies 2026(2):108-126. (DOI)
[5]
Nenadić, Luka; Rodriguez, David; Calandrino, Joseph A. (2026): "Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact", Proceedings on Privacy Enhancing Technologies 2026(4):703-723. (DOI)
[6]
Ahmad, Syed Suleman; Dar, Muhammad Daniyal; Zaffar, Muhammad Fareed; Vallina-Rodriguez, Narseo; Nithyanand, Rishab (2020): "Apophanies or Epiphanies? How Crawlers Impact Our Understanding of the Web", in: Proceedings of The Web Conference, pp. 271-280. (DOI)
[7]
Hantke, Florian; Snyder, Peter; Haddadi, Hamed; Stock, Ben (2025): "Web Execution Bundles: Reproducible, Accurate, and Archivable Web Measurements", in: Proceedings of the USENIX Security Symposium. (Link)
[8]
Raman, Ram Sundara; Shenoy, Prerana; Kohls, Katharina; Ensafi, Roya (2020): "Censored Planet: An Internet-wide, Longitudinal Censorship Observatory", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[9]
Niaki, Arian Akhavan; Cho, Shinyoung; Weinberg, Zachary; Hoang, Nguyen Phong; Razaghpanah, Abbas; Christin, Nicolas; Gill, Phillipa (2020): "ICLab: A Global, Longitudinal Internet Censorship Measurement Platform", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[10]
Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)
[11]
Dao, Ha; Shinde, Abhishek; Athar, Sana; Gosain, Devashish (2026): "Clicking into Exposure: Uncovering Privacy Risks of Google Click Identifier in YouTube Ads", Proceedings on Privacy Enhancing Technologies 2026(2):92-107. (DOI)
[12]
Lukić, Karlo; Papadopoulos, Lazaros (2026): "Privacy vs. Profit: The Impact of Google's Manifest Version 3 (MV3) Update on Ad Blocker Effectiveness", in: Proceedings on Privacy Enhancing Technologies. (Link)
[13]
Ahmad, Sumeer; Polychronakis, Michalis; Benson, Theophilus A.; Hoang, Nguyen Phong (2026): "More Space, Less Privacy? Measuring the Effectiveness of IP-based Website Fingerprinting in IPv6", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[14]
Rieder, Wolf; Raschke, Philip; Cory, Thomas (2025): "Beyond the Request: Harnessing HTTP Response Headers for Cross-Browser Web Tracker Detection in an Imbalanced Setting", in: Proceedings on Privacy Enhancing Technologies, pp. 100-117. (DOI)
[15]
Rasaii, Ali; Dao, Ha; Feldmann, Anja; Javid, Mohammadmahdi; Gasser, Oliver; Gosain, Devashish (2025): "Intractable Cookie Crumbs: Unveiling the Nexus of Stateful Banner Interaction and Tracking Cookies", in: Proceedings on Privacy Enhancing Technologies, pp. 429-445. (DOI)
[16]
Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[17]
Tang, Brian; Bui, Duc; Shin, Kang G. (2025): "Navigating Cookie Consent Violations Across the Globe", in: Proceedings of the USENIX Security Symposium. (Link)
[18]
Papadogiannakis, Emmanouil; Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2025): "Before & After: The Effect of EU's 2022 Code of Practice on Disinformation", in: Proceedings of the ACM Web Conference. (DOI)
[19]
Moti, Zahra; Senol, Asuman; Bostani, Hamid; Zuiderveen Borgesius, Frederik J.; Moonsamy, Veelasha; Mathur, Arunesh; Acar, Gunes (2024): "Targeted and Troublesome: Tracking and Advertising on Children's Websites", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[20]
Fouad, Imane; Santos, Cristiana; Laperdrix, Pierre (2024): "The Devil is in the Details: Detection, Measurement and Lawfulness of Server-Side Tracking on the Web", Proceedings on Privacy Enhancing Technologies 2024(4):450-465. (DOI)
[21]
Demir, Nurullah; Urban, Tobias; Pohlmann, Norbert; Wressnegger, Christian (2024): "A Large-Scale Study of Cookie Banner Interaction Tools and their Impact on Users' Privacy", in: Proceedings on Privacy Enhancing Technologies, pp. 5-20. (DOI)
[22]
Gavazzi, Anthony; Williams, Ryan; Kirda, Engin; Lu, Long; King, Andre; Davis, Andy; Leek, Tim (2023): "A Study of Multi-Factor and Risk-Based Authentication Availability", in: Proceedings of the USENIX Security Symposium. (Link)
[23]
Singanamalla, Sudheesh; Paracha, Muhammad Talha; Ahmad, Suleman; Hoyland, Jonathan; Valenta, Luke; Safronov, Yevgen; Wu, Peter; Galloni, Andrew; Heimerl, Kurtis; Sullivan, Nick; Wood, Christopher A.; Fayed, Marwan (2022): "Respect the ORIGIN!: a best-case evaluation of connection coalescing in the wild", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[24]
Kumar, Renuka; Virkud, Apurva; Sundara Raman, Ram; Prakash, Atul; Ensafi, Roya (2022): "A Large-scale Investigation into Geodifferences in Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)
[25]
Siby, Sandra; Juarez, Marc; Diaz, Claudia; Vallina-Rodriguez, Narseo; Troncoso, Carmela (2020): "Encrypted DNS -> Privacy? A Traffic Analysis Perspective", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[26]
Fouad, Imane; Bielova, Nataliia; Legout, Arnaud; Sarafijanovic-Djukic, Natasa (2020): "Missed by Filter Lists: Detecting Unknown Third-Party Trackers with Invisible Pixels", in: Proceedings on Privacy Enhancing Technologies, pp. 499-518. (DOI)
[27]
Urban, Tobias; Degeling, Martin; Holz, Thorsten; Pohlmann, Norbert (2020): "Beyond the Front Page:Measuring Third Party Dynamics in the Field", in: Proceedings of The Web Conference 2020, pp. 1275–1286. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[28]
Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[29]
Le, Huyen T.; Maragh, Raven; Ekdale, Brian; High, Andrew; Havens, Timothy; Shafiq, Zubair (2019): "Measuring Political Personalization of Google News Search", in: Proceedings of the ACM Web Conference. (DOI)
[30]
Lauinger, Tobias; Chaabane, Abdelberi; Arshad, Sajjad; Robertson, William; Wilson, Christo; Kirda, Engin (2017): "Thou Shalt Not Depend on Me: Analysing the Use of Outdated JavaScript Libraries on the Web", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[31]
Invernizzi, Luca; Thomas, Kurt; Kapravelos, Alexandros; Comanescu, Oxana; Picod, Jean-Michel; Bursztein, Elie (2016): "Cloak of Visibility: Detecting When Machines Browse a Different Web", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[32]
Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)
1)
https://api.github.com/repos/easylist/easylist/commits for easylist/easylist_general_block.txt, checked 2026-08-27.
2)
Checked 2026-08-27 against https://tranco-list.eu/api/lists/id/GVWK, which returns "available": true and a download URL. The same endpoint documents the configuration of any list id.
3)
All five rows read from the Tranco API on 2026-08-27: https://tranco-list.eu/api/lists/date/YYYYMMDD returns the list id and its configuration.providers for that day. Alexa's disappearance is forced — Amazon retired the Alexa Top Sites API on 15 December 2022, see Sampling.
4)
https://github.com/publicsuffix/list commit history for public_suffix_list.dat, checked 2026-08-27: ten commits in the eleven days to 19 August 2026, on three distinct days.
5)
Read from https://googlechromelabs.github.io/chrome-for-testing/known-good-versions-with-downloads.json on 2026-08-27; the last-known-good-versions.json endpoint on the same day reported Stable 152.0.7977.64.
6)
https://github.com/openwpm/OpenWPM/blob/master/scripts/install-firefox.sh pins TAG='9ce1ee6baeb9a3c326dbd180bdece65d8fc2eadc' # FIREFOX_154_0_RELEASE; release notes for v0.36.0 read “Firefox 152 → 154.0, with the Extension manifest's strict_min_version pinned to match”. Both checked 2026-08-27.
7)
https://raw.githubusercontent.com/microsoft/playwright/main/packages/playwright-core/browsers.json, checked 2026-08-27: {"name": "chromium", "revision": "1241", "browserVersion": "152.0.7977.54", "title": "Chrome for Testing"}.
8)
Header captured from https://easylist.to/easylist/easylist.txt on 2026-08-27; repository creation date from the GitHub API on the same day.
9)
https://httparchive.org/faq, checked 2026-08-27.
10)
FAQ link target https://github.com/HTTPArchive/httparchive/blob/main/docs/changelog.json returns 404; the file resolves on master, 12 entries, last dated 2017-06-01. Both checked 2026-08-27.
11)
Figures from data/extract/OVERVIEW.md; see Corpus.
You could leave a comment if you were logged in.
design/longitudinal.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki