User Tools

Site Tools


statistics:biases

This is an old revision of the document!


Biases

Every web measurement estimates a quantity about a population it could not observe. Between the two sits a chain of substitutions: the web becomes a ranking list, the list becomes the rows you could reach, the rows become the pages that loaded, and the pages become the ones that loaded from your machine, in your month, in your country. Each substitution has a measurable size, and for each one this literature contains a paper that measured it. Those measurements are the reason this page exists: the sizes are not small — they run from single-digit percentages up to a crawl missing 45% of the sites a real user's browser reveals.

This page is not about bias in general — a textbook covers that, and Hypothesis testing covers what non-independence does to a p-value. It is about the four substitutions above, in the form they take in a web measurement, with the effect sizes the field has actually measured for each: selection bias from a top list, survivorship in a repeated crawl, vantage-point bias, and denominator bias when the population you built is not the population you write about. It closes with a worked example of stating limitations honestly, using the biases of this site's own publication corpus.

All four biases have been measured by somebody in this literature. What is missing is not the knowledge — it is the reporting.

Of the 5,859 papers extracted from seven security, privacy and measurement venues (2010–2026):

  • Top lists: of the 857 papers that crawled the web, 479 (55.9%) draw from a popularity ranking, and 423 (49.4%) take a top-n slice without stratifying. Scheitle et al. [1Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] measured what that slice is: Alexa, Umbrella and Majestic agreed on only “99k out of 1M domains”, and against the general com/net/org population the lists over-represent IPv6 (“11-13%” against “4%”) and CDN use (“at least a factor of 2”, and a factor of 20 at top-1k). A top-n sample is a census of an unusual neighbourhood, not a sample of the web.
  • Survivorship: 250 of the 1,120 crawling papers (22.3%) crawl more than once. A full-text probe over all 250 finds 2 (0.8%) that use the words attrition or survivorship anywhere; widened to all 5,855 readable papers the rate is the same, 0.7%. The field runs longitudinal panels and does not name the bias they carry.
  • Vantage point: of the 3,908 papers that measured from somewhere, 2,773 (71.0%) give no location a fold can map, and only 603 (15.4%) name more than one country. Jueckstock et al. [2Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)] put a number on the cost: “Around 5% of content-providing domains show significant measurement bias across VP” — and a larger number on the axis a vantage point does not fix, the browser configuration, where among filter-list-flagged ad and tracker domains “nearly 20% of domains' traffic” is “strongly correlated to choice of BC.”
  • Denominator: 105 of 857 web-crawling papers (12.3%) sample both a site-level and a page-level unit, so the two are available to be silently swapped in a results sentence. Urban et al. [3Urban, Tobias; Degeling, Martin; Holz, Thorsten; Pohlmann, Norbert (2020): "Beyond the Front Page:Measuring Third Party Dynamics in the Field", in: Proceedings of The Web Conference 2020, pp. 1275–1286. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] show the size of the swap: “On average, 55 cookies were set when loading a landing page while 78 were set when a subsite was accessed.”
  • Stating limitations is becoming normal, slowly. A heading probe for Limitations or Threats to Validity over web-crawling papers rises from 3.8% (2010–2013) to 30.1% (2025–2026). Seven papers in ten still do not have the section.

Denominators differ per bullet because each question has its own population; none of these is a share of 5,859. Methodology, folds and the full query log are on biases.

What to Read First

Five entries, in the order that will save you the most time. Every one measures a bias rather than warning about one.

  • Scheitle et al. [1Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] (IMC 2018) — read this one first. The measurement of what a top list is: how much three lists disagree, how much they churn daily, how they differ from the domain population on IPv6, CDN and TLS adoption, and how cheaply Umbrella's rank can be manipulated (“10k probes at 1 query per day … achieve a rank of 38k”). Every “we crawled the top 1M” sentence in the field is scoped by this paper.
  • Ruth, Kumar, Wang, Valenta & Durumeric [4Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] (IMC 2022) — the update, and the one that makes the bias directional. Measured against Cloudflare's own resolver traffic: “of the 1,790 domains we measure in the Alexa top 10K, 70% of them are ranked by Cloudflare in a lower rank-magnitude bucket”; “Top lists exhibit noticeable and irregular geographic biases … all top lists poorly represent Japan”; “Top lists better approximate client behavior on desktop platforms than mobile platforms”; and “certain categories of websites (e.g., adult and gambling) are often underrepresented in top lists”. If your result is about mobile users, non-US users, or an under-represented category, this paper tells you which way your estimate is wrong.
  • Zeber et al. [5Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] (TheWebConf 2020) — the crawl-against-humans comparison. “crawler site visits issued requests to a median of 11.6 third-party domains, whereas for visits by humans, the median was 4.5 third parties”, and “the average similarity in third parties is low, with a median of 20%, and in most cases (87% of list domains) this is due to the crawler accessing more third parties.” A crawl is not a sample of browsing, and the direction of the error is known.
  • Jueckstock et al. [2Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)] (TheWebConf 2021) — the vantage point and browser configuration treated as experimental variables and varied on purpose. The standard citation for “where you measure from changes what you measure”, and the source of the refusenik concept: “sites that always failed to load from a particular vantage point (VP) or browser configuration (BC)”. See Crawling location for the full treatment.
  • Demir et al. [6Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)] (TheWebConf 2022) and [7Demir, Nurullah; Hörnemann, Jan; Große-Kampmann, Matteo; Urban, Tobias; Pohlmann, Norbert; Holz, Thorsten; Wressnegger, Christian (2023): "On the Similarity of Web Measurements Under Different Experimental Setups", in: Proceedings of the ACM Internet Measurement Conference, pp. 356-369. ACM DOI 10.1145/3618257.3624795 is listed by DBLP but was not registered with the DOI resolver as of 2026-08-12; the DOI above resolves to the authors' institutional record of the same paper (DOI)] (IMC 2023) — the two papers that measure how much of your result is your setup rather than the web. “when comparing two different profiles, 48% of the underlying data varies”, and “only 32% of the cookies appear in all profiles and 42% only in one profile.” Read these before you attribute a difference between your numbers and a prior paper's to a change on the web.

For the statistical framing rather than the web-specific measurement, Angrist & Pischke [8Angrist, Joshua D.; Pischke, Jörn-Steffen (2009): "Mostly Harmless Econometrics: An Empiricist's Companion". Princeton University Press. (DOI)] is the standard treatment of selection into a sample and what it does to an estimate.

1. Selection Bias: Your Sample Is a Ranking List

Website selection catalogues the lists and Sampling covers how to draw from one and how to record the draw. This section is about the part downstream of both: what a top-n sample licenses you to claim, and by how much it is known to be wrong.

A top list is a census of an unusual neighbourhood

The list is not a random sample of the web and it is not a random sample of anything. It is a ranked enumeration produced by one popularity proxy — browser telemetry, resolver queries, backlinks, panel traffic — and each proxy has its own blind spot. Scheitle et al. [1Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] quantified how unusual the resulting neighbourhood is by comparing top-list domains against all com/net/org:

Property Top lists General com/net/org population
IPv6 enabled 11–13% 4%
CDN use, Top 1M at least 2× the general population
CDN use, Top 1k at least 20× the general population
HTTP/2 adoption up to 26.6% (Alexa Top 1M) 7.84%
NXDOMAIN (does not resolve) 11.5% Umbrella, 2.7% Majestic 0.8%

The last row is the one that bites hardest: more than one in ten Umbrella domains does not resolve at all. If your denominator is “the top 1M” and your numerator is “sites doing X”, you have put non-existent domains in the denominator.

Ruth et al. [4Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] then measured the lists against ground truth — Cloudflare's own resolver traffic — and found the error is not noise but structure: rank magnitudes are wrong in a consistent direction (70% of measurable Alexa top-10K domains sit in a lower Cloudflare bucket), the lists track desktop behaviour better than mobile, geographic coverage is uneven and irregular (Japan worst), and adult and gambling sites are systematically under-represented. CrUX came out as the most accurate list across all their metrics.

Three claims a top-n sample does not license, and the sentence that fixes each:

  • “X% of websites do Y.” → ✓ “X% of the //n most popular sites in list L, version V, do Y.”// A top list has no defined relationship to “websites”; there is no sampling weight that gets you from one to the other.
  • “Y is becoming more common on the web.” → ✓ “Y is more common in list L in year 2 than in year 1.” The list's own composition changed — Scheitle et al. report “daily churn of up to 50% of domains” for some lists — so a trend over list snapshots is a trend in a moving population.
  • “our sample is representative.” → ✓ “our sample is the head of the web as measured by L, which over-represents CDN-fronted, IPv6-enabled, desktop-visited, non-adult sites [1Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)][4Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)].”

The list changed under the field, and Alexa is gone

Alexa's ranking service was retired on 1 May 2022, and its APIs on 15 December 2022.1) There is no successor at the same URL: alexa.com now redirects to Amazon's voice-assistant pages. The Tranco project states the consequence directly on its front page: “The Alexa ranking has been removed from the default Tranco list, as it is no longer available”, and that the Chrome UX Report and Cloudflare Radar rankings “have been integrated into the default Tranco list, starting from the daily updated list of August 1, 2023.”2)

This matters for reading the literature, not only for writing it. Any figure drawn from an Alexa list rests on a frame that no longer exists and cannot be re-drawn, so it cannot be replicated except from an archived copy of the list. In this corpus, of the 793 papers drawing from a ranking list, Alexa's share of each period falls from 95.9% to 10.8% while Tranco's rises to 71.2% — see Which list, and when below. The list changed under the field has the same transition measured on a web-unit population, and the two independent folds land within 3% of each other on Alexa (475 papers here, 463 there, on different populations by different methods), which is the best available check on either.

Rank-stratified sampling is the fix, and it is still rare

If you need a claim about the web rather than about the head of it, the standard correction is to stratify by rank: draw from rank bands and report per-band results, or weight. Of 857 web-crawling papers, 73 (8.5%) use a stratified draw and 423 (49.4%) take a top-n slice without one. Rank-stratified sampling is the standard fix, and it is rare covers how to do it, including a sampler that records its own provenance.

Stratifying does not remove the bias — the frame is still the list — but it makes the bias visible, because a per-band table shows the reader how fast the effect decays with rank. A single top-10k number hides exactly that.

2. Survivorship: The Sites That Stopped Answering

A crawl of n sites returns fewer than n results. What comes back are the survivors, and the gap is a population you know nothing about. Unlike the other three biases on this page, this one does not require a decision to commit: it is what happens if you write no code about it at all, which is why it is worth a section of its own.

Complete-case analysis is the default, and it is not a decision anyone made

The mechanics are ordinary: you crawl 100,000 sites, 91,000 load, and every figure in the paper is computed over the 91,000. That is a complete-case analysis, and it is only unbiased if dropping out is independent of what you measure. In web measurement it is not, and the dependence runs the wrong way — the sites that block datacenter IPs, sit behind bot management, or have been abandoned are also the sites whose tracking, consent and security behaviour differs. Attrition is a denominator, and it is usually invisible makes the same point about the first crawl; the longitudinal case compounds it every wave.

The honest version costs one sentence. Singh et al. [9Singh, Rachee; Nithyanand, Rishab; Afroz, Sadia; Pearce, Paul; Tschantz, Michael Carl; Gill, Phillipa; Paxson, Vern (2017): "Characterizing the Nature and Dynamics of Tor Exit Blocking", in: 26th USENIX Security Symposium (USENIX Security 17), pp. 325-341. USENIX Association. (Link)] write it: “several of the selected exit relays intermittently went offline, with a total of 0, 12, 19, and 28 offline during crawls 1-4, respectively. We account for the resulting page-load failures by excluding the failures from our analysis.” The exclusion is still an exclusion — but a reader can now bound what it could have done, which is the whole point.

Two survivorship effects, and they do not cancel

  • Drop-out. Sites present in wave 1 and gone by wave k. Your wave-k population is the sites that survived, which are disproportionately the well-maintained, commercially active ones. Anything correlated with being a going concern — a CMP, a tag manager, a current TLS configuration — will appear to have grown when what happened is that the sites without it left.
  • Drop-in via the frame. If you re-draw the list each wave instead of fixing a panel, new entrants arrive already adapted to whatever changed. Hils et al. [10Hils, Maximilian; Woods, Daniel W.; Böhme, Rainer (2021): "Privacy Preference Signals: Past, Present and Future", in: Proceedings on Privacy Enhancing Technologies. (DOI)] caught this and designed around it, choosing a deliberately older list “in order to avoid survivorship bias in our observations. Picking a later toplist would over-sample websites created post-2020 who are certain to adopt TCF 2.0 and de facto avoid a migration decision.” This is the clearest statement of the problem in the corpus and it is worth copying wholesale: the list you draw in wave k is post-selected on the very thing you are measuring.

The two arrive from opposite ends of the panel — one removes old sites, the other admits new ones — and both push an adoption trend the same way, upward. Fixing the panel removes the second and makes the first measurable; re-drawing removes the first and hides the second. So: fix the panel at wave 1 and report attrition, or re-draw and report both series — never re-draw silently.

The field runs the panels and does not name the bias

  • 250 of 1,120 crawling papers (22.3%) crawl more than once. Of those, 62 have a two-snapshot span, 77 a 3–5-snapshot span, and 37 a span of more than 52 — bands overlap, because a paper can record several spans.
  • Only 409 of the 1,120 (36.5%) state a snapshot count at all, so the panel population above is a floor.
  • A full-text probe over all 250 for attrition, survivorship or survivor bias returns 2 papers (0.8%). Widened to every paper with more than one snapshot regardless of study type (817 readable), 6 (0.7%). Over all 5,855 readable papers, 39 (0.7%).
  • The looser vocabulary is more common but still a minority: 61 of the 250 (24.4%) contain any of a seventeen-phrase attrition vocabulary, led by unreachable (19 papers) and churn (19).

A phrase hit is a mention, not a reported attrition rate, so 24.4% is an upper bound on papers that say anything about the sites they lost. The 0.8% is exact for what it measures — papers using those words — and is itself an upper bound on papers that analyse the bias rather than mentioning it. The gap between “22.3% run panels” and “0.8% name the bias” is this section's reason to exist.

Report three numbers per wave, not one. Sites attempted, sites that returned usable data, and sites present in every wave. Then compute your headline figure twice — over the balanced panel and over all complete cases — and say if they differ. If they do not, you have a robustness check for free; if they do, you have found something.

3. Vantage-Point Bias

Crawling location is the page for choosing and verifying a vantage point, and its figures are computed over the 1,120 crawling papers. This section states the bias in estimator terms and gives the exposure across the wider population — every paper that measured from anywhere, crawl or not.

The measured effect sizes

What varies Measured difference Source
Vantage point (cloud / residential / university) “Around 5% of content-providing domains show significant measurement bias across VP, clearly favoring non-cloud endpoints.” [2Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)]
Browser configuration (naive / stealth), from the most realistic vantage point “over 7% of domains and over 5% of total HTTP traffic volume”; among filter-list-flagged ad and tracker domains, “nearly 20% of domains' traffic strongly correlated to choice of BC” [2Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)]
Either axis: sites that never load at all Refuseniks, “sites that always failed to load from a particular vantage point (VP) or browser configuration (BC)”. By vantage point: 72 cloud, 11 residential, 2 university. By browser configuration: 69 naive, 30 stealth [2Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)]
Profile region Distinct trackers per page: “USA … (6.93; SD: 8.5), followed by Japan (5.6; SD: 6.39), and EU profiles (4.49; SD: 5.42)” [6Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)]
Country of the visitor “websites in 91% of the examined countries (21/23) embed trackers hosted in foreign nations” [11Singh, Sachin Kumar; Ricci, Robert; Gamero-Garrido, Alexander (2025): "Where in the World Are My Trackers? Mapping Web Tracking Flow Across Diverse Geographic Regions", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]
Crawl against real browsing Third-party domains per visit, median 11.6 (crawler) against 4.5 (humans); median third-party similarity 20% [5Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]
Crawl against real interaction “the real user browsing sessions detected 471 such fingerprinting websites, out of which the automated crawl missed 211 (45%)” [12Annamalai, Meenatchi Sundaram Muthu Selva; De Cristofaro, Emiliano; Bilogrevic, Igor (2025): "Beyond the Crawl: Unmasking Browser Fingerprinting in Real User Interactions", in: Proceedings of the ACM Web Conference. (DOI)]

The last two rows are not about geography at all, and they report the largest differences in the table. A crawl differs from real browsing by more than the vantage-point and profile-region effects above it — which is not a ranking of the biases, since these are different designs measuring different quantities, but it is a reason why “we crawled from AWS in Frankfurt” understates the problem: the vantage point is only one axis of a design that also fixes the browser, the interaction, the state and the identity.

The exposure, over every paper that measured from somewhere

Population: the 3,908 papers with at least one vantage tuple. Locations are free text and were folded through scripts/geo.mjs (12.1% of location strings unmapped; the residue is printed in full on the provenance page).

Vantage description Papers Share of 3,908
No location a fold can map — silence 2,773 71.0%
Exactly one country 532 13.6%
More than one country, or a region / multi-country claim 603 15.4%

And 1,764 of 3,908 (45.1%) never state an infrastructure type. Of those that do, research testbed (1,201) and university network (396) lead, while residential appears in 65 papers (1.7%) — the configuration Jueckstock et al. identify as the realism best case is the one almost nobody uses.

Counting locations by exact string is not an option. Folding vantage.locations through scripts/geo.mjs finds 639 papers naming the United States. Counting the exact string United States finds 238 — a 62.8% undercount, because “USA”, “US”, “California”, “US-East” and city names are separate values. Any table of “where the field measures from” built without a fold is wrong by roughly this much, in the direction that flatters the tail.

4. Denominator Bias: The Population You Built Is Not the One You Write About

The other three biases are about which units you observed. This one is about which units you counted, and it is the easiest to commit while being entirely honest, because nothing in the pipeline enforces that the denominator in your results sentence is the population you sampled.

Three ways the denominator slips

(a) The unit changes between sampling and reporting. You sampled sites and report per-page rates, or the reverse. 105 of 857 web-crawling papers (12.3%) sample both a site-or-domain unit and a page unit, so both denominators are in the dataset and the results sentence has to choose. Urban et al. [3Urban, Tobias; Degeling, Martin; Holz, Thorsten; Pohlmann, Norbert (2020): "Beyond the Front Page:Measuring Third Party Dynamics in the Field", in: Proceedings of The Web Conference 2020, pp. 1275–1286. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] measured what the choice is worth on the same sites: “subsites set considerably more (36 %) cookies than the respective landing pages. On average, 55 cookies were set when loading a landing page while 78 were set when a subsite was accessed.” Cite the paper's own 36% rather than deriving one — 78 against 55 is +42%, and the two numbers are not reconciled in the paper, which is itself a small lesson about reading a headline percentage next to the means it is supposed to summarise. They also report a 25% increase in fingerprinting on subsites, and “2.5% of the measured websites do not embed any trackers on the landing page but use trackers on subsites”, which a landing-page-only denominator scores as zero. Aqeel et al. [13Aqeel, Waqar; Chandrasekaran, Balakrishnan; Feldmann, Anja; Maggs, Bruce M. (2020): "On Landing and Internal Web Pages: The Strange Case of Jekyll and Hyde in Web Performance Measurement", in: Proceedings of the ACM Internet Measurement Conference, pp. 680-695. (DOI)] make the same distinction for performance measurement.

(b) Successes become the denominator. Covered in 2. Survivorship: The Sites That Stopped Answering; it is the same defect seen from the reporting end. “74% of sites” computed over pages that loaded is not “74% of sites”.

© A keyword- or query-derived population becomes “the web”. You build the population by searching — a search engine, a platform API, a full-text grep for a phrase — and then report a percentage of it. The percentage is a property of your query. It is the least examined of the four: 398 of 857 web-crawling papers (46.4%) draw at least one purposive population and 112 (13.1%) a convenience one, against 208 exhaustive. Not every purposive draw is keyword-derived — the enum does not separate them — so 46.4% bounds the exposure rather than measuring it.

A search API is a sampler you did not design and cannot inspect

The 2025 corpus contains the paper to cite here. Efstratiou [14Efstratiou, Alexandros (2025): "On YouTube Search API Use in Research", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] queried the YouTube Search API for six topics over twelve weeks and characterised what it returns:

  • API samples videos from empirical distributions, returning results based on the relative density of topical interest and even forcing zero videos to be returned when this relative density is adequately low.”
  • “The results suggest that drop-ins and drop-outs are the normative behavior” — the same item appears and disappears across repeated identical queries.
  • Return frequency is predicted by views, likes, duration, topic and channel: the API “tends to return shorter, more popular videos.”
  • ID-based endpoints, by contrast, are stable.

Read as a bias result, this says three things that generalise to any query-built population. The frame is not enumerable — you cannot know what your query did not return. It is not stable — a re-run is a different sample, so your “longitudinal” series mixes real change with sampler churn. And it is not uniform — the sampler prefers short, popular items, so your denominator is pre-weighted away from the long tail, which is usually where the phenomenon you went looking for lives. The same three properties hold for a search-engine-derived seed list and for a grep-derived corpus.

What to do instead

  • Report the query, not just the population. The exact query string, the date, the API version, the number of results requested and the number returned. Without those the population is not reconstructible even in principle.
  • Report the funnel. Candidates the query returned → candidates that passed your filters → units that produced data. Three numbers, and every percentage in the paper says which one it divides by.
  • Re-run the query. If a second identical run returns a materially different set, that instability is a result and belongs in the paper. Efstratiou's finding is that for at least one major API it will.
  • Never call it a sample of the platform. It is a sample of what the platform returned for your query, on those days, to your credentials.

If what you actually need is a population restricted to a class of site — news sites, e-commerce, sites in one sector — a keyword query is the wrong instrument for it, and there is now a documented alternative. Bozzolan, Calzavara & Cazzaro [15Bozzolan, Simone; Calzavara, Stefano; Cazzaro, Lorenzo (2026): "LLM-Assisted Web Measurements". arXiv:2510.08101, v3, 30 April 2026 (Link)] start from the observation that motivates this whole section: “existing top lists of popular websites are unlabeled and lack semantic information about the nature of the included websites, making targeted web measurements challenging, as researchers often rely on ad-hoc techniques to bias datasets toward specific website classes of interest.” They evaluate language models on the website-classification tasks web measurement actually needs, and propose a two-step pipeline that starts from Tranco and labels down to the class, rather than querying up to it. The frame stays Tranco — with all of 1. Selection Bias: Your Sample Is a Ranking List's biases and none of a query's unenumerability — and the class membership becomes a classification step you can validate. This is a 2026 arXiv preprint outside the seven venues, so treat it as the current pointer rather than as settled practice, and note that it moves the problem from the frame to the classifier: a model's own error profile then needs its own validation, which is an open question below.

Which Corrections Are Current

Dated, because a ranking of what the literature did is not advice about what to do now.

Correction Status Evidence
Fixing and citing a list version or list ID Current, and expected by reviewers. Tranco's list IDs exist for this. [16Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)]; 60.0% of the 857 web-crawling papers state a version, against 43.0% of the 5,712 that drew any population — a different population, not a trend
Rank-stratified sampling with per-band results Current best practice, still a minority. Grew to 8.9% of web sampling in 2022–2024 and plateaued. 8.5% of 857 web-crawling papers
Multi-vantage crawling Current. That page measures the EEA share of location-stating crawls tripling across the period, 19.0% to 59.6%. [2Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)]; Crawling location
CrUX as the accuracy reference for popularity Current. Most accurate list measured against resolver ground truth (2022), and now folded into Tranco's default list (since 2023-08-01).3) [4Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]
Reporting attempted / loaded / balanced-panel denominators Not superseded and not adopted. It has been correct practice throughout and 0.8% of panel papers name the bias, so there is no trend to date — only a gap. this page's probe
Real-user or in-the-field measurement as a validity check on a crawl Current, and the direction of travel — TheWebConf 2020, PETS 2022, TheWebConf 2025. Three papers is a direction, not a rate; the corpus has no field-wide count of in-the-field validation. [5Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)][17Cassel, Darion; Lin, Su-Chin; Buraggina, Alessio; Wang, William; Zhang, Andrew; Bauer, Lujo; Hsiao, Hsu-Chun; Jia, Limin; Libert, Timothy (2022): "OmniCrawl: Comprehensive Measurement of Web Tracking With Real Desktop and Mobile Browsers", in: Proceedings on Privacy Enhancing Technologies. (DOI)][12Annamalai, Meenatchi Sundaram Muthu Selva; De Cristofaro, Emiliano; Bilogrevic, Igor (2025): "Beyond the Crawl: Unmasking Browser Fingerprinting in Real User Interactions", in: Proceedings of the ACM Web Conference. (DOI)]
Alexa as a frame Superseded, and unavailable. Retired 2022-05-01; removed from Tranco's default list. Still named by 10.8% of 2025–2026 rank-list papers. this page's fold; Tranco front page
Treating a single US cloud vantage point as neutral Superseded. It is a configuration with a measured, directional effect. [2Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)][6Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)]
Reporting a single landing-page-only figure as a site-level rate Superseded where subsites are reachable. [3Urban, Tobias; Degeling, Martin; Holz, Thorsten; Pohlmann, Norbert (2020): "Beyond the Front Page:Measuring Third Party Dynamics in the Field", in: Proceedings of The Web Conference 2020, pp. 1275–1286. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)][13Aqeel, Waqar; Chandrasekaran, Balakrishnan; Feldmann, Anja; Maggs, Bruce M. (2020): "On Landing and Internal Web Pages: The Strange Case of Jekyll and Hyde in Web Performance Measurement", in: Proceedings of the ACM Internet Measurement Conference, pp. 680-695. (DOI)]

Two things this table cannot tell you, both because the 2025–2026 slice is the thinnest in the corpus (see Worked Example). No bias characterisation specific to LLM-based classification appears in the corpus's 2025–2026 material, although LLM classification itself is now common there: 166 of the 1,185 papers from 2025–2026 name a language model in a classification tuple (14.0%), and only 4 of those 166 mention bias or calibration anywhere — none of them about the selection or calibration behaviour of LLM labelling on web measurement data. So if your pipeline puts a model in the labelling step, its bias is an open question here, not a settled one. And no paper in this corpus characterises which sites drop out of a crawl and how that biases the result; it is listed as an open question on Open Questions and remains one.

Use in Publications

The figures below come from a structured extraction over 5,859 full-text papers from CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026. Each table names its own population; no figure on this page is a share of 5,859 unless it says so. Sentinel values (not-stated, none-mentioned) are counted as silence, never as answers, and papers are counted, never tuples. 2025 and 2026 are provisional venue-years — CCS and IMC 2026 have not been held and two further 2026 venue-years are incompletely indexed — so they are under-represented by construction, not by relevance.

The populations

Population Definition Papers
sampled drew at least one study population 5,712
crawled ran an automated web crawl 1,120
webCrawled crawled and measured the web platform 857
measuredFrom recorded at least one vantage point 3,908
repeated crawled with more than one snapshot 250

Which frame, and how exposed it is to list bias

Of 5,712 papers that drew a population, 793 (13.9%) name a web popularity ranking — but of the 857 webCrawled papers, 479 (55.9%) do. The corpus-wide figure is the wrong one to quote; the second is the exposure.

List Papers Share of the 793 using any ranking list
Alexa (retired 2022) 475 59.9%
Tranco 267 33.7%
Chrome UX Report (CrUX) 36 4.5%
Majestic 27 3.4%
Common Crawl 27 3.4%
Cisco Umbrella 27 3.4%
SimilarWeb 14 1.8%
Quantcast 8 1.0%
Statvoo / DomCop / other rank vendor 8 1.0%

Shares exceed 100% because a paper may use several lists. Names were matched by regular expression against a folded skeleton of population.sourceList; 10,011 distinct unmatched source skeletons remain and are not ranking lists (custom seed list 516 papers, Google Play 143, Prolific 122 lead them). Sixteen further strings match /alexa/ and are not the ranking list — skill stores, skill marketplaces, a smart-home integration set, and one author's surname — and are excluded by name. The residue head and every exclusion are on the provenance page.

Which list, and when

Denominator is the papers using any ranking list in that period, so this reads as composition rather than uptake.

Period Rank-list papers Alexa Tranco
2010–2013 49 95.9% 0.0%
2014–2017 130 97.7% 0.0%
2018–2021 227 83.7% 15.4%
2022–2024 248 38.7% 53.6%
2025–2026 (provisional) 139 10.8% 71.2%

Alexa's residue is the finding. Four years after retirement, one paper in ten that uses a ranking list still uses the one that no longer exists — and cannot be re-drawn by a replicator.

Top-//n// against stratified, among the 857 web-crawling papers

Draw Papers Share of 857
top-n 447 52.2%
stratified 73 8.5%
both 24 2.8%
top-n without stratification 423 49.4%

List versions: the bias you cannot even bound without one

Population N States a list version Share
sampled 5,712 2,455 43.0%
papers using a ranking list 793 510 64.3%
webCrawled 857 514 60.0%

Web crawlers report versions better than the corpus average, which is the right direction — but two papers in five that crawled the web still cannot have their frame reconstructed.

Repeated crawls, and what they say about the sites they lost

Population: the 250 repeated papers.

Snapshots Papers Share of 250
2 62 24.8%
3–5 77 30.8%
6–12 45 18.0%
13–52 48 19.2%
more than 52 37 14.8%

The rows sum to 269, not 250, and the column does not sum to 100%. A paper can record several temporal spans with different snapshot counts — a monthly crawl plus a one-off re-crawl — so it lands in every band it touches. The right reading is “how many papers ran a panel of this length”, not a partition.

Full-text probe over all 250 (all readable), whitespace-collapsed, case-insensitive:

Phrase Papers Share of 250
unreachable 19 7.6%
churn 19 7.6%
no longer available 12 4.8%
failed to load 9 3.6%
still alive 4 1.6%
could not be reached 3 1.2%
went offline 3 1.2%
survivorship 2 0.8%
attrition 0 0.0%
any of 17 phrases 61 24.4%

The word attrition appears in 0 of 250 repeated-crawl papers, and attrition or survivorship in 2 (0.8%). That rate does not change with the population: 6 of 817 (0.7%) papers with more than one snapshot of any study type, 8 of 1,120 (0.7%) crawling papers, 39 of 5,855 (0.7%) papers overall. A phrase hit is a mention and not a reported rate, so every figure here is an upper bound on papers that address the problem.

Vantage-point exposure

Reproduced from 3. Vantage-Point Bias so this section stands alone; population is the 3,908 measuredFrom papers, folded through scripts/geo.mjs.

Papers Share of 3,908
No mappable location 2,773 71.0%
Exactly one country 532 13.6%
More than one country, or a region 603 15.4%
No infrastructure type stated 1,764 45.1%

Top folded locations, as a share of the 1,135 papers with at least one mappable location: United States 639 (56.3%), Germany 181 (15.9%), multi-country 175 (15.4%), Europe 146 (12.9%), China 111 (9.8%), United Kingdom 104 (9.2%). Where from reports the same shape on the narrower crawling population.

Which unit, among the 857 web-crawling papers

Unit Papers Share of 857
websites 469 54.7%
other 355 41.4%
web pages 171 20.0%
domains 170 19.8%
documents 116 13.5%
both a site/domain unit and a page unit 105 12.3%

other at 41.4% is itself a finding: the extraction's unit vocabulary does not fit two web-crawling papers in five, which is a warning about reading any unit table — including this one — as exhaustive.

Stating limitations is becoming normal

A heading probe over decolumned full text for a line consisting of Limitations or Threats to Validity (with optional section number). A paper that discusses limitations inline without a heading counts as a miss, so every figure is a lower bound.

Population Papers read Has the heading Share
webCrawled 857 202 23.6%
crawled 1,120 255 22.8%
repeated 250 58 23.2%
all papers 5,855 1,197 20.4%
Period (webCrawled) Papers read Has the heading Share
2010–2013 80 3 3.8%
2014–2017 130 22 16.9%
2018–2021 241 55 22.8%
2022–2024 253 76 30.0%
2025–2026 (provisional) 153 46 30.1%

An eightfold rise, and a plateau at three papers in ten. This is the most encouraging trend on the page and still leaves the majority of web-crawling papers with no place a reader can look for the caveats.

Methodology and limitations of these figures

  • Every figure is produced by scripts/biases_report.mjs, whose complete unedited output — including the full location fold residue and every population definition — is on biases along with each query and its denominator.
  • Ranking-list matching is a regular expression over a folded skeleton, not a synonym fold. A list named in a way none of the nine patterns catch is counted as not a ranking list, so 793 is a floor. Amazon's voice assistant is excluded from the Alexa count by name; all fourteen excluded strings are on the provenance page. The independent fold on Sampling reaches 463 Alexa papers against 475 here on a different population by a different method — that agreement is the closest thing to a cross-check either page has, and it is worth knowing that it also held while this page's fold carried two bugs, so it is evidence and not proof.
  • The full-text probes are keyword probes. They find phrases, not concepts. A paper that reports attrition as “we successfully loaded 91,204 of 100,000 sites” and never uses the vocabulary is a miss; so is one whose limitations discussion has no heading. All probe figures are therefore bounds, in the direction each table states.
  • population.unit is an enum with an other escape that 41.4% of web-crawling papers land in, so the unit table describes the extraction's vocabulary as much as the field's practice.
  • The extraction's free-text fields agree run-to-run on roughly 20% of exact strings — measured on the previous corpus and not re-measured since — which is why sourceList and locations are reported as folded families and never as precise percentages. Enum fields (samplingMethod, unit, infrastructure) are stable and are quoted directly.
  • 32 quotes behind the figures and claims on this page were checked against the papers' own text by scripts/biases_quotecheck.mjs, with three invented-sentence negative controls; all 32 located and all three controls failed. Two of them were wrong when review started — a misquoted version number and a paraphrase typeset as a quote — and both were wrong precisely because they were not yet in the list. The check and its output, and both fixes, are on the provenance page.
  • The corpus is seven venues. EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent, and the last two are the human-factors venues — so nothing here is a claim about the field as a whole. See the worked example below.

Worked Example: The Biases of This Site's Own Corpus

The corpus behind every figure above has the same four biases, and writing them down is the exercise this page is asking you to do for your own study.

Selection bias, by venue. The frame is seven venues chosen for relevance, not sampled. So the corpus over-represents whatever those seven publish and cannot see the rest:

Venue Papers Share of 5,859 Ran a crawl
USENIX Security 1,410 24.1% 15.7%
CCS 990 16.9% 16.5%
TheWebConf 843 14.4% 28.7%
IEEE S&P 767 13.1% 14.3%
NDSS 701 12.0% 18.4%
IMC 638 10.9% 20.7%
PETS 510 8.7% 24.1%

EuroS&P, ACSAC, RAID, AsiaCCS, CHI and SOUPS are absent. CHI and SOUPS are where much of the field's user-study work is published, so a claim here about user studies is a claim about user studies at security venues — which is why Hypothesis testing anchors its statistical-practice claims on a SOUPS review rather than on this corpus.

Denominator bias. Only 1,120 of 5,859 papers (19.1%) ran a crawl and 857 (14.6%) crawled the web. “19.1% of papers ran a crawl” is a fact about a security-venue corpus, not about web measurement; the crawling percentages on this page are all against 1,120 or 857 for that reason.

Survivorship, in two forms. Retrieval: a paper with no obtainable full text is not in the corpus, and there is no reason to expect that loss to be random with respect to venue, year or artefact practice. Selection: papers were screened on abstracts, so a paper that measures the web without saying so in the abstract is systematically missing.

Provisional venue-years, which is truncation rather than a trend. CCS 2026 and IMC 2026 have not been held; IEEE S&P 2026 and TheWebConf 2026 abstracts are not yet in the indexing source that screening reads. The 2026 bucket is therefore under-represented by construction:

Year Papers Ran a crawl Share
2022 546 110 20.1%
2023 719 125 17.4%
2024 690 110 15.9%
2025 (provisional) 770 129 16.8%
2026 (provisional) 415 69 16.6%

Which is why no per-year trend on this site ends at 2026 without the label, and why the “current practice” rows in Which Corrections Are Current rest on the thinnest years in the corpus and say so. The full corpus-level caveats are on Corpus.

One bias that was fixed rather than caveated. Until August 2026 this corpus retrieved only 333 of the 780 IEEE S&P papers its screening had selected, and pages built on it carried an “IEEE S&P is only 43% retrieved” limitation. The retrieval gap was then closed and every figure computed before that became stale. 767 IEEE S&P records are in the extraction today — not 780, because retrieval and extraction are different steps and a handful of selected papers still yield no extractable record. That is the honest ordering: caveat what you cannot fix, fix what you can, re-derive every number when you do, and do not quote the selection count as if it were the extraction count.

What to Report

A checklist for a methods section. Each line exists because this page measured the field not doing it.

  1. The frame, with its version. Which list, which snapshot or list ID, which date. 40% of web-crawling papers do not.
  2. The draw. Top-n, stratified with the band boundaries, or random with the seed. If top-n, say what the n is a census of.
  3. The funnel, as three numbers. Rows in the frame → units attempted → units that produced data. Every percentage in the paper then says which of the three it divides by.
  4. For a panel: attrition per wave, and the balanced-panel size. Then your headline figure computed both ways.
  5. The vantage point: country, infrastructure type, and provider. And say whether you varied it — a single point is a choice, not a default.
  6. The unit, and the same unit in the sampling and the results sections. If you sampled sites and report pages, say the ratio.
  7. For a query-built population: the query. String, date, API version, results requested, results returned, and whether a re-run agrees.
  8. A Limitations section with a heading, naming which of the four biases above applies to your design and in which direction it moves your estimate. Seven papers in ten do not have one; having one is cheap and it is the section a reviewer looks for.

Direction matters more than magnitude. “Our top-10k frame over-represents CDN-fronted desktop-visited sites, so our tracking prevalence is likely an over-estimate for the web at large” is a usable sentence. “Our sample may not be representative” is not.

Open Questions

  • The drop-out set has not been characterised. No paper in this corpus measures which sites fail to load in a crawl and how their composition differs from the sites that load — the measurement that would let every crawling paper bound its own survivorship. An external sweep found one adjacent short paper on crawl refusals in Common Crawl, which is about server-side blocking generally rather than about how the lost set differs, so the question stands. Also open on Open Questions.
  • No bias analysis for LLM-based labelling inside the seven venues. LLM classification appears in the 2025–2026 corpus slice (166 of 1,185 papers name a model in a classification tuple) and a characterisation of its selection and calibration behaviour does not. The closest work is [15Bozzolan, Simone; Calzavara, Stefano; Cazzaro, Lorenzo (2026): "LLM-Assisted Web Measurements". arXiv:2510.08101, v3, 30 April 2026 (Link)], a 2026 arXiv preprint that benchmarks LLM website classification for exactly this purpose — it establishes that model and configuration choice matters a great deal, which is the premise of the question rather than its answer. What is missing is a per-class error and calibration profile a measurement paper could cite to bound its own labelling bias. If you know of one, please add it.
  • No post-2022 replacement for the top-list-against-ground-truth comparison. Ruth et al. [4Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] used Cloudflare resolver data in February 2022; CrUX has since been folded into Tranco's default list, so the thing they measured no longer exists in the same form and the comparison has not been re-run.
  • Attrition vocabulary. The probe on this page finds phrases. A hand-coded sample of repeated-crawl papers would establish how many report attrition numerically without using any of the vocabulary — the figure that would turn this page's upper bounds into estimates.
  • Representative sampling methods — how to draw, at what size, at what unit, and how to make the draw redrawable. The upstream page for 1. Selection Bias: Your Sample Is a Ranking List.
  • Website selection — the lists themselves, their provenance and their own documented limitations.
  • Crawling location — choosing and verifying a vantage point, with the full treatment of 3. Vantage-Point Bias on the crawling population.
  • Hypothesis testing — what non-independence and attrition do to a p-value once the sample is fixed.
  • Study preregistration and Pvalue corrections — fixing the population and the exclusion rule before the data arrives is the strongest available defence against the biases you would otherwise discover post hoc.
  • Archives — web archives sidestep the vantage point and add their own survivorship. Lerner et al. [18Lerner, Ada; Simpson, Anna Kornfeld; Kohno, Tadayoshi; Roesner, Franziska (2016): "Internet Jones and the Raiders of the Lost Trackers: An Archaeological Study of Web Tracking from 1996 to 2016", in: Proceedings of the USENIX Security Symposium. (Link)] measure the hole: “16.1% of all requests attempted to 'escape' … and were blocked by TrackingExcavator” — one request in six of what the historical page would have made is simply not in the archive.
  • Stateful stateless and Interaction — the two configuration axes behind the crawl-against-browsing gap that Zeber et al. [5Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] and Annamalai et al. [12Annamalai, Meenatchi Sundaram Muthu Selva; De Cristofaro, Emiliano; Bilogrevic, Igor (2025): "Beyond the Crawl: Unmasking Browser Fingerprinting in Real User Interactions", in: Proceedings of the ACM Web Conference. (DOI)] measure. Neither paper compares that gap against a geographic one; the observation that it is the larger number in this page's table is this page's, not theirs.
  • Corpus — corpus-level provenance for every figure on this site.
  • Provenance of this page's figures — every query with its denominator, the report script and its unedited output, the folds and their full residue, the quote checks, and what could not be established.

References

[1]
Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[2]
Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)
[3]
Urban, Tobias; Degeling, Martin; Holz, Thorsten; Pohlmann, Norbert (2020): "Beyond the Front Page:Measuring Third Party Dynamics in the Field", in: Proceedings of The Web Conference 2020, pp. 1275–1286. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[4]
Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[5]
Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[6]
Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)
[7]
Demir, Nurullah; Hörnemann, Jan; Große-Kampmann, Matteo; Urban, Tobias; Pohlmann, Norbert; Holz, Thorsten; Wressnegger, Christian (2023): "On the Similarity of Web Measurements Under Different Experimental Setups", in: Proceedings of the ACM Internet Measurement Conference, pp. 356-369. ACM DOI 10.1145/3618257.3624795 is listed by DBLP but was not registered with the DOI resolver as of 2026-08-12; the DOI above resolves to the authors' institutional record of the same paper (DOI)
[8]
Angrist, Joshua D.; Pischke, Jörn-Steffen (2009): "Mostly Harmless Econometrics: An Empiricist's Companion". Princeton University Press. (DOI)
[9]
Singh, Rachee; Nithyanand, Rishab; Afroz, Sadia; Pearce, Paul; Tschantz, Michael Carl; Gill, Phillipa; Paxson, Vern (2017): "Characterizing the Nature and Dynamics of Tor Exit Blocking", in: 26th USENIX Security Symposium (USENIX Security 17), pp. 325-341. USENIX Association. (Link)
[10]
Hils, Maximilian; Woods, Daniel W.; Böhme, Rainer (2021): "Privacy Preference Signals: Past, Present and Future", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[11]
Singh, Sachin Kumar; Ricci, Robert; Gamero-Garrido, Alexander (2025): "Where in the World Are My Trackers? Mapping Web Tracking Flow Across Diverse Geographic Regions", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[12]
Annamalai, Meenatchi Sundaram Muthu Selva; De Cristofaro, Emiliano; Bilogrevic, Igor (2025): "Beyond the Crawl: Unmasking Browser Fingerprinting in Real User Interactions", in: Proceedings of the ACM Web Conference. (DOI)
[13]
Aqeel, Waqar; Chandrasekaran, Balakrishnan; Feldmann, Anja; Maggs, Bruce M. (2020): "On Landing and Internal Web Pages: The Strange Case of Jekyll and Hyde in Web Performance Measurement", in: Proceedings of the ACM Internet Measurement Conference, pp. 680-695. (DOI)
[14]
Efstratiou, Alexandros (2025): "On YouTube Search API Use in Research", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[15]
Bozzolan, Simone; Calzavara, Stefano; Cazzaro, Lorenzo (2026): "LLM-Assisted Web Measurements". arXiv:2510.08101, v3, 30 April 2026 (Link)
[16]
Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)
[17]
Cassel, Darion; Lin, Su-Chin; Buraggina, Alessio; Wang, William; Zhang, Andrew; Bauer, Lujo; Hsiao, Hsu-Chun; Jia, Limin; Libert, Timothy (2022): "OmniCrawl: Comprehensive Measurement of Web Tracking With Real Desktop and Mobile Browsers", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[18]
Lerner, Ada; Simpson, Anna Kornfeld; Kohno, Tadayoshi; Roesner, Franziska (2016): "Internet Jones and the Raiders of the Lost Trackers: An Archaeological Study of Web Tracking from 1996 to 2016", in: Proceedings of the USENIX Security Symposium. (Link)
1)
Amazon's own retirement notice: “After two decades of helping you find, reach, and convert your digital audience, we made the difficult decision to retire Alexa.com on May 1, 2022 … The APIs will be retired on December 15, 2022.” support.alexa.com no longer resolves; read via the Internet Archive at https://web.archive.org/web/2022/https://support.alexa.com/hc/en-us/articles/4410503838999 on 2026-08-20.
2)
Verified against https://tranco-list.eu/ on 2026-08-20.
3)
The CrUX dataset is live and still updated monthly through BigQuery and the CrUX API. The standalone CrUX Dashboard product was retired in late 2025 — if a tutorial sends you there the data has not gone away, only that front end. Checked 2026-08-20.
You could leave a comment if you were logged in.
statistics/biases.1787256834.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki