Table of Contents
How Many Sites
Every crawl paper contains a number that nobody chose. It appears in the second paragraph of the methodology — we crawled the top 10,000 websites, we visited the Alexa top 1 million — and it is almost never derived from anything. Of 1,121 papers in our corpus of seven security and privacy venues that drew a population of websites, domains or web pages and stated its size, 6 (0.5%) derive that size from a stated precision, power or confidence requirement, and four of those six are sizing a hand-checked subsample rather than the crawl.1) The rest of the field inherits its n from the paper it read last.
This page is about the number. Sampling is about how you draw (which method, which unit, how to make the draw repeatable) and Hypothesis testing is about which test you run once you have the rows; neither owns the question how many, and it is the question a first crawl actually gets stuck on. It is also not a power-analysis tutorial: the closed-form formulas below are three lines of arithmetic, reproduced so the tables can be checked, and a textbook covers the derivations. What belongs here is what those formulas do to web data — where n buys you almost nothing, where it is the whole experiment, and why 100,000 sites are not 100,000 observations.
The one thing to take away. There is no single right n, because n answers three different questions with three opposite scaling laws:
- How common is a common thing? Precision improves with √n, so it saturates: at n = 10,000 a 95% interval on a 5% prevalence is already ±0.43 pp, and the classifier that produced the labels is wrong by more than that. More sites do not fix it. A few thousand is enough, and the budget belongs elsewhere.
- Does a rare thing exist, and how rare? Your expected count is n×p. At p = 0.01% you need 29,956 sites for a 95% chance of seeing even one, and 960,304 to pin it down to ±20% of itself. This — and only this — is what a million-site crawl is for, and if it is your reason, say so, because the reader cannot tell it apart from a scale boast.
- Do two groups differ? Here your n is not the n the test sees. Sites share tag managers, consent platforms, CMS templates and hosting; with clusters of 100 sites and an intra-class correlation of 0.05, a 1,000,000-site crawl carries the information of 168,067 independent ones. Adding sites from the same clusters buys almost nothing.
Whichever it is, the reader needs one sentence saying which — and the n you report next to it should be the one you analysed, not the one you drew.
What to Read First
Four entries. None of them is a statistics text; each is a web measurement whose result is about n itself.
- Sy, Burkert, Federrath & Fischer [2Sy, Erik; Burkert, Christian; Federrath, Hannes; Fischer, Mathias (2019): "A QUIC Look at Web Tracking", Proceedings on Privacy Enhancing Technologies 2019(3):255-266. (DOI)] (PoPETs 2019) — read this table first, and then never write “we measured the top n” without saying which n again. Measuring QUIC support at six nested list depths, they report: “Alexa Top 10 20.00% Alexa Top 100 21.00% Alexa Top 1K 8.10% Alexa Top 10K 1.69% Alexa Top 100K 0.19% Alexa Top 1M 0.02%”. The same phenomenon, the same crawler, the same list — and a thousand-fold spread in the headline percentage, entirely determined by where the list was cut. The absolute count is small either way: “We found 186 websites within the Alexa Top Million that support the QUIC protocol.”
- Oest et al. [3Oest, Adam; Safaei, Yeganeh; Doupé, Adam; Ahn, Gail-Joon; Wardman, Brad; Tyers, Kevin (2019): "PhishFarm: A Scalable Framework for Measuring the Effectiveness of Evasion Techniques against Browser Phishing Blacklists", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] (IEEE S&P 2019) and [4Oest, Adam; Safaei, Yeganeh; Zhang, Penghui; Wardman, Brad; Tyers, Kevin; Shoshitaishvili, Yan; Doupé, Adam; Ahn, Gail-Joon (2020): "PhishTime: Continuous Longitudinal Measurement of the Effectiveness of Anti-phishing Blacklists", in: Proceedings of the USENIX Security Symposium. (Link)] (USENIX Security 2020) — the only two papers in this corpus that size a web population from a power calculation. “Our goal was to obtain a power of 0.95 at the significance level of 0.05 in a one-way independent ANOVA test”, giving “a sample size of 384 phishing sites for each entity”. PhishFarm also records the attrition that did not happen — “All sites ultimately delivered 100% uptime during deployment, thus we ended up with an effective sample size of 396 per experiment” — and the follow-up, PhishTime, is the one that closes the loop by reporting the effect size it actually observed against the one it assumed. Read them for the shape of the argument, which transfers to any A/B crawl.
- Konoth et al. [5Konoth, Radhesh Krishnan; Vineti, Emanuele; Moonsamy, Veelasha; Lindorfer, Martina; Kruegel, Christopher; Bos, Herbert; Vigna, Giovanni (2018): "MineSweeper: An In-depth Look into Drive-by Cryptocurrency Mining and Its Defense", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] (CCS 2018) — the rare-event case, cleanly. Drive-by cryptomining on “1,735 websites” out of the “991,513” they managed to crawl from the Alexa top million: 0.18%. At n = 1,000 the expected count is 1.7 sites and the paper does not exist. This is the justification for scale, and it is arithmetic, not rhetoric.
- Hanley & Lippman-Hand [6Hanley, James A.; Lippman-Hand, Abby (1983): "If Nothing Goes Wrong, Is Everything All Right? Interpreting Zero Numerators", JAMA 249(13):1743-1745. (DOI)] (JAMA 1983) — outside this field and worth the three pages. The source of the rule of three: if you observe zero events in n trials, the 95% upper bound on the rate is about 3/n. It is the only correct way to report “we did not find any”, and it converts a null result into a bound, which is publishable.
For the two adjacent decisions, Sampling (how to draw, and rank-weighting as an alternative to scaling) and Website selection (which list, and what it measures) are the pages to read alongside this one. For what non-independence does to the test itself, Hypothesis testing.
The Number Is a Convention, and the Convention Is Measurable
1,000,000 is the modal answer
Over the 1,121 papers that drew at least one web population and stated a size, counting one number per paper — the largest web population it states:
| Largest stated web population | Papers | Share of 1,121 |
|---|---|---|
| 1,000,000 | 140 | 12.5% |
| 10,000 | 80 | 7.1% |
| 100,000 | 65 | 5.8% |
| 100 | 48 | 4.3% |
| 1,000 | 37 | 3.3% |
| 500 | 20 | 1.8% |
| 20,000 | 20 | 1.8% |
| 5,000 | 18 | 1.6% |
46.7% of these numbers are 1, 2 or 5 times a power of ten, and one paper in eight states exactly 1,000,000. The median is 20,000, the quartiles are 1,000 and 1,000,000, and 26.9% of the papers work with a thousand sites or fewer.2) A round number is not by itself wrong — 1,000,000 is a genuine list boundary, and so is a Tranco top-10k file — but a distribution this concentrated on powers of ten is a distribution of conventions, not of design decisions.
Design:Sampling reports a tuple-level median of 6,755 over 2,379 populations, and this page reports 20,000. Both are right and they answer different questions: that page describes every population a paper drew (including the small verification subsamples), this one describes the largest one per paper — the number that ends up in the abstract. Compare like with like before citing either.
The number stopped growing around 2018
| Years | Papers | Median n | p75 | n ≥ 1M | n ≤ 1,000 |
|---|---|---|---|---|---|
| 2010–2013 | 98 | 5,183 | 100,000 | 19.4% | 37.8% |
| 2014–2017 | 182 | 26,590 | 1,000,000 | 34.6% | 29.7% |
| 2018–2021 | 312 | 42,542 | 1,000,000 | 27.2% | 21.8% |
| 2022–2024 | 341 | 20,000 | 1,000,000 | 26.1% | 25.8% |
| 2025–2026 (provisional) | 188 | 20,000 | 1,000,000 | 29.3% | 29.3% |
Restricted to the 674 papers that also ran a crawl, the medians are 10,000 → 18,000 → 35,102 → 13,796 → 20,000. The median rose by roughly eight-fold to 2018–2021 and then fell back by half — to 20,000 over all papers and 13,800–20,000 among crawling ones — while storage, bandwidth and browser automation all got cheaper across the same window. The share of papers going to a million or more has been flat at a quarter to a third since 2014, so the fall is in the middle of the distribution, not at the top. Read the last row with care: 2025–2026 is provisional, because CCS and IMC 2026 have not been held and IEEE S&P and TheWebConf 2026 are incompletely selected (corpus).
Two readings are available and the corpus cannot choose between them: either the field discovered that the extra sites were not buying anything, or per-site analysis got more expensive — the next section is evidence for the second. What the table does establish is that there is no upward trend to keep up with. A number in the low tens of thousands is squarely inside what the last eight years published; the case for a larger one has to come from your question, not from the calendar.
What actually sets the number: the cost of what you do per site
Sorting the same 1,121 papers by what they do to each site produces a ranking that no statistical consideration does:
| Per-site analysis method | Papers | Median n | p25 | p75 |
|---|---|---|---|---|
| regex or signature | 111 | 132,798 | 4,000 | 1,000,000 |
| third-party service | 271 | 100,000 | 7,337 | 1,000,000 |
| heuristic rules | 605 | 93,427 | 2,846 | 1,000,000 |
| curated database | 228 | 93,284 | 3,806 | 1,000,000 |
| dynamic analysis | 66 | 74,216 | 5,000 | 1,000,000 |
| blocklist | 145 | 20,000 | 10,000 | 387,000 |
| supervised ML | 266 | 20,000 | 1,000 | 442,190 |
| static analysis | 50 | 15,000 | 1,899 | 1,000,000 |
| manual labelling | 307 | 10,000 | 361 | 906,731 |
| LLM | 29 | 10,000 | 2,892 | 90,000 |
The gap between a regex and a human is about an order of magnitude in n. The ordering is not perfectly monotone in cost — a blocklist lookup is cheap and its median sits at 20,000, alongside supervised ML — and classification.method is multi-valued and only moderately stable run-to-run (58%), so this is a ranking consistent with per-site cost rather than a measurement of it. Read it that way and the practical point survives: if you want a larger n, the lever is a cheaper per-site step, not a bigger machine.
The LLM row is the newest and the thinnest: 29 papers, 22 of them in the provisional 2025–2026 slice, with a median of 10,000 across all 29 (10,500 within 2025–2026) against 20,000 for non-LLM papers in the same years. That is the same order as manual labelling, which is what you would expect from a method with a real per-site price, and it is worth watching — but 29 papers is a ranking, not a rate, and this row should be re-derived when the 2025–2026 venue-years fill in.
Three Questions, Three Different Answers
A. Precision on a prevalence: the returns die at a few thousand
If the claim is “x% of sites do Y” and Y is not rare, the relevant quantity is the width of the interval on x. It shrinks with √n, which means the first thousand sites buy most of what you will ever get:
| n | p = 0.5 | p = 0.2 | p = 0.05 | p = 0.01 | p = 0.001 |
|---|---|---|---|---|---|
| 100 | ±9.80 pp | ±7.84 pp | ±4.27 pp | ±1.950 pp | ±0.619 pp |
| 1,000 | ±3.10 pp | ±2.48 pp | ±1.35 pp | ±0.617 pp | ±0.196 pp |
| 10,000 | ±0.98 pp | ±0.78 pp | ±0.43 pp | ±0.195 pp | ±0.062 pp |
| 100,000 | ±0.31 pp | ±0.25 pp | ±0.14 pp | ±0.062 pp | ±0.020 pp |
| 1,000,000 | ±0.10 pp | ±0.08 pp | ±0.04 pp | ±0.020 pp | ±0.006 pp |
Half-width of a 95% Wald interval, 1.96·√(p(1−p)/n). Nothing about the web is in this table; it is printed by the report script on the provenance page so the numbers below can be checked. Turned around, at the worst case p = 0.5: ±5 pp needs 385 sites, ±1 pp needs 9,604, and ±0.1 pp needs 960,400.
The decisive observation is not the saturation — it is what the interval is being compared against. Past a few thousand sites the sampling error is plausibly the smallest error in the pipeline, and the others do not shrink with n at all.3) The three that matter:
- Your labeller is wrong by more than the interval. A filter-list-based or classifier-based label on “is this a tracking cookie” that is 97% accurate can bias a prevalence by up to 3 percentage points — about seven times the ±0.43 pp sampling interval at n = 10,000 — and unlike the interval it does not shrink at all when you add sites. See Website classification and Filter lists for what those error rates actually are.
- Your frame is not the population. Scheitle et al. [7Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] measured top-list domains against the general
com/net/orgpopulation and found IPv6 support at 11–13% against 4%, and CDN use higher by at least a factor of two — a factor of twenty at the top 1,000. Ruth et al. [8Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] found that of the 1,790 Alexa top-10K domains they could measure, 70% sat in a lower rank-magnitude bucket in Cloudflare's own traffic data. Adding sites from the same list does not reduce coverage error at all — it only makes the census of the head more exact. Biases is the page for the sizes of those effects. - Your crawler is not a browser. Zeber et al. [9Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] found a crawler reaching a median of 11.6 third parties per visit against 4.5 for humans; Demir et al. [10Demir, Nurullah; Hörnemann, Jan; Große-Kampmann, Matteo; Urban, Tobias; Pohlmann, Norbert; Holz, Thorsten; Wressnegger, Christian (2023): "On the Similarity of Web Measurements Under Different Experimental Setups", in: Proceedings of the ACM Internet Measurement Conference, pp. 356-369. ACM DOI 10.1145/3618257.3624795 is listed by DBLP but was not registered with the DOI resolver as of 2026-08-12; the DOI above resolves to the authors' institutional record of the same paper (DOI)] found 48% of the underlying data varying between two crawler profiles. Neither gap narrows with n.
So for a prevalence question the sizing rule that follows is: pick the n at which the sampling interval is comfortably below the largest error you cannot remove, and spend the rest of the budget on a second frame, a second vantage point, or a validated labeller.
Worked, with the labeller as the binding error, because it is usually the largest one you can put a number on. A label that is 97% accurate leaves an irreducible ±3 pp. “Comfortably below” at a factor of five means a target half-width of 0.6 pp, which at the worst case p = 0.5 is n ≈ 26,700, and at p = 0.05 is n ≈ 5,100. Push the labeller to 99% and the same rule gives 240,000 and 46,000. Turn it round and the rule is easier to use: the n you need is set by how good your labels are, and a paper that reports a 97%-accurate classifier over a million sites has spent a hundred times the crawl budget to reduce an error it is not the bottleneck for.
That puts the defensible range for a prevalence question at roughly 5,000 to 50,000 sites, depending on your p and your labeller — which happens to bracket the field's median of 20,000. The agreement is worth noticing but it is not evidence: the field did not derive its median, as the rest of this page shows. The arithmetic gets there independently.
A top-n crawl usually has no sampling error to report at all. It is a census of its frame, so the binomial interval above is not the uncertainty in your number — the uncertainty is the extrapolation from that frame to the population you write about, and it is not binomial. Sampling states this at more length; the consequence for n is that buying more sites narrows an interval you were not entitled to compute. If you want a defensible interval, stratify and report per stratum.
B. Finding a rare thing: this is what a million sites is for
The other question inverts every conclusion above. If the phenomenon has prevalence p, your expected count is n·p, and no amount of statistical sophistication substitutes for having seen the thing:
| Prevalence p | Expected hits at n=10,000 | at n=100,000 | at n=1,000,000 | n for a 95% chance of ≥1 hit | n for a ±20% relative interval |
|---|---|---|---|---|---|
| 10% | 1,000 | 10,000 | 100,000 | 29 | 865 |
| 1% | 100 | 1,000 | 10,000 | 299 | 9,508 |
| 0.1% | 10 | 100 | 1,000 | 2,995 | 95,944 |
| 0.01% | 1 | 10 | 100 | 29,956 | 960,304 |
| 0.001% | 0.1 | 1 | 10 | 299,572 | 9,603,904 |
Two columns matter. The fifth is the existence threshold, roughly 3/p: below it your study can easily return zero and conclude nothing. The sixth is the estimation threshold: seeing three instances tells you the thing exists, not how common it is, and pinning a rate down to within a fifth of itself costs roughly 96/p sites for small p. The jump from column five to column six is the reason honest rare-phenomenon papers are large.
These are not hypothetical prevalences. Every figure below is a site-level rate this corpus reports, with the n the paper needed to get it:
| Phenomenon | Rate reported | Drawn n | Analysed n | Expected hits at n=10,000 |
|---|---|---|---|---|
| Drive-by cryptomining [5Konoth, Radhesh Krishnan; Vineti, Emanuele; Moonsamy, Veelasha; Lindorfer, Martina; Kruegel, Christopher; Bos, Herbert; Vigna, Giovanni (2018): "MineSweeper: An In-depth Look into Drive-by Cryptocurrency Mining and Its Defense", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] | 1,735 sites, 0.18% | Alexa top 1,000,000 | 991,513 | 18 |
| Cryptojacking, Internet-wide [11Bijmans, Hugo L.J.; Booij, Tim M.; Doerr, Christian (2019): "Inadvertently Making Cyber Criminals Rich: A Comprehensive Study of Cryptojacking Campaigns at Internet Scale", in: Proceedings of the USENIX Security Symposium. (Link)] | 5,190 domains, 0.011% — “one in every 9,090 websites” | ~20% of the Internet | 48,948,669 | 1.1 |
| QUIC support in the full top million [2Sy, Erik; Burkert, Christian; Federrath, Hannes; Fischer, Mathias (2019): "A QUIC Look at Web Tracking", Proceedings on Privacy Enhancing Technologies 2019(3):255-266. (DOI)] | 186 sites, 0.02% | Alexa top 1,000,000 | not reported | 2 |
| Server-Sent Events [12Murley, Paul; Ma, Zane; Mason, Joshua; Bailey, Michael D.; Kharraz, Amin (2021): "WebSocket Adoption and the Landscape of the Real-Time Web", in: Proceedings of the ACM Web Conference. (DOI)] | 0.05% of the top million | Tranco top 1,000,000 | 881,000 (88.1%) | 5 |
| Non-secure DNS dynamic updates [13Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | 1,877 domains, 0.065% | a 1% draw from 286,788,250 | 2,865,393 | 6.5 |
At n = 10,000 — the second most common value in this population after a million, and half the field's median — four of these five studies would have found single digits or nothing. That is the case for scale, and it is worth writing down explicitly, because it is invisible in the finished paper: the reader sees 1,735 sites and not the 989,778 that were visited and had nothing on them.
The two n columns are the page's own rule applied to its own examples, and applying it changed three of the five rows. Konoth et al. analysed 0.9% fewer sites than they drew and Murley et al. 11.9% fewer — both report it plainly, and neither gap changes a conclusion. Korczyński et al. is the one that matters: the structured extraction had recorded 286,788,250, the domain corpus the study drew its 1% sample from, which is a hundred times the population that was actually tested for the vulnerability. Sy et al. report no attrition figure at all, and Bijmans et al. have no drawn number to compare against — “~20% of the Internet” is a description, not a sampling frame. Only the analysed n belongs next to a prevalence, and the first draft of this table used the drawn n throughout — precisely the mistake the box at the top of this page warns against, caught by a reviewer reading the papers rather than by any check.
If you found nothing, report the bound, not the zero. Zero hits in n sites gives a 95% upper bound of about 3/n [6Hanley, James A.; Lippman-Hand, Abby (1983): "If Nothing Goes Wrong, Is Everything All Right? Interpreting Zero Numerators", JAMA 249(13):1743-1745. (DOI)] — 0.03% at n = 10,000, 0.003% at n = 100,000. “We did not observe X” is unreviewable; “X is present on at most 0.03% of the top 10,000 sites, 95% CI” is a result. The same arithmetic tells you in advance whether a null finding would be worth anything: if 3/n is above the prevalence anyone would care about, the crawl cannot answer the question at that size.
| n with zero hits | 95% upper bound | i.e. at most, per million sites |
|---|---|---|
| 100 | 3.0000% | 30,000 |
| 1,000 | 0.3000% | 3,000 |
| 10,000 | 0.0300% | 300 |
| 100,000 | 0.0030% | 30 |
| 1,000,000 | 0.0003% | 3 |
Where does p come from before you have measured it? From a pilot, and the rule of three is what makes a pilot informative. Crawl 10,000 sites first. If you find k instances, p ≈ k/10,000 and the table above tells you what n the full crawl needs; if you find none, you have a 95% bound of 0.03% and the same table says 29,956 sites for a coin-flip chance of a single hit and 960,304 to estimate the rate — at which point the honest options are to fund the large crawl, to reframe the question, or to publish the bound. A pilot costs a day and it is the difference between a sizing argument and a guess. It is also the one sizing move that survives the “you cannot know p in advance” objection, which is otherwise the reason nobody in this corpus does the calculation.
C. Comparing two groups: your //n// is not the //n// the test sees
The third question — EU against US, before against after consent, blocked against unblocked — is where the arithmetic in section A above becomes actively misleading, because a proportion's interval assumes independent draws and sites are not independent. They share tag-manager containers, consent platforms, CMS-plus-plugin stacks, hosting and operators, and each of those moves hundreds or thousands of sites together. Hypothesis testing covers what this does to a p-value and what to do about it; the consequence for n is a deflation.
For clusters of size m with intra-class correlation ρ, the design effect is 1 + (m−1)ρ and the effective sample size is n divided by it [14Killip, Shersten; Mahfoud, Ziyad; Pearce, Kevin (2004): "What Is an Intracluster Correlation Coefficient? Crucial Concepts for Primary Care Researchers", The Annals of Family Medicine 2(3):204-208. (DOI)]. Applied to a million-site crawl:
| Cluster size m | ρ = 0.02 | ρ = 0.05 | ρ = 0.20 |
|---|---|---|---|
| 10 | 847,458 | 689,655 | 357,143 |
| 100 | 335,570 | 168,067 | 48,077 |
| 1,000 | 47,664 | 19,627 | 4,980 |
A million sites in clusters of a thousand with ρ = 0.05 carries the information of 19,627 independent ones. That is a 51-fold deflation, and clusters of a thousand sites are not a stretch on the web: one consent-management platform, one analytics container or one popular WordPress plugin comfortably exceeds it.
Two honest caveats. First, nobody has measured ρ for any real web-measurement outcome — the open question on the neighbouring page says so, and this table therefore uses the same illustrative values as the simulation published there. Second, the deflation is a property of the outcome, not of the crawl: an outcome that genuinely varies site by site has ρ near zero and loses nothing. The practical consequence is not a formula but a habit:
- Adding sites drawn from the same clusters is the cheapest and least useful way to increase n. Ten thousand more sites all running the same CMS-plus-CMP stack add almost no information about that stack's behaviour.
- Count your clusters, not your sites, when you are sizing a comparison. If a comparison rests on twelve consent platforms, n is closer to twelve than to your site count, and going from 10,000 to 100,000 sites will not change that. Aggregating to the cluster and testing the clusters is the crude, always-valid version.
- Where the clusters themselves are the population, n is fixed by the world and no crawl size helps. There are only so many major CDNs.
The Tail Is a Different Population, Not More of the Same
The unspoken assumption behind “we can always add more sites” is that site 500,001 is like site 5,000. It is not. Extending a crawl down a popularity ranking does not enlarge your sample; it changes the estimand, and by amounts that dwarf every interval in section A above.
Nineteen papers that report the same measurement per rank band
A full-text probe over the 1,121 papers for a rank band, a cross-rank comparison and a percentage in one sentence returns 28 papers; reading all 28 leaves 19 that genuinely report the same measurement at more than one list depth.4) That is 1.7% of the population. It is still the most under-supplied table in this literature — Design:Sampling asks for a paper whose contribution is that table and there is none — but nineteen incidental results are enough to settle the direction question, and they settle it in a way that a single number would have got wrong.
| Paper | Measurement | Nearer the head | Deeper in the list |
|---|---|---|---|
| Falls with rank | |||
| [2Sy, Erik; Burkert, Christian; Federrath, Hannes; Fischer, Mathias (2019): "A QUIC Look at Web Tracking", Proceedings on Privacy Enhancing Technologies 2019(3):255-266. (DOI)] | QUIC support | 21.00% (top 100) | 0.02% (top 1M) |
| [12Murley, Paul; Ma, Zane; Mason, Joshua; Bailey, Michael D.; Kharraz, Amin (2021): "WebSocket Adoption and the Landscape of the Real-Time Web", in: Proceedings of the ACM Web Conference. (DOI)] | Server-Sent Events | 0.4% (top 1K) | 0.05% (top 1M) |
| [15Poteat, Tara; Li, Frank (2021): "Who You Gonna Call? An Empirical Evaluation of Website security.txt Deployment", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | security.txt adoption | 8–10% (top 1K) | ~1% (top 100K) |
| [16Al Roomi, Suood; Li, Frank (2023): "A Large-Scale Measurement of Website Login Policies", in: Proceedings of the USENIX Security Symposium. (Link)] | login rate limiting | 33.6% (top 10K) | 24.1% (top 1M) |
| [17Ikram, Muhammad; Masood, Rahat; Tyson, Gareth; Kaafar, Mohamed Ali; Loizon, Noha; Ensafi, Roya (2019): "The Chain of Implicit Trust: An Analysis of the Web Third-party Resources Loading", in: Proceedings of the ACM Web Conference. (DOI)] | third-party dependency chains | 55% (top 10K) | lower below, no second figure |
| [18Hantke, Florian; Calzavara, Stefano; Wilhelm, Moritz; Rabitti, Alvise; Stock, Ben (2023): "You Call This Archaeology? Evaluating Web Archives for Reproducible Web Security Measurements", in: Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, pp. 3168–3182. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] | web-archive hit rate | higher | lower; the bands are in a plot |
| Rises with rank | |||
| [19Aas, Josh; Barnes, Richard; Case, Benton; Durumeric, Zakir; Eckersley, Peter; Flores-López, Alan; Halderman, J. Alex; Hoffman-Andrews, Jacob; Kasten, James; Rescorla, Eric; Schoen, Seth D.; Warren, Brad (2019): "Let's Encrypt: An Automated Certificate Authority to Encrypt the Entire Web", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] | Let's Encrypt market share | 5% (top 1K) | 35% (top 1M) |
| [20Ashiq, Md. Ishtiaq; Li, Weitong; Fiebig, Tobias; Chung, Taejoong (2023): "You've Got Report: Measurement and Security Implications of DMARC Reporting", in: Proceedings of the USENIX Security Symposium. (Link)] | DMARC misconfiguration | ~10% (most popular 10K) | ~20% (least popular 10K) |
| [16Al Roomi, Suood; Li, Frank (2023): "A Large-Scale Measurement of Website Login Policies", in: Proceedings of the USENIX Security Symposium. (Link)] | passwords sent in the clear | 0.44% (top 10K) | 0.62% (top 1M) |
| [21Kashaf, Aqsa; Sekar, Vyas; Agarwal, Yuvraj (2020): "Analyzing Third Party Service Dependencies in Modern Web Services: Have We Learned from the Mirai-Dyn Incident?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | third-party CA use | 71% (top 100) | 77% (top 100K) |
| Flat | |||
| [13Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | non-secure DNS dynamic updates | 0.062% (Alexa top 1M) | 0.065% (random sample of 2.9M domains5)) |
| [22Singanamalla, Sudheesh; Jang, Esther Han Beol; Anderson, Richard; Kohno, Tadayoshi; Heimerl, Kurtis (2020): "Accept the Risk and Continue: Measuring the Long Tail of Government https Adoption", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | valid HTTPS on government sites | ~30% (top million) | ~30% (long tail) |
Twelve rows, eleven papers — Al Roomi and Li appear twice, in both directions — and ten of the twelve give a figure at each of two depths. The other two are directional only, and are marked as such rather than dropped, because the direction is the part this section is about.
The accounting, since the table and the probe are not the same set. Ten of the eleven tabled papers are among the nineteen. The eleventh, [13Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], is not: its comparison is between a ranking and a random sample of the whole zone rather than between two depths of one ranking, so the rank-band probe misses it. It was found through the rare-event query instead, and it is here because a genuinely flat result is the hardest of the three families to find. Nine of the nineteen are not tabled: two report something that is not a rank band of websites (popularity bands of trackers, and a 3-domain proxy-refusal rate), and seven either give a single figure plus a direction in the sentence the probe matched or put their bands in a table or plot the probe cannot read. All nine are named on the provenance page.
The direction is a property of the phenomenon, not of the tail
Read the table sideways: the tail is not systematically worse, or systematically emptier — it depends entirely on what you are measuring, and one paper can contain both signs. Al Roomi & Li [16Al Roomi, Suood; Li, Frank (2023): "A Large-Scale Measurement of Website Login Policies", in: Proceedings of the USENIX Security Symposium. (Link)] report that “0.44% of top 10K domains transmitted passwords in the clear, compared to 0.50% of top 100K domains and 0.62% of top 1M domains” — worse in the tail — and in the same results section that “33.6% of the domains in the top 10K and 26.7% of domains in the top 100K demonstrated rate limiting, compared to the 24.1% for domains in the top 1M” — also worse in the tail, but as a falling rate rather than a rising one. A single “the tail is worse” heuristic gets one of the two backwards, and a single “the tail is emptier” heuristic gets the other one backwards.
Three families are visible in the twelve rows. They are a description of what these papers found, not a rule you can apply in advance — the second family is full of host-side defaults, which is exactly where the third family says to expect no gradient, and no test on the corpus separates them.
- Adoption of anything new falls with rank. QUIC, Server-Sent Events,
security.txt, rate limiting. Popular sites have engineering teams; the tail runs whatever its host shipped. A crawl that stops at the top 1,000 will overstate the adoption of every emerging technology, sometimes by three orders of magnitude — Sy et al. [2Sy, Erik; Burkert, Christian; Federrath, Hannes; Fischer, Mathias (2019): "A QUIC Look at Web Tracking", Proceedings on Privacy Enhancing Technologies 2019(3):255-266. (DOI)] report “this share decreases for larger Top Alexa lists to only 0.0186% within the Alexa Top 1 Million” against 21% in the top 100. - Reliance on defaults, and plain neglect, rise with rank. Let's Encrypt [19Aas, Josh; Barnes, Richard; Case, Benton; Durumeric, Zakir; Eckersley, Peter; Flores-López, Alan; Halderman, J. Alex; Hoffman-Andrews, Jacob; Kasten, James; Rescorla, Eric; Schoen, Seth D.; Warren, Brad (2019): "Let's Encrypt: An Automated Certificate Authority to Encrypt the Entire Web", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] — “market share increases as site popularity decreases: 5% of the top 1K, 20% of the top 100K, and 35% of the top 1M”6) — third-party CAs [21Kashaf, Aqsa; Sekar, Vyas; Agarwal, Yuvraj (2020): "Analyzing Third Party Service Dependencies in Modern Web Services: Have We Learned from the Mirai-Dyn Incident?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], DMARC misconfiguration [20Ashiq, Md. Ishtiaq; Li, Weitong; Fiebig, Tobias; Chung, Taejoong (2023): "You've Got Report: Measurement and Security Implications of DMARC Reporting", in: Proceedings of the USENIX Security Symposium. (Link)], cleartext passwords [16Al Roomi, Suood; Li, Frank (2023): "A Large-Scale Measurement of Website Login Policies", in: Proceedings of the USENIX Security Symposium. (Link)].
- Some things genuinely do not vary. Korczyński et al. [13Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] found “1,877 (0.065%) and 587 (0.062%) of domains are vulnerable” in a random 2.9-million-domain sample and the Alexa top million respectively — the same rate to two significant figures. Singanamalla et al. [22Singanamalla, Sudheesh; Jang, Esther Han Beol; Anderson, Richard; Kohno, Tadayoshi; Heimerl, Kurtis (2020): "Accept the Risk and Continue: Measuring the Long Tail of Government https Adoption", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] found valid HTTPS on government sites at about 30% in the top million and in their long-tail dataset alike, noting only that “Though ranking does have an effect, overall valid https use” is comparable “in the long tail dataset, at” roughly the same level.7)
Because the families do not predict, the only reliable move is to measure two depths yourself, and it is cheap. Murley et al. [12Murley, Paul; Ma, Zane; Mason, Joshua; Bailey, Michael D.; Kharraz, Amin (2021): "WebSocket Adoption and the Landscape of the Real-Time Web", in: Proceedings of the ACM Web Conference. (DOI)] sized a stratified crawl at “the top 1000 sites, along with a random sample of 1000 sites from each of the top 10K, top 100K, and top 1M, for a total sample size of 4000 websites” — a fifth of the median crawl in this corpus, and it answers the question for your own measurement instead of borrowing an answer from someone else's phenomenon.
A fixed //n// forces you down the ranking, whether you meant to or not
The effect also runs backwards, and it is easy to miss. Bhuiyan et al. [23Bhuiyan, Masudul Hasan Masud; Varvello, Matteo; Staicu, Cristian-Alexandru; Zaki, Yasir (2025): "Digital Disparities: A Comparative Web Measurement Study Across Economic Boundaries", in: Proceedings of the ACM Web Conference. (DOI)] compared web behaviour across economies at a fixed per-country size, and record what that cost: “for countries such as Bangladesh, Pakistan, Nigeria, and the Philippines, less popular websites were included to meet the target sample size of 10,000 websites per country.” The number was held constant and the population moved — so a like-for-like comparison of countries became a comparison of “the top 10,000 of a large market” against “most of a small one”. Any design that fixes n across strata that are not all that large will do this. The alternatives are to fix the rank cut instead of the count, or to say plainly that the strata are not comparable at equal n.
What Extra Sites Actually Cost
The reason a plateau at 20,000 is defensible is that the marginal site is not free, and most of the costs are not the ones a proposal budgets for. Note that this section argues from mechanism, not from measurement: the corpus records no cost of any kind, because papers do not report crawl duration, bandwidth or money and the extraction has no field for them.
- The crawl parallelises; the human step does not. Adding machines buys you more site visits and not one more hand-checked page. Five of the seven sizing calculations in the next section are about a subsample for exactly this reason — the human step is the binding constraint at any n above a few thousand, and it is the cost that sets your real budget. The two that are not are experiments on sites the authors deployed themselves, where there is no crawl to hand-check.
- Attrition rises and is silently absorbed. More sites means more unreachable rows, more parked pages, more bot walls and more consent dialogues that do not match your handler. Sampling treats attrition as a denominator problem and Crawler detection as a tooling one; for sizing, the point is that the n you drew and the n you analysed diverge as you go deeper, and only the second one is the sample size. We could not measure the size of that divergence by rank: a probe for reachability figures reported per rank band returns four papers, three of them table fragments, so this is a mechanism argument and is flagged as one on the provenance page.
- Each site visit is a request on someone else's server. This is a real constraint and it is occasionally the stated reason for a smaller n: Bhaskar & Pearce [24Bhaskar, Abhishek; Pearce, Paul (2022): "Many Roads Lead To Rome: How Packet Headers Influence DNS Censorship Measurement", in: Proceedings of the USENIX Security Symposium. (Link)] reduced their domain set citing an ethical goal to minimise risk “combined with the likelihood of diminishing returns”. That is the correct shape of argument — the ethical cost is linear in n and the statistical benefit is not. Ethics is the page for the rest of it.
- Larger n makes a weaker labeller more tempting, which trades an error that shrinks with n for one that does not. The medians in the table above are partly this trade being made.
Sizing Calculations That Survived Review
Only 26 of the 1,121 papers contain a sentence that even mentions sample size, statistical power or a margin of error alongside a web unit and a number. Reading all 26 one by one:
| After reading the sentence | Papers | Share of 1,121 |
|---|---|---|
| derives n from a precision or power requirement | 6 | 0.5% |
| argues the size from design or cost, without a calculation | 5 | 0.4% |
| names a size and gives no argument for it at all | 1 | 0.1% |
| not about sizing a web population at all | 14 | 1.2% |
The last row is the reason a probe count is never the finding: those fourteen are confidence intervals on an already-collected dataset, thresholds for excluding small groups, a “largest studied to date” boast, two figure captions and a note that one rank band had too few domains. The provenance page prints all 26 sentences with the verdict on each, so the classification can be disagreed with.
Sizing a subsample: the calculation you probably do need
Four of the six calculations the probe found size a hand-verification subsample rather than the crawl, and so does the fifth paper the probe missed:
- Amjad et al. [25Amjad, Abdul Haddi; Shafiq, Zubair; Gulzar, Muhammad Ali (2023): "Blocking JavaScript Without Breaking the Web: An Empirical Investigation", in: Proceedings on Privacy Enhancing Technologies. (DOI)]: “manually inspected 383 websites, which is a statistically significant sample size for 100K websites” with ±5% margin of error.
- Le et al. [26Le, Hieu; Elmalaki, Salma; Markopoulou, Athina; Shafiq, Zubair (2023): "AutoFR: Automated Filter Rule Generation for Adblocking", in: Proceedings of the USENIX Security Symposium. (Link)]: “we randomly selected 272 sites (a sample size out of 933 sites to get a confidence level of 95% with a 5% confidence interval)”.
- Moore et al. [27Moore, Tyler; Leontiadis, Nektarios; Christin, Nicolas (2011): "Fashion Crimes: Trending-Term Exploitation on the Web", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], the earliest instance in the corpus: “we selected a statistically significant (95% confidence interval) random sample of 363 websites for manual inspection”.
- Chang et al. [28Chang, Li; Hsiao, Hsu-Chun; Jeng, Wei; Kim, Tiffany Hyun-Jin; Lin, Wei-Hsi (2017): "Security Implications of Redirection Trail in Popular Websites Worldwide", in: Proceedings of the ACM Web Conference. (DOI)]: a “domized sample of 2,000 websites (margin of error = 2.”8) — the same calculation at a larger population.
- Liao et al. [1Liao, Xiaojing; Liu, Chang; McCoy, Damon; Shi, Elaine; Hao, Shuang; Beyah, Raheem A. (2016): "Characterizing Long-tail SEO Spam on Cloud Web Hosting Services", in: Proceedings of the ACM Web Conference. (DOI)] use a Chernoff bound rather than a normal approximation for the same job, setting a trust interval of 0.01 and an error probability of 0.01 to arrive at n = 500 cloud directories to inspect by hand.9) This is the right tool when you want a guarantee on a proportion rather than an interval around it.
This is worth knowing because it is the one sample-size calculation almost every crawl paper needs and most do not do. You automated a label; a reviewer will ask how accurate it is; you will hand-check some sites. That subsample has genuine sampling error — it is a random draw from your own population, not a census of it — so it is exactly the case where the arithmetic in section A applies, and the answer is small: 385 sites for ±5 pp, about 1,068 for ±3 pp, and 9,604 for ±1 pp, almost independent of whether the population behind it is 10,000 sites or a million.
That “almost” explains the numbers in all three papers above, and it is worth one line because a student who copies 385 and then reads a paper saying 383 will wonder what they got wrong. The finite-population correction is n₀ / (1 + (n₀−1)/N), with n₀ = 385:
| Population N | Corrected size | The paper that reports it |
|---|---|---|
| 1,000,000 | 384 | — |
| 100,000 | 383 | Amjad et al.'s 383 [25Amjad, Abdul Haddi; Shafiq, Zubair; Gulzar, Muhammad Ali (2023): "Blocking JavaScript Without Breaking the Web: An Empirical Investigation", in: Proceedings on Privacy Enhancing Technologies. (DOI)] |
| 6,558 | 363 | Moore et al.'s 363 [27Moore, Tyler; Leontiadis, Nektarios; Christin, Nicolas (2011): "Fashion Crimes: Trending-Term Exploitation on the Web", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] |
| 933 | 272 | Le et al.'s 272 [26Le, Hieu; Elmalaki, Salma; Markopoulou, Athina; Shafiq, Zubair (2023): "AutoFR: Automated Filter Rule Generation for Adblocking", in: Proceedings of the USENIX Security Symposium. (Link)] |
All three are the same calculation. The correction is negligible above about 100,000 and starts to bite below about 10,000, which is why hand-check sizes in this literature cluster between 272 and 385 and never scale with the crawl. Reporting a hand-validation accuracy with no interval, or with a subsample of 50 sites, is an avoidable weakness and a cheap one to fix — five papers in 1,121 do the calculation, so the reviewer who asks for it is not asking for something unusual in the field so much as something almost absent from it.
The two that size the measured population itself
- A/B experiments on sites you control. Oest et al. [3Oest, Adam; Safaei, Yeganeh; Doupé, Adam; Ahn, Gail-Joon; Wardman, Brad; Tyers, Kevin (2019): "PhishFarm: A Scalable Framework for Measuring the Effectiveness of Evasion Techniques against Browser Phishing Blacklists", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] deployed phishing sites to measure blocklist response times, and sized the arms properly: “Our goal was to obtain a power of 0.95 at the significance level of 0.05 in a one-way independent ANOVA test” giving 384 sites per arm. The follow-up [4Oest, Adam; Safaei, Yeganeh; Zhang, Penghui; Wardman, Brad; Tyers, Kevin; Shoshitaishvili, Yan; Doupé, Adam; Ahn, Gail-Joon (2020): "PhishTime: Continuous Longitudinal Measurement of the Effectiveness of Anti-phishing Blacklists", in: Proceedings of the USENIX Security Symposium. (Link)] repeats the calculation — “To obtain a power of 0.95 at a p-value of 0.05, we initially assumed a medium effect size of 0.25” — and then reports the effect size it actually observed (0.36), which is the part almost nobody does and the part that makes the next study sizeable. If your design has arms, this is your template, and it applies unchanged to a browser-configuration A/B, a consent-action A/B or a vantage-point A/B.
- That appears to be the whole list. Two papers in 1,121 size a population they went on to measure from a statistical requirement, and both are experiments on sites the authors deployed themselves. We found no paper that derives the size of an observational crawl from a precision or power requirement — not through the sentence probe, and not through the structured
power-analysisfield, whose nine in-population papers are these two, Liao et al.'s Chernoff bound, and six participant power analyses in user studies. Two routes agreeing is not a census, and the probe's recall is a floor, so read this as “none that we could find” rather than “none”. That absence is not a scandal — for a census of a list there is no sampling error to compute a size from (see the box in section A) — but it does mean the field has no worked example to copy, and it is why the sizing argument this page recommends is a mechanism argument rather than a formula.
And two designs worth copying that are not calculations
- Empirical saturation. McDonald et al. [29McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] resampled “500 random combinations of different sample sizes” to see at what size a block page stops appearing. When you cannot write down p in advance, measuring where your own curve flattens is a defensible substitute, and it produces a figure a reviewer can read.
- Stratify instead of scaling. Murley et al. [12Murley, Paul; Ma, Zane; Mason, Joshua; Bailey, Michael D.; Kharraz, Amin (2021): "WebSocket Adoption and the Landscape of the Real-Time Web", in: Proceedings of the ACM Web Conference. (DOI)], above: 4,000 sites across four rank magnitudes answers more questions than 40,000 from the head. Sampling makes the general case, including rank-weighted estimators as the third option.
Which Practices Are Current
Dating matters here because the corpus rewards whatever was fashionable mid-window, and the sizing habits of 2015 are still being copied.
| Practice | Status | Evidence |
|---|---|---|
| The number 1,000,000, as a default population size | Current, and it outlived the list it came from. alexa.com was retired on 1 May 2022 and its APIs on 15 December 2022,10) but the round million did not go with it: papers stating exactly 1,000,000 run 7.1% → 15.4% → 11.5% → 12.6% → 13.8% across the five year buckets, and of the 26 such papers in 2025-2026, 25 draw a Tranco top 1M and one an Alexa list. The frame was replaced and the number was kept. Treat 1,000,000 as a list boundary you have chosen, not as a size you have justified. | this page, above; Website selection and [30Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)] for the list itself |
| A round number matched to the field's median (5k–50k) | Current and defensible. Since 2018 the median has stayed between 13,800 and 42,500 across the year buckets, and between 13,800 and 35,100 among crawling papers only. Say which question the n serves and you will not be challenged on the number itself. | this page, above |
| A million or more sites, for a rare phenomenon | Current and correct, when the prevalence justifies it. Between a quarter and a third of the population does this. The arithmetic supports a million at p around 0.01% or below; at p = 0.1% about a hundred thousand sites is already enough to estimate the rate to ±20% of itself. | [5Konoth, Radhesh Krishnan; Vineti, Emanuele; Moonsamy, Veelasha; Lindorfer, Martina; Kruegel, Christopher; Bos, Herbert; Vigna, Giovanni (2018): "MineSweeper: An In-depth Look into Drive-by Cryptocurrency Mining and Its Defense", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [11Bijmans, Hugo L.J.; Booij, Tim M.; Doerr, Christian (2019): "Inadvertently Making Cyber Criminals Rich: A Comprehensive Study of Cryptojacking Campaigns at Internet Scale", in: Proceedings of the USENIX Security Symposium. (Link)] |
| A million or more sites, for a common phenomenon | Superseded by stratification and weighting. At p = 5% the extra 990,000 sites narrow the interval from ±0.43 pp to ±0.04 pp, a precision no other component of the pipeline can support. | [31Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)], [8Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] |
| Rank-stratified sampling | Current best practice, still rare — 5.0% in 2018–2021, 8.9% in 2022–2024 and 6.3% in the provisional 2025–2026 bucket, against 6.0% of all 1,153 web-sampling papers, with no clear trend across those three buckets.11) It is the only design that answers the rank-tail question rather than assuming it away. | [12Murley, Paul; Ma, Zane; Mason, Joshua; Bailey, Michael D.; Kharraz, Amin (2021): "WebSocket Adoption and the Landscape of the Real-Time Web", in: Proceedings of the ACM Web Conference. (DOI)]; Sampling, whose report_sampling.mjs produces these shares |
| Rank weighting instead of a larger n | Current, and still essentially one paper's idea in this corpus. Prominence [31Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] makes an estimate comparable across studies with different n, which is exactly the problem the QUIC table illustrates. | Weight by rank instead of buying more sites |
| Traffic-weighted frames (CrUX) instead of a ranking | Current, and named directly by under 10% of papers — though many more use it indirectly, because Tranco has folded CrUX into its default inputs since 1 August 2023. Ruth et al. [8Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] found CrUX the most accurate list on every metric they tested in 2022. Accuracy of the frame substitutes for size of the sample. | [8Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)], [32Ruth, Kimberly; Fass, Aurore; Azose, Jonathan; Pearson, Mark; Thomas, Emma; Sadowski, Caitlin; Durumeric, Zakir (2022): "A World Wide View of Browsing the World Wide Web", in: Proceedings of the ACM Internet Measurement Conference, pp. 317-336. (DOI)]; Website selection for the list composition |
| Power analysis for a site count | Never established rather than superseded. 87 of 1,762 papers that ran inference report a power analysis, 74 of them for human participants; 6 also ran a crawl, and 3 of those derive a size for a web population. It remains a user-study habit. | below; [33Tang, Jenny; Bauer, Lujo; Christin, Nicolas (2025): "Misuse, Misreporting, Misinterpretation of Statistical Methods in Usable Privacy and Security Papers", in: Proceedings of the Twenty-First Symposium on Usable Privacy and Security (SOUPS), pp. 475-493. USENIX Association. (Link)] |
| Sizing a hand-verification subsample | Current, correct, and the gap worth closing. Done by five papers in 1,121, spread across 2011, 2016, 2017 and two in 2023 — never common, and never abandoned either. This is the calculation to copy. | [25Amjad, Abdul Haddi; Shafiq, Zubair; Gulzar, Muhammad Ali (2023): "Blocking JavaScript Without Breaking the Web: An Empirical Investigation", in: Proceedings on Privacy Enhancing Technologies. (DOI)], [26Le, Hieu; Elmalaki, Salma; Markopoulou, Athina; Shafiq, Zubair (2023): "AutoFR: Automated Filter Rule Generation for Adblocking", in: Proceedings of the USENIX Security Symposium. (Link)] |
| Per-site LLM analysis capping n around 10,000 | Emerging, and provisional. 29 papers, 22 of them in the incomplete 2025–2026 slice. Directionally consistent with manual labelling; re-derive before relying on it. | this page, above |
Use in Publications
Every figure on this page comes from one script over one extraction of the corpus. The population, the queries and the unedited output are on how_many_sites.
The population
| Population | Papers |
|---|---|
| corpus (seven venues, 2010–2026) | 5,859 |
| drew ≥1 population whose unit is a website, domain or web page | 1,153 |
| … and stated a size for at least one of them — this page's population | 1,121 |
| … of which also ran a crawl | 674 |
| … sampled the web without crawling (DNS, certificates, archives, passive data) | 447 |
97.2% of web-sampling papers state a size, so the 32 that do not are a rounding error rather than a reporting gap — this is the one methodological field in the schema that the field almost always reports. What it does not report is why.
Sizing, in one table
| Question | Papers | Share of 1,121 |
|---|---|---|
| states a size | 1,121 | 100% (by construction) |
| the size is 1, 2 or 5 times a power of ten | 523 | 46.7% |
| the size is exactly 1,000,000 | 140 | 12.5% |
| mentions sample size, power or a margin of error in one sentence with a web unit and a number | 26 | 2.3% |
| derives the size from a precision or power requirement | 6 | 0.5% |
| … and the thing sized is the crawl rather than a subsample | 2 | 0.2% |
| reports the same measurement at more than one rank depth | 19 | 1.7% |
Power analysis, over the population that could use one
| Population | Papers | Share |
|---|---|---|
| ran statistical inference (not descriptives only) | 1,762 | — |
| … reports a power analysis | 87 | 4.9% |
| … … of which recruited human participants | 74 | 85.1% of the 87 |
| … … of which also ran a crawl | 6 | 6.9% of the 87 |
| … … which derive a size for a web population — 2 from power, 1 from a Chernoff bound | 3 | 3.4% of the 87 |
Methodology and limitations of these figures
- How they were produced. One structured record per paper was extracted from full text; each tuple carries a verbatim evidence quote and its section, so every figure traces to a sentence. Sizes come from
population[].non tuples whoseunitiswebsites,domainsorweb-pages— the same web-unit definition Sampling uses, narrowed to papers that state a size. - Medians are rounded. Where a group has an even number of papers the median is the mean of the two central values; 26,590 and 42,542 in the year table are 26,589.5 and 42,541.5, and 13,796 is 13,795.5. The script prints the unrounded values.
- Paper-level, largest population. Except where a table says otherwise, one number per paper: the largest web population it states. Tuple-level distributions over the same field are on Sampling and are numerically different for that reason, not because either is wrong.
- The arithmetic is arithmetic. Sections A, B and C contain no measurement of the corpus. They are closed-form binomial and design-effect formulas, printed by the same script so they can be checked, and the ρ values in section C are illustrative — no measured intra-class correlation for a web-measurement outcome exists.
- Probe counts are not findings, and probe recall is a floor. The 26 sizing sentences and the 28 rank-band sentences were each read one by one and classified by hand; the verdicts, the residue and the full sentences are on the provenance page. Both probes returned substantially fewer hits before two bugs were fixed — a sentence splitter that cut at decimal points, and a rank pattern that required a digit after “top” — so the derived counts (6 and 19) are lower bounds on what the corpus contains, not exact censuses. Three of the quoted sentences are spliced by two-column reading order and are flagged on the provenance page.
classification.methodis a mid-band field (58% run-to-run agreement on the previous extraction run, not re-measured since) and multi-valued, so the medians by method are a ranking.population.nandpopulation.unitare among the stable fields.- Silence is not absence. “Does not derive its n” means the paper does not say so in a sentence the probe can see. A paper may have sized its crawl carefully and reported only the number.
- Venue coverage. Seven venues — CCS, IMC, NDSS, PoPETs, USENIX Security, TheWebConf, IEEE S&P. EuroS&P, ACSAC, RAID, AsiaCCS, WPES, CHI and SOUPS are absent, and SOUPS matters here because it is where careful sizing is most likely to appear. Every claim is a claim about those seven.
- 2025 and 2026 are provisional, labelled in every table where they appear. Trends should be read as ending in 2024.
- Every query, the report script and its unedited output are on how_many_sites; corpus-level caveats are on corpus.
What to Report
A sizing sentence a reviewer can accept names, in this order:
- Which of the three questions n is for — a prevalence, an existence or rate claim about something rare, or a comparison. One clause. It determines whether your number is generous or inadequate, and the reader cannot infer it.
- The n you drew and the n you analysed, with the reason for the gap. The second is the sample size; the first is a purchase order.
- For a prevalence: the interval, and one sentence saying which error dominates it. If the labeller's error exceeds the sampling error, say so — it is a stronger paper, not a weaker one.
- For a rare thing: the prevalence and the expected count, so the reader can see the design was capable of the finding. If you found none, the 3/n bound rather than the zero.
- For a comparison: what your independent unit is, and how many of them there are. Not how many sites.
- Whether the estimate varies with rank depth, from at least two depths of your own crawl if the mechanism suggests it might. Four rank bands of 1,000 sites cost less than one crawl of 10,000 and pre-empt the reviewer question.
- The size of any hand-checked subsample and how it was chosen, with the interval it supports. This is the sample-size calculation your paper most likely needs.
Open Questions
- The per-rank-band table still barely exists. Nineteen of 1,121 papers (1.7%) report the same measurement at more than one list depth, all of them incidentally, and none as its contribution. A paper whose whole result is “here is prevalence of X at ranks 1–1k, 1k–10k, 10k–100k and 100k–1M, for twenty X” would let every top-n paper in the field state what its cut costs. It is a by-product of any stratified crawl.
- No measured intra-class correlation exists for any web outcome. Section C is therefore arithmetic on assumed ρ. The same gap is open on the neighbouring page for the same reason, and one small study — the ICC of “sets a tracking cookie before consent”, by consent platform, by CMS and by hosting provider — would turn both pages' illustrative tables into real ones.
- Nobody has measured where the field's own curves flatten. McDonald et al. [29McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] did it for one phenomenon. A study that re-ran a handful of standard measurements at 1k, 10k, 100k and 1M and reported when the estimate stopped moving would answer “how many sites” empirically, which is how it should have been answered all along.
- Cost is not in any structured record. The corpus cannot say what a crawl of n sites costs in wall-clock, bandwidth or money, because papers do not report it and the extraction has no field for it. That absence is why the “what extra sites cost” section above argues from mechanism rather than from measurement.
- Do reviewers ever ask? Our impression from reading is that a small n draws a reviewer objection and a large one does not, regardless of which question is being asked. That asymmetry, if real, is the mechanism keeping the convention in place, and it is not measurable from published papers.
Related Pages
- Sampling — how you draw: method, unit, versioning, and rank weighting as the alternative to a larger n. Read together with this page; that one owns everything about the draw except the number.
- Website selection — which list, and what each ranking measures. A frame that is more accurate substitutes for a sample that is bigger.
- Hypothesis testing — what non-independence does to the test, once the rows exist. Section C here is the sizing consequence of that page's central error.
- Biases — coverage, selection and vantage bias: the errors that do not shrink with n, with the sizes the field has measured for each.
- Website classification — the labeller whose error rate caps what n can buy you.
- Longitudinal — a repeated crawl trades sites for time points, which is the same budget spent on a different axis.
- Crawler — attrition, timeouts and bot walls, which decide the gap between the n you drew and the n you analysed.
- Study preregistration — where a sizing decision is deposited before the data, if you want it to be credible.
References
- [1]
- Liao, Xiaojing; Liu, Chang; McCoy, Damon; Shi, Elaine; Hao, Shuang; Beyah, Raheem A. (2016): "Characterizing Long-tail SEO Spam on Cloud Web Hosting Services", in: Proceedings of the ACM Web Conference. (DOI)
- [2]
- Sy, Erik; Burkert, Christian; Federrath, Hannes; Fischer, Mathias (2019): "A QUIC Look at Web Tracking", Proceedings on Privacy Enhancing Technologies 2019(3):255-266. (DOI)
- [3]
- Oest, Adam; Safaei, Yeganeh; Doupé, Adam; Ahn, Gail-Joon; Wardman, Brad; Tyers, Kevin (2019): "PhishFarm: A Scalable Framework for Measuring the Effectiveness of Evasion Techniques against Browser Phishing Blacklists", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [4]
- Oest, Adam; Safaei, Yeganeh; Zhang, Penghui; Wardman, Brad; Tyers, Kevin; Shoshitaishvili, Yan; Doupé, Adam; Ahn, Gail-Joon (2020): "PhishTime: Continuous Longitudinal Measurement of the Effectiveness of Anti-phishing Blacklists", in: Proceedings of the USENIX Security Symposium. (Link)
- [5]
- Konoth, Radhesh Krishnan; Vineti, Emanuele; Moonsamy, Veelasha; Lindorfer, Martina; Kruegel, Christopher; Bos, Herbert; Vigna, Giovanni (2018): "MineSweeper: An In-depth Look into Drive-by Cryptocurrency Mining and Its Defense", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [6]
- Hanley, James A.; Lippman-Hand, Abby (1983): "If Nothing Goes Wrong, Is Everything All Right? Interpreting Zero Numerators", JAMA 249(13):1743-1745. (DOI)
- [7]
- Scheitle, Quirin; Hohlfeld, Oliver; Gamba, Julien; Jelten, Jonas; Zimmermann, Torsten; Strowes, Stephen D.; Vallina-Rodriguez, Narseo (2018): "A Long Way to the Top: Significance, Structure, and Stability of Internet Top Lists", in: Proceedings of the Internet Measurement Conference 2018, pp. 478–493. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
- [8]
- Ruth, Kimberly; Kumar, Deepak; Wang, Brandon; Valenta, Luke; Durumeric, Zakir (2022): "Toppling Top Lists: Evaluating the Accuracy of Popular Website Lists", in: Proceedings of the 22nd ACM Internet Measurement Conference, pp. 374–387. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
- [9]
- Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
- [10]
- Demir, Nurullah; Hörnemann, Jan; Große-Kampmann, Matteo; Urban, Tobias; Pohlmann, Norbert; Holz, Thorsten; Wressnegger, Christian (2023): "On the Similarity of Web Measurements Under Different Experimental Setups", in: Proceedings of the ACM Internet Measurement Conference, pp. 356-369. ACM DOI 10.1145/3618257.3624795 is listed by DBLP but was not registered with the DOI resolver as of 2026-08-12; the DOI above resolves to the authors' institutional record of the same paper (DOI)
- [11]
- Bijmans, Hugo L.J.; Booij, Tim M.; Doerr, Christian (2019): "Inadvertently Making Cyber Criminals Rich: A Comprehensive Study of Cryptojacking Campaigns at Internet Scale", in: Proceedings of the USENIX Security Symposium. (Link)
- [12]
- Murley, Paul; Ma, Zane; Mason, Joshua; Bailey, Michael D.; Kharraz, Amin (2021): "WebSocket Adoption and the Landscape of the Real-Time Web", in: Proceedings of the ACM Web Conference. (DOI)
- [13]
- Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [14]
- Killip, Shersten; Mahfoud, Ziyad; Pearce, Kevin (2004): "What Is an Intracluster Correlation Coefficient? Crucial Concepts for Primary Care Researchers", The Annals of Family Medicine 2(3):204-208. (DOI)
- [15]
- Poteat, Tara; Li, Frank (2021): "Who You Gonna Call? An Empirical Evaluation of Website security.txt Deployment", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [16]
- Al Roomi, Suood; Li, Frank (2023): "A Large-Scale Measurement of Website Login Policies", in: Proceedings of the USENIX Security Symposium. (Link)
- [17]
- Ikram, Muhammad; Masood, Rahat; Tyson, Gareth; Kaafar, Mohamed Ali; Loizon, Noha; Ensafi, Roya (2019): "The Chain of Implicit Trust: An Analysis of the Web Third-party Resources Loading", in: Proceedings of the ACM Web Conference. (DOI)
- [18]
- Hantke, Florian; Calzavara, Stefano; Wilhelm, Moritz; Rabitti, Alvise; Stock, Ben (2023): "You Call This Archaeology? Evaluating Web Archives for Reproducible Web Security Measurements", in: Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, pp. 3168–3182. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
- [19]
- Aas, Josh; Barnes, Richard; Case, Benton; Durumeric, Zakir; Eckersley, Peter; Flores-López, Alan; Halderman, J. Alex; Hoffman-Andrews, Jacob; Kasten, James; Rescorla, Eric; Schoen, Seth D.; Warren, Brad (2019): "Let's Encrypt: An Automated Certificate Authority to Encrypt the Entire Web", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [20]
- Ashiq, Md. Ishtiaq; Li, Weitong; Fiebig, Tobias; Chung, Taejoong (2023): "You've Got Report: Measurement and Security Implications of DMARC Reporting", in: Proceedings of the USENIX Security Symposium. (Link)
- [21]
- Kashaf, Aqsa; Sekar, Vyas; Agarwal, Yuvraj (2020): "Analyzing Third Party Service Dependencies in Modern Web Services: Have We Learned from the Mirai-Dyn Incident?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [22]
- Singanamalla, Sudheesh; Jang, Esther Han Beol; Anderson, Richard; Kohno, Tadayoshi; Heimerl, Kurtis (2020): "Accept the Risk and Continue: Measuring the Long Tail of Government https Adoption", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [23]
- Bhuiyan, Masudul Hasan Masud; Varvello, Matteo; Staicu, Cristian-Alexandru; Zaki, Yasir (2025): "Digital Disparities: A Comparative Web Measurement Study Across Economic Boundaries", in: Proceedings of the ACM Web Conference. (DOI)
- [24]
- Bhaskar, Abhishek; Pearce, Paul (2022): "Many Roads Lead To Rome: How Packet Headers Influence DNS Censorship Measurement", in: Proceedings of the USENIX Security Symposium. (Link)
- [25]
- Amjad, Abdul Haddi; Shafiq, Zubair; Gulzar, Muhammad Ali (2023): "Blocking JavaScript Without Breaking the Web: An Empirical Investigation", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [26]
- Le, Hieu; Elmalaki, Salma; Markopoulou, Athina; Shafiq, Zubair (2023): "AutoFR: Automated Filter Rule Generation for Adblocking", in: Proceedings of the USENIX Security Symposium. (Link)
- [27]
- Moore, Tyler; Leontiadis, Nektarios; Christin, Nicolas (2011): "Fashion Crimes: Trending-Term Exploitation on the Web", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [28]
- Chang, Li; Hsiao, Hsu-Chun; Jeng, Wei; Kim, Tiffany Hyun-Jin; Lin, Wei-Hsi (2017): "Security Implications of Redirection Trail in Popular Websites Worldwide", in: Proceedings of the ACM Web Conference. (DOI)
- [29]
- McDonald, Allison; Bernhard, Matthew; Valenta, Luke; VanderSloot, Benjamin; Scott, Will; Sullivan, Nick; Halderman, J. Alex; Ensafi, Roya (2018): "403 Forbidden: A Global View of CDN Geoblocking", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [30]
- Le Pochat, Victor; Van Goethem, Tom; Tajalizadehkhoob, Samaneh; Korczy´nski, Maciej; Joosen, Wouter (2019): "Tranco: A Research-Oriented Top Sites Ranking Hardened Against Manipulation", in: Proceedings of the 26th Annual Network and Distributed System Security Symposium. (DOI)
- [31]
- Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
- [32]
- Ruth, Kimberly; Fass, Aurore; Azose, Jonathan; Pearson, Mark; Thomas, Emma; Sadowski, Caitlin; Durumeric, Zakir (2022): "A World Wide View of Browsing the World Wide Web", in: Proceedings of the ACM Internet Measurement Conference, pp. 317-336. (DOI)
- [33]
- Tang, Jenny; Bauer, Lujo; Christin, Nicolas (2025): "Misuse, Misreporting, Misinterpretation of Statistical Methods in Usable Privacy and Security Papers", in: Proceedings of the Twenty-First Symposium on Usable Privacy and Security (SOUPS), pp. 475-493. USENIX Association. (Link)
., so 0.4% of the top thousand was split in half at the decimal point, and the rank-band pattern required a digit after top, so magnitudes written in words were invisible. Both are fixed in the published script, which is why the count here is higher than a reader of the earlier draft would expect — and why 19 is a floor, not a census. See provenance.report_sampling.mjs and are over that page's population of 1,153 papers with a web-unit population — slightly wider than this page's 1,121, which additionally requires a stated size. Bucket denominators are 322, 350 and 190 papers.