This is an old revision of the document!
Table of Contents
Provenance: Statistics — How Many Sites
Working log behind How many sites. Every query with its denominator, the report script and its unedited output, the two full-text probes and the bugs found in them, the hand classification of every probe hit, the quotes checked, the external sources verified and rejected, and what could not be established.
Corpus-level caveats — the venue scope, the selection funnel, the provisional 2025–2026 years — are on Corpus and are not restated here. Citation keys are shared with the content page and the single Bibliography; this page adds no bibliography entries of its own and carries no discussion block, following the convention set by hypothesis_testing that comments belong on the content page.
Voice here is a working log, not prose. It is for someone checking a number.
The run
| Item | Value |
|---|---|
| Date | 2026-09-02 |
| Corpus | data/extract/run1, 5,859 extracted papers, 7 venues, 2010–2026 |
| Page status | New page. statistics:how_many_sites was not a red link — nothing on the wiki promised it. It was commissioned because the choice of n had no page whose job it was |
| Scripts added | scripts/report_how_many_sites.mjs, scripts/hms_quotecheck.py, scripts/_mkprov_hms.py |
| Code published on the content page | none — the arithmetic tables are printed by the report script rather than embedded as a runnable file |
| Bibliography entries added | 10 |
| Models | Opus 5 (author, all queries, all probes, all hand classification, all quote checking); the review layer below |
| Neighbour edits | a cross-link box in Sampling's How Big section; a row in the Statistics hub table, with its “all 6 children” and page count moved to 7; a line in start |
| Publish order | Bibliography first (so no marker is ever red), then the content page, then the three neighbour edits, then this page last — so that every claim it makes about what was published is already true when it is saved |
Scope decision: create, not broaden
Two neighbours could have absorbed this material and both were read in full before the decision.
- Sampling (52 KB) already has a section titled How Big, and Why That Number and a Sample sizes are round numbers subsection. Broadening it was the serious alternative and was rejected on three grounds: (1) its subject is the draw — frame, method, unit, versioning — and n is orthogonal to all four, so the section sits there as a guest; (2) it is already one of the longest pages on the wiki and the material below adds roughly 45 KB; (3) the rank-tail, design-effect and rare-event arguments are statistical rather than design arguments and belong in the
statistics:namespace next to Hypothesis testing, whose central error — non-independence — has a direct sizing consequence that neither page previously stated. - Hypothesis testing owns non-independence and its effect on a p-value. It does not own, and should not own, the question of what n to pick before any test exists. Section C of the new page is explicitly the sizing corollary of that page's simulation and links to it rather than restating it.
Consequence, and the drift risk it creates. design:sampling's How Big section stays where it is; it was not carved out. That is deliberate — removing it would leave literal numbers stranded in the parts of that page that were not moved — but it means two pages now carry size figures. They are numerically different on purpose and each says so:
| Figure | design:sampling | statistics:how_many_sites |
|---|---|---|
| Unit of count | tuple (one row per population drawn) | paper (largest web population per paper) |
| Denominator | 2,379 web populations with a stated size | 1,121 papers with at least one such population |
| Median | 6,755 | 20,000 |
| Round-number share | 43.5% | 46.7% |
Both were re-derived on 2026-09-02 against the current run1, and design:sampling's published figures were confirmed unchanged, so the divergence is definitional and not drift. A tip box on the new page states this so a reader who meets both does not conclude one is stale.
Population and denominators
Every figure on the content page traces to one of these. Nothing on the page uses “of 5,859 papers”.
| Population | Definition | Papers |
|---|---|---|
| corpus | all extracted records | 5,859 |
| web-sampling | at least one population[] tuple whose unit is websites, domains or web-pages | 1,153 |
| page population | web-sampling and at least one of those tuples has a numeric n greater than 0 | 1,121 |
| … also crawled | crawlConfig is not null, or studyTypes includes automated-web-crawl | 674 |
| … web without a crawl | the complement (DNS, certificates, archives, passive data) | 447 |
| inferential | at least one statistics[] tuple whose kind is not descriptive-only | 1,762 |
The web-unit definition is copied from scripts/report_sampling.mjs on purpose, so the two pages' populations nest rather than cross. The narrowing from 1,153 to 1,121 discards 32 papers (2.8%) that name a web population and never give its size; they are this page's population residue and are counted in section 11 of the report output.
Paper-level, largest population. A paper that draws a 1,000,000-site crawl and a 300-site hand-check contributes 1,000,000. That is the number that reaches the abstract, and it is the number this page is about. The tuple-level view is Sampling's.
Queries, one per published figure
| Figure on the page | Where in the report output | Denominator |
|---|---|---|
| 1,121 / 674 / 447 | 1 | 1,153 web-sampling papers |
| modal-value table, 46.7% round, 12.5% at exactly 1M | 2 | 1,121 |
| median 20,000, quartiles 1,000 and 1,000,000 | 2 | 1,121 |
| year-bucket medians and the 1M-and-above shares | 2 | 1,121, split by bucket |
| crawling-only medians 10,000 through 20,000 | 2 | 674 |
| n by per-site analysis method | 3 | 1,121, one row per classification.method value |
| LLM row: 29 papers, 22 in 2025–2026 | 3 | 1,121, split by bucket |
| 87 power-analysis papers, 74 with participants, 6 crawling, 3 sizing a web population | 4 | 1,762 inferential |
| the nine power-analysis papers inside this page's population, printed in full | 4 | 1,121 |
| 26 sizing sentences, split 6 / 5 / 1 / 14 | 5 | 1,121 |
| precision table, and n-for-a-half-width table | 6 | arithmetic, no corpus population |
| rare-event table, rule-of-three table | 7 | arithmetic, no corpus population |
| design-effect and effective-n tables | 8 | arithmetic, no corpus population |
| 28 rank-band sentences, 19 per-band | 9 | 1,121 |
| the five rare-phenomenon rows in the page's table | 10, plus hand selection | 368 papers with n at or above 200,000 |
| 32 papers excluded; prevalence-string residue | 11 | 1,153 papers / 6,203 tuples |
Two figures on the page come from outside this report script and are flagged where they appear. Both come from scripts/report_sampling.mjs, re-run on 2026-09-02.
- The 2,379 / 6,755 / 43.5% comparison row above. Re-run and compared against the published
design:samplingfigures: identical. - The rank-stratified sampling shares in the currency table. Re-run: 1.0% (2010–2013), 4.8% (2014–2017), 5.0% (2018–2021), 8.9% (2022–2024), 6.3% (2025–2026, provisional), and 6.0% over the whole 1,153-paper web-sampling population (69 papers). That population is
report_sampling.mjs's, which is web-unit papers without the stated-size requirement, so it is 1,153 rather than this page's 1,121 — the currency table now says so in a footnote rather than leaving “the whole population” ambiguous between the two. A first draft of the currency table wrote this as “6–9% since 2018”, which is wrong at the low end — 2018–2021 is 5.0%, not 6%. It was caught by the figures reviewer re-running the script rather than trusting the prose, and it is recorded here because the same draft also claimed the figure had been “diffed with no change”, which was itself untrue. The lesson is the general one: a figure copied from a neighbouring page into a summary range is not a figure this page's script can check, and every such figure needs its own re-run.
Full-text probes: what was run, what broke, and what was rejected
Both probes read data/fulltext/<year>/<venue>/<slug>/paper.cols.txt — the repaired column reading order — never paper.norm.txt. Both collapse whitespace before matching, because a PDF line break inside a phrase silently loses the match.
Two bugs, both found by checking a paper the probe should have matched and did not
This is the part worth reading, because both bugs made a published count too low and both survived inspection.
- The sentence splitter cut at decimal points. The splitter treats the full stop in
0.4%as a sentence end. Murley et al.'s “We find them in use on only 0.4% of the top thousand and 0.05% of websites in the top million” — a textbook per-band result — was split into fragments none of which contained both bands. Fixed by masking digit-dot-digit, plus a short abbreviation list, with a placeholder byte before splitting and restoring it before printing. Effect: the sizing probe went from 24 hits to 26, and the rank-band probe from 15 to 20. - The rank-band pattern required a digit after “top”. It matched top 1K but not top thousand, top million or larger Top Alexa lists. Widened to include spelled-out magnitudes. Effect: the rank-band probe went from 20 hits to 28, and the per-band count from 10 to 19.
The published counts are therefore floors. A third under-count is known and not fixed: a sentence spliced across two columns can lose the join, which is how the sizing probe misses Liao et al.'s Chernoff calculation entirely — that paper was found through the structured statistics.kind = power-analysis field instead. The page reports 6 derived sizes from the probe and names Liao as a seventh found by another route.
Probe A — sizing sentences
One sentence containing all three of: sizing language, a web unit, and a number.
sizing language: sample size(s|d) | statistical power | power analysis/analyses |
powered to detect | margin of error | chernoff | hoeffding |
confidence interval | diminishing returns | law of large numbers
web unit: web site(s) | site(s) | domain(s) | web page(s) | page(s) | url(s) | host(s)
number: three or more digits with optional commas, or a decimal
followed by k / K / thousand / million / m / M
26 papers matched. All 26 matching sentences were read one by one and given one of four verdicts. The verdict is on the sentence the probe matched, not on the paper: a paper filed not-sizing may have sized its population carefully somewhere the probe cannot see. The verdicts are which live in a literal table inside the script so the script prints the judgement rather than hiding it. The script fails loudly with !! unread if a hit has no verdict, which is how the three new hits after the decimal-point fix were caught.
| Verdict | Papers | What it means |
|---|---|---|
derives-n | 6 | the size of a web population is derived from a stated precision, power or confidence requirement |
design-reason | 5 | the size is argued for, from a design or cost argument, without a calculation |
states-only | 1 | names the size and gives no argument at all |
not-sizing | 14 | the sentence is about the statistics of the analysis, a scale boast, a figure caption, or a caveat that a group was too small |
Every hit, its verdict and its full sentence are in the –sizing output reproduced below.
Probe B — rank-band sentences
One sentence containing a rank band, a phrase comparing across rank, and a percentage. 28 papers matched; all 28 were read and classified per-band (19) or not-rank (9). The nine rejected are: a long tail of vulnerable applications, of WHOIS service names, of FQDNs, of third-party domains to fix, and of URL popularity; popularity bands of browser extensions rather than websites; a figure legend whose comparison is not in the sentence; a figure fragment about the top-100 most informative features; and Ruth et al.'s measurement of list accuracy by rank bucket, which is a result about the list rather than a phenomenon measured per band.
The content page tables 12 of the 19, in 13 rows (Al Roomi and Li appear twice, in both directions), plus one paper the probe misses: [1Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], whose comparison is between a ranking and a random sample of the whole zone rather than between two depths of one ranking. It is included because a genuinely flat result is the hardest of the three families to find, and it was surfaced by the rare-event query in section 10 instead.
The seven per-band papers not tabled, all in section 9 of the report output with their sentences:
| Paper | Why not tabled |
|---|---|
| IMC/2015 Neither Snow Nor Rain Nor MITM | gives tail figures (82% TLS, 35% correctly configured across 700,000 long-tail SMTP servers) with no matching head figure in the sentence |
| USENIX/2018 O Single Sign-Off, Where Art Thou? | one figure — 10.8% SSO coverage in the top 100K — plus a direction |
| IMC/2022 A World Wide View of Browsing the World Wide Web | category mix across the ranking; directional, and the bands are in a figure |
| CCS/2025 PIIXEL Leaks | one figure, quoted for the 500K–1M band only |
| WWW/2025 Harmful Terms and Where to Find Them | one figure (42.06% of Tranco top-100K shopping sites) plus a direction |
| PETS/2026 Overcoming Language Barriers | has full per-rank-group data, but in a table the sentence probe cannot read |
| USENIX/2026 The State of Passkeys | direction stated in prose, figures in a plot |
Probes that were run and rejected
| Probe | Result | Why it is not on the page |
|---|---|---|
| attrition-by-rank: failure or unreachability language near a rank band | 4 papers, of which 3 are table fragments about parked domains and 1 is a footnote (“18% of Alexa top 100,000 websites were unreachable”) | Too thin to carry the claim “the tail is harder to reach”, which the page therefore argues from mechanism instead. Sampling owns attrition as a denominator problem |
detection[].prevalence percentages cross-tabulated against n | 1,008 papers parsed; the median paper-minimum prevalence fell from 13.67% at n at or below 1k to 5.67% at n at or above 1M, and the share reporting anything under 1% rose from 12.3% to 22.7% | Dropped. Reading 12 of the 1M-and-above rows showed the “minimum percentage” is frequently a false-negative rate, a validation-set share or a classifier metric rather than a site prevalence. The signal is real but the measure is contaminated; publishing it would have been a rate of nothing. What survives is the hand-picked example table in the page's section B, five rows each read in the source |
| “representative sample” claims | 93 papers, not pursued | It is a claim about generalisability, not about size. Biases is the page for it |
Hand-selected examples: the five rare-phenomenon rows
The page's rare-event table is a selection, not an aggregate, and the selection rule is written down here. From section 10 of the report output — prevalence strings naming a web unit, containing a percentage below 0.5%, in papers with n at or above 200,000 — keep rows where (a) the phenomenon is a property of a site or domain rather than of a classifier, (b) the paper's own n is stated, and © the source sentence checks out verbatim. Twenty-one tuples across 21 papers matched the mechanical filter; five were kept.
| Kept | Rate | Drawn n | Analysed n, from the paper |
|---|---|---|---|
| [2Konoth, Radhesh Krishnan; Vineti, Emanuele; Moonsamy, Veelasha; Lindorfer, Martina; Kruegel, Christopher; Bos, Herbert; Vigna, Giovanni (2018): "MineSweeper: An In-depth Look into Drive-by Cryptocurrency Mining and Its Defense", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] drive-by cryptomining | 1,735 sites, 0.18% | Alexa top 1,000,000 | 991,513 — “we managed to crawl 991,513 of them” |
| [3Bijmans, Hugo L.J.; Booij, Tim M.; Doerr, Christian (2019): "Inadvertently Making Cyber Criminals Rich: A Comprehensive Study of Cryptojacking Campaigns at Internet Scale", in: Proceedings of the USENIX Security Symposium. (Link)] cryptojacking | 5,190 domains, 0.011% | ~20% of the Internet | 48,948,669 — the table total |
| [4Sy, Erik; Burkert, Christian; Federrath, Hannes; Fischer, Mathias (2019): "A QUIC Look at Web Tracking", Proceedings on Privacy Enhancing Technologies 2019(3):255-266. (DOI)] QUIC support | 186 sites, 0.02% | Alexa top 1,000,000 | not reported; the paper probes “each Alexa-listed host” and states no attrition |
| [5Murley, Paul; Ma, Zane; Mason, Joshua; Bailey, Michael D.; Kharraz, Amin (2021): "WebSocket Adoption and the Landscape of the Real-Time Web", in: Proceedings of the ACM Web Conference. (DOI)] Server-Sent Events | 0.05% of the top million | Tranco top 1,000,000 | ~881,000 — “data for a total of 88.1% of websites in the top million” |
| [1Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] non-secure DNS dynamic updates | 1,877 domains, 0.065% | a 1% draw from 286,788,250 | 2,865,393 — “we randomly sampled 1%”, Table 1 gives 2,865,393 |
Rejected from the same 21 for failing (a): HPKP validity, CAA tag misconfiguration, DKIM l= tags, QUIC ECN validation, and several false-positive-rate rows. All 21 are printed by the report script under –rare.
The analysed-n column is a correction, and the largest single defect this run produced. The first draft of the table used population[].n from the structured extraction as “the n the paper needed”, which is the drawn n. The citations reviewer read the papers and found three wrong:
| Paper | First draft | Correct | Size of the error |
|---|---|---|---|
| [2Konoth, Radhesh Krishnan; Vineti, Emanuele; Moonsamy, Veelasha; Lindorfer, Martina; Kruegel, Christopher; Bos, Herbert; Vigna, Giovanni (2018): "MineSweeper: An In-depth Look into Drive-by Cryptocurrency Mining and Its Defense", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] | 1,000,000 | 991,513 | 0.9% — cosmetic, but it is the drawn number |
| [5Murley, Paul; Ma, Zane; Mason, Joshua; Bailey, Michael D.; Kharraz, Amin (2021): "WebSocket Adoption and the Landscape of the Real-Time Web", in: Proceedings of the ACM Web Conference. (DOI)] | 1,000,000 | ~881,000 | 12%, and the paper reports it plainly |
| [1Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | 286,788,250 | 2,865,393 | 100×. population[].n had recorded the corpus the study drew a 1% sample from; the 286-million set was never tested for the vulnerability |
None of these changes a rate or an expected-hit count, because the rates are the papers' own. All three change what n the reader is told the rate rests on, on a page whose central instruction is to report the analysed n and not the drawn one. population[].n is the drawn size and cannot be assumed to be the analysed size; the six new quote-check entries added afterwards pin each analysed figure to a sentence.
Quote checking
scripts/hms_quotecheck.py checks every quotation the content page prints against the corpus text the extractor read. Exact match after whitespace collapse and quote/dash normalisation; on failure the phrase is re-tried as overlapping five-word windows and the share found is printed, so a two-column splice shows as PARTIAL rather than as an absence.
42 entries, 41 distinct, 0 not located at all, 2 PARTIAL, and 0 page quotations uncovered by the audit. Both partials are two-column splices and are handled on the page rather than hidden:
| Quote | Result | What the page does |
|---|---|---|
| Liao et al., “to obtain the number of sampled cloud directories n = 500” | PARTIAL, 4 of 7 windows | The source sentence is interleaved with an adjacent one by column order. The page paraphrases the calculation and footnotes the splice |
| Poteat and Li, “This percentage decreases to 3-4% for the top 10K sites, and only a percent for the top 100K” | PARTIAL, 8 of 14 windows | Same cause. The page quotes only the intact run “a higher adoption rate for higher-ranked” and gives the figures as numbers |
Two coverage gaps in this script were found by the citations reviewer and are fixed. First, the two footnoted quotes from [6Moore, Tyler; Leontiadis, Nektarios; Christin, Nicolas (2011): "Fashion Crimes: Trending-Term Exploitation on the Web", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] and [7Chang, Li; Hsiao, Hsu-Chun; Jeng, Wei; Kim, Tiffany Hyun-Jin; Lin, Wei-Hsi (2017): "Security Implications of Redirection Trail in Popular Websites Worldwide", in: Proceedings of the ACM Web Conference. (DOI)] were not in the check list at all, so “every quotation the page prints” was not true when it was first written; both are now in it. Second, the normaliser collapsed whitespace but did not rejoin a word broken across a line by a hyphen. A de-hyphenating second pass was added, reported separately as EXACT-DH so it can never silently paper over a real absence.
Two honest qualifications on that second fix, both raised by the figures reviewer on a re-check. Chang et al.'s “ran-domized sample” was the example that motivated it, and the de-hyphenation pass does not rescue that quote: the real damage there is a two-column splice with an unrelated clause sitting between ran- and domized, which no adjacent-word rejoin can bridge. The page quotes the intact run starting after the hyphen instead. And no quote in the current list exercises the pass at all — the output contains zero EXACT-DH lines. It is a defence against a failure mode that has not yet occurred here, kept because it is cheap and because the failure mode is real in this corpus; it is not evidence that anything was fixed.
Third, and structurally the important one: curating the list by hand failed three times in a row, so it is no longer curated on trust. The script now audits itself against the content page: it extracts every //"..."// span from pages/statistics_how_many_sites.txt, requires each to be covered by a QUOTES entry, prints any that are not, and exits non-zero. That guard caught two further quotes while this page was being written — the McDonald “500 random combinations” quote the citations reviewer found, and a PhishFarm quote added minutes earlier. The provenance page's own page-quote spans are not audited, because most of them quote this page's superseded drafts rather than any paper.
A fourth splice, Chang et al.'s “domized sample of 2,000 websites (margin of error = 2.”, is quoted with its damage visible rather than repaired: the leading ran and the trailing digits are in the other column, and inventing them would have been a fabrication.
A third splice, Singanamalla et al.'s HTTPS long-tail sentence, was caught before it was quoted: the naive reconstruction reads “valid https use ment index scores”, which is nonsense. The page quotes the two intact fragments and paraphrases the join, and the check file records both fragments separately with a comment explaining why.
Four quotes were additionally chased by hand during drafting, because the extraction record and the page disagreed about what the number meant:
- [3Bijmans, Hugo L.J.; Booij, Tim M.; Doerr, Christian (2019): "Inadvertently Making Cyber Criminals Rich: A Comprehensive Study of Cryptojacking Campaigns at Internet Scale", in: Proceedings of the USENIX Security Symposium. (Link)] — the extraction gave “0.011% of all domains, or one in 9,090 websites”. The corpus text says “meaning that one in every 9,090 websites is cryptojacking”, and the page uses the source wording. Reading the surrounding paragraph also produced the head/tail comparison used in the page's rank-tail discussion — 0.065% in the Alexa top 1M against 0.011% in a random 49M-domain sample, “almost 6 times lower” — which no probe had surfaced.
- [4Sy, Erik; Burkert, Christian; Federrath, Hannes; Fischer, Mathias (2019): "A QUIC Look at Web Tracking", Proceedings on Privacy Enhancing Technologies 2019(3):255-266. (DOI)] — the extraction gave “186 websites … 0.02% in the Top 1M”. Reading the source found the full six-row table the content page now leads with, and a second figure (0.0186%) the paper gives to three decimals elsewhere. Both are quoted.
- [1Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] — checked because a “flat across rank” result is the one this page most wants to be true. It is in the abstract, verbatim.
External sources: verified, and rejected
Three references on the content page are outside the seven-venue corpus. Two are the external methodological citations below, both fetched on 2026-09-02 rather than recalled. The third is [9Tang, Jenny; Bauer, Lujo; Christin, Nicolas (2025): "Misuse, Misreporting, Misinterpretation of Statistical Methods in Usable Privacy and Security Papers", in: Proceedings of the Twenty-First Symposium on Usable Privacy and Security (SOUPS), pp. 475-493. USENIX Association. (Link)], which is SOUPS 2025 — a venue this corpus does not cover, and the one venue whose absence matters most for a page about statistical practice; it is cited from Hypothesis testing, which introduced the key.
| Source | Verification |
|---|---|
| [10Hanley, James A.; Lippman-Hand, Abby (1983): "If Nothing Goes Wrong, Is Everything All Right? Interpreting Zero Numerators", JAMA 249(13):1743-1745. (DOI)] Hanley and Lippman-Hand, If Nothing Goes Wrong, Is Everything All Right?, JAMA 249(13):1743–1745, 1983 | DOI 10.1001/jama.1983.03330370053031 resolved through the Crossref API: title, journal, volume 249, issue 13, first page 1743 and year 1983 all confirmed. Crossref lists only the first author, so the second was taken from PubMed 6827763, which gives Hanley JA, Lippman-Hand A and pages 1743–5 |
| [11Killip, Shersten; Mahfoud, Ziyad; Pearce, Kevin (2004): "What Is an Intracluster Correlation Coefficient? Crucial Concepts for Primary Care Researchers", The Annals of Family Medicine 2(3):204-208. (DOI)] Killip, Mahfoud and Pearce, What Is an Intracluster Correlation Coefficient?, Ann Fam Med 2(3):204–208, 2004 | DOI 10.1370/afm.141 resolved through Crossref: title, journal, volume 2, issue 3, pages 204–208 and year 2004 confirmed. Crossref again lists one author; the full list came from PubMed 15209195 |
Rejected. No industry or vendor source was used and none was sought. Sample-size choice is a methodological question with no vendor stake, and the one category of external material that would have been tempting — “how many sites should I crawl” blog posts and SEO tooling pages — is exactly the listicle failure mode the site's conventions warn about. No general statistics textbook is cited either, under the no-textbook rule: the three formulas the page uses are written out in the report script where they can be checked, and the two external citations above are each the primary source for one specific named rule rather than background reading.
Not established, and therefore not written: which software this field uses for a power calculation. Six papers is too few to say, and none of them names a tool.
Bibliography entries added
Ten, appended before the closing </bibtex> tag of Bibliography after a key and DOI and URL collision scan against a fresh export of the live page, not against a local copy. The scan caught seven would-be duplicates that a key-string check alone would have missed or mis-keyed:
| Wanted | Already present as | Caught by |
|---|---|---|
hantke2023_call | hantke2023you | DOI 10.1145/3576915.3616688 |
roomi2023_website | alroomi2023_login | landing URL |
bijmans2019_inadvertently | same key | key |
amjad2023_blocking | same key | key |
bhaskar2022_many | same key | key |
mcdonald2018_forbidden | same key | key |
chang2017_security | same key | key |
Added: sy2019_quic, korczynski2016_zone, ashiq2023_report, liao2016_characterizing, bhuiyan2025_digital, ikram2019_chain, singanamalla2020_accept, moore2011_fashion, hanley1983_nothing, killip2004_intracluster. PoPETs and USENIX index records carry no authors, so sy2019_quic and ashiq2023_report had their author lists fetched from the PoPETs and USENIX landing pages with curl and a browser User-Agent; bibgen.mjs derived the rest from OpenAlex DOIs. No entry contains a literal ASCII at-sign in a field, which silently drops an entry from bibtex4dw. One pre-existing duplicate was noticed and left alone: lerner2016internet and lerner2016_internet are the same paper under two keys. The page cites the underscored form, which matches the file's dominant style.
Two bugs in the published artefacts themselves
Both were introduced while fixing something else, and neither would have been visible on the rendered page until it was too late.
- A literal closing <file> tag inside a published <file> block.
hms_quotecheck.pystrips DokuWiki code blocks out of the page text before auditing it, so its source contained the literal string for a closing tag — and this page publishes that source inside such a block. DokuWiki would have ended the block at that line and rendered the remaining ~1,000 lines of the file as page markup. Caught by counting opening against closing tags after every generator run rather than by reading. The script now assembles the tag name from two fragments, with a comment saying why. - The report script's power-analysis section printed only the six papers that also crawled, while the page states a negative about all nine in the population. A negative claim whose audit trail shows a subset is not audited. It now prints all nine.
A mis-attribution caught before publication
Worth recording because it is the failure mode the site's conventions warn about most, and because no automated check would have caught it. An early draft described two of the subsample-sizing papers in a footnote rather than citing them, and named them from the slug and from memory: “Liao et al., Fashion Crimes … CCS 2011” and “Chen et al., Security Implications of Redirection Trail … TheWebConf 2017”. Both author attributions were wrong. The bibliographic index gives:
| Paper | Draft said | Index says |
|---|---|---|
| CCS 2011, Fashion Crimes: Trending-Term Exploitation on the Web | Liao et al. | Moore, Leontiadis and Christin, DOI 10.1145/2046707.2046761 |
| TheWebConf 2017, Security Implications of Redirection Trail in Popular Websites Worldwide | Chen et al. | Chang, Hsiao, Jeng, Kim and Lin, DOI 10.1145/3038912.3052698 |
“Liao” had leaked across from [8Liao, Xiaojing; Liu, Chang; McCoy, Damon; Shi, Elaine; Hao, Shuang; Beyah, Raheem A. (2016): "Characterizing Long-tail SEO Spam on Cloud Web Hosting Services", in: Proceedings of the ACM Web Conference. (DOI)], a different paper by a different group in the same section. It was found by looking both slugs up in data/corpus2/.meta/ rather than by any check — and no check would have found it. The quote check did not evaluate these two quotes at all (they were not in its list until afterwards), and even if it had, it would have passed them: the quotes were correct and only the names were invented, and no quote check reads a name. Both papers are now cited properly — chang2017_security was already in the bibliography, moore2011_fashion was added — and their quotes are in the check file.
The lesson for the next run: a name written in prose rather than resolved through bibgen.mjs or .meta/ has no verification path at all. Do not describe a paper in a footnote to avoid adding a bibliography entry.
What could not be established
- Whether the plateau in n is a plateau in ambition or in cost. Median n stopped growing around 2018 while storage and automation got cheaper. The corpus records no cost, so the page states the plateau and does not explain it.
- The intra-class correlation of any web-measurement outcome. No paper reports one, which is why section C of the page is arithmetic on assumed ρ. This is the same gap Hypothesis testing records; one small study would close both.
- Whether reviewers penalise a small n. The page states this as an impression, explicitly labelled as one, in its Open Questions box. It is not measurable from published papers, and it was tempting to assert.
- Whether the 2025–2026 LLM row is a real constraint or a small-sample artefact. 29 papers, 22 of them in provisional venue-years. Flagged in the page text as needing re-derivation.
- The recall of both probes. Two bugs were found by looking; a third source of loss — column splices — is known and unfixed. There is no estimate of how many further per-band or sizing papers exist, so the page says “floor” wherever a probe-derived count appears.
Judgement calls
- Create rather than broaden
design:sampling. Argued above. A reasonable person could have added 15 KB to that page instead and avoided two pages carrying size figures. - Leave
design:sampling's How Big section in place. Carving it out would have been cleaner but would have left literal numbers stranded in the parts of that page that were not moved. A cross-link and an explicit reconciliation table were chosen instead. - Publish the design-effect table with assumed ρ. The alternative was to omit section C for want of a measured ICC. It is kept because the shape of the deflation is the point and is not controversial, and because the assumption is stated three times — in the section, in the methodology bullet and in the open question. The ρ values match the sweep already published on Hypothesis testing, so the two pages cannot be read as disagreeing.
- Drop the prevalence-versus-n cross-tabulation. It was the most quotable finding this run produced, and it did not survive reading its own inputs. Recorded above so the next run does not rediscover and publish it.
- Call power analysis for site counts “never established” rather than “superseded”. 87 of 1,762 inferential papers report one and 74 are user studies; there is no earlier period in which crawl papers did this and then stopped. The distinction matters because “superseded” would imply something replaced it, and nothing did.
- Report the sizing verdicts as four categories rather than as “justified” against “not justified”. The
states-onlyanddesign-reasonrows exist because collapsing them would have made the headline figure look better and the classification less checkable. - Do not present the rank-tail direction as a rule. The nineteen papers split three ways and one paper contains two directions. The page gives a mechanism test instead of a heuristic.
- Quote the arithmetic tables from the report script rather than embedding a runnable file on the page. Sibling pages embed Python with a <file> block. Here the formulas are three lines each and a published script would have been mostly print statements, so the tables are printed by the report script instead and the page says where they come from.
Report script and its unedited output
node scripts/report_how_many_sites.mjs prints every figure on the content page with its denominator, plus the three arithmetic tables and the residues. –sizing, –ranktail and –rare add the full text of every probe hit. Runtime is about 20 seconds, dominated by reading 1,121 full texts twice.
- report_how_many_sites.mjs
#!/usr/bin/env node // Every figure on statistics:how_many_sites, with its denominator. // // node scripts/report_how_many_sites.mjs # the full report // node scripts/report_how_many_sites.mjs --wiki # the same tables as DokuWiki markup // node scripts/report_how_many_sites.mjs --sizing # all 24 sizing-sentence hits, in full // node scripts/report_how_many_sites.mjs --ranktail # all rank-band hits, in full // node scripts/report_how_many_sites.mjs --rare # rare site-level phenomena with their n // // Rules enforced here rather than remembered (data/extract/README.md): // * every figure names its own population; "of 5,859 papers" is never one // * sentinels are silence, never an answer // * papers are counted, never tuples, EXCEPT the size distribution, which is // tuple-level because one paper may draw several populations — it says so // * full-text probes print their hits so a count is never the claim; the // hand classification of the sizing hits is a literal table below, so the // script prints the judgement rather than hiding it // // Population: this page is about the SIZE of a web population, so a paper is in // scope if it drew at least one population whose unit is a web unit // (websites / domains / web-pages) AND stated a size for one of them. That is // the same web-unit definition scripts/report_sampling.mjs uses, narrowed to // papers with a stated n, because a page about the number cannot use a paper // that gives none. // // The arithmetic sections (§6, §7, §8) are NOT measurements of the corpus. // They are closed-form binomial and design-effect calculations printed here so // the page's tables are reproducible. They are labelled as such on the page. import fs from 'node:fs'; import path from 'node:path'; import { loadExtractions, dataRoot, pct, table, wikiTable } from './lib.mjs'; const WIKI = process.argv.includes('--wiki'); const SHOW_SIZING = process.argv.includes('--sizing'); const SHOW_RANKTAIL = process.argv.includes('--ranktail'); const SHOW_RARE = process.argv.includes('--rare'); const ROOT = dataRoot(); const rows = loadExtractions(); const WEB_UNITS = new Set(['websites', 'domains', 'web-pages']); const key = (p) => `${p.venue}/${p.year}/${p.slug}`; const crawled = (p) => p.crawlConfig !== null || p.studyTypes.includes('automated-web-crawl'); const webTuples = (p) => p.population.filter((t) => WEB_UNITS.has(t.unit)); const statedSizes = (p) => webTuples(p) .map((t) => t.n) .filter((n) => typeof n === 'number' && n > 0); const maxN = (p) => Math.max(...statedSizes(p)); const WEB_POP = rows.filter((p) => webTuples(p).length > 0); const POP = WEB_POP.filter((p) => statedSizes(p).length > 0); const N = POP.length; const out = []; const say = (s = '') => out.push(s); const h = (s) => { say(''); say('='.repeat(74)); say(s); say('='.repeat(74)); }; const t = (headers, body) => say(WIKI ? wikiTable(headers, body) : table(headers, body)); const med = (a) => { const s = [...a].sort((x, y) => x - y); if (!s.length) return null; return s.length % 2 ? s[(s.length - 1) / 2] : (s[s.length / 2 - 1] + s[s.length / 2]) / 2; }; const quant = (a, q) => { const s = [...a].sort((x, y) => x - y); return s[Math.min(s.length - 1, Math.floor(q * (s.length - 1)))]; }; const fmt = (n) => Number(n).toLocaleString('en-US'); const BUCKETS = [ ['2010–2013', 2010, 2013], ['2014–2017', 2014, 2017], ['2018–2021', 2018, 2021], ['2022–2024', 2022, 2024], ['2025–2026*', 2025, 2026], ]; // ------------------------------------------------------------------ §1 scope h('1. POPULATION — which papers this page is about'); say(`corpus ${fmt(rows.length)}`); say(`drew >=1 web-unit population ${fmt(WEB_POP.length)} ${pct(WEB_POP.length, rows.length)} of corpus`); say(`PAGE POPULATION: ... and stated a size for one ${fmt(N)} ${pct(N, WEB_POP.length)} of those`); say(` ... of which also ran a crawl ${fmt(POP.filter(crawled).length)} ${pct(POP.filter(crawled).length, N)}`); say(` ... sampled the web without crawling ${fmt(POP.filter((p) => !crawled(p)).length)} ${pct(POP.filter((p) => !crawled(p)).length, N)}`); say(`web-unit population tuples in scope ${fmt(POP.reduce((a, p) => a + webTuples(p).length, 0))}`); say(` ... of those, with a stated size ${fmt(POP.reduce((a, p) => a + statedSizes(p).length, 0))}`); // ------------------------------------------------------- §2 the number itself h('2. THE NUMBER — what sizes the field actually picks (paper-level, max stated web n)'); const allN = POP.map(maxN); say(`min ${fmt(Math.min(...allN))} p25 ${fmt(quant(allN, 0.25))} median ${fmt(med(allN))} p75 ${fmt(quant(allN, 0.75))} p90 ${fmt(quant(allN, 0.9))} max ${fmt(Math.max(...allN))}`); say(''); const freq = new Map(); for (const n of allN) freq.set(n, (freq.get(n) || 0) + 1); say('-- the 12 most common values, papers of ' + N + ' --'); t( ['Largest stated web population', 'Papers', 'Share of ' + N], [...freq.entries()] .sort((a, b) => b[1] - a[1] || a[0] - b[0]) .slice(0, 12) .map(([n, c]) => [fmt(n), c, pct(c, N)]) ); const isRound = (n) => { for (let e = 0; e <= 12; e++) for (const m of [1, 2, 5]) if (n === m * 10 ** e) return true; return false; }; const round = allN.filter(isRound).length; const exact1M = allN.filter((n) => n === 1e6).length; say(''); say(`sizes that are 1, 2 or 5 times a power of ten: ${round} ${pct(round, N)}`); say(`exactly 1,000,000: ${exact1M} ${pct(exact1M, N)}`); say(`at or above 1,000,000: ${allN.filter((n) => n >= 1e6).length} ${pct(allN.filter((n) => n >= 1e6).length, N)}`); say(`at or below 1,000: ${allN.filter((n) => n <= 1000).length} ${pct(allN.filter((n) => n <= 1000).length, N)}`); say(''); say('-- has the number grown? median largest stated web n, by four-year bucket --'); t( ['Years', 'Papers', 'Median n', 'p75', 'n >= 1M', 'n <= 1,000'], BUCKETS.map(([lab, a, b]) => { const ps = POP.filter((p) => p.year >= a && p.year <= b).map(maxN); return [lab, ps.length, fmt(med(ps)), fmt(quant(ps, 0.75)), pct(ps.filter((n) => n >= 1e6).length, ps.length), pct(ps.filter((n) => n <= 1000).length, ps.length)]; }) ); say('* 2025-2026 is provisional: CCS/IMC 2026 have not been held and IEEE S&P/WWW 2026 are'); say(' incompletely selected. See literature:corpus.'); say(''); say('-- the same, crawling papers only --'); t( ['Years', 'Crawling papers', 'Median n', 'n >= 1M'], BUCKETS.map(([lab, a, b]) => { const ps = POP.filter(crawled).filter((p) => p.year >= a && p.year <= b).map(maxN); return [lab, ps.length, fmt(med(ps)), pct(ps.filter((n) => n >= 1e6).length, ps.length)]; }) ); // -------------------------------------------------- §3 what sets the number h('3. WHAT SETS THE NUMBER — n by what the paper does to each site'); say('classification.method is a mid-band field (58% run-to-run agreement) and is'); say('multi-valued, so a paper appears in every row whose method it uses. Read the'); say('medians as a ranking, not as precise figures.'); say(''); const METHODS = ['heuristic-rules', 'regex-or-signature', 'blocklist', 'curated-database', 'third-party-service', 'supervised-ml', 'static-analysis', 'dynamic-analysis', 'manual-labelling', 'llm']; t( ['Per-site analysis method', 'Papers in scope', 'Median n', 'p25', 'p75'], METHODS.map((m) => { const ps = POP.filter((p) => p.classification.some((c) => c.method === m)); const ns = ps.map(maxN); return [m, ps.length, ns.length ? fmt(med(ns)) : '—', ns.length ? fmt(quant(ns, 0.25)) : '—', ns.length ? fmt(quant(ns, 0.75)) : '—']; }) ); const noCls = POP.filter((p) => p.classification.length === 0); say(''); say(`papers in scope with no classification tuple at all: ${noCls.length}, median n ${fmt(med(noCls.map(maxN)))}`); say(''); say('-- LLM classification by bucket, because the population is small and new --'); t( ['Years', 'LLM papers in scope', 'Median n', 'Non-LLM papers', 'Median n'], BUCKETS.map(([lab, a, b]) => { const inb = POP.filter((p) => p.year >= a && p.year <= b); const l = inb.filter((p) => p.classification.some((c) => c.method === 'llm')); const nl = inb.filter((p) => !p.classification.some((c) => c.method === 'llm')); return [lab, l.length, l.length ? fmt(med(l.map(maxN))) : '—', nl.length, fmt(med(nl.map(maxN)))]; }) ); // ------------------------------------------- §4 does anyone justify the number h('4. DOES ANYONE JUSTIFY THE NUMBER? — structured field'); const inferential = rows.filter((p) => p.statistics.some((s) => s.kind && s.kind !== 'descriptive-only')); const powerPapers = rows.filter((p) => p.statistics.some((s) => s.kind === 'power-analysis')); say(`papers running inference ('inferential') ${fmt(inferential.length)}`); say(` ... with a statistics tuple of kind power-analysis ${fmt(powerPapers.length)} ${pct(powerPapers.length, inferential.length)}`); say(` of those, recruited human participants ${powerPapers.filter((p) => p.participants.length > 0).length} ${pct(powerPapers.filter((p) => p.participants.length > 0).length, powerPapers.length)}`); say(` of those, ran a crawl ${powerPapers.filter(crawled).length}`); say(` of those, drew a web-unit population ${powerPapers.filter((p) => webTuples(p).length > 0).length}`); say(`power-analysis papers inside this page's ${N}-paper population ${powerPapers.filter((p) => POP.includes(p)).length}`); say(''); say('-- every power-analysis tuple in a paper INSIDE THIS PAGE\'S POPULATION --'); say(' (was: only the ones that also crawled. The page states a negative — no'); say(' observational crawl sized from power — so the audit trail has to show'); say(' every in-population paper the structured field returns, not a subset.)'); for (const p of powerPapers.filter((p) => POP.includes(p))) { say(` ${key(p)} participants=${p.participants.length} max web n=${statedSizes(p).length ? fmt(maxN(p)) : 'n/a'}`); for (const s of p.statistics.filter((s) => s.kind === 'power-analysis')) { say(` method: ${s.method}`); say(` detail: ${s.detail}`); say(` quote : ${s.evidence.quote}`); } } // -------------------------------------------------------- §5 full-text probe h('5. DOES ANYONE JUSTIFY THE NUMBER? — full-text probe over the ' + N + ' papers'); say('A sizing sentence is one sentence (<=400 chars, no terminal punctuation inside)'); say('that contains ALL THREE of: sample-size / precision language, a web unit, and a'); say('number. Text is whitespace-collapsed first, because a PDF line break inside a'); say('phrase silently loses the match. paper.cols.txt is read, not paper.norm.txt.'); say(''); const SIZE_RE = /\b(?:sample size[sd]?|statistical power|power analys[ei]s|powered to detect|margin of error|chernoff|hoeffding|confidence interval|diminishing returns|law of large numbers)\b/i; const UNIT_RE = /\b(?:web ?sites?|sites?|domains?|web ?pages?|pages?|urls?|hosts?)\b/i; const NUM_RE = /\b(?:\d[\d,]{2,}|\d+(?:\.\d+)?\s?(?:k|K|thousand|million|m|M))\b/; const SENT_RE = /[^.!?]{0,400}[.!?]/g; // Sentence splitting on [.!?] cuts "0.4% of the top thousand" in half at the // decimal point, which silently dropped murley2021_websocket from the rank-band // probe below. Decimal points, and periods inside common abbreviations, are // masked with \u0001 before splitting and restored before printing. Found // 2026-09-02 by checking a paper the probe should have matched and did not. const DOT = '\u0001'; const maskDots = (s) => s .replace(/(\d)\.(\d)/g, `$1${DOT}$2`) .replace(/\b(?:e\.g|i\.e|et al|cf|vs|Fig|Sec|approx|no)\./gi, (m) => m.slice(0, -1) + DOT); const unmaskDots = (s) => s.split(DOT).join('.'); const norm = (s) => maskDots(s.replace(/\s+/g, ' ')); // Hand classification of every hit, read one by one on 2026-09-02. The verdict // is a judgement and is printed so a reader can disagree with it. // derives-n the size of a web population is derived from a stated // precision, power or confidence requirement // design-reason the size is argued for, but from a design or cost argument // rather than a calculation // states-only the sentence names the size of a web population and gives no // argument for it at all // not-sizing the sentence is about the statistics of the analysis, or a // scale boast, or a caveat about a group being too small const VERDICT = { 'CCS/2011/fashion-crimes-trending-term-exploitation-on-the-web': ['derives-n', 'sized a 363-site manual-inspection subsample from a 95% confidence requirement'], 'IEEE-SP/2014/hunting-the-red-fox-online-understanding-and-detection-of-mass-redirect-script-i': ['not-sizing', 'dropped groups with <10 URLs; a filtering threshold, not a sample size'], 'IMC/2014/censorship-in-the-wild-analyzing-internet-filtering-in-syria': ['not-sizing', 'CI width on a proportion in an already-collected 32M-request dataset'], 'WWW/2016/remedying-web-hijacking-notification-effectiveness-and-webmaster-comprehension': ['not-sizing', '"largest studied to date" — a scale boast'], 'IMC/2017/email-typosquatting': ['not-sizing', 'CI on an extrapolated email volume, not on a sample size'], 'WWW/2017/security-implications-of-redirection-trail-in-popular-websites-worldwide': ['derives-n', 'sized a 2,000-site subsample from a stated margin of error'], 'IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking': ['design-reason', 'empirical saturation: resampled 500 combinations at varying sizes to see when a block page stops appearing'], 'IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni': ['derives-n', '384 phishing sites per arm from an ANOVA power calculation'], 'IMC/2019/a-longitudinal-analysis-of-the-ads-txt-standard': ['not-sizing', 'the population grew because the standard was adopted; no calculation'], 'WWW/2019/auditing-the-partisanship-of-google-search-snippets': ['not-sizing', 'a >100 threshold for including a website in a per-site test'], 'USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish': ['derives-n', 'deployment size from a power calculation; also reports the achieved effect size'], 'WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web': ['design-reason', '4,000 sites as 1,000 from each rank magnitude band — a stratification argument'], 'IMC/2022/a-world-wide-view-of-browsing-the-world-wide-web': ['design-reason', 'calls n=10,000 "conservative"; asserted, not derived'], 'USENIX/2022/many-roads-lead-to-rome-how-packet-headers-influence-dns-censorship-measurement': ['design-reason', 'reduced the domain set citing diminishing returns and minimising risk'], 'WWW/2022/et-tu-brute-privacy-analysis-of-government-websites-and-mobile-apps': ['not-sizing', 'a 100-site verification subsample described as limited; no calculation'], 'WWW/2022/leveraging-googles-publisher-specific-ids-to-detect-website-administration': ['not-sizing', 'a figure caption fragment about snapshot variance'], 'CCS/2023/read-between-the-lines-detecting-tracking-javascript-with-bytecode-classificatio': ['not-sizing', 'explains a recall difference by a smaller training set'], 'PETS/2023/blocking-javascript-without-breaking-the-web-an-empirical-investigation': ['derives-n', '383 sites hand-inspected, sized for 100K at +-5% margin of error'], 'USENIX/2023/autofr-automated-filter-rule-generation-for-adblocking': ['derives-n', '272 of 933 sites hand-inspected, sized for 95%/+-5%'], 'WWW/2024/a-study-of-gdpr-compliance-under-the-transparency-and-consent-framework': ['not-sizing', 'a caveat that one rank band had too few domains'], 'WWW/2025/digital-disparities-a-comparative-web-measurement-study-across-economic-boundari': ['design-reason', 'a fixed 10,000-per-country target, met by descending the ranking'], 'USENIX/2025/websites-global-privacy-control-compliance-at-scale-and-over-time': ['not-sizing', 'CI on an estimate derived from an already-fixed population'], 'PETS/2025/understanding-regional-filter-lists-efficacy-and-impact': ['not-sizing', 'diminishing returns of larger filter-rule sets, not of more sites'], 'PETS/2026/privacy-vs-profit-the-impact-of-googles-manifest-version-3-mv3-update-on-ad-bloc': ['not-sizing', 'a robustness remark about variation in sample size'], // added 2026-09-02 after masking decimal points in the sentence splitter 'PETS/2020/enhanced-performance-and-privacy-for-tls-over-tcp-fast-open': ['states-only', 'names a 30,000-hostname subsample of the Alexa top million with no derivation'], 'IMC/2021/web-censorship-measurements-of-http-3-over-quic': ['not-sizing', 'a table fragment: per-country replication counts after validation filtering'], 'WWW/2021/its-not-just-the-site-its-the-contents-intra-domain-fingerprinting-social-media': ['not-sizing', 'a figure caption about confidence intervals on burst sizes'], }; const sizingHits = []; for (const p of POP) { const f = path.join(ROOT, `fulltext/${p.year}/${p.venue}/${p.slug}/paper.cols.txt`); if (!fs.existsSync(f)) continue; const txt = norm(fs.readFileSync(f, 'utf8')); const keep = (txt.match(SENT_RE) || []).filter((s) => SIZE_RE.test(s) && UNIT_RE.test(s) && NUM_RE.test(s)); if (keep.length) sizingHits.push({ p, sents: keep }); } say(`papers with >=1 sizing sentence: ${sizingHits.length} ${pct(sizingHits.length, N)} of ${N}`); const unclassified = sizingHits.filter((x) => !VERDICT[key(x.p)]); say(`hand-read: ${sizingHits.length - unclassified.length} of ${sizingHits.length}` + (unclassified.length ? ' !! UNREAD HITS BELOW — the verdict table is stale' : '')); for (const x of unclassified) say(` !! unread: ${key(x.p)}`); const byVerdict = (v) => sizingHits.filter((x) => VERDICT[key(x.p)]?.[0] === v); say(''); t( ['Verdict after reading the sentence', 'Papers', 'Share of ' + N], [ ['derives n from a precision or power requirement', byVerdict('derives-n').length, pct(byVerdict('derives-n').length, N)], ['argues the size from design or cost, without a calculation', byVerdict('design-reason').length, pct(byVerdict('design-reason').length, N)], ['names a size and gives no argument for it', byVerdict('states-only').length, pct(byVerdict('states-only').length, N)], ['not about sizing a web population', byVerdict('not-sizing').length, pct(byVerdict('not-sizing').length, N)], ] ); say(''); say('-- the derives-n papers, and what the calculation sized --'); for (const x of byVerdict('derives-n')) { say(` ${key(x.p)}`); say(` n stated for the web population: ${fmt(maxN(x.p))}`); say(` ${VERDICT[key(x.p)][1]}`); } const subsample = byVerdict('derives-n').filter((x) => /subsample|hand-inspected|manual-inspection|hand/i.test(VERDICT[key(x.p)][1])); say(''); say(`of the ${byVerdict('derives-n').length} derives-n papers, ${subsample.length} size a HAND-CHECKED SUBSAMPLE`); say(`rather than the crawl itself; ${byVerdict('derives-n').length - subsample.length} size the measured population.`); if (SHOW_SIZING) { h('5b. EVERY SIZING-SENTENCE HIT, IN FULL'); for (const x of sizingHits) { const v = VERDICT[key(x.p)] ?? ['UNREAD', '']; say(''); say(`### ${key(x.p)} n=${fmt(maxN(x.p))} [${v[0]}] ${v[1]}`); for (const s of x.sents) say(' ' + unmaskDots(s).trim()); } } // ------------------------------------------------------- §6 precision, exact h('6. ARITHMETIC (not a corpus measurement) — precision on a prevalence'); say('Wald half-width of a 95% interval on a proportion: 1.96*sqrt(p(1-p)/n).'); say('Nothing about the web is in this table; it is here so the page can be checked.'); say(''); const Z = 1.96; const hw = (p, n) => Z * Math.sqrt((p * (1 - p)) / n); const NS = [100, 1000, 10000, 100000, 1000000]; t( ['n', 'p = 0.5', 'p = 0.2', 'p = 0.05', 'p = 0.01', 'p = 0.001'], NS.map((n) => [fmt(n), ...[0.5, 0.2, 0.05, 0.01, 0.001].map((p) => '+-' + (100 * hw(p, n)).toFixed(p <= 0.01 ? 3 : 2) + ' pp')]) ); say(''); say('-- n needed for a given half-width at p = 0.5, the worst case --'); t( ['Target half-width', 'n needed'], [0.05, 0.03, 0.01, 0.003, 0.001].map((w) => [`+-${(100 * w).toFixed(1)} pp`, fmt(Math.ceil((Z ** 2 * 0.25) / w ** 2))]) ); // ---------------------------------------------------- §7 rare-event capture h('7. ARITHMETIC (not a corpus measurement) — capturing a rare thing'); say('Expected count is n*p. P(at least one) = 1-(1-p)^n. The n for a 95% chance of'); say('seeing at least one is ceil(ln(0.05)/ln(1-p)), which is about 3/p. If you see'); say('zero in n trials, the 95% upper bound on p is about 3/n (the rule of three).'); say(''); const PS = [0.1, 0.01, 0.001, 0.0001, 0.00001]; t( ['Prevalence p', 'Expected hits at n=10,000', 'at n=100,000', 'at n=1,000,000', 'n for P(>=1 hit) >= 95%', 'n for a +-20% relative CI'], PS.map((p) => [ (100 * p).toFixed(3).replace(/0+$/, '').replace(/\.$/, '') + '%', (1e4 * p).toFixed(2), (1e5 * p).toFixed(1), (1e6 * p).toFixed(0), fmt(Math.ceil(Math.log(0.05) / Math.log(1 - p))), fmt(Math.ceil((Z ** 2 * (1 - p)) / (0.2 ** 2 * p))), ]) ); say(''); say('-- what a zero buys you: 95% upper bound on p after seeing no hits in n --'); t( ['n with zero hits', '95% upper bound on p', 'i.e. at most ... sites in a million'], NS.map((n) => [fmt(n), (100 * (3 / n)).toFixed(4) + '%', fmt(Math.round((3 / n) * 1e6))]) ); // ------------------------------------------------------- §8 the design effect h('8. ARITHMETIC (not a corpus measurement) — sites are not independent draws'); say('Design effect for equal clusters of size m with intra-class correlation rho:'); say('deff = 1 + (m-1)*rho, and the effective sample size is n/deff. The ICC of any'); say('real web-measurement outcome has never been measured (see the open question on'); say('statistics:hypothesis_testing), so these rho values are illustrative, chosen to'); say('match the sweep published on that page.'); say(''); const deff = (m, rho) => 1 + (m - 1) * rho; t( ['Cluster size m', 'rho = 0.02', 'rho = 0.05', 'rho = 0.20'], [10, 100, 1000].map((m) => [fmt(m), ...[0.02, 0.05, 0.2].map((r) => deff(m, r).toFixed(1) + 'x')]) ); say(''); say('-- effective n for a 1,000,000-site crawl --'); t( ['Cluster size m', 'rho = 0.02', 'rho = 0.05', 'rho = 0.20'], [10, 100, 1000].map((m) => [fmt(m), ...[0.02, 0.05, 0.2].map((r) => fmt(Math.round(1e6 / deff(m, r))))]) ); say(''); say('-- half-width on p=0.05 at n=1,000,000 once deflated --'); t( ['Cluster size m', 'rho = 0.02', 'rho = 0.05', 'rho = 0.20'], [10, 100, 1000].map((m) => [fmt(m), ...[0.02, 0.05, 0.2].map((r) => '+-' + (100 * hw(0.05, 1e6 / deff(m, r))).toFixed(3) + ' pp')]) ); // -------------------------------------------------------------- §9 rank tail h('9. THE RANK TAIL — papers reporting the same measurement per rank band'); say('Probe: one sentence containing a rank band, a phrase comparing across rank, and'); say('a percentage. Whitespace-collapsed. Every hit is printed with --ranktail; the'); say('count is not the claim.'); say(''); // Widened on 2026-09-02: the first version required a DIGIT after "top", so it // missed "0.4% of the top thousand and 0.05% of websites in the top million" // (murley2021_websocket) — a per-band result stated in words. Spelled-out // magnitudes and "Top Alexa lists" are now included. The narrow version found // 10 papers; the wide one is what the page reports. const BAND_RE = /\b(?:top[- ]?(?:\d[\d,]*\s?(?:k|K|m|M|thousand|million)?|ten|hundred|thousand|million)|larger top[- ]?\w+ lists?|rank(?:ed|ing|s)? (?:between|from) [\d,]+|\d[\d,]*\s?(?:k|K)\s?[–-]\s?\d[\d,]*\s?(?:k|K|m|M))\b/i; const DIR_RE = /\b(?:less popular|lower[- ]ranked|higher[- ]ranked|more popular|as (?:the )?rank(?:ing)? (?:increases|decreases|grows)|down the (?:list|ranking|tail)|long tail|popularity (?:decreases|declines|drops)|rank(?:ing)? bucket|rank(?:ing)? (?:band|strat|tier)|per[- ]rank|by rank|(?:share|rate|prevalence|adoption)[^.]{0,40}(?:decreases|increases|drops|rises)[^.]{0,30}(?:larger|lower|higher|less|more)|top thousand and[^.]{0,40}top million)\b/i; const PCT_RE = /\d+(?:\.\d+)?\s?%/; const rankHits = []; for (const p of POP) { const f = path.join(ROOT, `fulltext/${p.year}/${p.venue}/${p.slug}/paper.cols.txt`); if (!fs.existsSync(f)) continue; const txt = norm(fs.readFileSync(f, 'utf8')); const keep = (txt.match(SENT_RE) || []).filter((s) => BAND_RE.test(s) && DIR_RE.test(s) && PCT_RE.test(s)); if (keep.length) rankHits.push({ p, sents: keep }); } // Hand classification, read one by one on 2026-09-02. const RANK_VERDICT = { 'CCS/2019/lets-encrypt-an-automated-certificate-authority-to-encrypt-the-entire-web': ['per-band', 'CA market share RISES down the ranking: 5% of top 1K, 20% of top 100K, 35% of top 1M'], 'IMC/2021/who-you-gonna-call-an-empirical-evaluation-of-website-security-txt-deployment': ['per-band', 'security.txt adoption FALLS down the ranking: 3-4% of top 10K, ~1% of top 100K'], 'IMC/2020/analyzing-third-party-service-dependencies-in-modern-web-services-have-we-learne': ['per-band', 'third-party CA use 71% in top-100 against 77% in top-100K'], 'USENIX/2023/youve-got-report-measurement-and-security-implications-of-dmarc-reporting': ['per-band', 'DMARC misconfiguration ~10% in the most popular 10K against ~20% in the least popular 10K'], 'CCS/2023/you-call-this-archaeology-evaluating-web-archives-for-reproducible-web-security': ['per-band', 'archive hit rate plotted across Tranco rank buckets — the tail is less archived'], 'WWW/2019/the-chain-of-implicit-trust-an-analysis-of-the-web-third-party-resources-loading': ['per-band', 'dependency chains marginally more common in more popular sites (55% in the top 10K)'], 'WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web': ['per-band', 'Server-Sent Events on 0.4% of the top thousand against 0.05% of the top million — an 8x head/tail gap'], 'USENIX/2026/the-state-of-passkeys-studying-the-adoption-and-security-of-passkeys-on-the-web': ['per-band', 'lower-ranked and unranked sites less susceptible than the Tranco top 1K'], 'CCS/2025/piixel-leaks-passive-identification-of-personally-identifiable-information-leaka': ['per-band', 'PII leakage share quoted separately for the 500K-1M band'], 'WWW/2010/detection-and-analysis-of-drive-by-download-attacks-and-malicious-javascript-cod': ['not-rank', 'long tail of vulnerable applications, not of site rank'], 'IMC/2015/who-is-com-learning-to-parse-whois-records': ['not-rank', 'long tail of WHOIS service names'], 'IMC/2016/characterizing-website-behaviors-across-logged-in-and-not-logged-in-users': ['not-rank', 'long tail of FQDNs within the data, not a rank band'], 'USENIX/2016/internet-jones-and-the-raiders-of-the-lost-trackers-an-archaeological-study-of-w': ['per-band', 'Wayback coverage better for more popular trackers: 75% of the top 100 against 53% of all'], 'PETS/2021/the-cname-of-the-game-large-scale-analysis-of-dns-based-tracking-evasion': ['not-rank', 'a figure-legend fragment; the comparison is not in the sentence'], 'WWW/2019/unnecessarily-identifiable-quantifying-the-fingerprintability-of-browser-extensi': ['not-rank', 'popularity bands of browser EXTENSIONS, not of websites'], 'IMC/2022/a-world-wide-view-of-browsing-the-world-wide-web': ['per-band', 'category mix shifts across the ranking: News & Media above 15% of top-50 sites and dropping'], // added 2026-09-02 after masking decimal points in the sentence splitter widened // this probe from 15 hits to 28 'PETS/2019/a-quic-look-at-web-tracking': ['per-band', 'QUIC support at six nested depths: 21.00% (top 100) down to 0.02% (top 1M)'], 'USENIX/2023/a-large-scale-measurement-of-website-login-policies': ['per-band', 'BOTH directions in one paper: cleartext password submission 0.44%/0.50%/0.62% rising across top 10K/100K/1M, rate limiting 33.6%/26.7%/24.1% falling'], 'IMC/2020/accept-the-risk-and-continue-measuring-the-long-tail-of-government-https-adoptio': ['per-band', 'valid HTTPS on government sites ~30% in the top million and in the long tail alike — FLAT'], 'IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking': ['per-band', 'Luminati-protected domains 0.05% of the wide sample against 0.2% of the Alexa top 10K'], 'USENIX/2018/o-single-sign-off-where-art-thou-an-empirical-analysis-of-single-sign-on-account': ['per-band', 'SSO support higher in more popular sites, 10.8% coverage in the top 100K'], 'IMC/2015/neither-snow-nor-rain-nor-mitm-an-empirical-analysis-of-email-delivery-security': ['per-band', 'TLS support lags in the long tail of 700,000 SMTP servers behind the Alexa top million'], 'PETS/2026/overcoming-language-barriers-multilingual-analysis-of-the-2023-swiss-privacy-law': ['per-band', 'policy-generator use and disclosure rates tabulated by rank group; generators used by less popular sites'], 'WWW/2025/harmful-terms-and-where-to-find-them-measuring-and-modeling-unfavorable-financia': ['per-band', 'unfavorable financial terms more prevalent on less popular sites within the Tranco top 100K'], 'IMC/2022/toppling-top-lists-evaluating-the-accuracy-of-popular-website-lists': ['not-rank', 'measures the ACCURACY OF THE LIST by rank bucket, not a web phenomenon per band'], 'IMC/2011/towards-understanding-modern-web-traffic': ['not-rank', 'a figure fragment about URL popularity distributions'], 'CCS/2018/measuring-information-leakage-in-website-fingerprinting-attacks-and-defenses': ['not-rank', 'a figure fragment: top-100 most informative FEATURES by rank'], 'IMC/2025/towards-a-non-binary-view-of-ipv6-adoption': ['not-rank', 'long tail of third-party DOMAINS to fix, not a rank band of sites'], }; const rankUnread = rankHits.filter((x) => !RANK_VERDICT[key(x.p)]); say(`papers with >=1 rank-band sentence: ${rankHits.length} ${pct(rankHits.length, N)} of ${N}`); for (const x of rankUnread) say(` !! unread: ${key(x.p)}`); const perBand = rankHits.filter((x) => RANK_VERDICT[key(x.p)]?.[0] === 'per-band'); say(`after reading: ${perBand.length} report a measurement per rank band; ${rankHits.length - perBand.length - rankUnread.length} are a long tail of something else`); say(''); for (const x of perBand) say(` ${key(x.p)}\n ${RANK_VERDICT[key(x.p)][1]}`); if (SHOW_RANKTAIL) { h('9b. EVERY RANK-BAND HIT, IN FULL'); for (const x of rankHits) { const v = RANK_VERDICT[key(x.p)] ?? ['UNREAD', '']; say(''); say(`### ${key(x.p)} n=${fmt(maxN(x.p))} [${v[0]}] ${v[1]}`); for (const s of x.sents.slice(0, 2)) say(' ' + unmaskDots(s).trim()); } } // ------------------------------------------------------ §10 rare in practice h('10. RARE IN PRACTICE — site-level phenomena under 0.5%, in papers with n >= 200,000'); say('detection[].prevalence is free text. This is a PROBE, not an aggregate: it keeps'); say('tuples whose prevalence string names a web unit and contains a percentage below'); say('0.5, and prints them so each can be read. It is not a rate of anything.'); say(''); const UNIT_ONLY = /\b(sites?|websites?|domains?|pages?)\b/i; const PCT_G = /(\d+(?:\.\d+)?)\s*%/g; let rareTuples = 0; const rare = []; for (const p of POP.filter((p) => maxN(p) >= 200000)) { for (const d of p.detection) { if (!d.prevalence) continue; const s = String(d.prevalence); if (!UNIT_ONLY.test(s)) continue; PCT_G.lastIndex = 0; const vs = []; let m; while ((m = PCT_G.exec(s)) !== null) vs.push(parseFloat(m[1])); if (vs.length && Math.min(...vs) < 0.5) { rareTuples++; rare.push({ p, d, s, min: Math.min(...vs) }); } } } say(`papers in scope with n >= 200,000: ${POP.filter((p) => maxN(p) >= 200000).length}`); say(`tuples kept: ${rareTuples}, across ${new Set(rare.map((r) => key(r.p))).size} papers`); if (SHOW_RARE) for (const r of rare) say(` ${r.p.year} ${key(r.p)} n=${fmt(maxN(r.p))}\n ${r.d.phenomenon} :: ${r.s}`); else say('(run with --rare to print all of them)'); // -------------------------------------------------------------- §11 residues h('11. RESIDUES AND WHAT IS NOT COUNTED'); say(`web-unit papers with NO stated size, excluded from this page: ${WEB_POP.length - N} ${pct(WEB_POP.length - N, WEB_POP.length)} of ${WEB_POP.length}`); const anyPrev = POP.filter((p) => p.detection.some((d) => d.prevalence)); say(`papers in scope with >=1 detection prevalence string: ${anyPrev.length} ${pct(anyPrev.length, N)}`); let prevT = 0, prevUnit = 0, prevPctInUnit = 0; for (const p of POP) for (const d of p.detection) { if (!d.prevalence) continue; prevT++; const s = String(d.prevalence); if (!UNIT_ONLY.test(s)) continue; prevUnit++; if (/\d+(?:\.\d+)?\s*%/.test(s)) prevPctInUnit++; } say(`prevalence strings in scope: ${prevT}; naming a web unit: ${prevUnit} (${pct(prevUnit, prevT)}); of those with a %: ${prevPctInUnit} (${pct(prevPctInUnit, prevUnit)})`); say(`The ${pct(prevUnit - prevPctInUnit, prevUnit)} of unit-naming prevalence strings that carry no`); say('parseable percentage, and the percentages that turn out to be false-positive'); say('rates or classifier metrics rather than site shares,'); say('are why no aggregate of prevalence-versus-n is published. An earlier draft'); say('cross-tabulated them; it was dropped. See the provenance page.'); console.log(out.join('\n'));
Unedited output of node scripts/report_how_many_sites.mjs:
- report_how_many_sites-output.txt
========================================================================== 1. POPULATION — which papers this page is about ========================================================================== corpus 5,859 drew >=1 web-unit population 1,153 19.7% of corpus PAGE POPULATION: ... and stated a size for one 1,121 97.2% of those ... of which also ran a crawl 674 60.1% ... sampled the web without crawling 447 39.9% web-unit population tuples in scope 2,484 ... of those, with a stated size 2,379 ========================================================================== 2. THE NUMBER — what sizes the field actually picks (paper-level, max stated web n) ========================================================================== min 1 p25 1,000 median 20,000 p75 1,000,000 p90 7,341,165 max 250,000,000,000 -- the 12 most common values, papers of 1121 -- Largest stated web population Papers Share of 1121 ----------------------------- ------ ------------- 1,000,000 140 12.5% 10,000 80 7.1% 100,000 65 5.8% 100 48 4.3% 1,000 37 3.3% 500 20 1.8% 20,000 20 1.8% 5,000 18 1.6% 50 16 1.4% 20 14 1.2% 200 14 1.2% 10 12 1.1% sizes that are 1, 2 or 5 times a power of ten: 523 46.7% exactly 1,000,000: 140 12.5% at or above 1,000,000: 311 27.7% at or below 1,000: 302 26.9% -- has the number grown? median largest stated web n, by four-year bucket -- Years Papers Median n p75 n >= 1M n <= 1,000 ---------- ------ -------- --------- ------- ---------- 2010–2013 98 5,183 100,000 19.4% 37.8% 2014–2017 182 26,589.5 1,000,000 34.6% 29.7% 2018–2021 312 42,541.5 1,000,000 27.2% 21.8% 2022–2024 341 20,000 1,000,000 26.1% 25.8% 2025–2026* 188 20,000 1,000,000 29.3% 29.3% * 2025-2026 is provisional: CCS/IMC 2026 have not been held and IEEE S&P/WWW 2026 are incompletely selected. See literature:corpus. -- the same, crawling papers only -- Years Crawling papers Median n n >= 1M ---------- --------------- -------- ------- 2010–2013 49 10,000 16.3% 2014–2017 115 18,000 25.2% 2018–2021 191 35,102 21.5% 2022–2024 208 13,795.5 19.2% 2025–2026* 111 20,000 24.3% ========================================================================== 3. WHAT SETS THE NUMBER — n by what the paper does to each site ========================================================================== classification.method is a mid-band field (58% run-to-run agreement) and is multi-valued, so a paper appears in every row whose method it uses. Read the medians as a ranking, not as precise figures. Per-site analysis method Papers in scope Median n p25 p75 ------------------------ --------------- -------- ------ --------- heuristic-rules 605 93,427 2,846 1,000,000 regex-or-signature 111 132,798 4,000 1,000,000 blocklist 145 20,000 10,000 387,000 curated-database 228 93,283.5 3,806 1,000,000 third-party-service 271 100,000 7,337 1,000,000 supervised-ml 266 20,000 1,000 442,190 static-analysis 50 15,000 1,899 1,000,000 dynamic-analysis 66 74,215.5 5,000 1,000,000 manual-labelling 307 10,000 361 906,731 llm 29 10,000 2,892 90,000 papers in scope with no classification tuple at all: 101, median n 3,000 -- LLM classification by bucket, because the population is small and new -- Years LLM papers in scope Median n Non-LLM papers Median n ---------- ------------------- -------- -------------- -------- 2010–2013 0 — 98 5,183 2014–2017 0 — 182 26,589.5 2018–2021 0 — 312 42,541.5 2022–2024 7 10,000 334 20,000 2025–2026* 22 10,500 166 20,000 ========================================================================== 4. DOES ANYONE JUSTIFY THE NUMBER? — structured field ========================================================================== papers running inference ('inferential') 1,762 ... with a statistics tuple of kind power-analysis 87 4.9% of those, recruited human participants 74 85.1% of those, ran a crawl 6 of those, drew a web-unit population 9 power-analysis papers inside this page's 1121-paper population 9 -- every power-analysis tuple in a paper INSIDE THIS PAGE'S POPULATION -- (was: only the ones that also crawled. The page states a negative — no observational crawl sized from power — so the audit trail has to show every in-population paper the structured field returns, not a subset.) WWW/2016/characterizing-long-tail-seo-spam-on-cloud-web-hosting-services participants=0 max web n=1,073,642 method: Chernoff Bounds detail: trust interval δ = 0.01; error probability λ = 0.01; sample n = 500 quote : We set the trust interval δ = 0.01 and the error probability λ = 0.01 to obtain the number of sampled cloud directories n = 500. IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni participants=0 max web n=1,980 method: one-way independent ANOVA detail: power = 0.95; significance level = 0.05; assumed medium effect size f = 0.25 quote : Our goal was to obtain a power of 0.95 at the significance level of 0.05 in a one-way independent ANOVA test USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish participants=0 max web n=4,393 method: Statistical power analysis detail: power 0.95; p-value 0.05; assumed medium effect size 0.25; later estimated effect size 0.36 quote : To obtain a power of 0.95 at a p-value of 0.05, we initially assumed a medium effect size of 0.25. PETS/2021/managing-potentially-intrusive-practices-in-the-browser-a-user-centered-perspect participants=2 max web n=16 method: G*Power detail: At least 997 participants required for α = 0.05 and medium effect size 0.4. quote : We required at least 997 participants to achieve significance at α = 0.05 with a medium effect size (0.4). IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising-a-field-measurement participants=2 max web n=10 method: G*Power sample-size calculation detail: At least 126 participants for medium effects with 10 predictors quote : we calculated that we needed a sample size of at least 126 participants to detect medium effect sizes using a linear regression with 10 predictors PETS/2024/what-does-it-mean-to-be-creepy-responses-to-visualizations-of-personal-browsing participants=2 max web n=118,000 method: point biserial model power analysis detail: All comparisons had >=80% power quote : A power analysis, using the point biserial model, strengthens this hypothesis, as all comparisons between conditions had ≥80% power. IEEE-SP/2016/sending-out-an-sms-characterizing-the-security-of-the-sms-ecosystem-with-public participants=0 max web n=8 method: statistical power detail: all tests had statistical power of 0.98 or higher quote : Finally, we confirmed that all tests performed had a statistical power of 0.98 or higher NDSS/2025/the-kids-are-all-right-investigating-the-susceptibility-of-teens-and-adults-to-youtube-giveaway-scams participants=2 max web n=451 method: repeated measures ANOVA power analysis detail: effect size = 0.25, power = 0.95, α = 0.05, 2 measures, 12 groups, correlation = 0.5 quote : This calculation used a repeated measures ANOVA model, assuming an effect size of 0.25, power = 0.95, α = 0.05, 2 measures, 12 groups ... IEEE-SP/2025/lets-get-visual-testing-visual-analogies-and-metaphors-for-conveying-privacy-pol participants=4 max web n=10 method: a priori power analysis; post-hoc power analyses detail: minimum N=357; initial N=428; lowest post-hoc power .72 quote : An initial power analysis with the tool G*Power [78] suggested a minimum required sample size of N=357 participants ========================================================================== 5. DOES ANYONE JUSTIFY THE NUMBER? — full-text probe over the 1121 papers ========================================================================== A sizing sentence is one sentence (<=400 chars, no terminal punctuation inside) that contains ALL THREE of: sample-size / precision language, a web unit, and a number. Text is whitespace-collapsed first, because a PDF line break inside a phrase silently loses the match. paper.cols.txt is read, not paper.norm.txt. papers with >=1 sizing sentence: 26 2.3% of 1121 hand-read: 26 of 26 Verdict after reading the sentence Papers Share of 1121 ---------------------------------------------------------- ------ ------------- derives n from a precision or power requirement 6 0.5% argues the size from design or cost, without a calculation 5 0.4% names a size and gives no argument for it 1 0.1% not about sizing a web population 14 1.2% -- the derives-n papers, and what the calculation sized -- CCS/2011/fashion-crimes-trending-term-exploitation-on-the-web n stated for the web population: 6,558 sized a 363-site manual-inspection subsample from a 95% confidence requirement WWW/2017/security-implications-of-redirection-trail-in-popular-websites-worldwide n stated for the web population: 1,000,000 sized a 2,000-site subsample from a stated margin of error IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni n stated for the web population: 1,980 384 phishing sites per arm from an ANOVA power calculation USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish n stated for the web population: 4,393 deployment size from a power calculation; also reports the achieved effect size PETS/2023/blocking-javascript-without-breaking-the-web-an-empirical-investigation n stated for the web population: 100,000 383 sites hand-inspected, sized for 100K at +-5% margin of error USENIX/2023/autofr-automated-filter-rule-generation-for-adblocking n stated for the web population: 5,000 272 of 933 sites hand-inspected, sized for 95%/+-5% of the 6 derives-n papers, 4 size a HAND-CHECKED SUBSAMPLE rather than the crawl itself; 2 size the measured population. ========================================================================== 6. ARITHMETIC (not a corpus measurement) — precision on a prevalence ========================================================================== Wald half-width of a 95% interval on a proportion: 1.96*sqrt(p(1-p)/n). Nothing about the web is in this table; it is here so the page can be checked. n p = 0.5 p = 0.2 p = 0.05 p = 0.01 p = 0.001 --------- --------- --------- --------- ---------- ---------- 100 +-9.80 pp +-7.84 pp +-4.27 pp +-1.950 pp +-0.619 pp 1,000 +-3.10 pp +-2.48 pp +-1.35 pp +-0.617 pp +-0.196 pp 10,000 +-0.98 pp +-0.78 pp +-0.43 pp +-0.195 pp +-0.062 pp 100,000 +-0.31 pp +-0.25 pp +-0.14 pp +-0.062 pp +-0.020 pp 1,000,000 +-0.10 pp +-0.08 pp +-0.04 pp +-0.020 pp +-0.006 pp -- n needed for a given half-width at p = 0.5, the worst case -- Target half-width n needed ----------------- -------- +-5.0 pp 385 +-3.0 pp 1,068 +-1.0 pp 9,604 +-0.3 pp 106,712 +-0.1 pp 960,400 ========================================================================== 7. ARITHMETIC (not a corpus measurement) — capturing a rare thing ========================================================================== Expected count is n*p. P(at least one) = 1-(1-p)^n. The n for a 95% chance of seeing at least one is ceil(ln(0.05)/ln(1-p)), which is about 3/p. If you see zero in n trials, the 95% upper bound on p is about 3/n (the rule of three). Prevalence p Expected hits at n=10,000 at n=100,000 at n=1,000,000 n for P(>=1 hit) >= 95% n for a +-20% relative CI ------------ ------------------------- ------------ -------------- ----------------------- ------------------------- 10% 1000.00 10000.0 100000 29 865 1% 100.00 1000.0 10000 299 9,508 0.1% 10.00 100.0 1000 2,995 95,944 0.01% 1.00 10.0 100 29,956 960,304 0.001% 0.10 1.0 10 299,572 9,603,904 -- what a zero buys you: 95% upper bound on p after seeing no hits in n -- n with zero hits 95% upper bound on p i.e. at most ... sites in a million ---------------- -------------------- ----------------------------------- 100 3.0000% 30,000 1,000 0.3000% 3,000 10,000 0.0300% 300 100,000 0.0030% 30 1,000,000 0.0003% 3 ========================================================================== 8. ARITHMETIC (not a corpus measurement) — sites are not independent draws ========================================================================== Design effect for equal clusters of size m with intra-class correlation rho: deff = 1 + (m-1)*rho, and the effective sample size is n/deff. The ICC of any real web-measurement outcome has never been measured (see the open question on statistics:hypothesis_testing), so these rho values are illustrative, chosen to match the sweep published on that page. Cluster size m rho = 0.02 rho = 0.05 rho = 0.20 -------------- ---------- ---------- ---------- 10 1.2x 1.4x 2.8x 100 3.0x 6.0x 20.8x 1,000 21.0x 51.0x 200.8x -- effective n for a 1,000,000-site crawl -- Cluster size m rho = 0.02 rho = 0.05 rho = 0.20 -------------- ---------- ---------- ---------- 10 847,458 689,655 357,143 100 335,570 168,067 48,077 1,000 47,664 19,627 4,980 -- half-width on p=0.05 at n=1,000,000 once deflated -- Cluster size m rho = 0.02 rho = 0.05 rho = 0.20 -------------- ---------- ---------- ---------- 10 +-0.046 pp +-0.051 pp +-0.071 pp 100 +-0.074 pp +-0.104 pp +-0.195 pp 1,000 +-0.196 pp +-0.305 pp +-0.605 pp ========================================================================== 9. THE RANK TAIL — papers reporting the same measurement per rank band ========================================================================== Probe: one sentence containing a rank band, a phrase comparing across rank, and a percentage. Whitespace-collapsed. Every hit is printed with --ranktail; the count is not the claim. papers with >=1 rank-band sentence: 28 2.5% of 1121 after reading: 19 report a measurement per rank band; 9 are a long tail of something else IMC/2015/neither-snow-nor-rain-nor-mitm-an-empirical-analysis-of-email-delivery-security TLS support lags in the long tail of 700,000 SMTP servers behind the Alexa top million USENIX/2016/internet-jones-and-the-raiders-of-the-lost-trackers-an-archaeological-study-of-w Wayback coverage better for more popular trackers: 75% of the top 100 against 53% of all IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking Luminati-protected domains 0.05% of the wide sample against 0.2% of the Alexa top 10K USENIX/2018/o-single-sign-off-where-art-thou-an-empirical-analysis-of-single-sign-on-account SSO support higher in more popular sites, 10.8% coverage in the top 100K CCS/2019/lets-encrypt-an-automated-certificate-authority-to-encrypt-the-entire-web CA market share RISES down the ranking: 5% of top 1K, 20% of top 100K, 35% of top 1M PETS/2019/a-quic-look-at-web-tracking QUIC support at six nested depths: 21.00% (top 100) down to 0.02% (top 1M) WWW/2019/the-chain-of-implicit-trust-an-analysis-of-the-web-third-party-resources-loading dependency chains marginally more common in more popular sites (55% in the top 10K) IMC/2020/analyzing-third-party-service-dependencies-in-modern-web-services-have-we-learne third-party CA use 71% in top-100 against 77% in top-100K IMC/2020/accept-the-risk-and-continue-measuring-the-long-tail-of-government-https-adoptio valid HTTPS on government sites ~30% in the top million and in the long tail alike — FLAT IMC/2021/who-you-gonna-call-an-empirical-evaluation-of-website-security-txt-deployment security.txt adoption FALLS down the ranking: 3-4% of top 10K, ~1% of top 100K WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web Server-Sent Events on 0.4% of the top thousand against 0.05% of the top million — an 8x head/tail gap IMC/2022/a-world-wide-view-of-browsing-the-world-wide-web category mix shifts across the ranking: News & Media above 15% of top-50 sites and dropping CCS/2023/you-call-this-archaeology-evaluating-web-archives-for-reproducible-web-security archive hit rate plotted across Tranco rank buckets — the tail is less archived USENIX/2023/a-large-scale-measurement-of-website-login-policies BOTH directions in one paper: cleartext password submission 0.44%/0.50%/0.62% rising across top 10K/100K/1M, rate limiting 33.6%/26.7%/24.1% falling USENIX/2023/youve-got-report-measurement-and-security-implications-of-dmarc-reporting DMARC misconfiguration ~10% in the most popular 10K against ~20% in the least popular 10K CCS/2025/piixel-leaks-passive-identification-of-personally-identifiable-information-leaka PII leakage share quoted separately for the 500K-1M band PETS/2026/overcoming-language-barriers-multilingual-analysis-of-the-2023-swiss-privacy-law policy-generator use and disclosure rates tabulated by rank group; generators used by less popular sites WWW/2025/harmful-terms-and-where-to-find-them-measuring-and-modeling-unfavorable-financia unfavorable financial terms more prevalent on less popular sites within the Tranco top 100K USENIX/2026/the-state-of-passkeys-studying-the-adoption-and-security-of-passkeys-on-the-web lower-ranked and unranked sites less susceptible than the Tranco top 1K ========================================================================== 10. RARE IN PRACTICE — site-level phenomena under 0.5%, in papers with n >= 200,000 ========================================================================== detection[].prevalence is free text. This is a PROBE, not an aggregate: it keeps tuples whose prevalence string names a web unit and contains a percentage below 0.5, and prints them so each can be read. It is not a rate of anything. papers in scope with n >= 200,000: 368 tuples kept: 21, across 21 papers (run with --rare to print all of them) ========================================================================== 11. RESIDUES AND WHAT IS NOT COUNTED ========================================================================== web-unit papers with NO stated size, excluded from this page: 32 2.8% of 1153 papers in scope with >=1 detection prevalence string: 1114 99.4% prevalence strings in scope: 6203; naming a web unit: 1626 (26.2%); of those with a %: 837 (51.5%) The 48.5% of unit-naming prevalence strings that carry no parseable percentage, and the percentages that turn out to be false-positive rates or classifier metrics rather than site shares, are why no aggregate of prevalence-versus-n is published. An earlier draft cross-tabulated them; it was dropped. See the provenance page.
Unedited output of node scripts/report_how_many_sites.mjs –sizing –ranktail –rare — every probe hit, in full, with its hand verdict:
- report_how_many_sites-full-output.txt
========================================================================== 1. POPULATION — which papers this page is about ========================================================================== corpus 5,859 drew >=1 web-unit population 1,153 19.7% of corpus PAGE POPULATION: ... and stated a size for one 1,121 97.2% of those ... of which also ran a crawl 674 60.1% ... sampled the web without crawling 447 39.9% web-unit population tuples in scope 2,484 ... of those, with a stated size 2,379 ========================================================================== 2. THE NUMBER — what sizes the field actually picks (paper-level, max stated web n) ========================================================================== min 1 p25 1,000 median 20,000 p75 1,000,000 p90 7,341,165 max 250,000,000,000 -- the 12 most common values, papers of 1121 -- Largest stated web population Papers Share of 1121 ----------------------------- ------ ------------- 1,000,000 140 12.5% 10,000 80 7.1% 100,000 65 5.8% 100 48 4.3% 1,000 37 3.3% 500 20 1.8% 20,000 20 1.8% 5,000 18 1.6% 50 16 1.4% 20 14 1.2% 200 14 1.2% 10 12 1.1% sizes that are 1, 2 or 5 times a power of ten: 523 46.7% exactly 1,000,000: 140 12.5% at or above 1,000,000: 311 27.7% at or below 1,000: 302 26.9% -- has the number grown? median largest stated web n, by four-year bucket -- Years Papers Median n p75 n >= 1M n <= 1,000 ---------- ------ -------- --------- ------- ---------- 2010–2013 98 5,183 100,000 19.4% 37.8% 2014–2017 182 26,589.5 1,000,000 34.6% 29.7% 2018–2021 312 42,541.5 1,000,000 27.2% 21.8% 2022–2024 341 20,000 1,000,000 26.1% 25.8% 2025–2026* 188 20,000 1,000,000 29.3% 29.3% * 2025-2026 is provisional: CCS/IMC 2026 have not been held and IEEE S&P/WWW 2026 are incompletely selected. See literature:corpus. -- the same, crawling papers only -- Years Crawling papers Median n n >= 1M ---------- --------------- -------- ------- 2010–2013 49 10,000 16.3% 2014–2017 115 18,000 25.2% 2018–2021 191 35,102 21.5% 2022–2024 208 13,795.5 19.2% 2025–2026* 111 20,000 24.3% ========================================================================== 3. WHAT SETS THE NUMBER — n by what the paper does to each site ========================================================================== classification.method is a mid-band field (58% run-to-run agreement) and is multi-valued, so a paper appears in every row whose method it uses. Read the medians as a ranking, not as precise figures. Per-site analysis method Papers in scope Median n p25 p75 ------------------------ --------------- -------- ------ --------- heuristic-rules 605 93,427 2,846 1,000,000 regex-or-signature 111 132,798 4,000 1,000,000 blocklist 145 20,000 10,000 387,000 curated-database 228 93,283.5 3,806 1,000,000 third-party-service 271 100,000 7,337 1,000,000 supervised-ml 266 20,000 1,000 442,190 static-analysis 50 15,000 1,899 1,000,000 dynamic-analysis 66 74,215.5 5,000 1,000,000 manual-labelling 307 10,000 361 906,731 llm 29 10,000 2,892 90,000 papers in scope with no classification tuple at all: 101, median n 3,000 -- LLM classification by bucket, because the population is small and new -- Years LLM papers in scope Median n Non-LLM papers Median n ---------- ------------------- -------- -------------- -------- 2010–2013 0 — 98 5,183 2014–2017 0 — 182 26,589.5 2018–2021 0 — 312 42,541.5 2022–2024 7 10,000 334 20,000 2025–2026* 22 10,500 166 20,000 ========================================================================== 4. DOES ANYONE JUSTIFY THE NUMBER? — structured field ========================================================================== papers running inference ('inferential') 1,762 ... with a statistics tuple of kind power-analysis 87 4.9% of those, recruited human participants 74 85.1% of those, ran a crawl 6 of those, drew a web-unit population 9 power-analysis papers inside this page's 1121-paper population 9 -- every power-analysis tuple in a paper INSIDE THIS PAGE'S POPULATION -- (was: only the ones that also crawled. The page states a negative — no observational crawl sized from power — so the audit trail has to show every in-population paper the structured field returns, not a subset.) WWW/2016/characterizing-long-tail-seo-spam-on-cloud-web-hosting-services participants=0 max web n=1,073,642 method: Chernoff Bounds detail: trust interval δ = 0.01; error probability λ = 0.01; sample n = 500 quote : We set the trust interval δ = 0.01 and the error probability λ = 0.01 to obtain the number of sampled cloud directories n = 500. IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni participants=0 max web n=1,980 method: one-way independent ANOVA detail: power = 0.95; significance level = 0.05; assumed medium effect size f = 0.25 quote : Our goal was to obtain a power of 0.95 at the significance level of 0.05 in a one-way independent ANOVA test USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish participants=0 max web n=4,393 method: Statistical power analysis detail: power 0.95; p-value 0.05; assumed medium effect size 0.25; later estimated effect size 0.36 quote : To obtain a power of 0.95 at a p-value of 0.05, we initially assumed a medium effect size of 0.25. PETS/2021/managing-potentially-intrusive-practices-in-the-browser-a-user-centered-perspect participants=2 max web n=16 method: G*Power detail: At least 997 participants required for α = 0.05 and medium effect size 0.4. quote : We required at least 997 participants to achieve significance at α = 0.05 with a medium effect size (0.4). IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising-a-field-measurement participants=2 max web n=10 method: G*Power sample-size calculation detail: At least 126 participants for medium effects with 10 predictors quote : we calculated that we needed a sample size of at least 126 participants to detect medium effect sizes using a linear regression with 10 predictors PETS/2024/what-does-it-mean-to-be-creepy-responses-to-visualizations-of-personal-browsing participants=2 max web n=118,000 method: point biserial model power analysis detail: All comparisons had >=80% power quote : A power analysis, using the point biserial model, strengthens this hypothesis, as all comparisons between conditions had ≥80% power. IEEE-SP/2016/sending-out-an-sms-characterizing-the-security-of-the-sms-ecosystem-with-public participants=0 max web n=8 method: statistical power detail: all tests had statistical power of 0.98 or higher quote : Finally, we confirmed that all tests performed had a statistical power of 0.98 or higher NDSS/2025/the-kids-are-all-right-investigating-the-susceptibility-of-teens-and-adults-to-youtube-giveaway-scams participants=2 max web n=451 method: repeated measures ANOVA power analysis detail: effect size = 0.25, power = 0.95, α = 0.05, 2 measures, 12 groups, correlation = 0.5 quote : This calculation used a repeated measures ANOVA model, assuming an effect size of 0.25, power = 0.95, α = 0.05, 2 measures, 12 groups ... IEEE-SP/2025/lets-get-visual-testing-visual-analogies-and-metaphors-for-conveying-privacy-pol participants=4 max web n=10 method: a priori power analysis; post-hoc power analyses detail: minimum N=357; initial N=428; lowest post-hoc power .72 quote : An initial power analysis with the tool G*Power [78] suggested a minimum required sample size of N=357 participants ========================================================================== 5. DOES ANYONE JUSTIFY THE NUMBER? — full-text probe over the 1121 papers ========================================================================== A sizing sentence is one sentence (<=400 chars, no terminal punctuation inside) that contains ALL THREE of: sample-size / precision language, a web unit, and a number. Text is whitespace-collapsed first, because a PDF line break inside a phrase silently loses the match. paper.cols.txt is read, not paper.norm.txt. papers with >=1 sizing sentence: 26 2.3% of 1121 hand-read: 26 of 26 Verdict after reading the sentence Papers Share of 1121 ---------------------------------------------------------- ------ ------------- derives n from a precision or power requirement 6 0.5% argues the size from design or cost, without a calculation 5 0.4% names a size and gives no argument for it 1 0.1% not about sizing a web population 14 1.2% -- the derives-n papers, and what the calculation sized -- CCS/2011/fashion-crimes-trending-term-exploitation-on-the-web n stated for the web population: 6,558 sized a 363-site manual-inspection subsample from a 95% confidence requirement WWW/2017/security-implications-of-redirection-trail-in-popular-websites-worldwide n stated for the web population: 1,000,000 sized a 2,000-site subsample from a stated margin of error IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni n stated for the web population: 1,980 384 phishing sites per arm from an ANOVA power calculation USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish n stated for the web population: 4,393 deployment size from a power calculation; also reports the achieved effect size PETS/2023/blocking-javascript-without-breaking-the-web-an-empirical-investigation n stated for the web population: 100,000 383 sites hand-inspected, sized for 100K at +-5% margin of error USENIX/2023/autofr-automated-filter-rule-generation-for-adblocking n stated for the web population: 5,000 272 of 933 sites hand-inspected, sized for 95%/+-5% of the 6 derives-n papers, 4 size a HAND-CHECKED SUBSAMPLE rather than the crawl itself; 2 size the measured population. ========================================================================== 5b. EVERY SIZING-SENTENCE HIT, IN FULL ========================================================================== ### CCS/2011/fashion-crimes-trending-term-exploitation-on-the-web n=6,558 [derives-n] sized a 363-site manual-inspection subsample from a 95% confidence requirement To that ef- fect we selected a statistically significant (95% confidence interval) random sample of 363 websites for manual inspection. ### IEEE-SP/2014/hunting-the-red-fox-online-understanding-and-detection-of-mass-redirect-script-i n=1,558,690 [not-sizing] dropped groups with <10 URLs; a filtering threshold, not a sample size To avoid drawing any conclusion based upon such a small sample size, we ignored those with less than 10 unique URLs, which left us 213 types. ### IMC/2014/censorship-in-the-wild-analyzing-internet-filtering-in-syria n=28 [not-sizing] CI width on a proportion in an already-collected 32M-request dataset 2), requests from the same user accessing the same URL), some are for a sample size of n = 32M, the actual proportion in the full consistently denied, while others are sometimes or always allowed. ### WWW/2016/remedying-web-hijacking-notification-effectiveness-and-webmaster-comprehension n=313,190 [not-sizing] "largest studied to date" — a scale boast Second, our coverage of hijacked websites is biased towards threats caught by Google's pipelines, though we still capture a sample size of 760,935 incidents, the largest studied to date. ### IMC/2017/email-typosquatting n=1,000,000 [not-sizing] CI on an extrapolated email volume, not on a sample size Our model finds that the 1,211 typosquatting domains registered by others should receive approximately 260,514 emails per year, with a 95% confidence interval ranging between 22,577 and 905,174 emails per year. ### WWW/2017/security-implications-of-redirection-trail-in-popular-websites-worldwide n=1,000,000 [derives-n] sized a 2,000-site subsample from a stated margin of error We then selected a ran- (48.0%) of the reachable home pages (N = 9, 159) do not domized sample of 2,000 websites (margin of error = 2.19%) support HTTPS at all even though this analysis is based to represent the 1M sample, and we used the represented on the top 10K popular websites; only 24.7% provide secure sample to compare with the 10K websites. ### IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking n=87,000,000 [design-reason] empirical saturation: resampled 500 combinations at varying sizes to see when a block page stops appearing From each set of samples, we then selected 500 random combinations of different sample sizes to detect how many combinations would not yield a block page. ### IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni n=1,980 [derives-n] 384 phishing sites per arm from an ANOVA power calculation Sample Size Selection For our full tests, we chose a sample size of 384 phishing sites for each entity. All sites ultimately delivered 100% uptime during deployment, thus we ended up with an effective sample size of 396 per experiment. ### IMC/2019/a-longitudinal-analysis-of-the-ads-txt-standard n=240,000 [not-sizing] the population grew because the standard was adopted; no calculation Subsequently, our sample size grew from 100K websites on January 15, 2018 to 240K on April 1, 2019. ### WWW/2019/auditing-the-partisanship-of-google-search-snippets n=541,437 [not-sizing] a >100 threshold for including a website in a per-site test For each website that had a sample size >100, we use the Wilcoxon V signed-rank test to compare the snippets' Γ scores and webpages' Γ scores, and further applied a Boneferroni correction for multiple hypothesis testings. ### PETS/2020/enhanced-performance-and-privacy-for-tls-over-tcp-fast-open n=1,000,000 [states-only] names a 30,000-hostname subsample of the Alexa top million with no derivation Based upon a sample size of approx. 30 000 hostnames within the Alexa Top Million Sites, we then investigate to which extent changing server IP addresses affects the performance gains achievable by TFO. ### USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish n=4,393 [derives-n] deployment size from a power calculation; also reports the achieved effect size Despite our relatively large sample size for each deployment, we do not believe that the volume of URLs we reported hampered the anti-phishing ecosystem's 392 29th USENIX Security Symposium ability to mitigate real threats. ### IMC/2021/web-censorship-measurements-of-http-3-over-quic n=4,000 [not-sizing] a table fragment: per-country replication counts after validation filtering ations, TLS-hs-to overall Hosts Sample Size* China VPS, 69, 37.3% 25.9% 2.7% - (45090) 102 6706 Iran VPS, 36, 34.4% - 33.4% - - (62442) 120 3887 India PD, 2, 15.0% 7.5% - 4.5% (55836) 133 266 India VPS, 60, 16.3% - - - (14061) 133 7531 India PD, 1, 12.8% - - - (38266) 133 133 Kazakhstan VPN, 22, 3.2% - 3.2% - (9198) 82 1764 * final sample size of all replications after validation step filtering (c. ### WWW/2021/its-not-just-the-site-its-the-contents-intra-domain-fingerprinting-social-media n=120,000 [not-sizing] a figure caption about confidence intervals on burst sizes 5 244 22 Burst 4 242 Burst 4 20 240 0.4 238 1 2 Burst 3 Burst Index Burst Index 0.3 44 Amount of Transmitted Data Per Burst (MB) Amount of Transmitted Data Per Burst (MB) Burst 3 300 Burst 2 42 297.5 0.2 40 295 292.5 38 Burst 1 290 0.1 36 287.5 285 34 282.5 3 4 0.0 0.0 0.5 1.0 1.5 2.0 Burst Index Burst Index Web Page Loading Time (seconds) (d) Visualized the CDN bursts with 95% confidence interval. ### WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web n=1,000,000 [design-reason] 4,000 sites as 1,000 from each rank magnitude band — a stratification argument Our subset of sites for this crawl consisted of the top 1000 sites, along with a random sample of 1000 sites from each of the top 10K, top 100K, and top 1M, for a total sample size of 4000 websites. ### IMC/2022/a-world-wide-view-of-browsing-the-world-wide-web n=1,000,000 [design-reason] calls n=10,000 "conservative"; asserted, not derived We note that di�erent per-platform usage rates dataset of ranked lists does not provide raw tra�c volume to each site, we instead set of incognito mode (which is not captured in Chrome telemetry) a conservative sample size of # = 10, 000 (the number of sites under consideration). ### USENIX/2022/many-roads-lead-to-rome-how-packet-headers-influence-dns-censorship-measurement n=567 [design-reason] reduced the domain set citing diminishing returns and minimising risk Given our ethical goal to minimize risk combined 31st USENIX Security Symposium 455 with the likelihood of diminishing returns [31], we selected a shift would move an experiment into the consistency state, subset of domains from this list spread across categories. ### WWW/2022/et-tu-brute-privacy-analysis-of-government-websites-and-mobile-apps n=231,449 [not-sizing] a 100-site verification subsample described as limited; no calculation We manually verify our government website dataset (with a limited sample size of 100, selected randomly) to ensure false positives are eliminated. ### CCS/2023/read-between-the-lines-detecting-tracking-javascript-with-bytecode-classificatio n=50,000 [not-sizing] explains a recall difference by a smaller training set Bytecode classification using D3 dataset had lower recall compared to D1 and D2 datasets, which can be attributed to the smaller sample size due to only processing 1,500 websites. ### PETS/2023/blocking-javascript-without-breaking-the-web-an-empirical-investigation n=100,000 [derives-n] 383 sites hand-inspected, sized for 100K at +-5% margin of error We excluded a total of 117 websites and manually inspected 383 websites, which is a statistically significant sample size for 100K websites with ± 5% margin of error [16]. ### USENIX/2023/autofr-automated-filter-rule-generation-for-adblocking n=5,000 [derives-n] 272 of 933 sites hand-inspected, sized for 95%/+-5% To further confirm our results for AutoFR and EasyList, we randomly selected 272 sites (a sample size out of 933 sites to get a confidence level of 95% with a 5% confidence interval), and we visually inspected them. ### WWW/2024/a-study-of-gdpr-compliance-under-the-transparency-and-consent-framework n=2,230 [not-sizing] a caveat that one rank band had too few domains domains ranked between 2,000 and 4,000 which suffered from a Thus, TCF non-compliance is not only occurring on small domains small sample size). ### WWW/2025/digital-disparities-a-comparative-web-measurement-study-across-economic-boundari n=200,000 [design-reason] a fixed 10,000-per-country target, met by descending the ranking However, for countries such as Bangladesh, Pakistan, Nigeria, and the Philippines, less popular websites were included to meet the target sample size of 10,000 websites per country. ### USENIX/2025/websites-global-privacy-control-compliance-at-scale-and-over-time n=42,312 [not-sizing] CI on an estimate derived from an already-fixed population Based on this calculation we estimate that 4,465-5,929 of the 9,578 sites, 47-62% according to the 95% confidence interval, that had a request from an SAFG service in December 2023 also meet the traffic requirement per §3.2. ### PETS/2025/understanding-regional-filter-lists-efficacy-and-impact n=10,000 [not-sizing] diminishing returns of larger filter-rule sets, not of more sites Proceedings on Privacy Enhancing Technologies 2025(2) Larger rule sets offer diminishing returns, increasing runtime, and resource usage without proportionate benefits in blocked URLs. ### PETS/2026/privacy-vs-profit-the-impact-of-googles-manifest-version-3-mv3-update-on-ad-bloc n=924 [not-sizing] a robustness remark about variation in sample size 1) examined whether Guard MV3, Stands MV2, Stands MV3, uBlock MV2, uBlock our findings hold despite variations in sample size and website se-MV3) with the number of websites (924) yields the number lection. ========================================================================== 6. ARITHMETIC (not a corpus measurement) — precision on a prevalence ========================================================================== Wald half-width of a 95% interval on a proportion: 1.96*sqrt(p(1-p)/n). Nothing about the web is in this table; it is here so the page can be checked. n p = 0.5 p = 0.2 p = 0.05 p = 0.01 p = 0.001 --------- --------- --------- --------- ---------- ---------- 100 +-9.80 pp +-7.84 pp +-4.27 pp +-1.950 pp +-0.619 pp 1,000 +-3.10 pp +-2.48 pp +-1.35 pp +-0.617 pp +-0.196 pp 10,000 +-0.98 pp +-0.78 pp +-0.43 pp +-0.195 pp +-0.062 pp 100,000 +-0.31 pp +-0.25 pp +-0.14 pp +-0.062 pp +-0.020 pp 1,000,000 +-0.10 pp +-0.08 pp +-0.04 pp +-0.020 pp +-0.006 pp -- n needed for a given half-width at p = 0.5, the worst case -- Target half-width n needed ----------------- -------- +-5.0 pp 385 +-3.0 pp 1,068 +-1.0 pp 9,604 +-0.3 pp 106,712 +-0.1 pp 960,400 ========================================================================== 7. ARITHMETIC (not a corpus measurement) — capturing a rare thing ========================================================================== Expected count is n*p. P(at least one) = 1-(1-p)^n. The n for a 95% chance of seeing at least one is ceil(ln(0.05)/ln(1-p)), which is about 3/p. If you see zero in n trials, the 95% upper bound on p is about 3/n (the rule of three). Prevalence p Expected hits at n=10,000 at n=100,000 at n=1,000,000 n for P(>=1 hit) >= 95% n for a +-20% relative CI ------------ ------------------------- ------------ -------------- ----------------------- ------------------------- 10% 1000.00 10000.0 100000 29 865 1% 100.00 1000.0 10000 299 9,508 0.1% 10.00 100.0 1000 2,995 95,944 0.01% 1.00 10.0 100 29,956 960,304 0.001% 0.10 1.0 10 299,572 9,603,904 -- what a zero buys you: 95% upper bound on p after seeing no hits in n -- n with zero hits 95% upper bound on p i.e. at most ... sites in a million ---------------- -------------------- ----------------------------------- 100 3.0000% 30,000 1,000 0.3000% 3,000 10,000 0.0300% 300 100,000 0.0030% 30 1,000,000 0.0003% 3 ========================================================================== 8. ARITHMETIC (not a corpus measurement) — sites are not independent draws ========================================================================== Design effect for equal clusters of size m with intra-class correlation rho: deff = 1 + (m-1)*rho, and the effective sample size is n/deff. The ICC of any real web-measurement outcome has never been measured (see the open question on statistics:hypothesis_testing), so these rho values are illustrative, chosen to match the sweep published on that page. Cluster size m rho = 0.02 rho = 0.05 rho = 0.20 -------------- ---------- ---------- ---------- 10 1.2x 1.4x 2.8x 100 3.0x 6.0x 20.8x 1,000 21.0x 51.0x 200.8x -- effective n for a 1,000,000-site crawl -- Cluster size m rho = 0.02 rho = 0.05 rho = 0.20 -------------- ---------- ---------- ---------- 10 847,458 689,655 357,143 100 335,570 168,067 48,077 1,000 47,664 19,627 4,980 -- half-width on p=0.05 at n=1,000,000 once deflated -- Cluster size m rho = 0.02 rho = 0.05 rho = 0.20 -------------- ---------- ---------- ---------- 10 +-0.046 pp +-0.051 pp +-0.071 pp 100 +-0.074 pp +-0.104 pp +-0.195 pp 1,000 +-0.196 pp +-0.305 pp +-0.605 pp ========================================================================== 9. THE RANK TAIL — papers reporting the same measurement per rank band ========================================================================== Probe: one sentence containing a rank band, a phrase comparing across rank, and a percentage. Whitespace-collapsed. Every hit is printed with --ranktail; the count is not the claim. papers with >=1 rank-band sentence: 28 2.5% of 1121 after reading: 19 report a measurement per rank band; 9 are a long tail of something else IMC/2015/neither-snow-nor-rain-nor-mitm-an-empirical-analysis-of-email-delivery-security TLS support lags in the long tail of 700,000 SMTP servers behind the Alexa top million USENIX/2016/internet-jones-and-the-raiders-of-the-lost-trackers-an-archaeological-study-of-w Wayback coverage better for more popular trackers: 75% of the top 100 against 53% of all IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking Luminati-protected domains 0.05% of the wide sample against 0.2% of the Alexa top 10K USENIX/2018/o-single-sign-off-where-art-thou-an-empirical-analysis-of-single-sign-on-account SSO support higher in more popular sites, 10.8% coverage in the top 100K CCS/2019/lets-encrypt-an-automated-certificate-authority-to-encrypt-the-entire-web CA market share RISES down the ranking: 5% of top 1K, 20% of top 100K, 35% of top 1M PETS/2019/a-quic-look-at-web-tracking QUIC support at six nested depths: 21.00% (top 100) down to 0.02% (top 1M) WWW/2019/the-chain-of-implicit-trust-an-analysis-of-the-web-third-party-resources-loading dependency chains marginally more common in more popular sites (55% in the top 10K) IMC/2020/analyzing-third-party-service-dependencies-in-modern-web-services-have-we-learne third-party CA use 71% in top-100 against 77% in top-100K IMC/2020/accept-the-risk-and-continue-measuring-the-long-tail-of-government-https-adoptio valid HTTPS on government sites ~30% in the top million and in the long tail alike — FLAT IMC/2021/who-you-gonna-call-an-empirical-evaluation-of-website-security-txt-deployment security.txt adoption FALLS down the ranking: 3-4% of top 10K, ~1% of top 100K WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web Server-Sent Events on 0.4% of the top thousand against 0.05% of the top million — an 8x head/tail gap IMC/2022/a-world-wide-view-of-browsing-the-world-wide-web category mix shifts across the ranking: News & Media above 15% of top-50 sites and dropping CCS/2023/you-call-this-archaeology-evaluating-web-archives-for-reproducible-web-security archive hit rate plotted across Tranco rank buckets — the tail is less archived USENIX/2023/a-large-scale-measurement-of-website-login-policies BOTH directions in one paper: cleartext password submission 0.44%/0.50%/0.62% rising across top 10K/100K/1M, rate limiting 33.6%/26.7%/24.1% falling USENIX/2023/youve-got-report-measurement-and-security-implications-of-dmarc-reporting DMARC misconfiguration ~10% in the most popular 10K against ~20% in the least popular 10K CCS/2025/piixel-leaks-passive-identification-of-personally-identifiable-information-leaka PII leakage share quoted separately for the 500K-1M band PETS/2026/overcoming-language-barriers-multilingual-analysis-of-the-2023-swiss-privacy-law policy-generator use and disclosure rates tabulated by rank group; generators used by less popular sites WWW/2025/harmful-terms-and-where-to-find-them-measuring-and-modeling-unfavorable-financia unfavorable financial terms more prevalent on less popular sites within the Tranco top 100K USENIX/2026/the-state-of-passkeys-studying-the-adoption-and-security-of-passkeys-on-the-web lower-ranked and unranked sites less susceptible than the Tranco top 1K ========================================================================== 9b. EVERY RANK-BAND HIT, IN FULL ========================================================================== ### WWW/2010/detection-and-analysis-of-drive-by-download-attacks-and-malicious-javascript-cod n=115,706 [not-rank] long tail of vulnerable applications, not of site rank It is interesting to observe that even though a relatively high detection rate can be achieved with a small number of applications (about 90% with the top 5 applications), the detection curve is characterized by a long tail (one would have to install 22 applications to achieve 98% of detection). ### IMC/2011/towards-understanding-modern-web-traffic n=1,197 [not-rank] a figure fragment about URL popularity distributions .4 2008 86 2010 2010 % Requests % Requests 0.2 0.3 84 82 0.2 0.1 80 US 0.1 CN 78 FR BR 0 0 76 0 1 2 3 4 5 0 1 2 3 4 5 10 10 10 10 10 10 10 10 10 10 10 10 2006 2008 2010 URL Ranking URL Ranking Year (a) US: Top 100K URLs by % requests (b) CN: Top 100K URLs by % requests (c) % accessed once URLs (tail) Figure 12: URL popularity: The popular URLs grow, but the long tail of the content is also growing. ### IMC/2015/neither-snow-nor-rain-nor-mitm-an-empirical-analysis-of-email-delivery-security n=1,000,000 [per-band] TLS support lags in the long tail of 700,000 SMTP servers behind the Alexa top million However, such best practices continue to lag for the long tail of 700,000 SMTP servers associated with the Alexa Top Million: only 82% support TLS, of which a mere 35% are properly configured to allow server authentication. ### IMC/2015/who-is-com-learning-to-parse-whois-records n=102,077,202 [not-rank] long tail of WHOIS service names Although there is a long tail of service names, the top 10 account for 73% of protected domains. ### IMC/2016/characterizing-website-behaviors-across-logged-in-and-not-logged-in-users n=345 [not-rank] long tail of FQDNs within the data, not a rank band Limiting to the top 80% also ensures we could focus on categorizing the most prevalent FQDNs while avoiding the very long tail of FQDNs that were not frequently encountered. ### USENIX/2016/internet-jones-and-the-raiders-of-the-lost-trackers-an-archaeological-study-of-w n=500 [per-band] Wayback coverage better for more popular trackers: 75% of the top 100 against 53% of all In general, more popular trackers are better represented in Wayback data: 75% of the top 100 live trackers, compared to 53% of all live trackers. ### CCS/2018/measuring-information-leakage-in-website-fingerprinting-attacks-and-defenses n=2,000 [not-rank] a figure fragment: top-100 most informative FEATURES by rank Top 100 Most Informative Features (indexed by rank) Category Index (a) (b) 8 INFORMATION LEAKAGE IN WF DEFENSES Figure 9: Information Leakage Measurement Validation: 90% This section firstly gives the theoretical analysis on why accuracy Confidence Interval for the Measurement is not a reliable metric to validate a WF defense. ### IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking n=87,000,000 [per-band] Luminati-protected domains 0.05% of the wide sample against 0.2% of the Alexa top 10K This is only 0.05% of our sample, compared to 0.2% of Alexa Top 10K domains, indicating that more popular websites are more Total 5,462 238 (4.4%) protected by Luminati. ### USENIX/2018/o-single-sign-off-where-art-thou-an-empirical-analysis-of-single-sign-on-account n=1,000,000 [per-band] SSO support higher in more popular sites, 10.8% coverage in the top 100K We find that more popular websites are more likely to support SSO, as shown in Figure 4, with a 10.8% coverage in the top 100K, Cascading account compromise. ### CCS/2019/lets-encrypt-an-automated-certificate-authority-to-encrypt-the-entire-web n=223,000,000 [per-band] CA market share RISES down the ranking: 5% of top 1K, 20% of top 100K, 35% of top 1M The CA's market Let's Encrypt has seen rapidly growing adoption among top share increases as site popularity decreases: 5% of the top 1K, 20% million sites since its launch, while most other CAs have not (Figof the top 100K, and 35% of the top 1M sites with HTTPS use Let's ure 4b). ### PETS/2019/a-quic-look-at-web-tracking n=1,000,000 [per-band] QUIC support at six nested depths: 21.00% (top 100) down to 0.02% (top 1M) However, this share decreases for larger Top Alexa lists to only 0.0186% within the Alexa Top 1 Million. ### WWW/2019/the-chain-of-implicit-trust-an-analysis-of-the-web-third-party-resources-loading n=200,000 [per-band] dependency chains marginally more common in more popular sites (55% in the top 10K) It The propensity to form dependency chains is marginally higher has been commonly used in the academic literature to detect ma- in more popular websites; for example, 55% in the Alexa top 10K licious apps, executables, software and domains [7, 11, 13, 16, 17]. ### WWW/2019/unnecessarily-identifiable-quantifying-the-fingerprintability-of-browser-extensi n=50 [not-rank] popularity bands of browser EXTENSIONS, not of websites bloat with more than 9% of the top 5K extensions containing bloat, and approximately 4% of the less popular extensions. ### IMC/2020/analyzing-third-party-service-dependencies-in-modern-web-services-have-we-learne n=100,000 [per-band] third-party CA use 71% in top-100 against 77% in top-100K The use of third party CAs is also higher (77% in top-100K) in less popular websites as compared to more popular (71% in top-100) websites. ### IMC/2020/accept-the-risk-and-continue-measuring-the-long-tail-of-government-https-adoptio n=135,408 [per-band] valid HTTPS on government sites ~30% in the top million and in the long tail alike — FLAT not representative of the world, both having high human develop-Though ranking does have an effect, overall valid https use ment index scores (USA:15, ROK:22) and Internet adoption rates in government websites in the top million is similar to results (USA:90%, ROK:96%), among other unique factors, their relative in the long tail dataset, at ∼30%. ### IMC/2021/who-you-gonna-call-an-empirical-evaluation-of-website-security-txt-deployment n=100,000 [per-band] security.txt adoption FALLS down the ranking: 3-4% of top 10K, ~1% of top 100K We observe a higher adoption rate for higher-ranked percentage decreases to 3-4% for the top 10K sites, and only a per- websites, aligning with our observation that higher-ranked sites cent for the top 100K. ### PETS/2021/the-cname-of-the-game-large-scale-analysis-of-dns-based-tracking-evasion n=5,506,818 [not-rank] a figure-legend fragment; the comparison is not in the sentence CNAME-cloaking publisher domains in Tranco top 10k Less popular trackers in Tranco top 10k Change in # publishers 20% the list of IPs we found in October through a scan with zdns. ### WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web n=1,000,000 [per-band] Server-Sent Events on 0.4% of the top thousand against 0.05% of the top million — an 8x head/tail gap Higher-ranked sites more commonly use polling, with 19.8% of the top thousand sites leveraging the technique compared to 9.2% of sites in the top million. We find them in use on only 0.4% of the top thousand and 0.05% of websites in the top million. ### IMC/2022/a-world-wide-view-of-browsing-the-world-wide-web n=1,000,000 [per-band] category mix shifts across the ranking: News & Media above 15% of top-50 sites and dropping government services (26), and then a long tail of other types of sites Other categories are disproportionately represented in the middle that include gig economy (3), EdTech (6), ISPs and telecoms (9, 4 of of the range: News & Media peaks above 15% of top-50 sites and which provide TV service and 2 of which provide email), and job drops to less than 7% of top-10K sites. Top million lists capture well the vast majority of user tra�c (⇡95%), but studies that focus on the top million sites evenly are skewed heavily towards the long tail of the web. ### IMC/2022/toppling-top-lists-evaluating-the-accuracy-of-popular-website-lists n=1,000,000 [not-rank] measures the ACCURACY OF THE LIST by rank bucket, not a web phenomenon per band For instance, of the 1,790 domains we measure in the Alexa top 10K, 70% of them are ranked by Cloud�are in a lower rank-magnitude bucket, and 27.2% of them are ranked by Cloud�are in a bucket two or more orders of magnitude less popular. The CrUX list much more closely approximates the Cloud�are list by rank-magnitude movement: 47.1% of the 1410 domains in the CrUX top 10K are overranked compared to Cloud�are, and only 1% of them are overranked by two or more orders of magnitude. ### CCS/2023/you-call-this-archaeology-evaluating-web-archives-for-reproducible-web-security n=20,000 [per-band] archive hit rate plotted across Tranco rank buckets — the tail is less archived The only archive able to provide 3172 ACM CCS, November 26-30, 2023, Copenhagen Internet Archive Archive-It Arquivo Congress 100% 75% Hits 50% 25% 0% 100k 200k 300k 400k 500k 600k 700k 100% 75% Fresh Hits 50% 25% 0% Rank bucket Figure 3: Hit rates across popularity buckets fresh hits for a significant number of domains across the Tranco top 1M is again the IA. ### USENIX/2023/a-large-scale-measurement-of-website-login-policies n=1,000,000 [per-band] BOTH directions in one paper: cleartext password submission 0.44%/0.50%/0.62% rising across top 10K/100K/1M, rate limiting 33.6%/26.7%/24.1% falling As with HTTP-only login pages, we see a slight increase in HTTP password submission for lower-ranked: 0.44% of top 10K domains transmitted passwords in the clear, compared to 0.50% of top 100K domains and 0.62% of top 1M domains. Among domains successfully analyzed, we observe that rate limiting logins was more prevalent amongst higher-ranked domains; 33.6% of the domains in the top 10K and 26.7% of domains in the top 100K demonstrated rate limiting, compared to the 24.1% for domains in the top 1M. ### USENIX/2023/youve-got-report-measurement-and-security-implications-of-dmarc-reporting n=384,000,000 [per-band] DMARC misconfiguration ~10% in the most popular 10K against ~20% in the least popular 10K Interestingly, we also observe an increasing trend of such misconfiguration as the ranking increases; for example, almost 20% of the least 10K popular Alexa top-1M domains with DMARC that use external domains are misconfigured compared to about 10% for the most 10K popular domains. ### CCS/2025/piixel-leaks-passive-identification-of-personally-identifiable-information-leaka n=1,000,000 [per-band] PII leakage share quoted separately for the 500K-1M band The adoption rates follow a clear declining trend as website popularity decreases, starting at approximately 23% for the most popular sites (0-10K) and dropping to 17.24% for less popular sites (500K-1M). ### PETS/2026/overcoming-language-barriers-multilingual-analysis-of-the-2023-swiss-privacy-law n=90,000 [per-band] policy-generator use and disclosure rates tabulated by rank group; generators used by less popular sites Moreover, generators are mostly used by less popular websites, with less than 5% of the policies from Top 5k websites indicating use. .0% 52.6% 69.5% hum 34.6% 14.9% 22.6% 13.9% Average 74.7% 79.0% 67.3% 72.9% Total policies 1061 47 6224 1000 Proceedings on Privacy Enhancing Technologies 2026(4) Top 100k+ rank No generator Generator used 76.1% 89.3% 98.7% 100.0% 77.4% 82.1% 79.9% 82.2% 55.8% 67.7% 52.2% 67.1% 20.1% 16.3% 65.7% 72.1% 5028 973 Table 21: Disclosure rates per obligation by rank group and generator use (October 2023). ### WWW/2025/harmful-terms-and-where-to-find-them-measuring-and-modeling-unfavorable-financia n=100,000 [per-band] unfavorable financial terms more prevalent on less popular sites within the Tranco top 100K When applied to shopping websites from the Tranco top 100K, we find that 42.06% of these sites contain at least one unfavorable financial term, with such terms being more prevalent on less popular websites. ### IMC/2025/towards-a-non-binary-view-of-ipv6-adoption n=100,000 [not-rank] long tail of third-party DOMAINS to fix, not a rank band of sites The distribution exhibits a long tail: enabling IPv6 on just the top 500 (3.3%) IPv4-only third-party domains would allow over 25% of IPv6-partial websites to become IPv6-full. ### USENIX/2026/the-state-of-passkeys-studying-the-adoption-and-security-of-passkeys-on-the-web n=18,000,000 [per-band] lower-ranked and unranked sites less susceptible than the Tranco top 1K Figure 8b reveals a different picture: lower-ranked sites and even unranked ones are less susceptible than higher-ranked sites in 100% -5% the Tranco top 1k. ========================================================================== 10. RARE IN PRACTICE — site-level phenomena under 0.5%, in papers with n >= 200,000 ========================================================================== detection[].prevalence is free text. This is a PROBE, not an aggregate: it keeps tuples whose prevalence string names a web unit and contains a percentage below 0.5, and prints them so each can be read. It is not a rate of anything. papers in scope with n >= 200,000: 368 tuples kept: 21, across 21 papers 2015 USENIX/2015/cookies-lack-integrity-real-world-implications n=961,857 Secure cookies over HTTP :: 152 of 48,039 responding domains (0.32%) 2016 CCS/2016/predator-proactive-recognition-and-elimination-of-domain-abuse-at-time-of-regist n=12,824,401 Malicious .net domains :: 61% detection rate on .net domains under a 0.35% false positive rate 2016 IMC/2016/zone-poisoning-the-how-and-where-of-non-secure-dns-dynamic-updates n=286,788,250 non-secure DNS dynamic updates :: 1,877 (0.065%) of the random sample and 587 (0.062%) of Alexa top 1 million domains 2017 IMC/2017/millions-of-targets-under-attack-a-macroscopic-characterization-of-the-dos-ecosy n=210,000,000 attack intensity and migration delay :: 98.6% of top-0.1% intensity websites migrated within six days 2017 IMC/2017/mission-accomplished-https-security-after-diginotar n=193,000,000 HPKP deployment and validity :: 6181 domains (0.02%); 86.0% use HPKP correctly 2017 IMC/2017/understanding-the-role-of-registrars-in-dnssec-deployment n=118,147,199 Registrar policy over time :: PCExtreme increased from 0.44% to 98.3% DNSSEC-enabled domains in 10 days. 2018 CCS/2018/minesweeper-an-in-depth-look-into-drive-by-cryptocurrency-mining-and-its-defense n=1,000,000 drive-by cryptocurrency mining :: 1,735 websites (0.18%) 2018 IMC/2018/a-first-joint-look-at-dos-attacks-and-bgp-blackholing-in-the-wild n=228,100,000 Services in blackholed prefixes :: 0.33% of websites, 0.40% of MX names, and 0.13% of NS names 2018 IMC/2018/digging-into-browser-based-crypto-mining n=137,000,000 browser-based mining prevalence :: less than 0.08% of probed sites 2018 IMC/2018/from-deletion-to-re-registration-in-zero-seconds-domain-registrar-behaviour-duri n=4,599,802 malicious re-registrations :: 0.4 % of domains re-registered with a delay of 0 s were labelled malicious; fewer than 0.5 % overall. 2018 IMC/2018/needle-in-a-haystack-tracking-down-elite-phishing-domains-in-the-wild n=657,663 squatting phishing pages :: 1,175 confirmed domains, approximately 0.2% of 657,663 domains. 2019 PETS/2019/a-quic-look-at-web-tracking n=1,000,000 QUIC server deployment :: 186 websites in the Alexa Top Million supported QUIC; 0.02% in the Top 1M. 2019 USENIX/2019/inadvertently-making-cyber-criminals-rich-a-comprehensive-study-of-cryptojacking n=48,948,669 Internet-wide cryptojacking prevalence :: 0.011% of all domains, or one in 9,090 websites 2019 USENIX/2019/protecting-accounts-from-credential-stuffing-with-password-breach-alerting n=746,853 breached-credential reuse by domain category :: 0.2–0.3% for finance and government; 3.6–6.3% for entertainment and adult sites 2019 WWW/2019/who-watches-the-watchmen-exploring-complaints-on-the-web n=1,054,248,823 Domain content replicas :: Fewer than 0.01% of domains have any replicas 2021 WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web n=1,000,000 Server-Sent Events :: 0.4% of the top thousand and 0.05% of websites in the top million 2022 IMC/2022/zdns-a-fast-dns-toolkit-for-internet-measurement n=234,531,389 CAA tag configuration :: 459 domains (0.04%) are configured with invalid tags 2022 USENIX/2022/a-large-scale-and-longitudinal-measurement-study-of-dkim-deployment n=3,627,871 Insecure l= tags :: 6,860 domains (0.3%) 2023 IMC/2023/ecn-with-quic-challenges-in-the-wild n=183,280,000 QUIC ECN validation :: less than 2% of QUIC hosts, providing less than 0.3% of HTTP/3 websites, pass validation 2026 PETS/2026/cryptographically-secured-domain-validation n=400,800,000 CAA and DNSSEC readiness :: 689.3K domains (0.34%) 2026 NDSS/2026/phishlang-a-real-time-fully-client-side-phishing-detection-framework-using-mobilebert n=42,700,000 phishing websites in live stream :: 25,796 phishing URLs out of 42.7M domains (0.057%) ========================================================================== 11. RESIDUES AND WHAT IS NOT COUNTED ========================================================================== web-unit papers with NO stated size, excluded from this page: 32 2.8% of 1153 papers in scope with >=1 detection prevalence string: 1114 99.4% prevalence strings in scope: 6203; naming a web unit: 1626 (26.2%); of those with a %: 837 (51.5%) The 48.5% of unit-naming prevalence strings that carry no parseable percentage, and the percentages that turn out to be false-positive rates or classifier metrics rather than site shares, are why no aggregate of prevalence-versus-n is published. An earlier draft cross-tabulated them; it was dropped. See the provenance page.
Quote-check script and its unedited output
- hms_quotecheck.py
#!/usr/bin/env python3 """Quote check for statistics:how_many_sites. Every quotation the page publishes is listed here with the paper it is attributed to. The check is done on WHITESPACE-COLLAPSED text, because a PDF line break inside a phrase makes an exact search fail for no good reason, and against paper.cols.txt (the repaired column reading order), never paper.norm.txt. A quote that fails the exact test is retried as a set of five-word windows, and the share of windows found is printed, so a two-column splice shows up as a partial rather than as an absence. python3 scripts/hms_quotecheck.py """ import os import re import sys ROOTS = ["/workspace/publications_dataset/data", "/workspace/publications_dataset"] ROOT = next(r for r in ROOTS if os.path.isdir(os.path.join(r, "fulltext"))) # (paper key, the phrase as the page prints it) QUOTES = [ ("WWW/2016/characterizing-long-tail-seo-spam-on-cloud-web-hosting-services", "to obtain the number of sampled cloud directories n = 500"), ("IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni", "we chose a sample size of 384 phishing sites for each entity"), ("IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni", "Our goal was to obtain a power of 0.95 at the significance level of 0.05 in a one-way independent ANOVA test"), ("IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni", "All sites ultimately delivered 100% uptime during deployment, thus we ended up with an effective sample size of 396 per experiment"), ("USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish", "To obtain a power of 0.95 at a p-value of 0.05, we initially assumed a medium effect size of 0.25"), ("PETS/2023/blocking-javascript-without-breaking-the-web-an-empirical-investigation", "manually inspected 383 websites, which is a statistically significant sample size for 100K websites"), ("USENIX/2023/autofr-automated-filter-rule-generation-for-adblocking", "we randomly selected 272 sites (a sample size out of 933 sites to get a confidence level of 95% with a 5% confidence interval)"), ("CCS/2018/minesweeper-an-in-depth-look-into-drive-by-cryptocurrency-mining-and-its-defense", "1,735 websites"), ("PETS/2019/a-quic-look-at-web-tracking", "We found 186 websites within the Alexa Top Million that support the QUIC protocol"), ("PETS/2019/a-quic-look-at-web-tracking", "Alexa Top 10 20.00% Alexa Top 100 21.00% Alexa Top 1K 8.10% Alexa Top 10K 1.69% Alexa Top 100K 0.19% Alexa Top 1M 0.02%"), ("PETS/2019/a-quic-look-at-web-tracking", "this share decreases for larger Top Alexa lists to only 0.0186% within the Alexa Top 1 Million"), ("WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web", "a random sample of 1000 sites from each of the top 10K, top 100K, and top 1M, for a total sample size of 4000 websites"), ("CCS/2019/lets-encrypt-an-automated-certificate-authority-to-encrypt-the-entire-web", "5% of the top 1K, 20%"), ("USENIX/2019/inadvertently-making-cyber-criminals-rich-a-comprehensive-study-of-cryptojacking", "meaning that one in every 9,090 websites is cryptojacking"), ("USENIX/2019/inadvertently-making-cyber-criminals-rich-a-comprehensive-study-of-cryptojacking", "In the Alexa Top 1M, 0.065% of the websites was actively cryptojacking, in this random sample only 0.011% of the websites, which is almost 6 times lower"), ("USENIX/2023/youve-got-report-measurement-and-security-implications-of-dmarc-reporting", "increasing trend of such misconfiguration as the ranking increases"), ("IMC/2021/who-you-gonna-call-an-empirical-evaluation-of-website-security-txt-deployment", "higher adoption rate for higher-ranked"), ("USENIX/2022/many-roads-lead-to-rome-how-packet-headers-influence-dns-censorship-measurement", "the likelihood of diminishing returns"), ("WWW/2025/digital-disparities-a-comparative-web-measurement-study-across-economic-boundari", "less popular websites were included to meet the target sample size of 10,000 websites per country"), ("CCS/2023/you-call-this-archaeology-evaluating-web-archives-for-reproducible-web-security", "Hit rates across popularity buckets"), ("IMC/2022/a-world-wide-view-of-browsing-the-world-wide-web", "a conservative sample size of"), ("IMC/2016/zone-poisoning-the-how-and-where-of-non-secure-dns-dynamic-updates", "We analyze a random sample of 2.9 million domains and the Alexa top 1 million domains and find that at least 1,877 (0.065%) and 587 (0.062%) of domains are vulnerable, respectively"), ("IMC/2021/who-you-gonna-call-an-empirical-evaluation-of-website-security-txt-deployment", "This percentage decreases to 3-4% for the top 10K sites, and only a percent for the top 100K"), ("IMC/2020/analyzing-third-party-service-dependencies-in-modern-web-services-have-we-learne", "The use of third party CAs is also higher (77% in top-100K) in less popular websites as compared to more popular (71% in top-100) websites"), ("USENIX/2023/youve-got-report-measurement-and-security-implications-of-dmarc-reporting", "almost 20% of the least 10K popular Alexa top-1M domains with DMARC that use external domains are misconfigured compared to about 10% for the most 10K popular domains"), ("WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web", "We find them in use on only 0.4% of the top thousand and 0.05% of websites in the top million"), ("CCS/2018/minesweeper-an-in-depth-look-into-drive-by-cryptocurrency-mining-and-its-defense", "1,735 websites"), ("USENIX/2016/internet-jones-and-the-raiders-of-the-lost-trackers-an-archaeological-study-of-w", "more popular trackers are better represented in Wayback data"), ("USENIX/2023/a-large-scale-measurement-of-website-login-policies", "0.44% of top 10K domains transmitted passwords in the clear, compared to 0.50% of top 100K domains and 0.62% of top 1M domains"), ("USENIX/2023/a-large-scale-measurement-of-website-login-policies", "33.6% of the domains in the top 10K and 26.7% of domains in the top 100K demonstrated rate limiting, compared to the 24.1% for domains in the top 1M"), # This paper's sentence is spliced by the two-column reading order; the page # quotes only the two intact fragments and paraphrases the join. ("IMC/2020/accept-the-risk-and-continue-measuring-the-long-tail-of-government-https-adoptio", "Though ranking does have an effect, overall valid https use"), ("IMC/2020/accept-the-risk-and-continue-measuring-the-long-tail-of-government-https-adoptio", "in the long tail dataset, at"), # Added after the citations review: the ANALYSED n of each rare-event row. ("CCS/2018/minesweeper-an-in-depth-look-into-drive-by-cryptocurrency-mining-and-its-defense", "we managed to crawl 991,513 of them"), ("IMC/2016/zone-poisoning-the-how-and-where-of-non-secure-dns-dynamic-updates", "From the total 286,788,250 unique domains in the set, we randomly sampled 1%"), ("IMC/2016/zone-poisoning-the-how-and-where-of-non-secure-dns-dynamic-updates", "1% Sample Alexa 1M Domains 2,865,393 947,823"), ("WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web", "we obtained data for a total of 88.1% of websites in the top million"), ("USENIX/2019/inadvertently-making-cyber-criminals-rich-a-comprehensive-study-of-cryptojacking", "Total 48,948,669 5,190 (0.011)%"), ("PETS/2019/a-quic-look-at-web-tracking", "to send a client hello message to UDP port 443 of each Alexa-listed host"), ("CCS/2011/fashion-crimes-trending-term-exploitation-on-the-web", "we selected a statistically significant (95% confidence interval) random sample of 363 websites for manual inspection"), # Damaged by column order: the page quotes the intact run only. ("WWW/2017/security-implications-of-redirection-trail-in-popular-websites-worldwide", "domized sample of 2,000 websites (margin of error = 2."), ("IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking", "500 random combinations of different sample sizes"), ("IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking", "This is only 0.05% of our sample, compared to 0.2% of Alexa Top 10K domains, indicating that more popular websites are more"), ] def collapse(s: str) -> str: s = s.replace("’", "'").replace("‘", "'") s = s.replace("“", '"').replace("”", '"') s = s.replace("–", "-").replace("—", "-").replace("−", "-") return re.sub(r"\s+", " ", s) def dehyphenate(s: str) -> str: """Rejoin a word broken across a PDF line by a hyphen. "ran-\ndomized" collapses to "ran- domized", which no exact search for "randomized" will match. A reviewer found this: a quote that IS in the paper would have been reported MISSING. Applied only as a second pass, and reported as such, because it also joins genuine hyphenated compounds. """ return re.sub(r"(\w)-\s+(\w)", r"\1\2", s) def load(k: str) -> str: venue, year, slug = k.split("/", 2) p = os.path.join(ROOT, "fulltext", year, venue, slug, "paper.cols.txt") with open(p, encoding="utf8", errors="replace") as fh: return collapse(fh.read()) # ---------------------------------------------------------------- coverage # Curating QUOTES by hand failed twice: two footnoted quotes and one McDonald # quote were printed on the page and never checked, while the docstring and the # provenance page both claimed "every quotation". So the list is now AUDITED # against the pages rather than trusted: every //"..."// span on either page # must be covered by some QUOTES entry, and an uncovered one fails the run. # Only the CONTENT page is audited. The provenance page's //"..."// spans are # mostly quotations of this page's own earlier drafts and of extraction records # ("the same crawler, the same week", "96/(p.4)"), which have no paper to check # against; the paper quotes it repeats are already covered here. PAGES = ["pages/statistics_how_many_sites.txt"] # A DokuWiki italic run (//X//) inside a quotation would otherwise let the span # run from one quotation into the next, so reject any match containing //. WIKI_QUOTE = re.compile(r'//"([^"]+?)"//') def page_quotes(): seen = [] for rel in PAGES: path = os.path.join("/workspace/artifacts/wiki", rel) if not os.path.exists(path): continue text = open(path, encoding="utf8").read() # Strip the wiki's downloadable-code blocks: those hold this script's # own output, not page prose. The tag name is assembled rather than # written literally, because this file is itself published inside one # of those blocks and a literal closing tag here would end it early. tag = "fi" + "le" text = re.sub(rf"<{tag}[^>]*>.*?</{tag}>", "", text, flags=re.S) for mm in WIKI_QUOTE.finditer(text): if "//" in mm.group(1): continue seen.append((rel, mm.group(1))) return seen def audit_coverage(): listed = [collapse(q).lower() for _, q in QUOTES] missing = [] for rel, q in page_quotes(): qq = collapse(q).lower() if any(qq in l or l in qq for l in listed): continue missing.append((rel, q)) return missing uncovered = audit_coverage() if uncovered: print("!! QUOTES on a page that this script does not check:") for rel, q in uncovered: print(f" {rel}: {q}") print() fails = 0 for k, q in QUOTES: txt = load(k) qq = collapse(q) if qq.lower() in txt.lower(): print(f"EXACT {k}\n {q}") continue if collapse(dehyphenate(qq)).lower() in dehyphenate(txt).lower(): print(f"EXACT-DH {k} (matched only after rejoining a hyphenated line break)\n {q}") continue words = qq.split() wins = [" ".join(words[i:i + 5]) for i in range(max(1, len(words) - 4))] found = sum(1 for w in wins if w.lower() in txt.lower()) verdict = "PARTIAL" if found else "MISSING" if not found: fails += 1 print(f"{verdict} {k} ({found}/{len(wins)} five-word windows)\n {q}") distinct = len({(k, collapse(q).lower()) for k, q in QUOTES}) print(f"\n{len(QUOTES)} entries ({distinct} distinct), {fails} not located at all.") print(f"page-quote coverage audit: {len(uncovered)} uncovered") sys.exit(1 if (fails or uncovered) else 0)
- hms_quotecheck-output.txt
PARTIAL WWW/2016/characterizing-long-tail-seo-spam-on-cloud-web-hosting-services (4/7 five-word windows) to obtain the number of sampled cloud directories n = 500 EXACT IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni we chose a sample size of 384 phishing sites for each entity EXACT IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni Our goal was to obtain a power of 0.95 at the significance level of 0.05 in a one-way independent ANOVA test EXACT IEEE-SP/2019/phishfarm-a-scalable-framework-for-measuring-the-effectiveness-of-evasion-techni All sites ultimately delivered 100% uptime during deployment, thus we ended up with an effective sample size of 396 per experiment EXACT USENIX/2020/phishtime-continuous-longitudinal-measurement-of-the-effectiveness-of-anti-phish To obtain a power of 0.95 at a p-value of 0.05, we initially assumed a medium effect size of 0.25 EXACT PETS/2023/blocking-javascript-without-breaking-the-web-an-empirical-investigation manually inspected 383 websites, which is a statistically significant sample size for 100K websites EXACT USENIX/2023/autofr-automated-filter-rule-generation-for-adblocking we randomly selected 272 sites (a sample size out of 933 sites to get a confidence level of 95% with a 5% confidence interval) EXACT CCS/2018/minesweeper-an-in-depth-look-into-drive-by-cryptocurrency-mining-and-its-defense 1,735 websites EXACT PETS/2019/a-quic-look-at-web-tracking We found 186 websites within the Alexa Top Million that support the QUIC protocol EXACT PETS/2019/a-quic-look-at-web-tracking Alexa Top 10 20.00% Alexa Top 100 21.00% Alexa Top 1K 8.10% Alexa Top 10K 1.69% Alexa Top 100K 0.19% Alexa Top 1M 0.02% EXACT PETS/2019/a-quic-look-at-web-tracking this share decreases for larger Top Alexa lists to only 0.0186% within the Alexa Top 1 Million EXACT WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web a random sample of 1000 sites from each of the top 10K, top 100K, and top 1M, for a total sample size of 4000 websites EXACT CCS/2019/lets-encrypt-an-automated-certificate-authority-to-encrypt-the-entire-web 5% of the top 1K, 20% EXACT USENIX/2019/inadvertently-making-cyber-criminals-rich-a-comprehensive-study-of-cryptojacking meaning that one in every 9,090 websites is cryptojacking EXACT USENIX/2019/inadvertently-making-cyber-criminals-rich-a-comprehensive-study-of-cryptojacking In the Alexa Top 1M, 0.065% of the websites was actively cryptojacking, in this random sample only 0.011% of the websites, which is almost 6 times lower EXACT USENIX/2023/youve-got-report-measurement-and-security-implications-of-dmarc-reporting increasing trend of such misconfiguration as the ranking increases EXACT IMC/2021/who-you-gonna-call-an-empirical-evaluation-of-website-security-txt-deployment higher adoption rate for higher-ranked EXACT USENIX/2022/many-roads-lead-to-rome-how-packet-headers-influence-dns-censorship-measurement the likelihood of diminishing returns EXACT WWW/2025/digital-disparities-a-comparative-web-measurement-study-across-economic-boundari less popular websites were included to meet the target sample size of 10,000 websites per country EXACT CCS/2023/you-call-this-archaeology-evaluating-web-archives-for-reproducible-web-security Hit rates across popularity buckets EXACT IMC/2022/a-world-wide-view-of-browsing-the-world-wide-web a conservative sample size of EXACT IMC/2016/zone-poisoning-the-how-and-where-of-non-secure-dns-dynamic-updates We analyze a random sample of 2.9 million domains and the Alexa top 1 million domains and find that at least 1,877 (0.065%) and 587 (0.062%) of domains are vulnerable, respectively PARTIAL IMC/2021/who-you-gonna-call-an-empirical-evaluation-of-website-security-txt-deployment (8/14 five-word windows) This percentage decreases to 3-4% for the top 10K sites, and only a percent for the top 100K EXACT IMC/2020/analyzing-third-party-service-dependencies-in-modern-web-services-have-we-learne The use of third party CAs is also higher (77% in top-100K) in less popular websites as compared to more popular (71% in top-100) websites EXACT USENIX/2023/youve-got-report-measurement-and-security-implications-of-dmarc-reporting almost 20% of the least 10K popular Alexa top-1M domains with DMARC that use external domains are misconfigured compared to about 10% for the most 10K popular domains EXACT WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web We find them in use on only 0.4% of the top thousand and 0.05% of websites in the top million EXACT CCS/2018/minesweeper-an-in-depth-look-into-drive-by-cryptocurrency-mining-and-its-defense 1,735 websites EXACT USENIX/2016/internet-jones-and-the-raiders-of-the-lost-trackers-an-archaeological-study-of-w more popular trackers are better represented in Wayback data EXACT USENIX/2023/a-large-scale-measurement-of-website-login-policies 0.44% of top 10K domains transmitted passwords in the clear, compared to 0.50% of top 100K domains and 0.62% of top 1M domains EXACT USENIX/2023/a-large-scale-measurement-of-website-login-policies 33.6% of the domains in the top 10K and 26.7% of domains in the top 100K demonstrated rate limiting, compared to the 24.1% for domains in the top 1M EXACT IMC/2020/accept-the-risk-and-continue-measuring-the-long-tail-of-government-https-adoptio Though ranking does have an effect, overall valid https use EXACT IMC/2020/accept-the-risk-and-continue-measuring-the-long-tail-of-government-https-adoptio in the long tail dataset, at EXACT CCS/2018/minesweeper-an-in-depth-look-into-drive-by-cryptocurrency-mining-and-its-defense we managed to crawl 991,513 of them EXACT IMC/2016/zone-poisoning-the-how-and-where-of-non-secure-dns-dynamic-updates From the total 286,788,250 unique domains in the set, we randomly sampled 1% EXACT IMC/2016/zone-poisoning-the-how-and-where-of-non-secure-dns-dynamic-updates 1% Sample Alexa 1M Domains 2,865,393 947,823 EXACT WWW/2021/websocket-adoption-and-the-landscape-of-the-real-time-web we obtained data for a total of 88.1% of websites in the top million EXACT USENIX/2019/inadvertently-making-cyber-criminals-rich-a-comprehensive-study-of-cryptojacking Total 48,948,669 5,190 (0.011)% EXACT PETS/2019/a-quic-look-at-web-tracking to send a client hello message to UDP port 443 of each Alexa-listed host EXACT CCS/2011/fashion-crimes-trending-term-exploitation-on-the-web we selected a statistically significant (95% confidence interval) random sample of 363 websites for manual inspection EXACT WWW/2017/security-implications-of-redirection-trail-in-popular-websites-worldwide domized sample of 2,000 websites (margin of error = 2. EXACT IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking 500 random combinations of different sample sizes EXACT IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking This is only 0.05% of our sample, compared to 0.2% of Alexa Top 10K domains, indicating that more popular websites are more 42 entries (41 distinct), 0 not located at all. page-quote coverage audit: 0 uncovered
Review log
Four reviewers, all told explicitly that the author's context may not be exhaustive, all handed the page text, the report script, its unedited output and these notes. The three focused passes ran in parallel first; the generic pass ran after their findings were acted on. Rejections are listed as well as fixes, because they are the only record of whether a reviewer slot earns its place.
Author's own pass, before the reviewers
Eight things found by re-reading rather than by any check, listed because they are the class of error no automated check can see:
| Found | Action |
|---|---|
| Two author attributions written in a footnote from memory were both wrong (see the section above) | Fixed; both papers now cited properly |
| “the same crawler, the same week” of Sy et al. — the paper does not say the measurements were a week apart | Changed to “the same list” |
| Hanley and Lippman-Hand described as “four pages”; it is 1743–1745 | Changed to three |
| “two orders of magnitude larger” for a 97%-accurate labeller against a ±0.43 pp interval; it is about seven times | Rewritten with the arithmetic shown |
| n for a ±20% relative interval described as “96/(p·4)”; it is roughly 96/p | Fixed |
| “the most common crawl size in this literature” for n = 10,000; 1,000,000 is the mode over the whole population | Changed to “the second most common value in this population after a million” |
| Murley et al.'s stratified-design quote truncated so that the arithmetic in it (three samples of 1,000 summing to 4,000) did not add up | Full quote restored |
| The attrition-rises-at-the-tail bullet asserted as fact what the corpus does not support | Rewritten as an explicit mechanism argument, with the failed probe named |
''sonnet'' — figures against the script
Re-ran all three scripts, diffed the committed outputs, recomputed the arithmetic tables independently in Python, and re-derived the design:sampling cross-check.
| Finding | Verdict |
|---|---|
| The currency table's “6–9% stratify since 2018” is wrong at the low end: 2018–2021 is 5.0%, and this page's claim that the figure had been “diffed with no change” was itself untrue | Accepted. The only substantive defect it found, and a real one. Both the page row and this page's certification are rewritten above, with all five bucket shares given rather than a range |
| Committed outputs byte-identical to a fresh run; all three scripts reproduce | Confirmation, no action |
| Every arithmetic table reproduces from the stated formula, independently recomputed: Wald half-widths, ceil(ln(0.05)/ln(1−p)), ceil(Z²(1−p)/(0.04p)), 1+(m−1)ρ, the 50.95× deflation and the ±0.305 pp deflated half-width | Confirmation, no action |
| Every corpus figure checked matches the script, including the disclosed .5-rounding | Confirmation, no action |
| Independently recomputed the mode restricted to the 674 crawling papers: 10,000 with 67 papers against 1,000,000 with 62 | Noted. The page's wording is about the full 1,121-paper population, where 1,000,000 is the mode, and is left as it stands |
''sonnet'' — citations and quotes
Read every {[key]} marker against a fresh export of the live bibliography, re-verified the two external references against JAMA/PMC full text as well as Crossref and PubMed, and re-checked twelve of the fourteen rank-band rows against paper.cols.txt by hand.
| Finding | Verdict |
|---|---|
| The rare-event table's n for [1Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] is 286,788,250 but the 0.065% rate comes from a 2,865,393-domain 1% sample of that corpus — an error of two orders of magnitude | Accepted. Verified in the source; the table now carries drawn and analysed n separately and the correction is documented above |
| [2Konoth, Radhesh Krishnan; Vineti, Emanuele; Moonsamy, Veelasha; Lindorfer, Martina; Kruegel, Christopher; Bos, Herbert; Vigna, Giovanni (2018): "MineSweeper: An In-depth Look into Drive-by Cryptocurrency Mining and Its Defense", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] is paired with the drawn n of 1,000,000; the paper analysed 991,513 — the exact drawn-versus-analysed error the page tells readers not to make | Accepted. Fixed in the table and in the “What to Read First” entry. Checking the other three rows for the same defect found a fourth: [5Murley, Paul; Ma, Zane; Mason, Joshua; Bailey, Michael D.; Kharraz, Amin (2021): "WebSocket Adoption and the Landscape of the Real-Time Web", in: Proceedings of the ACM Web Conference. (DOI)] analysed 88.1% of the top million |
hms_quotecheck.py does not check every quotation the page prints: the two footnoted quotes were not in its list, and its normaliser cannot rejoin a hyphenated line break, so a quote that IS present would have been reported MISSING | Accepted. Both quotes added before the review returned; the de-hyphenation pass added after it, reported separately as EXACT-DH. The check is now 40 quotes |
| The rank-band table files [1Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]'s random 2.9M sample under a column headed “Deeper in the list”, which it is not — it is an independent unranked frame | Accepted as a wording fix. The cell now carries a footnote saying so. The row stays, because a flat result is the hardest of the three families to find |
[12Ikram, Muhammad; Masood, Rahat; Tyson, Gareth; Kaafar, Mohamed Ali; Loizon, Noha; Ensafi, Roya (2019): "The Chain of Implicit Trust: An Analysis of the Web Third-party Resources Loading", in: Proceedings of the ACM Web Conference. (DOI)] and [13Singanamalla, Sudheesh; Jang, Esther Han Beol; Anderson, Richard; Kohno, Tadayoshi; Heimerl, Kurtis (2020): "Accept the Risk and Continue: Measuring the Long Tail of Government https Adoption", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] will not resolve — absent from the bibliography and from the additions file | Rejected: false positive, caused by the author. The reviewer was pointed at /tmp/newbib.bib, a superseded 12-entry scratch file, rather than at scripts/bib_additions_how_many_sites.bib, which contains both. Both keys resolve against the file that was actually appended. The author's error, not the reviewer's — a review brief that names the wrong artefact produces a confident finding about nothing, and this one cost a real reviewer slot |
| Both external references verified against primary full text, not only Crossref: the rule of three and the design-effect formula each say what the page attributes to them | Confirmation, no action |
| All seven bibliography additions the reviewer could see check out on authors, year, venue and DOI; no duplicate keys beyond the five already excluded | Confirmation, no action |
| Twelve of fourteen rank-band rows re-verified with correct head/tail assignment; the three column splices confirmed genuine and correctly handled | Confirmation, no action |
''sonnet'' — external currency, as of 2026-09-02
Fetched rather than recalled: alexa.com and alexa.com/topsites, tranco-list.eu and its /latest_list endpoint, developer.chrome.com/docs/crux and its methodology page, the zakird/crux-top-lists repository, the IETF datatracker entry for RFC 9116, and both external DOIs through Crossref and doi.org.
| Finding | Verdict |
|---|---|
| No factual error in any externally-verifiable claim | Confirmation |
| Alexa still retired: both alexa.com and alexa.com/topsites 302 to the unrelated Amazon voice-assistant product | Confirmation |
| Tranco live and generating lists as of 2026-09-02; since August 2023 its default inputs include CrUX and Cloudflare Radar | Accepted as an addition. The CrUX row of the currency table said “under-used”; it now says CrUX is named directly by under 10% of papers but reaches many more indirectly through Tranco, and points at Website selection for the list composition |
| CrUX still published monthly; the “most accurate list” finding is not superseded | Confirmation |
| Both external DOIs resolve and match their cited metadata. JAMA's own site returns 403 to a bot; Crossref is the authoritative check | Confirmation |
| No literal URL anywhere on either page, so no link rot to check beyond the DOIs | Confirmation |
RFC 9116 (security.txt) still current, not obsoleted | Confirmation |
| No 2025–2026 popularity list or dataset worth adding. Prefix Top Lists Reloaded (PAM 2026) reconstructs prefix and AS rankings from Tranco, Umbrella, CrUX and Radar rather than replacing them, and PAM is outside the seven venues | Rejected as an addition, on the reviewer's own reasoning |
| No primary source published in 2025–2026 on choosing web-measurement sample size. Checked Measuring What the Crawler Sees (arXiv 2607.13636), which models longitudinal URL discovery rather than cross-sectional sizing, and SoK: State of the Krawlers (USENIX Security 2024), which is about crawler design | Accepted as a confirmation of absence. The page's Open Questions box stands unchanged; the searched-and-found-nothing result is recorded here so the next run does not repeat the search |
''fable'' — generic
No checklist. Read both pages end to end against the neighbours. 19 findings; 14 accepted, 4 rejected or deferred, 1 already fixed.
| Finding | Verdict |
|---|---|
| The bolded rank-tail rule contradicts the table above it. “Decided by its host, its registrar or its CMS vendor, expect none” — but the rises-with-rank family is entirely host-side defaults and has the steepest gradients on the page | Accepted. The rule is deleted. The three families are kept as a description of what these papers found, with an explicit sentence saying they do not predict and that no test on the corpus separates them. The advice that replaces it is Murley et al.'s: measure two depths yourself |
| “Plateau” is not what the table shows — the median fell 42,542 to 20,000 between consecutive buckets, a halving described as staying put; and “you will not be penalised for meeting it” is an opinion the page's own Open Questions calls unmeasurable | Accepted. Rewritten as “rose by roughly eight-fold to 2018-2021 and then fell back by half”, with the two available readings named and neither chosen. The penalty sentence is deleted; its twin already lives in Open Questions |
| The page gives three different answers to its own question — “a few thousand”, “5,000 to 50,000”, and the currency table's “5k-50k” — and the 5k-50k was fitted to the field's median rather than derived | Accepted, and it produced the best paragraph on the page. Section A now derives the number from the labeller error: 97% accurate leaves +-3 pp, “comfortably below” at a factor of five is a 0.6 pp target, giving n = 26,700 at p = 0.5 and 5,100 at p = 0.05; at 99% accuracy, 240,000 and 46,000. The range is now an output, and the page says explicitly that its agreement with the field's median is a coincidence rather than evidence |
| “Every one of the sizing calculations in the next section is about a subsample” — two of seven are not | Accepted. “Five of the seven” |
| PhishFarm does not report an achieved effect; it reports an effective sample size of 396. Only PhishTime reports 0.36 | Accepted, verified in the source. The entry now attributes the effective-sample-size figure to PhishFarm and the achieved effect to PhishTime |
| Absolute negatives are drawn from a probe whose recall the page itself calls a floor | Accepted. “No paper in this corpus derives…” became “we found no paper”, with both routes named. The report script now prints all nine in-population power-analysis papers rather than only the six that also crawled, so the second route is visible in the audit trail |
| “12.5% still states exactly 1,000,000” uses a 2010-2026 denominator for a claim about persistence, and conflates the dead list with the surviving number | Accepted, and the finding got better for it. By bucket the figure is 7.1 / 15.4 / 11.5 / 12.6 / 13.8%, and of the 26 papers stating exactly 1,000,000 in 2025-2026, 25 draw a Tranco top 1M and one an Alexa list. The frame was replaced and the number was kept — which is a sharper claim than the one it replaced |
| Scheitle misquoted: “tens of percentage points on IPv6 and CDN adoption”. IPv6 is 11-13% against 4%, and CDN is a ratio | Accepted. Both figures now given as the paper gives them |
| Three weak rows in the rank-band table: Lerner et al. rank trackers, not sites; McDonald et al.'s 0.05% is 3 domains out of 6,180 that a proxy network declined to fetch, against an independent frame; Ikram et al. and Hantke et al. carry no second number | Accepted. Lerner and McDonald are dropped from the table and counted among the nine per-band papers that are not tabled. Ikram and Hantke stay, marked as directional. The table is now twelve rows over eleven papers, ten of which give a figure at each of two depths, and the page says so |
| “Explains far more of the variance” — no variance decomposition exists, and blocklist (cheap) sits with supervised ML | Accepted. Now “a ranking consistent with per-site cost rather than a measurement of it”, with the non-monotone row named |
| The finite-population footnote is wrong in a way that hides the useful fact | Accepted, and this was the single most useful finding for a reader. 385 with the finite-population correction gives 383 at N = 100,000, 363 at 6,558 and 272 at 933 — which are, exactly, the three numbers the three papers report. The footnote became a four-row table showing that all three are the same calculation |
| Small overstatements: “present continuously from 2011 to 2023” for five papers; “the trend is up and slow” for 5.0 to 8.9 to 6.3 | Accepted. Both softened to what the numbers support |
| Design:Sampling hedges “coverage error typically dominates” as an explicit judgement; this page bolded the same claim as fact | Accepted. The claim now carries the same footnote hedge, in the same words, so the two pages agree about how confident they are |
| Nowhere says where p comes from for the rare-event calculation | Accepted. A paragraph on pilots now closes the loop: crawl 10,000, and either k/10,000 gives you p or the rule of three gives you a 0.03% bound, and the table tells you what the full crawl needs. It is also the answer to the “you cannot know p in advance” objection that is otherwise why nobody does this |
| Structure: swap “The Number Is a Convention” and “Three Questions” so the page answers first | Rejected. The takeaway box already answers first, in three bullets, above everything else. The convention section is the page's evidence that the number is unexamined, which is what licenses the rest; moving it after the answer would make the page argue with a reader who has not yet been shown there is a problem. Recorded here because it is a defensible call the other way |
| The rank-tail section is the longest and overlaps Statistics:Biases and Design:Sampling | Partly accepted. The predictive rule is cut and two weak rows removed, which is most of the overlap. The remainder is the per-band table, which neither neighbour has |
| The currency table restates other sections with “see this page” as evidence in four rows | Accepted, trimmed |
| “Sizing, in one table” repeats the intro's numbers | Rejected. It is the section a reader lands on from a search or a link, and a summary table that cannot stand alone is worse than a repeated figure |
| The provenance page claims the bibliography and the neighbour edits as done | Accepted. They are done, but they were not when the reviewer read the file. The publish order is now stated in the run table, and this page is saved last so that every claim it makes about what was published is true when it is saved |
What the review layer was worth
Four reviewers, seven substantive defects between them, and one false positive that was the author's fault for naming the wrong file in a brief.
| Pass | Substantive defects found | Would the page have been wrong without it? |
|---|---|---|
| author's own re-read | 8, including two invented author attributions | Yes. Two fabricated surnames would have shipped |
sonnet figures | 1 (the stratification range, plus this page's false certification of it) | Yes, and it was the only pass that would have caught it, because it re-ran the neighbouring script instead of trusting the prose |
sonnet figures, re-check | 2 (the de-hyphenation claim being false; the drawn-versus-analysed tip miscounting its own rows) | Yes. Both were introduced by fixing something else, which is the argument for re-running a reviewer whose findings you acted on |
sonnet citations | 4 (the 100x Korczyński n; the Konoth drawn-versus-analysed n; quote-check coverage; a mislabelled column) | Yes. The Korczyński error was the largest single defect in the run and no script could have found it — it needed someone to read the paper |
sonnet citations, re-check | 3 (a spurious diacritic in an author name; a third uncovered quote; a count that did not match its own table) | Marginal individually, but the third uncovered quote is what forced the coverage guard to be written rather than the list patched again |
sonnet currency | 0 defects, 1 accepted addition, 1 rejected addition, and one useful confirmed absence | No, but its value here is the recorded negative: the next run does not have to search again for a 2025-2026 sizing source |
fable generic | 19 findings, of which 14 accepted | Yes. It was the only pass that noticed the page gave three different answers to its own question, that a bolded predictive rule contradicted the table above it, and that “plateau” described a median that had halved |
The pattern worth carrying forward: the focused passes caught wrong numbers, and the generic pass caught wrong arguments. Every defect that changed what the page advises — as opposed to what it reports — came from the unstructured pass or from re-reading. A review layer of only the three focused reviewers would have shipped a numerically correct page that answered its own question three different ways.
Related
- How many sites — the content page this log is behind.
- Corpus — corpus-level caveats: venues, funnel, provisional years.
- sampling — the neighbouring draw-side log, and the source of the tuple-level size figures reconciled above.
- hypothesis_testing — the source of the ICC sweep whose ρ values section C reuses.
References
- [1]
- Korczyński, Maciej; Król, Michal; van Eeten, Michel (2016): "Zone Poisoning: The How and Where of Non-Secure DNS Dynamic Updates", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [2]
- Konoth, Radhesh Krishnan; Vineti, Emanuele; Moonsamy, Veelasha; Lindorfer, Martina; Kruegel, Christopher; Bos, Herbert; Vigna, Giovanni (2018): "MineSweeper: An In-depth Look into Drive-by Cryptocurrency Mining and Its Defense", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [3]
- Bijmans, Hugo L.J.; Booij, Tim M.; Doerr, Christian (2019): "Inadvertently Making Cyber Criminals Rich: A Comprehensive Study of Cryptojacking Campaigns at Internet Scale", in: Proceedings of the USENIX Security Symposium. (Link)
- [4]
- Sy, Erik; Burkert, Christian; Federrath, Hannes; Fischer, Mathias (2019): "A QUIC Look at Web Tracking", Proceedings on Privacy Enhancing Technologies 2019(3):255-266. (DOI)
- [5]
- Murley, Paul; Ma, Zane; Mason, Joshua; Bailey, Michael D.; Kharraz, Amin (2021): "WebSocket Adoption and the Landscape of the Real-Time Web", in: Proceedings of the ACM Web Conference. (DOI)
- [6]
- Moore, Tyler; Leontiadis, Nektarios; Christin, Nicolas (2011): "Fashion Crimes: Trending-Term Exploitation on the Web", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [7]
- Chang, Li; Hsiao, Hsu-Chun; Jeng, Wei; Kim, Tiffany Hyun-Jin; Lin, Wei-Hsi (2017): "Security Implications of Redirection Trail in Popular Websites Worldwide", in: Proceedings of the ACM Web Conference. (DOI)
- [8]
- Liao, Xiaojing; Liu, Chang; McCoy, Damon; Shi, Elaine; Hao, Shuang; Beyah, Raheem A. (2016): "Characterizing Long-tail SEO Spam on Cloud Web Hosting Services", in: Proceedings of the ACM Web Conference. (DOI)
- [9]
- Tang, Jenny; Bauer, Lujo; Christin, Nicolas (2025): "Misuse, Misreporting, Misinterpretation of Statistical Methods in Usable Privacy and Security Papers", in: Proceedings of the Twenty-First Symposium on Usable Privacy and Security (SOUPS), pp. 475-493. USENIX Association. (Link)
- [10]
- Hanley, James A.; Lippman-Hand, Abby (1983): "If Nothing Goes Wrong, Is Everything All Right? Interpreting Zero Numerators", JAMA 249(13):1743-1745. (DOI)
- [11]
- Killip, Shersten; Mahfoud, Ziyad; Pearce, Kevin (2004): "What Is an Intracluster Correlation Coefficient? Crucial Concepts for Primary Care Researchers", The Annals of Family Medicine 2(3):204-208. (DOI)
- [12]
- Ikram, Muhammad; Masood, Rahat; Tyson, Gareth; Kaafar, Mohamed Ali; Loizon, Noha; Ensafi, Roya (2019): "The Chain of Implicit Trust: An Analysis of the Web Third-party Resources Loading", in: Proceedings of the ACM Web Conference. (DOI)
- [13]
- Singanamalla, Sudheesh; Jang, Esther Han Beol; Anderson, Richard; Kohno, Tadayoshi; Heimerl, Kurtis (2020): "Accept the Risk and Continue: Measuring the Long Tail of Government https Adoption", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
