User Tools

Site Tools


design:sampling

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

design:sampling [2026/08/12 21:28] – New page: sampling designs for web measurement — frame vs target population, top-n as a census of the head, rank-stratified sampling, sample size, unit (site/page/visit), attrition, and versioning the draw. Figures from 1,153 corpus papers that sampled a karel.kubicek.claudedesign:sampling [2026/08/12 21:40] (current) – Generic review pass: soften three overstatements in the lead, attach the Alexa unknown-provenance claim to the 34 undated papers rather than all 66, add sections on purposive sampling and on turning a list row into a URL, date the Scheitle figures, hedge karel.kubicek.claude
Line 7: Line 7:
 The short version, from 1,153 papers in our corpus of seven security and privacy venues that drew a population of websites, domains or web pages (see [[#Use in Publications]]): The short version, from 1,153 papers in our corpus of seven security and privacy venues that drew a population of websites, domains or web pages (see [[#Use in Publications]]):
  
-  * **57.6% take a top-//n//**, and the share is still rising. Only **6.0%** stratify by rank. +  * **57.6% take a top-//n//.** It has not been supersededit has consolidated — the share has sat between 58% and 60% since 2018, on the year-buckets that are complete. Only **6.0%** stratify at all, by rank or by anything else
-  * **99.1% name the list, 97.2% give a size, and 51.7% say which version** — so half of the field's samples cannot be redrawn.+  * **99.1% name the list, 97.2% give a size, and 51.7% say which version** — so for half of them a reader cannot tell which draw was made. About a third of the rest publish the drawn list instead, which closes the gap; the remainder is genuinely unrecoverable.
   * The list itself changed under the field: papers naming Alexa fell from **47.1% of the population in 2021 to 8.3% in 2025**, while Tranco rose from 31.0% to 56.4% over the same years. **66 papers published in 2023–2026 still sample from "the Alexa top //n//"**, a list that has not been published since 2022, and **51.5% of them give no version at all**.   * The list itself changed under the field: papers naming Alexa fell from **47.1% of the population in 2021 to 8.3% in 2025**, while Tranco rose from 31.0% to 56.4% over the same years. **66 papers published in 2023–2026 still sample from "the Alexa top //n//"**, a list that has not been published since 2022, and **51.5% of them give no version at all**.
  
Line 39: Line 39:
 That is **8.3% of the drawn sample gone before analysis**, and the paper's percentages mean something precise because it says so. A 2025 example at the other scale: {[annamalai2025_beyond]}((Cited here for its reporting of attrition; see the paper for its actual subject.)) reports that "the automated crawler failed to visit 15 out of the 3K websites (≈ 1%)", and attributes most of them to bot detection — hedged as "we believe", on the evidence that 86.7% of those sites returned 4XX errors. Even at 1%, that attrition is itself a finding, because it is not random with respect to what is being measured. That is **8.3% of the drawn sample gone before analysis**, and the paper's percentages mean something precise because it says so. A 2025 example at the other scale: {[annamalai2025_beyond]}((Cited here for its reporting of attrition; see the paper for its actual subject.)) reports that "the automated crawler failed to visit 15 out of the 3K websites (≈ 1%)", and attributes most of them to bot detection — hedged as "we believe", on the evidence that 86.7% of those sites returned 4XX errors. Even at 1%, that attrition is itself a finding, because it is not random with respect to what is being measured.
  
-Attrition is not random. Sites that block datacenter IPs, sites behind bot management, and sites in the tail are all more likely to drop out, and all three correlate with the things privacy and security papers measure. Report both denominators.+Attrition is not random. Sites that block datacenter IPs, sites behind bot management, and sites in the tail are all more likely to drop out, and all three plausibly correlate with the things privacy and security papers measure.((Plausibly, not measurably: we found no study that characterises which sites drop out of a crawl and how that biases the result. Jueckstock et al. {[jueckstock2021_realistic]} show that the vantage point changes what a crawl sees, and {[annamalai2025_beyond]} attributes its own failures to bot detection, but neither measures the composition of the lost set. It is an open question — see [[#Open Questions]].)) Report both denominators.
  
 ===== The Methods, and What Each Licenses ===== ===== The Methods, and What Each Licenses =====
Line 61: Line 61:
 The dominant design in the field is "we crawled the Tranco top 10k". It has real advantages — cheap, comparable across papers, weighted towards sites that matter to users — and one property that is constantly forgotten: **it involves no sampling at all**. There is no randomness, so a confidence interval computed on it is answering a question nobody asked. If 6.4% of your top-10k sites do //X//, that 6.4% is an exact census figure for those 10,000 sites, and the uncertainty that matters is entirely about how unlike the rest of the web those 10,000 sites are. The dominant design in the field is "we crawled the Tranco top 10k". It has real advantages — cheap, comparable across papers, weighted towards sites that matter to users — and one property that is constantly forgotten: **it involves no sampling at all**. There is no randomness, so a confidence interval computed on it is answering a question nobody asked. If 6.4% of your top-10k sites do //X//, that 6.4% is an exact census figure for those 10,000 sites, and the uncertainty that matters is entirely about how unlike the rest of the web those 10,000 sites are.
  
-They are, measurably, very unlike it. Scheitle et al. {[scheitle2018_long]} compared list domains against the general ''com''/''net''/''org'' population and found the head is a technological outlier: HTTP/2 adoption reached **26.6% for the Alexa Top 1M against 7.84% for com/net/org overall**, IPv6 enablement 11–13% against about 4%, and CDN use on the Top 1M lists exceeded the general population "by at least a factor of two". Any measurement of adoption, modernity or professionalism run on a top list is biased upward, and the size of the bias is of the same order as the effect most papers report.+They are, measurably, very unlike it. Scheitle et al. {[scheitle2018_long]} compared list domains against the general ''com''/''net''/''org'' population and found the head is a technological outlier. Their 2018 figures: HTTP/2 adoption **26.6% for the Alexa Top 1M against 7.84% for com/net/org overall**, IPv6 enablement 11–13% against about 4%, and CDN use on the Top 1M lists exceeding the general population "by at least a factor of two". The absolute numbers are eight years old and every one of them has moved since; the //gap// is the durable finding. Any measurement of adoption, modernity or professionalism run on a top list is biased upward, and on these three indicators the bias was a factor of two to three — large enough to swamp the effect a typical paper reports, though we know of no study that compares the two systematically.
  
 The head is also a small part of the web but most of the //browsing//. Ruth et al.'s study of Chrome telemetry {[ruth2022_world]} finds that "the top million sites capture over 95% of all page loads and time spent online, but they do so extremely unequally": six sites account for 25% of page loads, and one site alone for 17% of desktop page loads globally. The consequence they draw is the one that should change your design — a study "calculated directly from a simple set of the top million sites place[s] disproportionate focus on the long tail of the web", because treating a million ranks as a million equal observations weights the rank-900,000 site exactly like Google. The head is also a small part of the web but most of the //browsing//. Ruth et al.'s study of Chrome telemetry {[ruth2022_world]} finds that "the top million sites capture over 95% of all page loads and time spent online, but they do so extremely unequally": six sites account for 25% of page loads, and one site alone for 17% of desktop page loads globally. The consequence they draw is the one that should change your design — a study "calculated directly from a simple set of the top million sites place[s] disproportionate focus on the long tail of the web", because treating a million ranks as a million equal observations weights the rank-900,000 site exactly like Google.
Line 86: Line 86:
  
 **Seed-and-crawl** (100 papers) follows links out from a seed set. It is a snowball sample on the link graph, with the biases of one: well-linked pages are over-represented, islands are unreachable, and the sample depends on crawl order and depth. It is the only way to reach pages no list enumerates, and it should never be described as random. **Seed-and-crawl** (100 papers) follows links out from a seed set. It is a snowball sample on the link graph, with the biases of one: well-linked pages are over-represented, islands are unreachable, and the sample depends on crawl order and depth. It is the only way to reach pages no list enumerates, and it should never be described as random.
 +
 +==== Purposive samples, and the claim they do not license ====
 +
 +Purposive sampling — picking sites because they have a property — is the field's **second most common design (319 papers, 27.7%)** and the one most often described in a way a reader cannot act on. Two things have to be on the page.
 +
 +**The selection rule, precisely enough to re-run.** "Sites with a cookie banner", "European news sites", "e-commerce sites" are not rules; they are the output of one. What did the rule take as input — a categoriser, a filter list, a manual pass, a search? Over what frame? With what threshold? If a categoriser drew the boundary, **its error rate is now your sample's composition error**, and it needs reporting as such: at 90% precision, one site in ten of your "news sites" is not one, and any difference you measure between categories is attenuated toward zero by that mixing. See [[Design:Website classification]] for what those services actually achieve.
 +
 +**And no prevalence claim about anything larger.** A purposive sample of 200 sites that use a given CMP supports "of the 200 sites we selected, //x//% did //y//" and supports nothing at all about how common //y// is on the web, or on the top 10k, or among CMP users generally — because the selection was not a draw. This is the same denominator discipline as everywhere else on this page, and it is easiest to lose here, because a purposive sample //feels// like a sample. In the corpus, purposive populations are small by nature (median 102, against 10,000 for top-//n//), which is a good reason to treat them as case studies and describe them as such.
  
 ==== Reusing someone else's sample ==== ==== Reusing someone else's sample ====
Line 104: Line 112:
 ==== Weight by rank instead of buying more sites ==== ==== Weight by rank instead of buying more sites ====
  
-There is a third option that the size argument usually misses, and the canonical example is in {[englehardt2016online]}. Rather than count how many of //n// sites a third party appears on — a figure that moves whenever you change //n//, because the sites you add are always less popular than the ones you had — they define **prominence**: sum 1/rank(//s//) over the sites //s// where the third party appears. That is a rank-weighted prevalence, and it is motivated exactly as a sampling fix. Their motivation is a sampling one: a plain prevalence rank is sensitive to how many websites the measurement visited, and so "different studies may rank third parties differently" for no reason but their //n//. Prominence instead "de-emphasizes obscure sites, and hence can be adequately … approximated by relatively small-scale measurements".+There is a third option that the size argument usually misses, and the canonical example is in {[englehardt2016online]}. Rather than count how many of //n// sites a third party appears on — a figure that moves whenever you change //n//, because the sites you add are always less popular than the ones you had — they define **prominence**: sum 1/rank(//s//) over the sites //s// where the third party appears. That is a rank-weighted prevalence, and their motivation for it is a sampling one: a plain prevalence rank is sensitive to how many websites the measurement visited, and so "different studies may rank third parties differently" for no reason but their //n//. Prominence instead "de-emphasizes obscure sites, and hence can be adequately … approximated by relatively small-scale measurements".
  
 The general form is the useful part. If your frame is a ranking, you already have a weight for every row, and a weighted estimate is both more comparable across studies and cheaper to compute at a given //n// than an unweighted one is to stabilise. The weights need a stated model — prominence assumes the power-law visit distribution that {[ruth2022_world]} later measured directly — and, like every weighting scheme, it has to be reported as one, because a weighted and an unweighted percentage over the same crawl are different numbers. The general form is the useful part. If your frame is a ranking, you already have a weight for every row, and a weighted estimate is both more comparable across studies and cheaper to compute at a given //n// than an unweighted one is to stabilise. The weights need a stated model — prominence assumes the power-law visit distribution that {[ruth2022_world]} later measured directly — and, like every weighting scheme, it has to be reported as one, because a weighted and an unweighted percentage over the same crawl are different numbers.
Line 114: Line 122:
 ===== Site, Page, or Visit ===== ===== Site, Page, or Visit =====
  
-The unit is a sampling decision, and it is the one most often made by default. In the corpus723 papers sample **websites**, 393 **domains** and 251 **web pages** — and among the 492 papers whose unit is ''websites'' and which ran a crawl, **24.8% visit only the landing page** and just **32.4% of the crawling papers in this population visit more than one page per site**.+The unit is a sampling decision, and it is the one most often made by default. Of the 1,153 papers on this page, 723 sample **websites**, 393 **domains** and 251 **web pages**. Of the 680 that also crawled, just **32.4% visit more than one page per site**; and of the 492 whose unit is explicitly ''websites'' and which crawled, **24.8% visit only the landing page**
 + 
 +==== A list row is not yet a URL ==== 
 + 
 +Before any of that, there is a step the lists do not do for you and most methods sections skip. **Tranco rows are registrable domains** (''example.com'' — no scheme, no host); **CrUX rows are origins** (''https://www.example.com''). Getting from a row to something a browser visits means deciding: 
 + 
 +  * **Scheme.** ''http://'' or ''https://'' first, and what to do when only one answers. A crawl that tries HTTPS only measures a different population from one that falls back. 
 +  * **''www'' or apex.** They can serve different content, different cookies and different consent configurations, and a domain-based list does not tell you which the site prefers. 
 +  * **Where the redirect lands.** If ''example.com'' redirects to ''example.co.uk'', you have measured a different eTLD+1 from the one you sampled. Decide in advance whether that counts as the sampled unit or as attrition, and say which. 
 +  * **Duplicates.** The same site under several rows — country domains, redirect chains, ''www'' and apex both listed — silently weights it more heavily than the design says. De-duplicating by eTLD+1 is the common fix and it changes your //n//, so report the number before and after. 
 + 
 +None of these is exotic and all of them change the denominator. They are the first source of attrition, they happen before the crawler is ever blocked by anything, and a reader cannot reconstruct your population without them.
  
 ==== A landing page is not a sample of a site ==== ==== A landing page is not a sample of a site ====
  
-The strongest result here is Aqeel et al. {[aqeel2020_landing]}, who characterised landing versus internal pages of 1,000 sites and then went back through the literature to see how much it mattered. Reviewing web-performance papers from five networking venues 2015–2019, they scored the 119 that used a top list: **41 (34.5%) needed no revision, 48 (40.3%) a minor revision, and 30 (25.2%) a major revision** for their claims to apply to internal pages. Nearly two-thirds of published claims about "websites" were really claims about one page per website.+The strongest result here is Aqeel et al. {[aqeel2020_landing]}, who characterised landing versus internal pages of 1,000 sites and then went back through the literature to see how much it mattered. Reviewing web-performance papers from five networking venues 2015–2019, they scored the 119 that used a top list: **41 (34.5%) needed no revision, 48 (40.3%) a minor revision, and 30 (25.2%) a major revision** for their claims to apply to internal pages. All 119 were claims about one page per website; for nearly two-thirds of them, that turned out to matter.
  
 Their incidental figure is worth repeating too: of those 119 papers, **only 10 (8.4%) used a top list other than Alexa** — a monoculture that the corpus data below shows has since broken up, but that shaped a decade of results. Their incidental figure is worth repeating too: of those 119 papers, **only 10 (8.4%) used a top list other than Alexa** — a monoculture that the corpus data below shows has since broken up, but that shaped a decade of results.
Line 163: Line 182:
  
 <WRAP important> <WRAP important>
-**"The Alexa top //n//" is no longer a frame.** Amazon retired ''alexa.com'' on **1 May 2022** and the Alexa Top Sites and Web Information Service APIs on **15 December 2022**.((Alexa Support, "We retired Alexa.com on May 1, 2022", archived at [[https://web.archive.org/web/2023/https://support.alexa.com/hc/en-us/articles/4410503838999-We-retired-Alexa-com-on-May-1-2022|web.archive.org]]: "we made the difficult decision to retire Alexa.com on May 1, 2022 … The APIs will be retired on December 15, 2022.")) A 2025 paper that says it crawled "the Alexa top 10,000" is describing a copy of unknown provenance — a mirror, a cached CSV, or a list inherited from an earlier paper — and **66 papers published 2023–2026 in this corpus do exactly that, 51.5% of them without any version or date.** If you are reusing an archived Alexa listsay so and say which snapshot; that is a defensible longitudinal choice. Citing it as though it were live is not.+**"The Alexa top //n//" is no longer a frame.** Amazon retired ''alexa.com'' on **1 May 2022** and the Alexa Top Sites and Web Information Service APIs on **15 December 2022**.((Alexa Support, "We retired Alexa.com on May 1, 2022", archived at [[https://web.archive.org/web/2023/https://support.alexa.com/hc/en-us/articles/4410503838999-We-retired-Alexa-com-on-May-1-2022|web.archive.org]]: "we made the difficult decision to retire Alexa.com on May 1, 2022 … The APIs will be retired on December 15, 2022.")) A 2025 paper that says it crawled "the Alexa top 10,000" with no date is describing a copy of unknown provenance — a mirror, a cached CSV, or a list inherited from an earlier paper**66 papers published 2023–2026 in this corpus still draw from Alexaand 34 of them (51.5%) give no version or date at all**; those 34 are the ones a reader cannot pin to any snapshot. The other 32 do date it, several explicitly as an archive ("March 2018 snapshot", "Alexa archived 2022-01-01"), which is a defensible longitudinal choice — say so and say which snapshot. Citing it as though it were live is not.
  
 Note also that **Tranco is not what it was in 2019.** Its published list of 2026-08-01 was built from CrUX, Farsight, Majestic, Cloudflare Radar and Umbrella — Alexa is long gone from the inputs — over a 30-day window with the Dowdall rule.((Read from Tranco's own API for the daily list, ''https://tranco-list.eu/api/lists/date/2026-08-01'', on 2026-08-12: ''providers: [crux, farsight, majestic, radar, umbrella]'', ''combinationMethod: dowdall'', ''startDate: 2026-07-03'', ''endDate: 2026-08-01''. Tranco's prose methodology page still says "all four providers".)) Describing Tranco's composition from the 2019 paper is a small error that a reviewer who works on lists will catch. Note also that **Tranco is not what it was in 2019.** Its published list of 2026-08-01 was built from CrUX, Farsight, Majestic, Cloudflare Radar and Umbrella — Alexa is long gone from the inputs — over a 30-day window with the Dowdall rule.((Read from Tranco's own API for the daily list, ''https://tranco-list.eu/api/lists/date/2026-08-01'', on 2026-08-12: ''providers: [crux, farsight, majestic, radar, umbrella]'', ''combinationMethod: dowdall'', ''startDate: 2026-07-03'', ''endDate: 2026-08-01''. Tranco's prose methodology page still says "all four providers".)) Describing Tranco's composition from the 2019 paper is a small error that a reviewer who works on lists will catch.
Line 170: Line 189:
 ==== A sampler that records its own provenance ==== ==== A sampler that records its own provenance ====
  
-This draws a rank-stratified sample from a live Tranco or CrUX list and writes a manifest containing the list id, the download URL, the SHA-256 of the frame bytes, the strata, the seed, and a ready-made sentence for your methods section. Tested against both sources on 2026-08-12.+This draws a rank-stratified sample from a live Tranco or CrUX list and writes a manifest containing the list identity, the download URL, the SHA-256 of the frame bytes, the strata, the seed, and a ready-made sentence for your methods section. For CrUX it resolves the newest **dated** monthly snapshot (''202607'' when run on 2026-08-12) rather than ''current.csv.gz'', because an unversioned URL is the thing this page is arguing against. Tested against both sources on 2026-08-12.
  
 <file python stratified_sample.py> <file python stratified_sample.py>
Line 180: Line 199:
 emits both as a manifest next to the sample. emits both as a manifest next to the sample.
  
-    # current Tranco list, 200 sites per decade of rank, seed 20260812+    # current Tranco list, 200 sites per rank stratum, seed 20260812
     python3 stratified_sample.py --source tranco --per-stratum 200 --seed 20260812     python3 stratified_sample.py --source tranco --per-stratum 200 --seed 20260812
  
Line 187: Line 206:
  
     # CrUX, whose ranks are magnitude buckets rather than positions     # CrUX, whose ranks are magnitude buckets rather than positions
-    python3 stratified_sample.py --source crux --per-stratum 200+    python3 stratified_sample.py --source crux --crux-month 202607 --per-stratum 200
  
-Writes sample.csv (rank,domain,stratum) and sample.manifest.json. Paste the+Writes sample.csv (rank,name,stratum) and sample.manifest.json. Paste the
 manifest's `cite` line into your methods section. manifest's `cite` line into your methods section.
 +
 +Note what the rows are, because it decides your unit: Tranco rows are
 +registrable domains (no scheme, no host), CrUX rows are origins (scheme + host).
 +Turning either into a URL a browser visits is a further sampling decision -- see
 +the page this script is published on.
 """ """
 import argparse, csv, gzip, hashlib, io, json, random, sys, urllib.request, zipfile import argparse, csv, gzip, hashlib, io, json, random, sys, urllib.request, zipfile
Line 199: Line 223:
 TRANCO_DOWNLOAD = "https://tranco-list.eu/download/{list_id}/1000000" TRANCO_DOWNLOAD = "https://tranco-list.eu/download/{list_id}/1000000"
 TRANCO_PERMALINK = "https://tranco-list.eu/list/{list_id}" TRANCO_PERMALINK = "https://tranco-list.eu/list/{list_id}"
-CRUX_CURRENT = "https://raw.githubusercontent.com/zakird/crux-top-lists/main/data/global/current.csv.gz"+CRUX_MONTHLY = "https://raw.githubusercontent.com/zakird/crux-top-lists/main/data/global/{month}.csv.gz
 +CRUX_INDEX = "https://api.github.com/repos/zakird/crux-top-lists/contents/data/global"
  
 # Decades of rank. The head is where the popular-site literature lives and the # Decades of rank. The head is where the popular-site literature lives and the
Line 236: Line 261:
  
  
-def load_crux(): +def latest_crux_month(): 
-    raw = fetch(CRUX_CURRENT)+    """Newest YYYYMM snapshot in the mirror. Never `current.csv.gz`: that URL is 
 +    unversioned, which is the anti-pattern this script exists to avoid.""" 
 +    listing = json.loads(fetch(CRUX_INDEX)) 
 +    months = sorted(f["name"][:6] for f in listing if f["name"][:6].isdigit()) 
 +    return months[-1] 
 + 
 + 
 +def load_crux(month): 
 +    if month is None: 
 +        month = latest_crux_month() 
 +    raw = fetch(CRUX_MONTHLY.format(month=month))
     text = gzip.decompress(raw).decode()     text = gzip.decompress(raw).decode()
     rows = list(csv.DictReader(io.StringIO(text)))     rows = list(csv.DictReader(io.StringIO(text)))
Line 246: Line 281:
     return ranked, {     return ranked, {
         "source": "crux",         "source": "crux",
-        "list_id": None+        "list_id": month
-        "download": CRUX_CURRENT,+        "download": CRUX_MONTHLY.format(month=month),
         "permalink": "https://github.com/zakird/crux-top-lists",         "permalink": "https://github.com/zakird/crux-top-lists",
         "bytes_sha256": hashlib.sha256(raw).hexdigest(),         "bytes_sha256": hashlib.sha256(raw).hexdigest(),
-        "rank_semantics": "magnitude bucket (1000/10000/100000/1000000), random order within bucket",+        "rank_semantics": "magnitude bucket (1000/10000/100000/1000000), random order within bucket; strata below are bucket membership, not positions",
     }     }
  
Line 258: Line 293:
     ap.add_argument("--source", choices=["tranco", "crux"], default="tranco")     ap.add_argument("--source", choices=["tranco", "crux"], default="tranco")
     ap.add_argument("--list-id", help="Tranco list id; default is today's list")     ap.add_argument("--list-id", help="Tranco list id; default is today's list")
 +    ap.add_argument("--crux-month", help="CrUX snapshot as YYYYMM; default is the newest published")
     ap.add_argument("--per-stratum", type=int, default=200)     ap.add_argument("--per-stratum", type=int, default=200)
     ap.add_argument("--seed", type=int, default=0)     ap.add_argument("--seed", type=int, default=0)
Line 263: Line 299:
     args = ap.parse_args()     args = ap.parse_args()
  
-    ranked, prov = load_tranco(args.list_id) if args.source == "tranco" else load_crux()+    ranked, prov = load_tranco(args.list_id) if args.source == "tranco" else load_crux(args.crux_month)
     by_rank = {}     by_rank = {}
     for rank, name in ranked:     for rank, name in ranked:
Line 298: Line 334:
             f"We drew {len(sample)} sites from the {prov['source']} list "             f"We drew {len(sample)} sites from the {prov['source']} list "
             f"{prov['list_id'] or '(dated snapshot, see manifest)'} "             f"{prov['list_id'] or '(dated snapshot, see manifest)'} "
-            f"({prov['permalink']}), {args.per_stratum} per rank decade over "+            f"({prov['permalink']}), {args.per_stratum} per rank stratum over "
             f"{', '.join(s['stratum'] for s in strata_report)}, "             f"{', '.join(s['stratum'] for s in strata_report)}, "
             f"simple random sampling without replacement within each stratum, seed {args.seed}."             f"simple random sampling without replacement within each stratum, seed {args.seed}."
Line 313: Line 349:
 </file> </file>
  
-Real output, ''--source tranco --per-stratum 200 --seed 20260812'', run on 2026-08-12 (the four ''strata'' blocks are abbreviated to one line each here; everything else is verbatim):+Real output, ''--source tranco --per-stratum 200 --seed 20260812'', run on 2026-08-12 (the four ''strata'' objects are reflowed to one line each here; everything else is verbatim):
  
 <code> <code>
 { {
-  "drawn_at_utc": "2026-08-12T21:04:41+00:00",+  "drawn_at_utc": "2026-08-12T21:36:44+00:00",
   "frame": {   "frame": {
     "source": "tranco",     "source": "tranco",
Line 329: Line 365:
   "design": "rank-stratified, simple random sample without replacement within each stratum",   "design": "rank-stratified, simple random sample without replacement within each stratum",
   "strata": [   "strata": [
-    { "stratum": "1-1000",         "in_frame": 1000,   "drawn": 200 }, +    { "stratum": "1-1000",        "in_frame": 1000,   "drawn": 200 }, 
-    { "stratum": "1001-10000",     "in_frame": 9000,   "drawn": 200 }, +    { "stratum": "1001-10000",    "in_frame": 9000,   "drawn": 200 }, 
-    { "stratum": "10001-100000",   "in_frame": 90000,  "drawn": 200 },+    { "stratum": "10001-100000",  "in_frame": 90000,  "drawn": 200 },
     { "stratum": "100001-1000000", "in_frame": 900000, "drawn": 200 }     { "stratum": "100001-1000000", "in_frame": 900000, "drawn": 200 }
   ],   ],
Line 337: Line 373:
   "n": 800,   "n": 800,
   "sample_sha256": "e3c38a539111c9872b62682f8d289f5d573248d95fafe6ce3bf49ea02bf92249",   "sample_sha256": "e3c38a539111c9872b62682f8d289f5d573248d95fafe6ce3bf49ea02bf92249",
-  "cite": "We drew 800 sites from the tranco list 645KX (https://tranco-list.eu/list/645KX), 200 per rank decade over 1-1000, 1001-10000, 10001-100000, 100001-1000000, simple random sampling without replacement within each stratum, seed 20260812."+  "cite": "We drew 800 sites from the tranco list 645KX (https://tranco-list.eu/list/645KX), 200 per rank stratum over 1-1000, 1001-10000, 10001-100000, 100001-1000000, simple random sampling without replacement within each stratum, seed 20260812."
 } }
 </code> </code>
Line 444: Line 480:
 Median 6,755; quartiles 283 and 103,541; the 90th percentile is 1,067,968. **43.5% of stated sizes are 1, 2 or 5 times a power of ten**, and 198 populations are exactly 1,000,000. Median size by method: top-//n// 10,000 (879 populations), exhaustive 708,000 (235), seed-and-crawl 47,847 (116), pre-existing-dataset 24,000 (346), stratified 4,084 (76), random 2,000 (237), purposive 102 (406). Median 6,755; quartiles 283 and 103,541; the 90th percentile is 1,067,968. **43.5% of stated sizes are 1, 2 or 5 times a power of ten**, and 198 populations are exactly 1,000,000. Median size by method: top-//n// 10,000 (879 populations), exhaustive 708,000 (235), seed-and-crawl 47,847 (116), pre-existing-dataset 24,000 (346), stratified 4,084 (76), random 2,000 (237), purposive 102 (406).
  
-==== Half the field's samples cannot be redrawn ====+==== Half the field never says which version it drew ====
  
 ^ Reported for the same web population ^ Papers ^ Share of 1,153 ^ ^ Reported for the same web population ^ Papers ^ Share of 1,153 ^
design/sampling.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki