User Tools

Site Tools


design

This is an old revision of the document!


Design

This namespace is for choices you make before the crawler runs — which list stands in for the web, how you draw from it, where you appear to be, whether you fetch live or from an archive, what you hold fixed between waves, and when the object of study is an app or a platform rather than a URL. It is not a tutorial on experimental design, and it is not the crawler, the classifier, or the test. Those live in Programming, Privacy / Security, and Statistics. The publication corpus behind these pages is seven venues (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026, 5,859 extracted papers). Only 1,120 of them ran a crawl. Each child names its own population.

A namespace page outlines the pages inside it rather than carrying its own content. 1) All 11 pages start used to promise here are written, plus DNS, added on 2026-08-27 to give scan-branch DNS measurement a home of its own — Automated measurements had been routing it to TLS certificates and IP classification, which answer different questions. The four per-company platform pages start used to promise (Facebook, Twitter, TikTok, Amazon) were assessed against the corpus and deliberately not written; Platforms carries the reasoning and the paper counts behind it.

The pages

Page What a student needs it for On the wiki
Website selection Which list substitutes for the web (Tranco, CrUX, Alexa residue). written
Sampling How you draw from a list you already chose (top-n vs stratified, size, version). written
Website classification Topic / industry / company labels — not popularity. written
IP classification Turning an observed address into a defendable claim (ASN, geo, network type). written
Crawling location The vantage point you control: country, datacenter vs residential, bot treatment. written
Archives Live crawl vs Wayback / Common Crawl — what an archive can and cannot answer. written
Longitudinal Pinning list, browser, vantage and classifier so wave two is comparable to wave one. written
Mobile and app measurement Store scraping, static vs dynamic, and whether pinning hid the traffic. written
Automated measurements Orient between crawling, scanning and app analysis. written
Existing datasets Measuring from somebody else's scans: which snapshot, which query, which join key, and the denominator you inherit with the data. written
User studies Participants vs annotators. Crowdworkers labelling data are annotation, not a user study. written
Platforms One platform instead of a sample of the web: which access routes exist in 2026, what denominator each one hands you, rate limits, ToS and bans. written
DNS Measuring names rather than certificates: which resolver answered, whether the answer was manipulated, and why a single-vantage resolution is a result about one path. written
Ad archives The one platform access route with its own instrument page: what each of the thirteen DSA ad repositories covers, what an archive structurally cannot contain, and the two-sided error problem. written 2026-09-08
Blocking and geodifference What your vantage point was not allowed to see: geoblocking, GDPR walls, censorship and network interference — and what a blocking claim has to contain before a reviewer accepts it. written 2026-09-10

Read Website selection and Sampling together: one is which list, the other is how you draw. Crawling location is the address you control; IP classification is everyone else's. Platforms is the case where none of that applies, because the platform, not a ranking, defines the population. DNS is the step before all of them: the resolution that turns a name in your list into the address you connect to, and what that step alone can get wrong. Ad archives is the only child page in this namespace: it sits under Platforms because it is one route in, not a namespace of its own. Blocking and geodifference is the pair to Crawling location: that page is how you choose where to appear from, this one is what that choice may not be allowed to see, and the phenomenon was split out of it on 2026-09-10 because a 30 KB page about choosing a vantage point is the wrong home for the instruments and error rates of censorship measurement.

Two of these pages describe how the corpus name-folds interact: “Alexa” in a paper's stated population is the retired ranking list of Website selection in 412 papers and an Amazon platform in 39, and “Amazon” is User studies' Mechanical Turk in 179. Platforms publishes that fold in full.

Where this namespace stops

  • Crawler — which browser and which control channel. Design decides whether to crawl; programming decides with what.
  • Tranco / Cloudflare Radar / CrUXAPI and construction notes the selection page points at, not a second copy of “which list”.
  • Traffic files — the recording a crawl leaves. Linked from Archives because an archive is someone else's recording.
  • Biases — the effect sizes of a top-n frame, a vantage, a missing denominator. Design chooses the frame; statistics names what that does to the number.
  • Privacy / Security — what you classify once you have the bytes.
  • Artifacts — publishing the pinned list. Pinning is Longitudinal; the deposit is artifacts.

Methodology and limitations of these figures

The 5,859 / 1,120 are paper counts from the 5,859-paper extraction (seven venues, 2010–2026). 2025–2026 venue-years are provisional — see corpus. The 15 is the number of rows in the table above, counted on 2026-09-10. Two corrections were made that day: the count read 12 against a 13-row table (stale since Ad archives was added on 2026-09-08), and Existing datasets — a live page in this namespace since before either — was missing from the table altogether, found by a review pass on Blocking and geodifference. The 11 in the box is the number of pages start used to promise in this namespace; start no longer lists individual pages, so that number is a historical count and not a live one. Queries and the inventory: design.

1)
contributing, “Namespace and page structure”.
You could leave a comment if you were logged in.
design.1789084068.txt.gz · Last modified: by karel.kubicek.claude