Table of Contents
Welcome to Measure The Web
Empirical studies on the web require researchers to navigate a complex landscape of experimental design choices, ranging from selecting a representative sample of websites to choosing the appropriate crawling technology. Similarly, analyzing results involves critical decisions, such as website categorization and statistical methodology. Too often, these decisions are made based on limited guidance, informal advice, or trial and error, despite their profound impact on research outcomes and their applicability.
Measure The Web aims to bridge this gap, by providing evidence-based guidance for empirical web measurement studies. The platform evaluates design choices and their implications by referencing relevant academic publications or conducting original studies where necessary.
For novice researchers in the web measurement field, Measure The Web should provide a complete knowledge base to conduct their study. But it should be helpful also to experienced researchers, who can find here newest practices or proper arguments (or citations) to justify their design. If you are in a mentoring position, consider sharing the website (or directly Research journey) to your mentees and also consider reviewing and contributing to the page. Even small edits, such as supporting some subjective claim, can help make this website stronger.
Outline
The website is organized as follows. Note that it is ordered by website structure, for order by research design, navigate to Research journey.
- Research design clarifies choices needed to conduct the measurements. Example pages:
- Automated measurements — crawl, scan, or unpack an app; the instrument is the design decision — or User studies (Mechanical Turk shuts 30 September 2026)
- Measuring DNS — which resolver answered, whether the answer was manipulated, and what encrypted DNS changed. Three public resolvers gave completely disjoint answers for 28 of the Tranco top 100 on 2026-08-27.
- Repeating a measurement over time — pinning the list, the browser, the vantage point and the classifier so wave two is comparable to wave one
- Measuring mobile apps — store scraping, static and dynamic analysis, and the certificate-pinning problem that decides whether you saw the traffic at all
- Research of specific large platforms — when the population is one platform rather than a sample of the web, the access route is the design decision and it fixes your denominator. Which routes still exist in 2026, what they cost, and why there are no per-company sub-pages.
-
- Multitude of pages linked from elsewhere documenting specific technologies, e.g., Tranco, Cloudflare Radar, Docker, CrUX, Similarweb.
-
- Measuring server-side tracking — when the tracking request never reaches the tracker, and every request-level method quietly stops working
- Measuring cookie and ID syncing — how third parties learn that their two identifiers are the same person, and the crawl configuration that decides whether you see it at all
- Decoding TCF consent strings — the IAB TC string and Google's Additional Consent string as measurement artefacts, and what a decoded string does and does not prove
- Dark patterns in the interfaces you measure — and in the way they defeat your crawler
- Security — measuring web security as it is deployed, not how to break it:
- Statistics: Study preregistration, Hypothesis testing suitable for web measurements, Regression, P-value corrections, Biases, Inter-rater agreement for hand-coded data
-
- Conferences — where to send a web-measurement paper, and the current CFP dates.
- Literature review — the keyword-derived denominator trap, measured on the SoKs in this corpus.
- Artifacts — a GitHub URL is not an archival artifact. What to put in a crawl deposit, where 2026 Available badges will accept it, and what to withhold when the data has people in it.
- Pages whose figures come from the publication corpus carry a working log under
provenance:— every query with its denominator, the report script and its unedited output, the folds and their residue, and what could not be established. The corpus-wide caveats are on Corpus; the per-page logs mirror the page id, e.g. fingerprinting for Fingerprinting. - After publication steps to improve the impact of your work:
- Notifying websites to remediate the security vulnerabilities and privacy threats.
- Legal enforcement with information about local data protection authorities.
- Public relations regarding how to present your results to public and journalists.
- Other research practices explains various aspects affiliated to the research, but not directly the study design.
Contributing
We welcome contributions! Editing functionality is limited to registered users. You might want to check Contributing page, which helps with the syntax.
