data/extract/README.md, table “Which fields can carry a published percentage”. Multi-valued fields are scored by exact set equality, so one extra label flips the paper. This is an old revision of the document!
Table of Contents
Automated Measurements
You are about to measure something without recruiting humans. That is one of two top-level design branches on this site — the other is User studies — and it is not one method. It is three instruments that see different things, fail in different ways, and already have specialised pages. This page exists to pick the instrument. It does not replace those pages.
| If you need to observe… | The instrument is | Then read |
|---|---|---|
| a rendered page, cookies, JavaScript, banners, third-party requests | a crawl | Crawler, Stateful stateless, Interaction, Crawling location, Consent |
| open ports, certificates, protocol banners, DNS, IPv4/IPv6 hosts | a scan (or a search engine over someone else's scans) | TLS certificates, IP classification, Ethics (scanning checklist), Notifying websites |
| an APK or IPA, a store listing, an app's runtime | app analysis | Mobile and app measurement |
The instrument is the design decision. “We measured tracking” does not say whether you loaded pages, probed addresses, or unpacked binaries. A crawl cannot directly observe tracking that remains entirely server-side (Server side tracking); a port scan cannot see cookies; unpacking an APK cannot see a website. Papers that pick the wrong instrument still get published — they answer a different question from the one in the title.
Everything below about “the literature” is a claim about seven venues — CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026, 5,859 papers with extracted full text. The extractor field this page is built on, studyTypes, agrees with itself run-to-run on 57% of papers in a 100-paper sample that has not been re-measured on the 5,859-paper run.1) Treat every studyTypes share on this page as a ranking, not as a precise percentage. Venue counts, year buckets and reporting rates are paper counts over other fields; they still have the usual extraction error, but they are not the 57% field. See Corpus.
Which of the three
Of 5,859 papers, 2,198 (37.5%) are tagged with at least one of the three automated-measurement types. Almost all of those papers are only one type. The overlap is small and it is a design, not a rounding error:
| Slice | Papers |
|---|---|
only crawl (automated-web-crawl) | 792 |
only scan (network-scan-or-probe) | 780 |
only app (mobile-app-analysis) | 402 |
| crawl and scan, not app | 97 |
| crawl and app, not scan | 74 |
| scan and app, not crawl | 49 |
| all three | 4 |
| union | 2,198 |
The three branches are close in size — 967 crawl, 930 scan, 529 app — and they are not a partition of empirical work. 2,615 papers analyse an existing dataset; 1,149 are tagged user-study. 234 of the 2,198 (10.6%) are also tagged user-study: automated measurement and a user study are not exclusive, they are two instruments that co-occur when a paper both crawls and then asks people about what it found. Crowdworkers who only label data are annotators, not participants — see User studies when that page exists.
Two traps sit under those headlines.
- The site's “crawled” population is 1,120, not 967. Crawler and the other crawl pages count every paper with a
crawlConfigor theautomated-web-crawltag. All 967 tagged crawls are inside that 1,120. The extra 153 have a crawl configuration and were tagged as something else — 125 of them as a system-or-defence-proposal, 44 as mobile-app-analysis, 30 as a scan. They ran a crawler; the extractor did not treat the crawl as the study type. When this page says “crawl papers” it means the 967. When a child page says “crawling papers” it means the 1,120. Both numbers are real; they answer different questions. - Almost half of “scans” also analyse an existing dataset (429 of 930, 46.1%). Naming Censys or Shodan is often not running ZMap. The specialised TLS page already has to separate “we scanned” from “we queried a scan search engine”. Do that split before you count.
Crawling
A crawl drives a browser (or an HTTP client) at URLs and records what the site does in response. It is the instrument for cookies, scripts, banners, third-party requests, and anything else that only exists after a page is rendered. It is a bad instrument for a question about hosts, ports, or binaries.
Read the child pages, do not re-derive them here. Crawler is the library comparison (browser vs control channel). Stateful stateless is the profile. Interaction is what you do on the page. Crawling location is where from. Consent is the banner. Website selection is which URLs. Longitudinal is how to do it twice.
What this page adds is the routing evidence:
- Of the 967, 640 (66.2%) name a crawler-framework or browser-automation tool they used or produced. Selenium is the most-named family (214, 22.1%), then Puppeteer (71), OpenWPM (57), Playwright (31), PhantomJS and other legacy scriptable browsers (27). Those are papers, not a recommendation.
- Currency, dated 2026-08-27. Playwright is the library that arrived in this corpus in 2022–2024 (15 of 294 crawl papers, 5.1%) and is 16 of 169 (9.5%) in the provisional 2025–2026* window. Selenium is still the most-named family in that same window (44 of 169, 26.0%). PhantomJS is gone: 0 of 294 in 2022–2024. npm
playwrightis 1.62.1 today. “Use Playwright” is current advice for a new crawl; “the literature used Selenium” is a statement about the literature. They are not the same sentence. The comparison itself is Crawler. - A crawl is not a user. Zeber et al. [1Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] compared crawls to Mozilla telemetry and found that for a majority of the top domains they inspected, “the number of third parties reported by the crawler is well into the right tail” of the human distribution. Ahmad et al. [2Ahmad, Syed Suleman; Dar, Muhammad Daniyal; Zaffar, Muhammad Fareed; Vallina-Rodriguez, Narseo; Nithyanand, Rishab (2020): "Apophanies or Epiphanies? How Crawlers Impact Our Understanding of the Web", in: Proceedings of The Web Conference, pp. 271-280. (DOI)] drove several crawlers at the same task and report “a variation of over 16% in the number of successful page loads”. If your claim is about people, a crawl is a surrogate with a known direction of bias; if your claim is about what a site serves a browser, pick the browser on purpose and report it.
Scanning
A scan probes addresses, ports, or names. It is the instrument for “what does this host speak”, “which certificates are live”, “who is scanning the Internet”. It does not render a page. Mixing a crawl and a scan in one paper is a legitimate design (97 papers here do crawl-and-scan without also doing apps); mixing them in one count is a mistake.
Current instruments, dated 2026-08-27.
- ZMap is still an actively maintained IPv4 scanner. Durumeric, Wustrow and Halderman introduced it at USENIX Security 2013 [3Durumeric, Zakir; Wustrow, Eric; Halderman, J. Alex (2013): "ZMap: Fast Internet-wide Scanning and Its Security Applications", in: Proceedings of the USENIX Security Symposium. (Link)] — that paper is not in this extraction (it is missing from both the USENIX 2013 index and
data/fulltext/2013/USENIX/), which is why the paper to read in-corpus is the ten-year retrospective [4Durumeric, Zakir; Adrian, David; Stephens, Phillip; Wustrow, Eric; Halderman, J. Alex (2024): "Ten Years of ZMap", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]: “more than 300 research papers have used ZMap”, and “over 33% of all Internet-wide IPv4 scan traffic can be fingerprinted as coming from ZMap”. The 300 is that paper's count of the literature, not a count in these 5,859. In this corpus, 93 of 930 scan papers name ZMap as a used or producednetwork-scanner. The GitHub repozmap/zmapis not archived and was pushed on 2026-08-26;zmap/zgrab2was pushed on 2026-07-27. “The” scanner would overclaim what a repo date can prove; the 33% traffic figure is the usage evidence, and it is one paper's measurement of scan traffic, not of this corpus. - Censys is often a dataset, not a scan you ran. Durumeric et al. [5Durumeric, Zakir; Adrian, David; Mirian, Ariana; Bailey, Michael D.; Halderman, J. Alex (2015): "A Search Engine Backed by Internet-Wide Scanning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] introduced it as “a public search engine and data processing facility backed by data collected from ongoing Internet-wide scans”. 72 papers in the corpus name Censys as a used or produced tool; 63 of those are in the scan branch; 43 of those 63 are also tagged
existing-dataset-analysis. That tag does not prove they only queried Censys — a paper can scan and re-analyse — but it is why you must not count the name as “they ran ZMap”. Of scan papers, Censys is named by about 10% in 2018–2024 (26 of 255, then 24 of 245). The 2025–2026* cell is 5 of 168 (3.0%) and is too thin to read as a decline.censys.comserved HTTP 200 on 2026-08-27;search.censys.ioreturns 403 to automated clients, including with a browser User-Agent — a bot wall, not an outage. - nmap, traceroute, RIPE Atlas are the other named families (40, 31, 25 papers in the scan branch). They answer different questions than ZMap does (host discovery vs path vs volunteer vantage points). 167 scan papers name a
network-scannerthe fold did not map — one-off research scanners, listed on the provenance page. That residue is the scanning analogue of the home-grown crawler row on Crawler.
Scanning ethics is not this page. Durumeric, Bailey and Halderman [6Durumeric, Zakir; Bailey, Michael; Halderman, J. Alex (2014): "An Internet-Wide View of Internet-Wide Scanning", in: Proceedings of the USENIX Security Symposium. (Link)] measured who was already scanning the Internet; the operational checklist (identify yourself, publish an opt-out, rate-limit) lives on Ethics, and telling the operator lives on Notifying websites. A scan from cloud address space is also a vantage-point decision.
Certificate and TLS measurement is TLS certificates. Turning the addresses you collected into a claim is IP classification.
App analysis
An app is not a URL. Getting the binary, driving it, and knowing whether you saw its traffic are the three problems Mobile and app measurement is for. This page only records that the branch exists and how large it is: 529 papers, 402 of them only this type, 99.2% on the mobile platform. PETS (12.9% of venue output) and NDSS (11.7%) are where this work is densest as a share; USENIX Security (143 papers) and CCS (106) are where it is largest as a count. IMC is scan-heavy; app analysis is less common there (32 papers, 5.0% of venue output) but it is not absent.
Certificate pinning, Frida, AndroZoo, and the coverage number you owe the reader are on that page. Start with Pradeep et al. [7Pradeep, Amogh; Paracha, Muhammad Talha; Bhowmick, Protick; Davanian, Ali; Razaghpanah, Abbas; Chung, Taejoong; Lindorfer, Martina; Vallina-Rodriguez, Narseo; Levin, Dave; Choffnes, David (2022): "A Comparative Analysis of Certificate Pinning in Android & iOS", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] if you are about to intercept app traffic.
What this is not
- A user study. User studies is the other top-level Design branch (red link until written). 1,149 papers are tagged
user-study. Do not put Mechanical Turk in a crawlConfig. - An existing-dataset analysis by itself. Re-analysing Censys, a CT log, AndroZoo, or last year's crawl is empirical work and it is most of the corpus (2,615 papers). It is not running a measurement. The design questions that remain — which snapshot, which query, which join key — belong on the child page for that artefact (Tranco, Archives, TLS certificates, Mobile and app measurement).
- A manual audit, a code analysis, or a system paper. Those are the three larger
studyTypes(1,826 / 1,484 / 3,967). A system paper that also crawled is in the 153-paper gap above. - A tutorial on HTTP, DNS, or TLS. Those are specs. This page is which instrument the field uses to observe them.
What to report
The same four columns, asked of each branch. Location, tool version and an artifact URL are comparable. robots.txt is a crawl/scan-ethics question — the app row is 3 of 529 (0.6%) because the question barely applies, not because app papers forgot a robots file. The methods-section checklist after the table names fields this table does not measure; those live on the child pages.
| Population | N | Vantage location stated | Any used/produced tool version | Authors-own artifact URL | robots.txt stated |
|---|---|---|---|---|---|
crawl (automated-web-crawl) | 967 | 271 (28.0%) | 446 (46.1%) | 581 (60.1%) | 48 (5.0%) |
scan (network-scan-or-probe) | 930 | 487 (52.4%) | 401 (43.1%) | 514 (55.3%) | 22 (2.4%) |
app (mobile-app-analysis) | 529 | 117 (22.1%) | 207 (39.1%) | 334 (63.1%) | 3 (0.6%) |
| crawled population (the 1,120) | 1,120 | 301 (26.9%) | 506 (45.2%) | 684 (61.1%) | 53 (4.7%) |
A crawl methods section that does not name the browser, the library and its version, the list pin, the vantage point, and the consent action is incomplete — those pages exist because the field does not report them. A scan methods section that does not name the scanner and version, the port set, the source addresses, and the opt-out story is incomplete for the same reason. An app methods section that does not report the share of the app set whose traffic you actually decrypted is incomplete — that sentence is the one Mobile and app measurement is built on.
What to read first
| Paper | Why it is on this page |
|---|---|
| Zeber et al., TheWebConf 2020, The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing [1Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] | Whether a crawl is a stand-in for people. It is not, and the bias has a direction. |
| Ahmad et al., TheWebConf 2020, Apophanies or Epiphanies? [2Ahmad, Syed Suleman; Dar, Muhammad Daniyal; Zaffar, Muhammad Fareed; Vallina-Rodriguez, Narseo; Nithyanand, Rishab (2020): "Apophanies or Epiphanies? How Crawlers Impact Our Understanding of the Web", in: Proceedings of The Web Conference, pp. 271-280. (DOI)] | The crawler is an experimental variable. “a variation of over 16% in the number of successful page loads”. |
| Durumeric, Wustrow and Halderman, USENIX Security 2013, ZMap: Fast Internet-wide Scanning and Its Security Applications [3Durumeric, Zakir; Wustrow, Eric; Halderman, J. Alex (2013): "ZMap: Fast Internet-wide Scanning and Its Security Applications", in: Proceedings of the USENIX Security Symposium. (Link)] | The instrument paper for IPv4 scanning. Not in this extraction; cited from the USENIX page, which was live on 2026-08-27. |
| Durumeric et al., IMC 2024, Ten Years of ZMap [4Durumeric, Zakir; Adrian, David; Stephens, Phillip; Wustrow, Eric; Halderman, J. Alex (2024): "Ten Years of ZMap", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | That the instrument is still current, what changed, and the 33% traffic figure. |
| Durumeric et al., CCS 2015, A Search Engine Backed by Internet-Wide Scanning [5Durumeric, Zakir; Adrian, David; Mirian, Ariana; Bailey, Michael D.; Halderman, J. Alex (2015): "A Search Engine Backed by Internet-Wide Scanning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] | Censys: the scan you query instead of running. |
| Pradeep et al., IMC 2022, A Comparative Analysis of Certificate Pinning in Android & iOS [7Pradeep, Amogh; Paracha, Muhammad Talha; Bhowmick, Protick; Davanian, Ali; Razaghpanah, Abbas; Chung, Taejoong; Lindorfer, Martina; Vallina-Rodriguez, Narseo; Levin, Dave; Choffnes, David (2022): "A Comparative Analysis of Certificate Pinning in Android & iOS", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | The app-branch question a web-trained intuition misses: did your instrumentation see the traffic? |
Englehardt and Narayanan [8Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] is the canonical million-site crawl (“based on a crawl of the top 1 million websites”) and lives on OpenWPM. Jueckstock et al. [9Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)] is the vantage-point experiment and lives on Crawling location. Demir et al. [10Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)] is the repeat-run noise floor and lives on Longitudinal. Durumeric, Bailey and Halderman [6Durumeric, Zakir; Bailey, Michael; Halderman, J. Alex (2014): "An Internet-Wide View of Internet-Wide Scanning", in: Proceedings of the USENIX Security Symposium. (Link)] is who was already scanning you, and lives next to Ethics.
Use in publications
Figures below are from scripts/report_automated_measurements.mjs over the 5,859-paper extraction. studyTypes is multi-valued: a paper can be in more than one row, so the type table does not sum to 5,859. 2025–2026 venue-years are provisional — CCS and IMC 2026 have not been held, and IEEE S&P and WWW 2026 abstracts are under-selected by construction. Do not read the last year-bucket as a completed year.
The type ranking
studyTypes value | Papers | Share of 5,859 |
|---|---|---|
| system-or-defence-proposal | 3,967 | 67.7% |
| existing-dataset-analysis | 2,615 | 44.6% |
| manual-audit | 1,826 | 31.2% |
| code-or-binary-analysis | 1,484 | 25.3% |
| user-study | 1,149 | 19.6% |
| automated-web-crawl | 967 | 16.5% |
| network-scan-or-probe | 930 | 15.9% |
| interview-or-survey | 755 | 12.9% |
| mobile-app-analysis | 529 | 9.0% |
| simulation-or-theory-only | 335 | 5.7% |
14 papers have an empty studyTypes. They are listed on the provenance page; they are not forced into a branch.
Where the three branches publish
Share of that venue's own output, not share of the branch. IMC is the scanning venue. TheWebConf is the crawling venue. IMC is not the crawling venue — that claim was false last time someone made it from memory, and the numbers here are why.
| Venue | Papers | crawl | crawl share | scan | scan share | app | app share |
|---|---|---|---|---|---|---|---|
| CCS | 990 | 137 | 13.8% | 132 | 13.3% | 106 | 10.7% |
| IEEE S&P | 767 | 89 | 11.6% | 87 | 11.3% | 63 | 8.2% |
| IMC | 638 | 124 | 19.4% | 296 | 46.4% | 32 | 5.0% |
| NDSS | 701 | 106 | 15.1% | 117 | 16.7% | 82 | 11.7% |
| PETS | 510 | 112 | 22.0% | 49 | 9.6% | 66 | 12.9% |
| USENIX Security | 1,410 | 176 | 12.5% | 199 | 14.1% | 143 | 10.1% |
| TheWebConf | 843 | 223 | 26.5% | 50 | 5.9% | 37 | 4.4% |
Over the years
| Window | Papers | crawl | scan | app | crawl share | scan share | app share |
|---|---|---|---|---|---|---|---|
| 2010–2013 | 511 | 93 | 89 | 30 | 18.2% | 17.4% | 5.9% |
| 2014–2017 | 769 | 142 | 173 | 113 | 18.5% | 22.5% | 14.7% |
| 2018–2021 | 1,439 | 269 | 255 | 146 | 18.7% | 17.7% | 10.1% |
| 2022–2024 | 1,955 | 294 | 245 | 150 | 15.0% | 12.5% | 7.7% |
| 2025–2026* | 1,185 | 169 | 168 | 90 | 14.3% | 14.2% | 7.6% |
Crawl and scan shares are below their 2014–2017 / 2010–2013 peaks; app analysis is below its 2014–2017 peak (14.7%) but above its earliest share (5.9% → 7.6%). Absolute counts are up because the venues are bigger. The 2025–2026* column is starred.
Posters
138 of 5,859 papers are posters (poster- slug or Poster: title); 251 records are 4 pages or fewer. Scan papers are the most poster-heavy of the three (34 of 930, 3.7%). Dropping posters does not move the ranking of the three branches. Posters omit methods more often, which raises silence rates and lowers the reporting rates in the table above — those rates may understate how often full papers state the same things.
Methodology and limitations of these figures
Every number above is a count of papers, from the 5,859-paper extraction, with the denominator in the same sentence or table header. Sentinels (not-stated, none-mentioned) are never answers. Free-text tool names are folded before counting; the scanner fold and its 205 unmapped strings / 167 residue papers are on automated_measurements. studyTypes is the least stable field in the schema (57% on a 100-paper sample, not re-measured on this run) — rankings only. 2025–2026 are provisional; see corpus. The report script is scripts/report_automated_measurements.mjs. It exits 1 if the corpus size, the three OVERVIEW.md type counts, the crawled population, or the Venn arithmetic disagree with the contracts it encodes.
Related pages
- User studies — the other top-level Design branch (not yet written).
- Crawler — which library, which browser, which version.
- Crawling location — the address you crawl from.
- Website selection / Sampling — which URLs.
- Longitudinal — doing it twice.
- Mobile and app measurement — the app branch in full.
- TLS certificates — the scan branch when the question is the web PKI.
- IP classification — the addresses a crawl or a scan left you with.
- Ethics / Notifying websites — scanning and crawling other people's machines.
- Corpus — venue scope, funnel, provisional years.
- [1]
- Zeber, David; Bird, Sarah; Oliveira, Camila; Rudametkin, Walter; Segall, Ilana; Wolls´en, Fredrik; Lopatka, Martin (2020): "The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing", in: Proceedings of The Web Conference 2020, pp. 167–178. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
- [2]
- Ahmad, Syed Suleman; Dar, Muhammad Daniyal; Zaffar, Muhammad Fareed; Vallina-Rodriguez, Narseo; Nithyanand, Rishab (2020): "Apophanies or Epiphanies? How Crawlers Impact Our Understanding of the Web", in: Proceedings of The Web Conference, pp. 271-280. (DOI)
- [3]
- Durumeric, Zakir; Wustrow, Eric; Halderman, J. Alex (2013): "ZMap: Fast Internet-wide Scanning and Its Security Applications", in: Proceedings of the USENIX Security Symposium. (Link)
- [4]
- Durumeric, Zakir; Adrian, David; Stephens, Phillip; Wustrow, Eric; Halderman, J. Alex (2024): "Ten Years of ZMap", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [5]
- Durumeric, Zakir; Adrian, David; Mirian, Ariana; Bailey, Michael D.; Halderman, J. Alex (2015): "A Search Engine Backed by Internet-Wide Scanning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [6]
- Durumeric, Zakir; Bailey, Michael; Halderman, J. Alex (2014): "An Internet-Wide View of Internet-Wide Scanning", in: Proceedings of the USENIX Security Symposium. (Link)
- [7]
- Pradeep, Amogh; Paracha, Muhammad Talha; Bhowmick, Protick; Davanian, Ali; Razaghpanah, Abbas; Chung, Taejoong; Lindorfer, Martina; Vallina-Rodriguez, Narseo; Levin, Dave; Choffnes, David (2022): "A Comparative Analysis of Certificate Pinning in Android & iOS", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [8]
- Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
- [9]
- Jueckstock, Jordan; Sarker, Shaown; Snyder, Peter; Beggs, Aidan; Papadopoulos, Panagiotis; Varvello, Matteo; Livshits, Benjamin; Kapravelos, Alexandros (2021): "Towards Realistic and Reproducible Web Crawl Measurements", in: Proceedings of the ACM Web Conference. (DOI)
- [10]
- Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)
