temporal and population evidence quotes; located verbatim in data/fulltext/2018/IMC/is-the-web-ready-for-ocsp-must-staple/paper.cols.txt.Sometimes the measurement is a query, not a crawl. You ask Censys which hosts speak TLS on a port, you pull the AndroZoo APKs for a release window, you re-analyse the crawl a 2022 paper released, you join a CT log against a domain list. Nothing about your study visits a website. This page is about that instrument: which snapshot, which query, which join key, and what you are entitled to claim when the data was collected by somebody else, from their vantage point, on their schedule, for their purpose.
It is not a directory of datasets. It is the four decisions a reviewer will ask about and the field mostly does not report.
| If your question is… | Then read |
|---|---|
| whether an archive can answer a question about the past web (Wayback, Common Crawl, HTTP Archive as recordings) | Archives |
| Certificate Transparency, scan data or CT logs as an instrument for the web PKI | TLS certificates |
| which list of websites to sample from (Tranco, CrUX, Alexa) | Website selection and Sampling |
| turning IP addresses you obtained into a claim about who owns them | IP classification |
| turning domains into categories, or trusting a classification service's labels | Website classification |
| getting APKs and driving them | Mobile and app measurement |
| whether to crawl, scan or unpack at all | Automated measurements |
| any of the above, but you did not collect the data yourself | this page |
The one thing to take away. A dataset you queried has a producer's denominator, not yours. Censys's IPv4 view is the set of ports Censys chose to scan, at the rate Censys chose to scan them, from the addresses Censys scans from. When you write “X % of hosts”, the population in that sentence is not “the Internet” — it is “hosts as seen by that producer's configuration at that moment”, and the two differ by more than measurement error. Izhikevich et al. showed that only 3% of HTTP and 6% of TLS services run on ports 80 and 443 [izhikevich2021_identifying]; Wu et al. measured what the search engines actually scan, from a year of honeypot traffic, and found that they “do not scan all ports of the entire IPv4 space once a day, and make trade-offs between different ports” — the bulk of each engine's traffic lands on a few dozen ports, and which few dozen differs by engine [2Wu, Mengying; Hong, Geng; Chen, Jinsong; Liu, Qi; Tang, Shujun; Li, Youhao; Liu, Baojun; Duan, Haixin; Yang, Min (2025): "Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Consideration", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. Both facts are about the producer, and both silently become facts about your result.
This distinction is litigated on every page that touches it, so here is the measurement. Three independent fields in the publication corpus each mean “this paper analysed data somebody else collected”, and they disagree with each other:
| Field | What it means | Papers | Denominator |
|---|---|---|---|
temporal.mode = existing-dataset | the paper's data provenance is a dataset, not a collection it ran | 2,534 | 5,342 papers with any provenance tuple |
population.samplingMethod = pre-existing-dataset | the study population was drawn from a dataset that already existed | 2,029 | 5,712 papers with any population tuple |
studyTypes contains existing-dataset-analysis | the paper reads as a re-analysis | 2,615 | 5,859 papers |
| Slice | Papers |
|---|---|
| all three fields agree — the conservative core | 1,292 |
| exactly one of the three | 892 |
| union — the inclusive reuse population | 3,389 |
The core is 38.1% of the union: 61.9% of the papers that any one of these fields calls a re-analysis are not called that by both of the others. That is not extraction noise alone — studyTypes is the least stable field in the schema — it is that “we used an existing dataset” is genuinely a spectrum from “our entire dataset is theirs” to “we ran a crawl and also joined one public list”. Read 1,292–3,389 as a range, not 3,389 as a number, and never as a share of the whole corpus.
The overlap the parent page warns about is real and it is not just scanning:
| Population | N | Also in the reuse union | Share |
|---|---|---|---|
scan-tagged (network-scan-or-probe) | 930 | 518 | 55.7% |
| the crawled population | 1,120 | 644 | 57.5% |
| all papers (the base rate) | 5,859 | 3,389 | 57.8% |
Read that table with the third row, not without it. More than half of scan papers and more than half of crawl papers also analyse data they did not collect — but so does the literature as a whole, at the same rate. Reuse is not something that happens to scanning; it is the base rate. The scan-specific figure the parent page carries is narrower: 429 of 930 scan papers (46.1%) are also tagged existing-dataset-analysis.
Going the other way is where the number bites: 1,563 of the 3,389 (46.1%) record no live crawl, no active probe and no passive collection at all — for those the query is the paper's only collection of measurement data. The methodological consequence is simple and is the reason this page exists — a sentence about a dataset you queried needs the producer's parameters in it, and a sentence about your own crawl needs yours. A paper that mixes both and reports one methods section for the pair has hidden the join.
Censys and Shodan are the clearest case because they look like an oracle and are not. Censys was introduced as “a public search engine and data processing facility backed by data collected from ongoing Internet-wide scans” [3Durumeric, Zakir; Adrian, David; Mirian, Ariana; Bailey, Michael D.; Halderman, J. Alex (2015): "A Search Engine Backed by Internet-Wide Scanning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] — a facility, with a configuration.
| The claim you want to make | What a scan search engine can actually support |
|---|---|
| “N % of hosts on the Internet run X” | N % of hosts in that service's index, on the ports it scans, at the freshness it happens to have |
| “we found N hosts running X” | you found the ones on the ports it scans. Services on unexpected ports are disproportionately insecure [izhikevich2021_identifying], so the direction of the bias is towards a rosier result |
| “N hosts are vulnerable to CVE-Y” | what the service's detection method says. For banner-based tags, Huang et al. found Nuclei “contradicts over 95% of the Shodan banner-based detections” for 18 of 21 CVEs [huang2025_trust] |
| “N hosts are not vulnerable” | nothing. The same study found 52.07% of the vulnerable endpoints its own scans identified were not reported by Shodan [huang2025_trust] |
| “this is what the Internet looked like on date” | what the index held on the date you queried, which is not the date the host was last probed |
| “our result is reproducible” | only if you say which snapshot and which query — see below |
Four properties of the producer become properties of your result, and none of them is under your control:
Coverage is not a property you can assume, and it is cheap to check. Sethuraman et al. compared the ANT/ISI IPv4 census with Censys for July 2021: 118M addresses in both, but 268M unique to the ANT census and 92M unique to Censys — the two agree on only about a quarter of the addresses either sees [sethuraman2022_ipv4]. VanderSloot et al. did the equivalent for certificates and found that aggregated CT logs and Censys snapshots “encompass over 99% of all certificates found by any of these techniques” while still missing 1.5% [6VanderSloot, Benjamin; Amann, Johanna; Bernhard, Matthew; Durumeric, Zakir; Bailey, Michael D.; Halderman, J. Alex (2016): "Towards a Complete View of the Certificate Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] — the lesson there is not the 99% but that they had to compare four perspectives to know it.
The cheap defence. Query two sources for a slice you can afford, and report the disagreement. It takes an afternoon, it is a publishable sentence, and it converts “we used Censys” into “we used Censys, and on a 10,000-host sample it agreed with our own scan on N %”. No reviewer rejects that; plenty reject the first.
A dataset is not a thing, it is a sequence of things. “We used Censys” identifies a company, not data. What identifies data is producer + product + snapshot date (or version) + the query.
Of the 2,534 papers whose provenance is an existing dataset, this is what they state — with papers that ran a live crawl as the comparison:
| Stated on the provenance tuple | existing-dataset (n=2,534) | Share | live-crawl (n=1,261) | Share |
|---|---|---|---|---|
| First date of the data | 1,353 | 53.4% | 796 | 63.1% |
| Last date of the data | 1,358 | 53.6% | 791 | 62.7% |
| How many snapshots were used | 377 | 14.9% | 386 | 30.6% |
| Cadence, or how a snapshot was chosen | 364 | 14.4% | 316 | 25.1% |
1,097 of the 2,534 (43.3%) give neither a first nor a last date for data they did not collect. This is the reverse of the archive result on Archives, where papers date their data better than live-crawl papers because the date is the point. Here the date belongs to somebody else, so it goes unrecorded — and it is the one thing a reader cannot recover.
It matters because reused data is old. Taking the venue year minus the last date of the data, for the papers that give a parsable year:
| Age of the newest data at publication | existing-dataset (n=1,329) | Share | live-crawl (n=779) | Share |
|---|---|---|---|---|
| same year | 239 | 18.0% | 172 | 22.1% |
| 1 year | 579 | 43.6% | 466 | 59.8% |
| 2 years | 275 | 20.7% | 120 | 15.4% |
| 3–4 years | 131 | 9.9% | 20 | 2.6% |
| 5–9 years | 85 | 6.4% | 1 | 0.1% |
| 10 years or more | 20 | 1.5% | 0 | 0.0% |
17.8% of dateable reuse papers describe a world at least three years older than their venue year, against 2.7% of live-crawl papers. Much of that is legitimate — a longitudinal or historical study should have old data. The problem is that you cannot tell which is which for the 43.3% that give no dates at all, and a three-year-old snapshot presented in the present tense is a different paper from one that says so.
Three from the corpus, all Censys, all checkable — and one counter-example:
censys.io — and, for its accuracy simulation, the window and the cadence: monthly snapshots from 09/2015 to 03/2016, seven of them. Range and count, so a reader knows what one data point is.2)And the counter-example, from the same corpus and the same service: “We analyzed TLS handshakes using the data provided at censys.io.”4) That sentence cannot be reproduced, compared or dated. It is not an outlier: 43.3% of the 2,534 existing-dataset papers have no first or last date recorded for their data at all.
The second irreproducible half. A snapshot plus a query is a dataset; a snapshot alone is a shelf.
SELECT, a CT log filter, a jq expression over a JSONL dump — these are short, and they are the method. Put them in the artifact, not only in prose.443.https.tls.version, or its v3 equivalent — is not. Name the schema generation as well as the field, because the two do not use the same names.(IP, port), service, certificate, eTLD+1, package name — these give different Ns from the same query, and the difference is often larger than the effect you are reporting.https://api.platform.censys.io/v3/, and the older Search API is being retired;5) code written against the old Search API in a 2019 paper does not run today.Only 42.9% of the 2,029 papers that drew their population from a pre-existing dataset name a version for it — indistinguishable from the 43.0% corpus-wide rate for any population list, and below the 57.1% of papers that took an exhaustive population and the 49.3% that took a top-n. Versioning is not harder here; it is just nobody's job.
population.samplingMethod | Papers | Names a version | Share |
|---|---|---|---|
| exhaustive | 1,230 | 702 | 57.1% |
| top-n | 950 | 468 | 49.3% |
| pre-existing-dataset | 2,029 | 871 | 42.9% |
| stratified | 240 | 83 | 34.6% |
| seed-and-crawl | 328 | 108 | 32.9% |
| random | 1,002 | 277 | 27.6% |
| purposive | 2,744 | 744 | 27.1% |
| convenience | 1,222 | 185 | 15.1% |
| snowball | 89 | 9 | 10.1% |
Almost every reuse study is a join: this scan corpus against that domain list, this APK set against those store listings, these certificates against that ranking. The join key is where the study silently changes population, and the corpus cannot measure it for you — so this section is practice, not a figure.
(IP, port), domain, eTLD+1, registered domain, host, certificate SHA-256, APK SHA-256, package name, ASN. A one-to-many join inflates your N; the inflated N is the one that ends up in the abstract.www. stripping, case, port defaults, IPv6 zero-compression. Two spellings of the same key are a silent left-anti-join.The commonest existing dataset in this literature is not a service, it is somebody's released crawl or scan. It is also the one with no documentation team behind it, and the corpus shows two reasons to treat it carefully: 9 of the 60 residue strings hand-audited in the next section name a prior paper's dataset by citation rather than by name — the reader has to chase a reference to find out what the data even is — and only 57.7% of the 3,389 reuse papers release an artifact of their own (against 61.1% of the crawled population), so the chain usually stops with you.
Before you build on one, establish five things, and put the answers in your methods section:
third_party encodes somebody's definition of third party. If you do not reproduce their definition, you are measuring a different thing under their column name.And the thing a re-analysis cannot recover: re-running their measurement is not the same as re-analysing their file. Demir et al. ran a web measurement over 4.5 million pages with 24 different setups and conclude that “slight differences in the experimental setup directly affect the overall results” [7Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)]. A released dataset is one setup's output. You inherit it whole — its browser, its vantage point, its interaction depth — and no amount of re-analysis reveals how much of your result is that setup. If your claim depends on that, you need your own collection. Longitudinal is the page for doing it twice on purpose.
Of the 3,389 reuse papers, 2,919 (86.1%) name at least one source. Those names are free text, so they are folded into families before counting and counted by paper; the fold rules and the full residue are on existing_datasets.
| Internet, web and security data source | Papers | Share of 3,389 | Raw spellings folded |
|---|---|---|---|
| App store or app corpus (AndroZoo, Google Play, F-Droid) | 109 | 3.2% | 54 |
| RIPE and RIR data other than Atlas/RIS | 93 | 2.7% | 100 |
| CAIDA (Ark, ITDK, telescope, AS relationships) | 85 | 2.5% | 109 |
| Popularity ranking list (Tranco, Alexa, CrUX, Umbrella) | 77 | 2.3% | 61 |
| Vulnerability database (NVD, CVE, KEV, Snyk) | 76 | 2.2% | 78 |
| RouteViews | 67 | 2.0% | 60 |
| Passive or active DNS (CZDS, zone files, DGArchive) | 56 | 1.7% | 63 |
| Traffic trace archive (MAWI, CRAWDAD, CIC, DARPA) | 52 | 1.5% | 73 |
| VirusTotal | 51 | 1.5% | 33 |
| Censys | 51 | 1.5% | 35 |
| Threat feed or blocklist (PhishTank, OpenPhish, APWG, Safe Browsing) | 48 | 1.4% | 51 |
| RIPE RIS | 43 | 1.3% | 40 |
| Malware corpus (VirusShare, Drebin, Genome, EMBER, Koodous) | 40 | 1.2% | 40 |
| Other public data service (OpenStreetMap, OpenSky, WiGLE, Google Trends) | 34 | 1.0% | 19 |
| Common Crawl | 30 | 0.9% | 21 |
| Bug tracker / fuzzing corpus (syzbot, OSS-Fuzz, Bugzilla) | 29 | 0.9% | 22 |
| RIPE Atlas | 28 | 0.8% | 23 |
| Certificate Transparency | 28 | 0.8% | 27 |
| Tor Metrics / CollecTor | 25 | 0.7% | 17 |
| Farsight DNSDB / SIE | 23 | 0.7% | 24 |
| Rapid7 / scans.io | 21 | 0.6% | 23 |
| Internet Archive / Wayback | 17 | 0.5% | 10 |
| Censorship list (Citizen Lab, OONI) | 17 | 0.5% | 20 |
| Website-fingerprinting trace set (Wang, AWF, DF, BigEnough) | 17 | 0.5% | 20 |
| HTTP Archive | 15 | 0.4% | 4 |
| OpenINTEL | 14 | 0.4% | 11 |
| M-Lab | 13 | 0.4% | 16 |
| BGPStream / BGPmon | 13 | 0.4% | 7 |
| IPv6 Hitlist | 13 | 0.4% | 11 |
| Shodan | 9 | 0.3% | 8 |
| Filter list / tracker database | 9 | 0.3% | 11 |
| iPlane / DIMES / mrinfo | 5 | 0.1% | 5 |
| Address-space census (LANDER, Trinocular) | 5 | 0.1% | 4 |
| PeeringDB | 5 | 0.1% | 4 |
| Other scan search engine | 3 | 0.1% | 5 |
Four further scopes exist in the fold and are counted but not tabulated here, because they are not what this page is about: adjacent reused data — social and UGC dumps (216 papers), blockchains (99), code hosts and package registries (92), leaked-credential corpora (39), scholarly indexes (33), the Enron email corpus (27), official statistics (25), privacy-policy corpora (14) and underground-forum corpora (10); benchmarks — ML datasets (307) and software or fuzzing suites (60); explicit industry or partner data (55); and authors' own or unnamed sources (79). All of them, with their raw-spelling counts, are in the report output on existing_datasets.
Three things in that table are worth more than the ranking itself.
Nothing is big. The largest family is named by 3.2% of reuse papers. That table is the complete Internet/web/security scope of the fold — nothing is truncated off the bottom — and it accounts for only 853 of 3,389 papers (25.2%); 367 (10.8%) name an ML or software benchmark suite instead, which is what a broad security corpus looks like. There is no shared substrate here comparable to Tranco for website selection or EasyList for tracker labelling.
Censys is named by 51 papers and Shodan by 9, which is a fact about publication, not about quality — Shodan is a commercial product with an academic tier, Censys began as a research project and runs a research-access programme. Do not read the ratio as an endorsement.
The tail is one-off and mostly not reusable. Across the 2,919 papers that name something, there are 5,103 distinct raw strings, 4,782 of which appear exactly once, and only 59 appear in five or more papers. After folding, 3,192 strings (62.6%) match no family at all, and 3,162 of those are named by a single paper. A hand-audit of a 60-string random sample of that residue (full classification on the provenance page) came out:
| What the residue string names | of 60 |
|---|---|
| a public, citable data source the fold simply has no family for | 19 |
| an ML, NLP or software benchmark, out of this page's scope | 18 |
| industry, partner or internal data a reader cannot obtain | 10 |
| a prior paper's dataset, referenced by citation rather than by name | 9 |
| too vague to identify, or not a dataset at all (“public encrypted-traffic datasets”) | 4 |
Split those last two carefully, because they are not the same problem. 10 of 60 (17%) name data a reader cannot obtain at all — commercial telemetry, a partner's logs, an anonymised industry feed. A further 9 (15%) are identifiable only by chasing a citation: they may be perfectly downloadable once you find them, but the paper's own text does not say what the data is. Data reuse in this literature is not the same thing as shared public infrastructure. Plan for that: when you build on a dataset, check that a reader can still obtain the exact thing you used, and if they cannot, deposit what you can.
That obligation is not being met either. 1,954 of the 3,389 reuse papers (57.7%) release an artifact link of their own, against 61.1% of the crawled population and 56.7% of the corpus. Re-analysis is not passing more data on than crawling does.
Checked on 2026-08-28 by fetching each primary source. A 2018–2022 methods section is not a working recipe; several of these routes have closed or changed hands.
What is in this table and what is not. It lists sources from the table above whose access route has changed since roughly 2020, plus Pushshift — which is out of this page's scope as a data source but is the clearest example of a whole method going away. A source is absent because its access route is unchanged, not because it is unimportant: VirusTotal (51 papers) and the vulnerability databases (76) are absent for that reason, and label quality for the first is Website classification. This is not a directory.
| Source | Status on 2026-08-28 | What changed |
|---|---|---|
| Censys | Alive; censys.io now redirects to censys.com | Free “Research Access to Censys Data” gives verified researchers “the same access to our data as our highest-tiered commercial customers” across the Universal Internet Dataset, certificates and deprecated IPv4 scans.6) Current interface is the Platform API v3. search.censys.io returns HTTP 403 to automated clients behind a Cloudflare interstitial — verified with a browser User-Agent — so scripted scraping of the web UI does not work |
| Shodan | Alive; academic tier unchanged in substance | Free membership upgrade for academic email addresses: “100 query credits per month”, “100 scan credits per month”, monitoring up to 16 IPs.7) Those credit limits are a real constraint on a large study and belong in your methods section |
scans.io | Alive but a different service | Now the “Stanford Internet Research Data Repository”, hosted by the Stanford Empirical Security Research Group, “restricted to non-commercial use”, indexing Censys datasets and paper artifacts.8) The University of Michigan scan dumps a 2015–2019 paper cites are not what you get by following that URL today |
| Rapid7 Project Sonar | Restricted | opendata.rapid7.com is now sonardata.rapid7.com, which “provides commercial access to data from Project Sonar” and offers only “Sign In (existing accounts only)”.9) A paper whose method is “we downloaded Sonar's SSL dataset” cannot be reproduced by a new researcher |
| PhishTank | Restricted | “New user registration temporarily disabled.”10) An existing key may still work; a new one cannot be minted, so this is a dead route for a new study |
| Pushshift (Reddit) | Gone for public use | https://api.pushshift.io/reddit/search/submission returns HTTP 403 with {"detail":"Not authenticated"} to an unauthenticated request.11) Any “we used Pushshift” method from 2018–2023 is unreproducible |
| AndroZoo | Alive and growing | “Current number of APKs: 27,616,457” on 2026-08-28.12) Access is by application from an institutional address, co-signed by a permanent-position faculty member; “API keys will automatically expire after a period of 6 months” and “The maximum number of successful APK downloads is limited to 500,000 APKs during this 6-month period”.13) A paper reporting more than 500,000 APKs used several key periods, and should say so |
| OpenINTEL | Alive | “308 million domains measured on a daily basis”, “13.6 trillion data points collected since the start in 2015”.14) |
| CAIDA | Alive; catalogue moved | www.caida.org/catalog/datasets/ now redirects to catalog.caida.org, so dataset URLs in older methods sections resolve to a search UI rather than the page they named |
| RouteViews, RIPE RIS, RIPE Atlas | Alive | RIPE Atlas still gates custom measurements behind its credit system while published results stay free; RouteViews has added API and looking-glass interfaces beyond raw MRT dumps |
| Farsight DNSDB | Alive, renamed by acquisition | DomainTools acquired Farsight in 2021; academic access is case-by-case rather than a standing programme. Cite it as “Farsight DNSDB (now DomainTools)” |
| Certificate Transparency | Alive, log API changing | Let's Encrypt shut down its RFC 6962 logs on 2026-02-28;15) CT-fetching code written against the classic endpoints is on a deprecation path. The log ecosystem, the current API and what CT can be used to measure are TLS certificates, not this page |
If you are about to deposit the snapshot you queried, do not use an OSF Project. The Center for Open Science has announced that “Starting November 16, 2026, no new projects or child components of existing projects can be created on OSF” and that “After February 19, 2027, all public and private OSF projects will become read-only”; existing DOIs keep resolving, and OSF Registries and Preprints are unaffected.16) Zenodo and institutional Dataverse instances are the going-forward options for a data deposit. Preregistration on OSF is not affected — see Study preregistration — and what belongs in a deposit is Artifacts.
A methods paragraph a reader can act on, in this order. The first four are the ones the corpus says are usually missing.
dex_date in 2024”.| Paper | Why it is on this page |
|---|---|
| Durumeric et al., CCS 2015, A Search Engine Backed by Internet-Wide Scanning [3Durumeric, Zakir; Adrian, David; Mirian, Ariana; Bailey, Michael D.; Halderman, J. Alex (2015): "A Search Engine Backed by Internet-Wide Scanning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] | The instrument paper: what a scan search engine is, in its designers' own words |
| Huang et al., USENIX Security 2025, Trust but Verify: An Assessment of Vulnerability Tagging Services [huang2025_trust] | The labels are a classifier with a precision. Banner-based CVE tags “consist almost completely out of false positives”, and 52.07% of the vulnerable endpoints its own scans found went unreported by Shodan |
| Wu et al., NDSS 2025, Revealing the Black Box of Device Search Engine [2Wu, Mengying; Hong, Geng; Chen, Jinsong; Liu, Qi; Tang, Shujun; Li, Youhao; Liu, Baojun; Duan, Haixin; Yang, Min (2025): "Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Consideration", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] | What Censys, Shodan, FOFA and ZoomEye actually scan, from a year of honeypot data. The port sets, the vantage points, and the ethics |
| Izhikevich, Teixeira and Durumeric, USENIX Security 2021, LZR: Identifying Unexpected Internet Services [izhikevich2021_identifying] | Why a fixed port set biases towards a rosier security posture: only 3% of HTTP and 6% of TLS is on the assigned port |
| Sethuraman, Bischof and Dainotti, IMC 2022, Analysis of IPv4 Address Space Utilization with ANT ISI dataset and Censys [sethuraman2022_ipv4] | Two reputable IPv4 datasets, one address space, ~25% agreement. Two pages, and it will change how you write your coverage sentence |
| VanderSloot et al., IMC 2016, Towards a Complete View of the Certificate Ecosystem [6VanderSloot, Benjamin; Amann, Johanna; Bernhard, Matthew; Durumeric, Zakir; Bailey, Michael D.; Halderman, J. Alex (2016): "Towards a Complete View of the Certificate Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | The method for establishing coverage at all: compare every perspective you can get and report what each one misses |
| Durumeric et al., IMC 2024, Ten Years of ZMap [8Durumeric, Zakir; Adrian, David; Stephens, Phillip; Wustrow, Eric; Halderman, J. Alex (2024): "Ten Years of ZMap", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | What a decade of that scan traffic became, and how the assumptions behind it were revised |
| Demir et al., TheWebConf 2022, Reproducibility and Replicability of Web Measurement Studies [7Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)] | If the dataset you are about to reuse is somebody's released crawl: 4.5 million pages, 24 setups, and “slight differences in the experimental setup directly affect the overall results” |
For an archive rather than a scan corpus, start from Hantke et al. [9Hantke, Florian; Calzavara, Stefano; Wilhelm, Moritz; Rabitti, Alvise; Stock, Ben (2023): "You Call This Archaeology? Evaluating Web Archives for Reproducible Web Security Measurements", in: Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, pp. 3168–3182. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] and Singh et al. [10Singh, Sachin Kumar; Mahmud, Faisal; Ricci, Robert; Siby, Sandra (2026): "The Empire Strikes Back (at Your Privacy): An Archaeology of Tracking on Government Websites", Proceedings on Privacy Enhancing Technologies 2026(2):108-126. (DOI)] on Archives — the anachronism trap described there (today's classifier over yesterday's data) applies to every dataset on this page, not only to archives.
Figures are from scripts/report_existing_datasets.mjs over the 5,859-paper extraction of CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026. Every claim is a claim about those seven venues. 2025–2026 venue-years are provisional — CCS and IMC 2026 have not been held, and IEEE S&P and WWW 2026 are under-selected by construction.
| Window | Papers | Reuse union | Share | Core (all three fields) | Share |
|---|---|---|---|---|---|
| 2010–2013 | 511 | 278 | 54.4% | 96 | 18.8% |
| 2014–2017 | 769 | 439 | 57.1% | 178 | 23.1% |
| 2018–2021 | 1,439 | 850 | 59.1% | 334 | 23.2% |
| 2022–2024 | 1,955 | 1,127 | 57.6% | 431 | 22.0% |
| 2025–2026 (provisional) | 1,185 | 695 | 58.6% | 253 | 21.4% |
Flat. Re-analysis is not a trend, it is a constant of this literature — which is an argument for treating it as a first-class design with its own reporting standard, rather than as the thing you say when you did not crawl.
Share of that venue's own output:
| Venue | Papers | Reuse union | Share |
|---|---|---|---|
| TheWebConf | 843 | 625 | 74.1% |
| IMC | 638 | 448 | 70.2% |
| NDSS | 701 | 432 | 61.6% |
| USENIX Security | 1,410 | 763 | 54.1% |
| CCS | 990 | 509 | 51.4% |
| IEEE S&P | 767 | 379 | 49.4% |
| PETS | 510 | 233 | 45.7% |
studyTypes is the least stable field in the schema — 57% run-to-run on a 100-paper sample, measured on the previous corpus and not re-measured on this one. It is one of the three fields in the union, so the union inherits that instability, and the core does too since it requires all three. Read both as ranking-grade. Of the three, only temporal.mode has a measured stability figure (69% on that same sample); population.samplingMethod is an enum but was not in the stability comparison, so nothing is claimed about it.not-stated and none-mentioned are counted as silence, and silence is reported as itself.paper.cols.txt; all 40 located (24 exact after whitespace collapse, 14 via an 8-word run, 2 via a 5-word run).spanEnd, so it is granular to a year and inherits the venue-year convention. 28 existing-dataset papers had a spanEnd with no parsable year and 1 had a year after its venue year; both are excluded and reported rather than absorbed.temporal and population evidence quotes; located verbatim in data/fulltext/2018/IMC/is-the-web-ready-for-ocsp-must-staple/paper.cols.txt.population.dataset and temporal.evaluation evidence quotes. Both sentences are interleaved with a table in paper.cols.txt, so they are paraphrased here rather than quoted; the extraction's verbatim quotes are reproduced on existing_datasets.population.methodology evidence quote.temporal.methodology evidence quote.https://docs.censys.com/docs/platform-api-transition-guide, fetched 2026-08-28, HTTP 200. The guide maps Search API v2 endpoints to v3 and publishes no shutdown date; it carries no v1 mapping at all, v1 having been retired earlier.https://docs.censys.com/docs/research-access-to-censys-data, fetched 2026-08-28, HTTP 200.https://help.shodan.io/the-basics/academic-upgrade, fetched 2026-08-28, HTTP 200.https://scans.io, fetched 2026-08-28, HTTP 200.https://sonardata.rapid7.com/about/, fetched 2026-08-28, HTTP 200.https://www.phishtank.com/register.php, fetched 2026-08-28, HTTP 200 with a browser User-Agent. The site rate-limits automated clients: a second fetch from a different address the same day returned HTTP 429 with retry-after: 86372. The string is also present in a Wayback capture of 2026-05-12.https://redditinc.com/policies/data-api-terms, fetched 2026-08-28, HTTP 200, most recent revision dated 2026-07-01 — and require “a separate agreement with Reddit” for “research in excess of rate limits”. The researcher-programme help pages under support.reddithelp.com refuse automated clients (403), so the programme's own eligibility text could not be checked here.https://androzoo.uni.lu/, fetched 2026-08-28, HTTP 200.https://androzoo.uni.lu/access, fetched 2026-08-28, HTTP 200.https://www.openintel.nl/, fetched 2026-08-28, HTTP 200.https://letsencrypt.org/2025/08/14/rfc-6962-logs-eol, fetched 2026-08-28, HTTP 200. Read-only from 2025-11-30.https://www.cos.io/osf-changes, fetched 2026-08-28, HTTP 200.