User Tools

Site Tools


design:existing_datasets

Measuring from an existing dataset

Sometimes the measurement is a query, not a crawl. You ask Censys which hosts speak TLS on a port, you pull the AndroZoo APKs for a release window, you re-analyse the crawl a 2022 paper released, you join a CT log against a domain list. Nothing about your study visits a website. This page is about that instrument: which snapshot, which query, which join key, and what you are entitled to claim when the data was collected by somebody else, from their vantage point, on their schedule, for their purpose.

It is not a directory of datasets. It is the four decisions a reviewer will ask about and the field mostly does not report.

If your question is… Then read
whether an archive can answer a question about the past web (Wayback, Common Crawl, HTTP Archive as recordings) Archives
Certificate Transparency, scan data or CT logs as an instrument for the web PKI TLS certificates
which list of websites to sample from (Tranco, CrUX, Alexa) Website selection and Sampling
turning IP addresses you obtained into a claim about who owns them IP classification
turning domains into categories, or trusting a classification service's labels Website classification
getting APKs and driving them Mobile and app measurement
whether to crawl, scan or unpack at all Automated measurements
any of the above, but you did not collect the data yourself this page

The one thing to take away. A dataset you queried has a producer's denominator, not yours. Censys's IPv4 view is the set of ports Censys chose to scan, at the rate Censys chose to scan them, from the addresses Censys scans from. When you write “X % of hosts”, the population in that sentence is not “the Internet” — it is “hosts as seen by that producer's configuration at that moment”, and the two differ by more than measurement error. Izhikevich et al. showed that only 3% of HTTP and 6% of TLS services run on ports 80 and 443 [izhikevich2021_identifying]; Wu et al. measured what the search engines actually scan, from a year of honeypot traffic, and found that they “do not scan all ports of the entire IPv4 space once a day, and make trade-offs between different ports” — the bulk of each engine's traffic lands on a few dozen ports, and which few dozen differs by engine [2Wu, Mengying; Hong, Geng; Chen, Jinsong; Liu, Qi; Tang, Shujun; Li, Youhao; Liu, Baojun; Duan, Haixin; Yang, Min (2025): "Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Consideration", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. Both facts are about the producer, and both silently become facts about your result.

"We scanned" versus "we queried"

This distinction is litigated on every page that touches it, so here is the measurement. Three independent fields in the publication corpus each mean “this paper analysed data somebody else collected”, and they disagree with each other:

Field What it means Papers Denominator
temporal.mode = existing-dataset the paper's data provenance is a dataset, not a collection it ran 2,534 5,342 papers with any provenance tuple
population.samplingMethod = pre-existing-dataset the study population was drawn from a dataset that already existed 2,029 5,712 papers with any population tuple
studyTypes contains existing-dataset-analysis the paper reads as a re-analysis 2,615 5,859 papers
Slice Papers
all three fields agree — the conservative core 1,292
exactly one of the three 892
union — the inclusive reuse population 3,389

The core is 38.1% of the union: 61.9% of the papers that any one of these fields calls a re-analysis are not called that by both of the others. That is not extraction noise alone — studyTypes is the least stable field in the schema — it is that “we used an existing dataset” is genuinely a spectrum from “our entire dataset is theirs” to “we ran a crawl and also joined one public list”. Read 1,292–3,389 as a range, not 3,389 as a number, and never as a share of the whole corpus.

The overlap the parent page warns about is real and it is not just scanning:

Population N Also in the reuse union Share
scan-tagged (network-scan-or-probe) 930 518 55.7%
the crawled population 1,120 644 57.5%
all papers (the base rate) 5,859 3,389 57.8%

Read that table with the third row, not without it. More than half of scan papers and more than half of crawl papers also analyse data they did not collect — but so does the literature as a whole, at the same rate. Reuse is not something that happens to scanning; it is the base rate. The scan-specific figure the parent page carries is narrower: 429 of 930 scan papers (46.1%) are also tagged existing-dataset-analysis.

Going the other way is where the number bites: 1,563 of the 3,389 (46.1%) record no live crawl, no active probe and no passive collection at all — for those the query is the paper's only collection of measurement data. The methodological consequence is simple and is the reason this page exists — a sentence about a dataset you queried needs the producer's parameters in it, and a sentence about your own crawl needs yours. A paper that mixes both and reports one methods section for the pair has hidden the join.

What you cannot claim from a search engine over someone else's scans

Censys and Shodan are the clearest case because they look like an oracle and are not. Censys was introduced as “a public search engine and data processing facility backed by data collected from ongoing Internet-wide scans” [3Durumeric, Zakir; Adrian, David; Mirian, Ariana; Bailey, Michael D.; Halderman, J. Alex (2015): "A Search Engine Backed by Internet-Wide Scanning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] — a facility, with a configuration.

The claim you want to make What a scan search engine can actually support
N % of hosts on the Internet run X N % of hosts in that service's index, on the ports it scans, at the freshness it happens to have
“we found N hosts running X you found the ones on the ports it scans. Services on unexpected ports are disproportionately insecure [izhikevich2021_identifying], so the direction of the bias is towards a rosier result
N hosts are vulnerable to CVE-Y what the service's detection method says. For banner-based tags, Huang et al. found Nuclei “contradicts over 95% of the Shodan banner-based detections” for 18 of 21 CVEs [huang2025_trust]
N hosts are not vulnerable” nothing. The same study found 52.07% of the vulnerable endpoints its own scans identified were not reported by Shodan [huang2025_trust]
“this is what the Internet looked like on date what the index held on the date you queried, which is not the date the host was last probed
“our result is reproducible” only if you say which snapshot and which query — see below

Four properties of the producer become properties of your result, and none of them is under your control:

  • The port set. Wu et al. instrumented 28 honeypots for a year and traced 1,407 scanner IPs: “20% of Shodan and ZoomEye's traffic targets 29 ports, whereas Censys scans 49 ports, and FOFA targets seven ports” [2Wu, Mengying; Hong, Geng; Chen, Jinsong; Liu, Qi; Tang, Shujun; Li, Youhao; Liu, Baojun; Duan, Haixin; Yang, Min (2025): "Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Consideration", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. Against Izhikevich et al.'s finding that protocol deployment is diffuse [izhikevich2021_identifying], a fixed port set is a systematic, one-directional undercount of exactly the insecure tail.
  • The vantage point. The same study traced every Censys scanner address it found to the United States, on enterprise ISPs. The others sit elsewhere and on different kinds of network: 67% of ZoomEye's and 72% of FOFA's addresses are in China, FOFA on cloud ISPs and ZoomEye on consumer ISPs, while Shodan mixes enterprise and cloud — 24.69% of Shodan's addresses are in the Netherlands and 19.17% of FOFA's in Finland [2Wu, Mengying; Hong, Geng; Chen, Jinsong; Liu, Qi; Tang, Shujun; Li, Youhao; Liu, Baojun; Duan, Haixin; Yang, Min (2025): "Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Consideration", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]. A dataset collected from one country inherits every geographic blocking and geo-differentiation effect on Crawling location — you simply cannot see them, because you have one vantage point and no control.
  • The cadence. “Not scanned all ports once a day” means the record you read has an age, and the age varies by port. If your claim is temporal, you are reading a smear, not a snapshot.
  • The labels. A CVE tag, a device type, a “known software” label is a classifier the producer runs, and it has a precision. In a ground-truth experiment with 43 deployed Docker instances, Huang et al. measured Shodan at 85.71% precision and 75.00% recall and ONYPHE at 80.00% precision for the CVEs each tracks [huang2025_trust]. Their conclusion is directed at us: “the widespread use of banner-based vulnerability detection in academic work is problematic”.

Coverage is not a property you can assume, and it is cheap to check. Sethuraman et al. compared the ANT/ISI IPv4 census with Censys for July 2021: 118M addresses in both, but 268M unique to the ANT census and 92M unique to Censys — the two agree on only about a quarter of the addresses either sees [sethuraman2022_ipv4]. VanderSloot et al. did the equivalent for certificates and found that aggregated CT logs and Censys snapshots “encompass over 99% of all certificates found by any of these techniques” while still missing 1.5% [6VanderSloot, Benjamin; Amann, Johanna; Bernhard, Matthew; Durumeric, Zakir; Bailey, Michael D.; Halderman, J. Alex (2016): "Towards a Complete View of the Certificate Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] — the lesson there is not the 99% but that they had to compare four perspectives to know it.

The cheap defence. Query two sources for a slice you can afford, and report the disagreement. It takes an afternoon, it is a publishable sentence, and it converts “we used Censys” into “we used Censys, and on a 10,000-host sample it agreed with our own scan on N %”. No reviewer rejects that; plenty reject the first.

Which snapshot

A dataset is not a thing, it is a sequence of things. “We used Censys” identifies a company, not data. What identifies data is producer + product + snapshot date (or version) + the query.

Of the 2,534 papers whose provenance is an existing dataset, this is what they state — with papers that ran a live crawl as the comparison:

Stated on the provenance tuple existing-dataset (n=2,534) Share live-crawl (n=1,261) Share
First date of the data 1,353 53.4% 796 63.1%
Last date of the data 1,358 53.6% 791 62.7%
How many snapshots were used 377 14.9% 386 30.6%
Cadence, or how a snapshot was chosen 364 14.4% 316 25.1%

1,097 of the 2,534 (43.3%) give neither a first nor a last date for data they did not collect. This is the reverse of the archive result on Archives, where papers date their data better than live-crawl papers because the date is the point. Here the date belongs to somebody else, so it goes unrecorded — and it is the one thing a reader cannot recover.

It matters because reused data is old. Taking the venue year minus the last date of the data, for the papers that give a parsable year:

Age of the newest data at publication existing-dataset (n=1,329) Share live-crawl (n=779) Share
same year 239 18.0% 172 22.1%
1 year 579 43.6% 466 59.8%
2 years 275 20.7% 120 15.4%
3–4 years 131 9.9% 20 2.6%
5–9 years 85 6.4% 1 0.1%
10 years or more 20 1.5% 0 0.0%

17.8% of dateable reuse papers describe a world at least three years older than their venue year, against 2.7% of live-crawl papers. Much of that is legitimate — a longitudinal or historical study should have old data. The problem is that you cannot tell which is which for the 43.3% that give no dates at all, and a three-year-old snapshot presented in the present tense is a different paper from one that says so.

What a snapshot statement looks like

Three from the corpus, all Censys, all checkable — and one counter-example:

  • “We use the Censys snapshot of this dataset that was collected on April 24th, 2018, which contains 489,580,002 certificates” — producer, product, exact date, size.1)
  • IMC 2016's Towards Better Internet Citizenship gives both the volume and the count — 4.1 TB derived from 28 full IPv4 scans obtained from censys.io — and, for its accuracy simulation, the window and the cadence: monthly snapshots from 09/2015 to 03/2016, seven of them. Range and count, so a reader knows what one data point is.2)
  • “We queried Censys on February 13, 2017.” — the query date, which is what you actually control.3)

And the counter-example, from the same corpus and the same service: “We analyzed TLS handshakes using the data provided at censys.io.”4) That sentence cannot be reproduced, compared or dated. It is not an outlier: 43.3% of the 2,534 existing-dataset papers have no first or last date recorded for their data at all.

Which query

The second irreproducible half. A snapshot plus a query is a dataset; a snapshot alone is a shelf.

  1. Publish the query text, not a description of it. A Censys or Shodan search string, a BigQuery SELECT, a CT log filter, a jq expression over a JSONL dump — these are short, and they are the method. Put them in the artifact, not only in prose.
  2. Say which fields you read and which you ignored. “Hosts with TLS” is ambiguous across a producer's schema versions; a fully qualified field path — Censys's legacy 443.https.tls.version, or its v3 equivalent — is not. Name the schema generation as well as the field, because the two do not use the same names.
  3. Say what you deduplicated on and at what level. Host, IP, (IP, port), service, certificate, eTLD+1, package name — these give different Ns from the same query, and the difference is often larger than the effect you are reporting.
  4. Say what you excluded. Cloud ranges, CDNs, parked domains, non-2XX responses, apps below a download threshold. Exclusions move headline percentages more than most modelling choices.
  5. Pin the version of the producer's schema or API. Censys's current interface is the Platform API v3 at https://api.platform.censys.io/v3/, and the older Search API is being retired;5) code written against the old Search API in a 2019 paper does not run today.

Only 42.9% of the 2,029 papers that drew their population from a pre-existing dataset name a version for it — indistinguishable from the 43.0% corpus-wide rate for any population list, and below the 57.1% of papers that took an exhaustive population and the 49.3% that took a top-n. Versioning is not harder here; it is just nobody's job.

population.samplingMethod Papers Names a version Share
exhaustive 1,230 702 57.1%
top-n 950 468 49.3%
pre-existing-dataset 2,029 871 42.9%
stratified 240 83 34.6%
seed-and-crawl 328 108 32.9%
random 1,002 277 27.6%
purposive 2,744 744 27.1%
convenience 1,222 185 15.1%
snowball 89 9 10.1%

Which join key

Almost every reuse study is a join: this scan corpus against that domain list, this APK set against those store listings, these certificates against that ranking. The join key is where the study silently changes population, and the corpus cannot measure it for you — so this section is practice, not a figure.

  • Name the key and its cardinality. IP, (IP, port), domain, eTLD+1, registered domain, host, certificate SHA-256, APK SHA-256, package name, ASN. A one-to-many join inflates your N; the inflated N is the one that ends up in the abstract.
  • Both sides must be normalised the same way. Trailing dots, IDN vs punycode, www. stripping, case, port defaults, IPv6 zero-compression. Two spellings of the same key are a silent left-anti-join.
  • The join is a filter. Report the size of both inputs and the size of the result. “We joined A with B” without three numbers hides an arbitrary sample.
  • A domain is not a site and an IP is not a host. Turning either into an organisation or a category is its own instrument with its own error rate: IP classification for addresses, Website classification for domains. Do not treat a lookup as a fact.
  • Time is a join dimension. Joining a 2024 scan to a 2021 domain list produces a population that existed at neither moment. If the two sides have different dates, say so and say which one your claim is about.
  • Do not join on a label produced by the same vendor on both sides. Correlation between two of a producer's own derived fields tells you about their pipeline, not about the world.

When the dataset is another paper's artifact

The commonest existing dataset in this literature is not a service, it is somebody's released crawl or scan. It is also the one with no documentation team behind it, and the corpus shows two reasons to treat it carefully: 9 of the 60 residue strings hand-audited in the next section name a prior paper's dataset by citation rather than by name — the reader has to chase a reference to find out what the data even is — and only 57.7% of the 3,389 reuse papers release an artifact of their own (against 61.1% of the crawled population), so the chain usually stops with you.

Before you build on one, establish five things, and put the answers in your methods section:

  1. Does the deposit still resolve, and is it the version the paper used? A GitHub repository moves; a DOI does not. If the artifact is a repository, record the commit, not the URL.
  2. Is it their whole population or their filtered one? Papers routinely release the analysed subset rather than the collected set. Your denominator is then their exclusion rules, which are in their prose and not in their tarball.
  3. What are the collection dates of the data, as opposed to the paper's publication year? They are usually only in the paper. Carry them forward — this is the field that section Which snapshot is about, and it is lost at exactly this hand-off.
  4. Are the field semantics written down anywhere but the paper? A column called third_party encodes somebody's definition of third party. If you do not reproduce their definition, you are measuring a different thing under their column name.
  5. Are you allowed to redistribute it? Many corpora are licensed for the recipient only, which decides whether your own artifact can include the join or only the code that performs it.

And the thing a re-analysis cannot recover: re-running their measurement is not the same as re-analysing their file. Demir et al. ran a web measurement over 4.5 million pages with 24 different setups and conclude that “slight differences in the experimental setup directly affect the overall results” [7Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)]. A released dataset is one setup's output. You inherit it whole — its browser, its vantage point, its interaction depth — and no amount of re-analysis reveals how much of your result is that setup. If your claim depends on that, you need your own collection. Longitudinal is the page for doing it twice on purpose.

What the field actually queries

Of the 3,389 reuse papers, 2,919 (86.1%) name at least one source. Those names are free text, so they are folded into families before counting and counted by paper; the fold rules and the full residue are on existing_datasets.

Internet, web and security data source Papers Share of 3,389 Raw spellings folded
App store or app corpus (AndroZoo, Google Play, F-Droid) 109 3.2% 54
RIPE and RIR data other than Atlas/RIS 93 2.7% 100
CAIDA (Ark, ITDK, telescope, AS relationships) 85 2.5% 109
Popularity ranking list (Tranco, Alexa, CrUX, Umbrella) 77 2.3% 61
Vulnerability database (NVD, CVE, KEV, Snyk) 76 2.2% 78
RouteViews 67 2.0% 60
Passive or active DNS (CZDS, zone files, DGArchive) 56 1.7% 63
Traffic trace archive (MAWI, CRAWDAD, CIC, DARPA) 52 1.5% 73
VirusTotal 51 1.5% 33
Censys 51 1.5% 35
Threat feed or blocklist (PhishTank, OpenPhish, APWG, Safe Browsing) 48 1.4% 51
RIPE RIS 43 1.3% 40
Malware corpus (VirusShare, Drebin, Genome, EMBER, Koodous) 40 1.2% 40
Other public data service (OpenStreetMap, OpenSky, WiGLE, Google Trends) 34 1.0% 19
Common Crawl 30 0.9% 21
Bug tracker / fuzzing corpus (syzbot, OSS-Fuzz, Bugzilla) 29 0.9% 22
RIPE Atlas 28 0.8% 23
Certificate Transparency 28 0.8% 27
Tor Metrics / CollecTor 25 0.7% 17
Farsight DNSDB / SIE 23 0.7% 24
Rapid7 / scans.io 21 0.6% 23
Internet Archive / Wayback 17 0.5% 10
Censorship list (Citizen Lab, OONI) 17 0.5% 20
Website-fingerprinting trace set (Wang, AWF, DF, BigEnough) 17 0.5% 20
HTTP Archive 15 0.4% 4
OpenINTEL 14 0.4% 11
M-Lab 13 0.4% 16
BGPStream / BGPmon 13 0.4% 7
IPv6 Hitlist 13 0.4% 11
Shodan 9 0.3% 8
Filter list / tracker database 9 0.3% 11
iPlane / DIMES / mrinfo 5 0.1% 5
Address-space census (LANDER, Trinocular) 5 0.1% 4
PeeringDB 5 0.1% 4
Other scan search engine 3 0.1% 5

Four further scopes exist in the fold and are counted but not tabulated here, because they are not what this page is about: adjacent reused data — social and UGC dumps (216 papers), blockchains (99), code hosts and package registries (92), leaked-credential corpora (39), scholarly indexes (33), the Enron email corpus (27), official statistics (25), privacy-policy corpora (14) and underground-forum corpora (10); benchmarksML datasets (307) and software or fuzzing suites (60); explicit industry or partner data (55); and authors' own or unnamed sources (79). All of them, with their raw-spelling counts, are in the report output on existing_datasets.

Three things in that table are worth more than the ranking itself.

Nothing is big. The largest family is named by 3.2% of reuse papers. That table is the complete Internet/web/security scope of the fold — nothing is truncated off the bottom — and it accounts for only 853 of 3,389 papers (25.2%); 367 (10.8%) name an ML or software benchmark suite instead, which is what a broad security corpus looks like. There is no shared substrate here comparable to Tranco for website selection or EasyList for tracker labelling.

Censys is named by 51 papers and Shodan by 9, which is a fact about publication, not about quality — Shodan is a commercial product with an academic tier, Censys began as a research project and runs a research-access programme. Do not read the ratio as an endorsement.

The tail is one-off and mostly not reusable. Across the 2,919 papers that name something, there are 5,103 distinct raw strings, 4,782 of which appear exactly once, and only 59 appear in five or more papers. After folding, 3,192 strings (62.6%) match no family at all, and 3,162 of those are named by a single paper. A hand-audit of a 60-string random sample of that residue (full classification on the provenance page) came out:

What the residue string names of 60
a public, citable data source the fold simply has no family for 19
an ML, NLP or software benchmark, out of this page's scope 18
industry, partner or internal data a reader cannot obtain 10
a prior paper's dataset, referenced by citation rather than by name 9
too vague to identify, or not a dataset at all (“public encrypted-traffic datasets”) 4

Split those last two carefully, because they are not the same problem. 10 of 60 (17%) name data a reader cannot obtain at all — commercial telemetry, a partner's logs, an anonymised industry feed. A further 9 (15%) are identifiable only by chasing a citation: they may be perfectly downloadable once you find them, but the paper's own text does not say what the data is. Data reuse in this literature is not the same thing as shared public infrastructure. Plan for that: when you build on a dataset, check that a reader can still obtain the exact thing you used, and if they cannot, deposit what you can.

That obligation is not being met either. 1,954 of the 3,389 reuse papers (57.7%) release an artifact link of their own, against 61.1% of the crawled population and 56.7% of the corpus. Re-analysis is not passing more data on than crawling does.

Currency: what you can still get in 2026

Checked on 2026-08-28 by fetching each primary source. A 2018–2022 methods section is not a working recipe; several of these routes have closed or changed hands.

What is in this table and what is not. It lists sources from the table above whose access route has changed since roughly 2020, plus Pushshift — which is out of this page's scope as a data source but is the clearest example of a whole method going away. A source is absent because its access route is unchanged, not because it is unimportant: VirusTotal (51 papers) and the vulnerability databases (76) are absent for that reason, and label quality for the first is Website classification. This is not a directory.

Source Status on 2026-08-28 What changed
Censys Alive; censys.io now redirects to censys.com Free “Research Access to Censys Data” gives verified researchers “the same access to our data as our highest-tiered commercial customers” across the Universal Internet Dataset, certificates and deprecated IPv4 scans.6) Current interface is the Platform API v3. search.censys.io returns HTTP 403 to automated clients behind a Cloudflare interstitial — verified with a browser User-Agent — so scripted scraping of the web UI does not work
Shodan Alive; academic tier unchanged in substance Free membership upgrade for academic email addresses: “100 query credits per month”, “100 scan credits per month”, monitoring up to 16 IPs.7) Those credit limits are a real constraint on a large study and belong in your methods section
scans.io Alive but a different service Now the “Stanford Internet Research Data Repository”, hosted by the Stanford Empirical Security Research Group, “restricted to non-commercial use”, indexing Censys datasets and paper artifacts.8) The University of Michigan scan dumps a 2015–2019 paper cites are not what you get by following that URL today
Rapid7 Project Sonar Restricted opendata.rapid7.com is now sonardata.rapid7.com, which “provides commercial access to data from Project Sonar” and offers only “Sign In (existing accounts only)”.9) A paper whose method is “we downloaded Sonar's SSL dataset” cannot be reproduced by a new researcher
PhishTank Restricted “New user registration temporarily disabled.”10) An existing key may still work; a new one cannot be minted, so this is a dead route for a new study
Pushshift (Reddit) Gone for public use https://api.pushshift.io/reddit/search/submission returns HTTP 403 with {"detail":"Not authenticated"} to an unauthenticated request.11) Any “we used Pushshift” method from 2018–2023 is unreproducible
AndroZoo Alive and growing “Current number of APKs: 27,616,457” on 2026-08-28.12) Access is by application from an institutional address, co-signed by a permanent-position faculty member; API keys will automatically expire after a period of 6 months” and “The maximum number of successful APK downloads is limited to 500,000 APKs during this 6-month period”.13) A paper reporting more than 500,000 APKs used several key periods, and should say so
OpenINTEL Alive “308 million domains measured on a daily basis”, “13.6 trillion data points collected since the start in 2015”.14)
CAIDA Alive; catalogue moved www.caida.org/catalog/datasets/ now redirects to catalog.caida.org, so dataset URLs in older methods sections resolve to a search UI rather than the page they named
RouteViews, RIPE RIS, RIPE Atlas Alive RIPE Atlas still gates custom measurements behind its credit system while published results stay free; RouteViews has added API and looking-glass interfaces beyond raw MRT dumps
Farsight DNSDB Alive, renamed by acquisition DomainTools acquired Farsight in 2021; academic access is case-by-case rather than a standing programme. Cite it as “Farsight DNSDB (now DomainTools)”
Certificate Transparency Alive, log API changing Let's Encrypt shut down its RFC 6962 logs on 2026-02-28;15) CT-fetching code written against the classic endpoints is on a deprecation path. The log ecosystem, the current API and what CT can be used to measure are TLS certificates, not this page

If you are about to deposit the snapshot you queried, do not use an OSF Project. The Center for Open Science has announced that “Starting November 16, 2026, no new projects or child components of existing projects can be created on OSF” and that “After February 19, 2027, all public and private OSF projects will become read-only”; existing DOIs keep resolving, and OSF Registries and Preprints are unaffected.16) Zenodo and institutional Dataverse instances are the going-forward options for a data deposit. Preregistration on OSF is not affected — see Study preregistration — and what belongs in a deposit is Artifacts.

What to report

A methods paragraph a reader can act on, in this order. The first four are the ones the corpus says are usually missing.

  1. Producer, product and snapshot. Not “Censys” but “the Censys Universal Internet Dataset snapshot of 2026-03-14”. Not “AndroZoo” but “the AndroZoo index as of 2026-03, apps with dex_date in 2024”.
  2. The query, verbatim, in the artifact. Search string, SQL, filter expression, API call with parameters.
  3. The producer's parameters that bound your claim — port set, vantage point, scan cadence, list source, inclusion rule — as far as the producer documents them, and a sentence saying which ones you could not establish.
  4. Dates: collection and query. When the producer collected it, and when you pulled it. They are different, and only the second is under your control.
  5. The join: key, normalisation, size of each input, size of the result.
  6. Deduplication and exclusion rules, with the N before and after each.
  7. A validation slice. Something you checked against a second source or against your own collection, with the agreement rate. One number, and it is the difference between “we used Censys” and a defensible claim.
  8. Where your derived data lives, so the next paper joins to yours rather than re-deriving it. See the OSF warning above and Artifacts.

What to read first

Paper Why it is on this page
Durumeric et al., CCS 2015, A Search Engine Backed by Internet-Wide Scanning [3Durumeric, Zakir; Adrian, David; Mirian, Ariana; Bailey, Michael D.; Halderman, J. Alex (2015): "A Search Engine Backed by Internet-Wide Scanning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] The instrument paper: what a scan search engine is, in its designers' own words
Huang et al., USENIX Security 2025, Trust but Verify: An Assessment of Vulnerability Tagging Services [huang2025_trust] The labels are a classifier with a precision. Banner-based CVE tags “consist almost completely out of false positives”, and 52.07% of the vulnerable endpoints its own scans found went unreported by Shodan
Wu et al., NDSS 2025, Revealing the Black Box of Device Search Engine [2Wu, Mengying; Hong, Geng; Chen, Jinsong; Liu, Qi; Tang, Shujun; Li, Youhao; Liu, Baojun; Duan, Haixin; Yang, Min (2025): "Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Consideration", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] What Censys, Shodan, FOFA and ZoomEye actually scan, from a year of honeypot data. The port sets, the vantage points, and the ethics
Izhikevich, Teixeira and Durumeric, USENIX Security 2021, LZR: Identifying Unexpected Internet Services [izhikevich2021_identifying] Why a fixed port set biases towards a rosier security posture: only 3% of HTTP and 6% of TLS is on the assigned port
Sethuraman, Bischof and Dainotti, IMC 2022, Analysis of IPv4 Address Space Utilization with ANT ISI dataset and Censys [sethuraman2022_ipv4] Two reputable IPv4 datasets, one address space, ~25% agreement. Two pages, and it will change how you write your coverage sentence
VanderSloot et al., IMC 2016, Towards a Complete View of the Certificate Ecosystem [6VanderSloot, Benjamin; Amann, Johanna; Bernhard, Matthew; Durumeric, Zakir; Bailey, Michael D.; Halderman, J. Alex (2016): "Towards a Complete View of the Certificate Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] The method for establishing coverage at all: compare every perspective you can get and report what each one misses
Durumeric et al., IMC 2024, Ten Years of ZMap [8Durumeric, Zakir; Adrian, David; Stephens, Phillip; Wustrow, Eric; Halderman, J. Alex (2024): "Ten Years of ZMap", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] What a decade of that scan traffic became, and how the assumptions behind it were revised
Demir et al., TheWebConf 2022, Reproducibility and Replicability of Web Measurement Studies [7Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)] If the dataset you are about to reuse is somebody's released crawl: 4.5 million pages, 24 setups, and “slight differences in the experimental setup directly affect the overall results”

For an archive rather than a scan corpus, start from Hantke et al. [9Hantke, Florian; Calzavara, Stefano; Wilhelm, Moritz; Rabitti, Alvise; Stock, Ben (2023): "You Call This Archaeology? Evaluating Web Archives for Reproducible Web Security Measurements", in: Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, pp. 3168–3182. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] and Singh et al. [10Singh, Sachin Kumar; Mahmud, Faisal; Ricci, Robert; Siby, Sandra (2026): "The Empire Strikes Back (at Your Privacy): An Archaeology of Tracking on Government Websites", Proceedings on Privacy Enhancing Technologies 2026(2):108-126. (DOI)] on Archives — the anachronism trap described there (today's classifier over yesterday's data) applies to every dataset on this page, not only to archives.

Use in Publications

Figures are from scripts/report_existing_datasets.mjs over the 5,859-paper extraction of CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010–2026. Every claim is a claim about those seven venues. 2025–2026 venue-years are provisional — CCS and IMC 2026 have not been held, and IEEE S&P and WWW 2026 are under-selected by construction.

Over the years

Window Papers Reuse union Share Core (all three fields) Share
2010–2013 511 278 54.4% 96 18.8%
2014–2017 769 439 57.1% 178 23.1%
2018–2021 1,439 850 59.1% 334 23.2%
2022–2024 1,955 1,127 57.6% 431 22.0%
2025–2026 (provisional) 1,185 695 58.6% 253 21.4%

Flat. Re-analysis is not a trend, it is a constant of this literature — which is an argument for treating it as a first-class design with its own reporting standard, rather than as the thing you say when you did not crawl.

By venue

Share of that venue's own output:

Venue Papers Reuse union Share
TheWebConf 843 625 74.1%
IMC 638 448 70.2%
NDSS 701 432 61.6%
USENIX Security 1,410 763 54.1%
CCS 990 509 51.4%
IEEE S&P 767 379 49.4%
PETS 510 233 45.7%

Methodology and limitations of these figures

  • Populations. Every table names its own denominator. “All 5,859 papers” is used only where the field applies to every paper. The reuse population (3,389) is the union of three fields and the core (1,292) is their intersection; both are given because the difference between them is a finding, not a rounding error.
  • studyTypes is the least stable field in the schema — 57% run-to-run on a 100-paper sample, measured on the previous corpus and not re-measured on this one. It is one of the three fields in the union, so the union inherits that instability, and the core does too since it requires all three. Read both as ranking-grade. Of the three, only temporal.mode has a measured stability figure (69% on that same sample); population.samplingMethod is an enum but was not in the stability comparison, so nothing is claimed about it.
  • Sentinels are never answers. not-stated and none-mentioned are counted as silence, and silence is reported as itself.
  • Counts are of papers, never tuples. A paper naming Censys four times is one paper.
  • Free-text dataset names were folded into families before counting, by an all-matches rule so that one string naming four sources counts for all four. The fold and its 3,192-string residue are published in full on existing_datasets, along with the hand-audit of the residue sample.
  • Quotes were spot-checked. 40 evidence quotes behind the dataset-family and old-data figures were matched against paper.cols.txt; all 40 located (24 exact after whitespace collapse, 14 via an 8-word run, 2 via a 5-word run).
  • Data age is approximate. It is the venue year minus a 4-digit year parsed from spanEnd, so it is granular to a year and inherits the venue-year convention. 28 existing-dataset papers had a spanEnd with no parsable year and 1 had a year after its venue year; both are excluded and reported rather than absorbed.
  • Every query, the report script and its unedited output are on existing_datasets; corpus-level caveats are on Corpus.

Open Questions

  • Nobody has measured how often a reused dataset is still obtainable. The residue audit suggests roughly a third of unplaced source strings are not something a reader can download, but that is a 60-string sample classified by one person, and the residue is not the whole name universe. A systematic link-and-access check over the named sources in this corpus is an afternoon of scripting and a real result.
  • No study has measured how much a snapshot choice moves a published result. The archive literature has this for snapshot selection [9Hantke, Florian; Calzavara, Stefano; Wilhelm, Moritz; Rabitti, Alvise; Stock, Ben (2023): "You Call This Archaeology? Evaluating Web Archives for Reproducible Web Security Measurements", in: Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, pp. 3168–3182. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]; the scan-corpus literature does not. Re-running one published Censys-based analysis over twelve consecutive monthly snapshots would answer it.
  • The coverage comparisons are ad hoc and ageing. Sethuraman et al. [sethuraman2022_ipv4] is a two-page poster on one month of 2021; VanderSloot et al. [6VanderSloot, Benjamin; Amann, Johanna; Bernhard, Matthew; Durumeric, Zakir; Bailey, Michael D.; Halderman, J. Alex (2016): "Towards a Complete View of the Certificate Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] is from 2016. There is no current, systematic agreement measurement across the scan search engines a 2026 paper would actually query.
  • Reuse does not pass data on. 57.7% of reuse papers release an artifact of their own, below the crawled population's 61.1%. Whether the derived datasets that are released are joinable to each other — same keys, same normalisation — is unmeasured and is the precondition for this literature compounding rather than repeating.
[2]
Wu, Mengying; Hong, Geng; Chen, Jinsong; Liu, Qi; Tang, Shujun; Li, Youhao; Liu, Baojun; Duan, Haixin; Yang, Min (2025): "Revealing the Black Box of Device Search Engine: Scanning Assets, Strategies, and Ethical Consideration", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[3]
Durumeric, Zakir; Adrian, David; Mirian, Ariana; Bailey, Michael D.; Halderman, J. Alex (2015): "A Search Engine Backed by Internet-Wide Scanning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[6]
VanderSloot, Benjamin; Amann, Johanna; Bernhard, Matthew; Durumeric, Zakir; Bailey, Michael D.; Halderman, J. Alex (2016): "Towards a Complete View of the Certificate Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[7]
Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)
[8]
Durumeric, Zakir; Adrian, David; Stephens, Phillip; Wustrow, Eric; Halderman, J. Alex (2024): "Ten Years of ZMap", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[9]
Hantke, Florian; Calzavara, Stefano; Wilhelm, Moritz; Rabitti, Alvise; Stock, Ben (2023): "You Call This Archaeology? Evaluating Web Archives for Reproducible Web Security Measurements", in: Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, pp. 3168–3182. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[10]
Singh, Sachin Kumar; Mahmud, Faisal; Ricci, Robert; Siby, Sandra (2026): "The Empire Strikes Back (at Your Privacy): An Archaeology of Tracking on Government Websites", Proceedings on Privacy Enhancing Technologies 2026(2):108-126. (DOI)
1)
IMC 2018, Is the Web Ready for OCSP Must-Staple?, temporal and population evidence quotes; located verbatim in data/fulltext/2018/IMC/is-the-web-ready-for-ocsp-must-staple/paper.cols.txt.
2)
population.dataset and temporal.evaluation evidence quotes. Both sentences are interleaved with a table in paper.cols.txt, so they are paraphrased here rather than quoted; the extraction's verbatim quotes are reproduced on existing_datasets.
3)
USENIX Security 2017, Measuring HTTPS Adoption on the Web, population.methodology evidence quote.
4)
IMC 2017, Large-Scale Scanning of TCP's Initial Window, temporal.methodology evidence quote.
5)
https://docs.censys.com/docs/platform-api-transition-guide, fetched 2026-08-28, HTTP 200. The guide maps Search API v2 endpoints to v3 and publishes no shutdown date; it carries no v1 mapping at all, v1 having been retired earlier.
6)
https://docs.censys.com/docs/research-access-to-censys-data, fetched 2026-08-28, HTTP 200.
7)
https://help.shodan.io/the-basics/academic-upgrade, fetched 2026-08-28, HTTP 200.
8)
https://scans.io, fetched 2026-08-28, HTTP 200.
9)
https://sonardata.rapid7.com/about/, fetched 2026-08-28, HTTP 200.
10)
https://www.phishtank.com/register.php, fetched 2026-08-28, HTTP 200 with a browser User-Agent. The site rate-limits automated clients: a second fetch from a different address the same day returned HTTP 429 with retry-after: 86372. The string is also present in a Wayback capture of 2026-05-12.
11)
Fetched 2026-08-28. The bare domain answers 405/307; the 403 is on the data endpoints. Reddit's own Data API Terms are the governing document — https://redditinc.com/policies/data-api-terms, fetched 2026-08-28, HTTP 200, most recent revision dated 2026-07-01 — and require “a separate agreement with Reddit” for “research in excess of rate limits”. The researcher-programme help pages under support.reddithelp.com refuse automated clients (403), so the programme's own eligibility text could not be checked here.
12)
https://androzoo.uni.lu/, fetched 2026-08-28, HTTP 200.
13)
https://androzoo.uni.lu/access, fetched 2026-08-28, HTTP 200.
14)
https://www.openintel.nl/, fetched 2026-08-28, HTTP 200.
15)
https://letsencrypt.org/2025/08/14/rfc-6962-logs-eol, fetched 2026-08-28, HTTP 200. Read-only from 2025-11-30.
16)
https://www.cos.io/osf-changes, fetched 2026-08-28, HTTP 200.
You could leave a comment if you were logged in.
design/existing_datasets.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki