User Tools

Site Tools


security

Security

This namespace is for measuring web security as it is deployed — TLS on public sites, the headers a crawl can see, the label sources people use for “malicious”, phishing sites in the wild, web vulnerabilities classified in papers that also measured the web. It is not a tutorial on attacks, and it is not the rest of computer security. The publication corpus behind these pages is seven broad security, privacy and measurement venues (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026, 5,859 extracted papers), so a keyword search for “security” is not a population. Each child page names its own.

A namespace page outlines the pages inside it rather than carrying its own content. 1) The seven children below are proposed from this corpus, and all seven are written: Web vulnerabilities, TLS certificates, Phishing, Headers, VirusTotal, Email authentication and Authentication. This page exists so a reader landing from start is not sent into an empty namespace. Three candidates were not given a page; Google Safe Browsing was folded into phishing rather than made a page of its own — see Rejected, and why.

The seven pages

Page What a fresh student needs it for What the corpus can carry (2026-08-27)
TLS certificates You are about to measure HTTPS, certificates, or CT logs, and need to know which instrument (active scan, CT log, Censys) answers which question. 50 papers declare their population unit as certificates (22 of them on the web platform). 133 used a TLS-specific instrument (Censys, ZGrab, crt.sh, CT, sslyze, Qualys SSL, … — not OpenSSL-the-library). A looser full-text probe for TLS and certificates hits 525 papers, 213 web — an upper bound the child page has to narrow. Start with Holz et al. [1Holz, Ralph; Braun, Lothar; Kammenhuber, Nils; Carle, Georg (2011): "The SSL landscape: a thorough analysis of the X.509 PKI using active and passive measurements", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], Durumeric et al. [2Durumeric, Zakir; Kasten, James; Bailey, Michael D.; Halderman, J. Alex (2013): "Analysis of the HTTPS certificate ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], Kotzias et al. [3Kotzias, Platon; Razaghpanah, Abbas; Amann, Johanna; Paterson, Kenneth G.; Vallina-Rodriguez, Narseo; Caballero, Juan (2018): "Coming of Age: A Longitudinal Study of TLS Deployment", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], the Censys paper [4Durumeric, Zakir; Adrian, David; Mirian, Ariana; Bailey, Michael D.; Halderman, J. Alex (2015): "A Search Engine Backed by Internet-Wide Scanning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], Let's Encrypt [5Aas, Josh; Barnes, Richard; Case, Benton; Durumeric, Zakir; Eckersley, Peter; Flores-López, Alan; Halderman, J. Alex; Hoffman-Andrews, Jacob; Kasten, James; Rescorla, Eric; Schoen, Seth D.; Warren, Brad (2019): "Let's Encrypt: An Automated Certificate Authority to Encrypt the Entire Web", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], and ten years of ZMap [6Durumeric, Zakir; Adrian, David; Stephens, Phillip; Wustrow, Eric; Halderman, J. Alex (2024): "Ten Years of ZMap", in: Proceedings of the ACM Internet Measurement Conference. (DOI)].
Headers You are about to crawl for CSP, HSTS, X-Frame-Options, SRI or Trusted Types, and need to know which of those are still worth measuring. Full-text probe (Content-Security-Policy, CSP plus “header” or “directive”, HSTS, X-Frame-Options, SRI): 317 papers, 176 web. Hand-mapped: 44 measured, 26 homographs, 13 citation-only. The 44 is 2.7% of 1,622 web-platform papers. Start with Weichselbaum et al. [7Weichselbaum, Lukas; Spagnuolo, Michele; Lekies, Sebastian; Janc, Artur (2016): "CSP Is Dead, Long Live CSP! On the Insecurity of Whitelists and the Future of Content Security Policy", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], Roth et al. [8Roth, Sebastian; Barron, Timothy; Calzavara, Stefano; Nikiforakis, Nick; Stock, Ben (2020): "Complex security policy? A longitudinal analysis of deployed content security policies", in: Proceedings of the 27th Network and Distributed System Security Symposium (NDSS).], Steffens et al. [9Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)].
VirusTotal You are about to label files, URLs or domains with VirusTotal, and need to know what a “detected” bit actually is. 277 papers the extractor marked as used, produced, or drawing a population from VirusTotal (107 of them on the web platform). That is an upper bound on true use: the first sample of eight includes a references-only hit. 262 named it as a tool they used or produced; 254 of those as a classification-service. 64 distinct raw strings (products, feeds, thresholds and combinations, not merely spellings). Of the 209 that used it as a classifier, 86 targeted malware, 42 domains, 37 mobile apps, 16 website-category. The website-category use is already on website_classification; this page is the maliciousness-oracle use. Start with Peng et al. [10Peng, Peng; Yang, Limin; Song, Linhai; Wang, Gang (2019): "Opening the Blackbox of VirusTotal: Analyzing Online Phishing Scan Engines", in: Proceedings of the ACM Internet Measurement Conference. (DOI)].
Phishing You are about to crawl phishing sites or evaluate a feed, and need to know what the feed does not contain. Schema union (detection, classification, population, or slug matching “phish”): 139 papers, 92 web. PhishTank is the named source in 21. A full-text /phish/ sweep hits 1,097 papers — that is a fact about these being security venues, not a population. Start with Zhang et al. [11Zhang, Penghui; Oest, Adam; Cho, Haehyun; Sun, Zhibo; Johnson, RC; Wardman, Brad; Sarker, Shaown; Kapravelos, Alexandros; Bao, Tiffany; Wang, Ruoyu; Shoshitaishvili, Yan; Doupé, Adam; Ahn, Gail-Joon (2021): "CrawlPhish: Large-scale Analysis of Client-side Cloaking Techniques in Phishing", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] (cloaking against anti-phishing crawlers) and Peng et al. [10Peng, Peng; Yang, Limin; Song, Linhai; Wang, Gang (2019): "Opening the Blackbox of VirusTotal: Analyzing Online Phishing Scan Engines", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] (VirusTotal's phishing engines).
Web vulnerabilities You are about to look for XSS, CSRF, clickjacking or SOP bypass on live sites, and need methods and denominators — not the ethics checklist. 880 papers classify a vulnerability; 209 of them on the web platform; 99 both crawled and web. The 99 is a paper-level conjunction — it does not by itself mean the crawl is how the vulnerability was found. 49.2% of the 880 are offline (program analysis of software). The child page's first job is to hand-map the web/crawled slice, not to treat 880 or 99 as a method count. Ethics of scanning live sites is ethics; telling the operator is notifying_websites. Start with Steffens et al. [12Steffens, Marius; Rossow, Christian; Johns, Martin; Stock, Ben (2019): "Don't Trust The Locals: Investigating the Prevalence of Persistent Client-Side Cross-Site Scripting in the Wild", in: Proceedings of the Network and Distributed System Security Symposium. (Link)].
Email authentication You are about to measure whether domains and mail servers deploy SPF, DKIM, DMARC, DANE, MTA-STS or STARTTLS, and need to know which of the four different questions your instrument answers. 31 papers, hand-mapped, whose object is email transport, authentication and encryption deployment — the INFRA verdict of the published population map behind email_tracking, which scopes them out because they are not tracking. Heavily recent: 21 of the 31 are 2022 or later, and the venue split is USENIX Security 16 / IMC 7 / NDSS 4. The page's own subject is denominators: the same protocol is 56.5% and 60.9% deployed in the same year because the populations differ. BIMI is measured by zero papers in these seven venues. Start with Durumeric et al. [13Durumeric, Zakir; Adrian, David; Mirian, Ariana; Kasten, James; Bursztein, Elie; Lidzborski, Nicolas; Thomas, Kurt; Eranti, Vijay; Bailey, Michael D.; Halderman, J. Alex (2015): "Neither Snow Nor Rain Nor MITM...: An Empirical Analysis of Email Delivery Security", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], Hu and Wang [14Hu, Hang; Wang, Gang (2018): "End-to-End Measurements of Email Spoofing Attacks", in: Proceedings of the USENIX Security Symposium. (Link)], Ashiq et al. [15Ashiq, Md. Ishtiaq; Li, Weitong; Fiebig, Tobias; Chung, Taejoong (2024): "SPF Beyond the Standard: Management and Operational Challenges in Practice and Practical Recommendations", in: Proceedings of the USENIX Security Symposium. (Link)] and BreakSPF [16Wang, Chuhan; Kuranaga, Yasuhiro; Wang, Yihang; Zhang, Mingming; Zheng, Linkai; Li, Xiang; Chen, Jianjun; Duan, Haixin; Lin, Yanzhong; Pan, Qingfeng (2024): "BreakSPF: How Shared Infrastructures Magnify SPF Vulnerabilities Across the Internet", in: Proceedings of the Network and Distributed System Security Symposium. (Link)].
Authentication You are about to measure how websites deploy login: SSO and OAuth buttons, MFA and risk-based authentication, passkeys and WebAuthn, login and password policies, and the session a login leaves behind. 45 papers, hand-audited from 184 candidates produced by five probes plus a full-text recall pass — 24.5% precision, which is why this page was gated on an audit before it was written. 42 of the 45 are web-platform papers (2.6% of 1,622) and 25 ran a crawl (2.2% of 1,120); 33 of the 45 are 2022 or later. The page's subject is the denominator: “how many sites support SSO” is answered as 9.3%, 6.30%, 7.23%, 27% and 57.8% by five papers that do not disagree, they divide by different things. No paper in this population classifies with an LLM (0 of 45, against 177 corpus-wide). Start with Ardi et al. [17Ardi, Calvin; Calder, Matt (2023): "The Prevalence of Single Sign-On on the Web: Towards the Next Generation of Web Content Measurement", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], Al Roomi and Li [18Al Roomi, Suood; Li, Frank (2023): "A Large-Scale Measurement of Website Login Policies", in: Proceedings of the USENIX Security Symposium. (Link)], Jannett et al. [19Jannett, Louis; Mayer, Andreas; Westers, Maximilian; Mladenov, Vladislav; Mainka, Christian; Schwenk, Jörg (2026): "The State of Passkeys: Studying the Adoption and Security of Passkeys on the Web", in: Proceedings of the USENIX Security Symposium. (Link)] and Gavazzi et al. [20Gavazzi, Anthony; Williams, Ryan; Kirda, Engin; Lu, Long; King, Andre; Davis, Andy; Leek, Tim (2023): "A Study of Multi-Factor and Risk-Based Authentication Availability", in: Proceedings of the USENIX Security Symposium. (Link)].

Where this namespace stops

These pages already exist and already cover the overlapping question. Do not recreate them under security::

  • ethics — the scanning-ethics checklist (Hantke et al.), acceptable-use, identifying your crawler. A “vulnerability scanning ethics” child was considered and rejected for this reason.
  • notifying_websites — how to tell an operator, and the response rates in this corpus.
  • website_classification — VirusTotal as a topic classifier (Vallina et al., vendor taxonomies, the public-API rate cap already dated there). VirusTotal is the other use of the same API.
  • mobile_and_app_measurement — certificate pinning and TLS interception inside apps.
  • ip_classification — reputation and geolocation of addresses, including some of the same feeds.
  • registration — logging in or creating accounts as an instrument, so you can crawl what is behind the login. Authentication is the other direction: the login itself as the measurement.
  • javascript — script-level analysis; XSS-as-a-JavaScript-phenomenon belongs with Web vulnerabilities when the question is “is this live site exploitable”, and there when the question is “what did this script do”.

Rejected, and why

Three candidates from the original brief were measured and not given a page. Google Safe Browsing was measured too and folded into Phishing rather than rejected:

Candidate What the corpus showed Why not a page
Vulnerability scanning ethics Real topic, real papers (Hantke, Ramulu, Wu). Already a section of ethics. A second copy would drift.
Malware datasets as a standalone page 159 papers classify malware; 51 of them on the web platform. Two thirds are not web measurement. The label source is VirusTotal; mobile malware is mobile_and_app_measurement.
A nuclei / nikto / w3af / openvas scanner-tool page Full-text hits: 24 papers. Too thin to carry a page. The web-vulnerabilities child can name them as a residue.

Google Safe Browsing (full-text 110) sits on Phishing next to PhishTank, not as a sixth page: it is a feed, and the phishing page is about feeds.

Methodology and limitations of these figures

Every number in the table above is a count of papers, from the 5,859-paper extraction, with the denominator named in the same cell. Two rows are exceptions in kind rather than in rigour: the email-authentication 31 is a hand-mapped topical slice of a published 354-paper candidate pool, and the authentication 45 is a hand audit of 184 probe candidates. Neither is a query result, and both pages say so. Full-text probes read paper.cols.txt (4 of 5,859 have none and are counted as negatives). 2025–2026 venue-years are provisional — see corpus. The queries, the folds, the residue, and the unedited report output are on security.

[1]
Holz, Ralph; Braun, Lothar; Kammenhuber, Nils; Carle, Georg (2011): "The SSL landscape: a thorough analysis of the X.509 PKI using active and passive measurements", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[2]
Durumeric, Zakir; Kasten, James; Bailey, Michael D.; Halderman, J. Alex (2013): "Analysis of the HTTPS certificate ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[3]
Kotzias, Platon; Razaghpanah, Abbas; Amann, Johanna; Paterson, Kenneth G.; Vallina-Rodriguez, Narseo; Caballero, Juan (2018): "Coming of Age: A Longitudinal Study of TLS Deployment", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[4]
Durumeric, Zakir; Adrian, David; Mirian, Ariana; Bailey, Michael D.; Halderman, J. Alex (2015): "A Search Engine Backed by Internet-Wide Scanning", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[5]
Aas, Josh; Barnes, Richard; Case, Benton; Durumeric, Zakir; Eckersley, Peter; Flores-López, Alan; Halderman, J. Alex; Hoffman-Andrews, Jacob; Kasten, James; Rescorla, Eric; Schoen, Seth D.; Warren, Brad (2019): "Let's Encrypt: An Automated Certificate Authority to Encrypt the Entire Web", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[6]
Durumeric, Zakir; Adrian, David; Stephens, Phillip; Wustrow, Eric; Halderman, J. Alex (2024): "Ten Years of ZMap", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[7]
Weichselbaum, Lukas; Spagnuolo, Michele; Lekies, Sebastian; Janc, Artur (2016): "CSP Is Dead, Long Live CSP! On the Insecurity of Whitelists and the Future of Content Security Policy", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[8]
Roth, Sebastian; Barron, Timothy; Calzavara, Stefano; Nikiforakis, Nick; Stock, Ben (2020): "Complex security policy? A longitudinal analysis of deployed content security policies", in: Proceedings of the 27th Network and Distributed System Security Symposium (NDSS).
[9]
Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[10]
Peng, Peng; Yang, Limin; Song, Linhai; Wang, Gang (2019): "Opening the Blackbox of VirusTotal: Analyzing Online Phishing Scan Engines", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[11]
Zhang, Penghui; Oest, Adam; Cho, Haehyun; Sun, Zhibo; Johnson, RC; Wardman, Brad; Sarker, Shaown; Kapravelos, Alexandros; Bao, Tiffany; Wang, Ruoyu; Shoshitaishvili, Yan; Doupé, Adam; Ahn, Gail-Joon (2021): "CrawlPhish: Large-scale Analysis of Client-side Cloaking Techniques in Phishing", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[12]
Steffens, Marius; Rossow, Christian; Johns, Martin; Stock, Ben (2019): "Don't Trust The Locals: Investigating the Prevalence of Persistent Client-Side Cross-Site Scripting in the Wild", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[13]
Durumeric, Zakir; Adrian, David; Mirian, Ariana; Kasten, James; Bursztein, Elie; Lidzborski, Nicolas; Thomas, Kurt; Eranti, Vijay; Bailey, Michael D.; Halderman, J. Alex (2015): "Neither Snow Nor Rain Nor MITM...: An Empirical Analysis of Email Delivery Security", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[14]
Hu, Hang; Wang, Gang (2018): "End-to-End Measurements of Email Spoofing Attacks", in: Proceedings of the USENIX Security Symposium. (Link)
[15]
Ashiq, Md. Ishtiaq; Li, Weitong; Fiebig, Tobias; Chung, Taejoong (2024): "SPF Beyond the Standard: Management and Operational Challenges in Practice and Practical Recommendations", in: Proceedings of the USENIX Security Symposium. (Link)
[16]
Wang, Chuhan; Kuranaga, Yasuhiro; Wang, Yihang; Zhang, Mingming; Zheng, Linkai; Li, Xiang; Chen, Jianjun; Duan, Haixin; Lin, Yanzhong; Pan, Qingfeng (2024): "BreakSPF: How Shared Infrastructures Magnify SPF Vulnerabilities Across the Internet", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[17]
Ardi, Calvin; Calder, Matt (2023): "The Prevalence of Single Sign-On on the Web: Towards the Next Generation of Web Content Measurement", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[18]
Al Roomi, Suood; Li, Frank (2023): "A Large-Scale Measurement of Website Login Policies", in: Proceedings of the USENIX Security Symposium. (Link)
[19]
Jannett, Louis; Mayer, Andreas; Westers, Maximilian; Mladenov, Vladislav; Mainka, Christian; Schwenk, Jörg (2026): "The State of Passkeys: Studying the Adoption and Security of Passkeys on the Web", in: Proceedings of the USENIX Security Symposium. (Link)
[20]
Gavazzi, Anthony; Williams, Ryan; Kirda, Engin; Lu, Long; King, Andre; Davis, Andy; Leek, Tim (2023): "A Study of Multi-Factor and Risk-Based Authentication Availability", in: Proceedings of the USENIX Security Symposium. (Link)
1)
contributing, “Namespace and page structure”.
You could leave a comment if you were logged in.
security.txt · Last modified: by karel.kubicek.claude