User Tools

Site Tools


security

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
security [2026/08/27 12:54] – Create security namespace outline: 5 proposed children with corpus support counts, 3 rejected, boundaries. Authored by Claude. karel.kubicek.claudesecurity [2026/09/11 07:14] (current) – Add Security:Authentication as the seventh child (login deployment: SSO, MFA/RBA, passkeys, login and password policy); note the programming:registration boundary. Authored by Claude karel.kubicek.claude
Line 4: Line 4:
  
 <WRAP important> <WRAP important>
-**A namespace page outlines the pages inside it rather than carrying its own content.** (([[contributing]], "Namespace and page structure".)) The five children below are proposed from this corpus: four have a schema population, and [[Security:Headers]] has only a full-text upper bound until a later sitting hand-maps itThey are red links until written; this page exists so a reader landing from [[start]] is not sent into an empty namespace. Three candidates were not given a page; Google Safe Browsing was folded into phishing rather than made a sixth — see [[#Rejected, and why]].+**A namespace page outlines the pages inside it rather than carrying its own content.** (([[contributing]], "Namespace and page structure".)) The **seven** children below are proposed from this corpus, and all seven are written: [[Security:Web vulnerabilities]], [[Security:TLS certificates]], [[Security:Phishing]], [[Security:Headers]], [[Security:VirusTotal]], [[Security:Email authentication]] and [[Security:Authentication]]This page exists so a reader landing from [[start]] is not sent into an empty namespace. Three candidates were not given a page; Google Safe Browsing was folded into phishing rather than made a page of its own — see [[#Rejected, and why]].
 </WRAP> </WRAP>
  
-===== The five pages =====+===== The seven pages =====
  
 ^ Page ^ What a fresh student needs it for ^ What the corpus can carry (2026-08-27) ^ ^ Page ^ What a fresh student needs it for ^ What the corpus can carry (2026-08-27) ^
 | [[Security:TLS certificates]] | You are about to measure HTTPS, certificates, or CT logs, and need to know which instrument (active scan, CT log, Censys) answers which question. | **50** papers declare their population unit as certificates (22 of them on the web platform). **133** used a TLS-specific instrument (Censys, ZGrab, crt.sh, CT, sslyze, Qualys SSL, … — **not** OpenSSL-the-library). A looser full-text probe for TLS and certificates hits **525** papers, **213** web — an upper bound the child page has to narrow. Start with Holz et al. {[holz2011_landscape]}, Durumeric et al. {[durumeric2013_https]}, Kotzias et al. {[kotzias2018_coming]}, the Censys paper {[durumeric2015_search]}, Let's Encrypt {[aas2019_encrypt]}, and ten years of ZMap {[durumeric2024_years]}. | | [[Security:TLS certificates]] | You are about to measure HTTPS, certificates, or CT logs, and need to know which instrument (active scan, CT log, Censys) answers which question. | **50** papers declare their population unit as certificates (22 of them on the web platform). **133** used a TLS-specific instrument (Censys, ZGrab, crt.sh, CT, sslyze, Qualys SSL, … — **not** OpenSSL-the-library). A looser full-text probe for TLS and certificates hits **525** papers, **213** web — an upper bound the child page has to narrow. Start with Holz et al. {[holz2011_landscape]}, Durumeric et al. {[durumeric2013_https]}, Kotzias et al. {[kotzias2018_coming]}, the Censys paper {[durumeric2015_search]}, Let's Encrypt {[aas2019_encrypt]}, and ten years of ZMap {[durumeric2024_years]}. |
-| [[Security:Headers]] | You are about to crawl for CSP, HSTS, X-Frame-Options, SRI or Trusted Types, and need to know which of those are still worth measuring. | Full-text probe (Content-Security-Policy, CSP plus "header" or "directive", HSTS, X-Frame-Options, SRI): **317** papers, **176** web. **That 176 is an upper bound, not a population** — it still contains bibliography hits and homographs. A schema match on header-ish tokens is not usable here — see the provenance page. The child page has to hand-map before any prevalence figure. Start with Weichselbaum et al. {[weichselbaum2016_dead]}, Roth et al. {[roth2020complex]}, Steffens et al. {[steffens2021_blockparty]}. |+| [[Security:Headers]] | You are about to crawl for CSP, HSTS, X-Frame-Options, SRI or Trusted Types, and need to know which of those are still worth measuring. | Full-text probe (Content-Security-Policy, CSP plus "header" or "directive", HSTS, X-Frame-Options, SRI): **317** papers, **176** web. Hand-mapped: **44** measured, 26 homographs, 13 citation-only. The 44 is 2.7% of 1,622 web-platform papers. Start with Weichselbaum et al. {[weichselbaum2016_dead]}, Roth et al. {[roth2020complex]}, Steffens et al. {[steffens2021_blockparty]}. |
 | [[Security:VirusTotal]] | You are about to label files, URLs or domains with VirusTotal, and need to know what a "detected" bit actually is. | **277** papers the extractor marked as used, produced, or drawing a population from VirusTotal (107 of them on the web platform). That is an upper bound on true use: the first sample of eight includes a references-only hit. **262** named it as a tool they used or produced; **254** of those as a classification-service. **64** distinct raw strings (products, feeds, thresholds and combinations, not merely spellings). Of the 209 that used it as a classifier, **86** targeted malware, 42 domains, 37 mobile apps, 16 website-category. The website-category use is already on [[design:website_classification]]; this page is the maliciousness-oracle use. Start with Peng et al. {[peng2019_opening]}. | | [[Security:VirusTotal]] | You are about to label files, URLs or domains with VirusTotal, and need to know what a "detected" bit actually is. | **277** papers the extractor marked as used, produced, or drawing a population from VirusTotal (107 of them on the web platform). That is an upper bound on true use: the first sample of eight includes a references-only hit. **262** named it as a tool they used or produced; **254** of those as a classification-service. **64** distinct raw strings (products, feeds, thresholds and combinations, not merely spellings). Of the 209 that used it as a classifier, **86** targeted malware, 42 domains, 37 mobile apps, 16 website-category. The website-category use is already on [[design:website_classification]]; this page is the maliciousness-oracle use. Start with Peng et al. {[peng2019_opening]}. |
 | [[Security:Phishing]] | You are about to crawl phishing sites or evaluate a feed, and need to know what the feed does not contain. | Schema union (detection, classification, population, or slug matching "phish"): **139** papers, **92** web. PhishTank is the named source in **21**. A full-text /phish/ sweep hits **1,097** papers — that is a fact about these being security venues, not a population. Start with Zhang et al. {[zhang2021_crawlphish]} (cloaking against anti-phishing crawlers) and Peng et al. {[peng2019_opening]} (VirusTotal's phishing engines). | | [[Security:Phishing]] | You are about to crawl phishing sites or evaluate a feed, and need to know what the feed does not contain. | Schema union (detection, classification, population, or slug matching "phish"): **139** papers, **92** web. PhishTank is the named source in **21**. A full-text /phish/ sweep hits **1,097** papers — that is a fact about these being security venues, not a population. Start with Zhang et al. {[zhang2021_crawlphish]} (cloaking against anti-phishing crawlers) and Peng et al. {[peng2019_opening]} (VirusTotal's phishing engines). |
 | [[Security:Web vulnerabilities]] | You are about to look for XSS, CSRF, clickjacking or SOP bypass on live sites, and need methods and denominators — not the ethics checklist. | **880** papers classify a vulnerability; **209** of them on the web platform; **99** both crawled and web. The 99 is a paper-level conjunction — it does not by itself mean the crawl is how the vulnerability was found. **49.2%** of the 880 are ''offline'' (program analysis of software). The child page's first job is to hand-map the web/crawled slice, not to treat 880 or 99 as a method count. Ethics of scanning live sites is [[practices:ethics]]; telling the operator is [[practices:notifying_websites]]. Start with Steffens et al. {[steffens2019_dont]}. | | [[Security:Web vulnerabilities]] | You are about to look for XSS, CSRF, clickjacking or SOP bypass on live sites, and need methods and denominators — not the ethics checklist. | **880** papers classify a vulnerability; **209** of them on the web platform; **99** both crawled and web. The 99 is a paper-level conjunction — it does not by itself mean the crawl is how the vulnerability was found. **49.2%** of the 880 are ''offline'' (program analysis of software). The child page's first job is to hand-map the web/crawled slice, not to treat 880 or 99 as a method count. Ethics of scanning live sites is [[practices:ethics]]; telling the operator is [[practices:notifying_websites]]. Start with Steffens et al. {[steffens2019_dont]}. |
 +| [[Security:Email authentication]] | You are about to measure whether domains and mail servers deploy SPF, DKIM, DMARC, DANE, MTA-STS or STARTTLS, and need to know which of the four different questions your instrument answers. | **31** papers, hand-mapped, whose object is email transport, authentication and encryption **deployment** — the ''INFRA'' verdict of the published population map behind [[privacy:email_tracking]], which scopes them out because they are not tracking. Heavily recent: 21 of the 31 are 2022 or later, and the venue split is USENIX Security 16 / IMC 7 / NDSS 4. The page's own subject is denominators: the same protocol is 56.5% and 60.9% deployed in the same year because the populations differ. **BIMI is measured by zero papers** in these seven venues. Start with Durumeric et al. {[durumeric2015_neither]}, Hu and Wang {[hu2018_measurements]}, Ashiq et al. {[ashiq2024_beyond]} and BreakSPF {[wang2024_breakspf]}. | 
 +| [[Security:Authentication]] | You are about to measure how websites deploy **login**: SSO and OAuth buttons, MFA and risk-based authentication, passkeys and WebAuthn, login and password policies, and the session a login leaves behind. | **45** papers, hand-audited from **184** candidates produced by five probes plus a full-text recall pass — 24.5% precision, which is why this page was gated on an audit before it was written. 42 of the 45 are web-platform papers (2.6% of 1,622) and 25 ran a crawl (2.2% of 1,120); 33 of the 45 are 2022 or later. The page's subject is the denominator: "how many sites support SSO" is answered as 9.3%, 6.30%, 7.23%, 27% and 57.8% by five papers that do not disagree, they divide by different things. **No paper in this population classifies with an LLM** (0 of 45, against 177 corpus-wide). Start with Ardi et al. {[ardi2023_prevalence]}, Al Roomi and Li {[alroomi2023_login]}, Jannett et al. {[jannett2026_passkeys]} and Gavazzi et al. {[gavazzi2023_multi]}. |
 ===== Where this namespace stops ===== ===== Where this namespace stops =====
  
Line 25: Line 26:
   * **[[design:mobile_and_app_measurement]]** — certificate pinning and TLS interception inside apps.   * **[[design:mobile_and_app_measurement]]** — certificate pinning and TLS interception inside apps.
   * **[[design:ip_classification]]** — reputation and geolocation of addresses, including some of the same feeds.   * **[[design:ip_classification]]** — reputation and geolocation of addresses, including some of the same feeds.
 +  * **[[programming:registration]]** — logging in or creating accounts **as an instrument**, so you can crawl what is behind the login. [[Security:Authentication]] is the other direction: the login itself as the measurement.
   * **[[privacy:javascript]]** — script-level analysis; XSS-as-a-JavaScript-phenomenon belongs with [[Security:Web vulnerabilities]] when the question is "is this live site exploitable", and there when the question is "what did this script do".   * **[[privacy:javascript]]** — script-level analysis; XSS-as-a-JavaScript-phenomenon belongs with [[Security:Web vulnerabilities]] when the question is "is this live site exploitable", and there when the question is "what did this script do".
  
Line 40: Line 42:
 ===== Methodology and limitations of these figures ===== ===== Methodology and limitations of these figures =====
  
-Every number in the table above is a count of **papers**, from the 5,859-paper extraction, with the denominator named in the same cell. Full-text probes read ''paper.cols.txt'' (4 of 5,859 have none and are counted as negatives). 2025–2026 venue-years are provisional — see [[literature:corpus]]. The queries, the folds, the residue, and the unedited report output are on [[provenance:security]].+Every number in the table above is a count of **papers**, from the 5,859-paper extraction, with the denominator named in the same cell. Two rows are exceptions in kind rather than in rigour: the email-authentication 31 is a hand-mapped topical slice of a published 354-paper candidate pool, and the authentication 45 is a hand audit of 184 probe candidates. Neither is a query result, and both pages say so. Full-text probes read ''paper.cols.txt'' (4 of 5,859 have none and are counted as negatives). 2025–2026 venue-years are provisional — see [[literature:corpus]]. The queries, the folds, the residue, and the unedited report output are on [[provenance:security]].
  
 <bibtex bibliography></bibtex> <bibtex bibliography></bibtex>
security.1787835259.txt.gz · Last modified: by karel.kubicek.claude