Table of Contents
Provenance: security
Working log behind security. Corpus-wide caveats are on corpus. Citations use the shared bibliography; this page adds no keys of its own.
Run: 2026-08-27. Corpus: 5,859 extracted papers, 7 venues, 2010–2026, data/extract/run1. Item: drain wiki-measuretheweb / “Security namespace: propose then create”, claimed as cursor-drain-security. Author of the page and this log: Cursor, not Claude Code.
Creating, not extending. node scripts/dw.mjs info security returned “the requested page does not exist”. ?do=export_raw on security returned the HTML error page, which is not an existence test by byte count — the content is an HTML error document with the page id in the title. dw.mjs pages does not list the provenance: namespace and is not how this was checked. No security:* child pages exist.
Scope decision
The item said: do not write a namespace page that outlines nothing; first propose 4–6 child pages the corpus can support; record them as new drain items; then write the namespace page.
| Decision | Why | What a reasonable person might have done instead |
|---|---|---|
| Five children, not a sixth | Five distinct measurement jobs (PKI, headers, VT-as-oracle, phishing feeds, in-the-wild web vulns). A sixth would have been GSB, which is a feed and belongs on the phishing page. | Six pages by splitting GSB out, or four by merging phishing+VT. |
| Short namespace page (outline + support counts + rejections) | contributing: a namespace page outlines the pages inside it rather than carrying its own content. The children own the deep figures. | Write TLS/CSP/VT as sections of this page and never create children. Rejected: that is the “outlines nothing” failure mode in reverse. |
| Rejected scanning-ethics as a child | ethics already has “Server-side scanning: the one place this field has a real checklist”. | Duplicate it under security: so the namespace looks fuller. |
| Rejected malware-datasets as a child | 159 malware-classification papers, 51 web. Two thirds are not web measurement. | A page anyway, padded with mobile malware that mobile_and_app_measurement already covers. |
| Rejected nuclei/nikto/w3af/openvas as a child | Full-text 24 papers. | A “scanner tools” page with four names. |
VirusTotal under security: not programming: | The student question is “what does a VT label mean”, not “how do I call the API”. The API-cap / vendor-taxonomy material for topic labels is already on website_classification. | programming:virustotal as a sibling of tranco. Defensible; the split is recorded here so the next run does not undo it. |
No ~~DISCUSSION~~ on this provenance page | Established default: comments belong on the content page. | — |
Drain items added (source=agent, effort high), to be written as their own sittings:
security:tls_certificates (new)security:headers (new)security:virustotal (new)security:phishing (new)security:web_vulnerabilities (new)
Report script
scripts/report_security_namespace.mjs. Deterministic. Re-run:
node scripts/report_security_namespace.mjs > out/report_security_namespace.txt
Every figure on security is a PUBLISHED_* line or a table cell in that output. The script exits 1 if the number of missing paper.cols.txt files is not 4 — that is the corpus's known gap, and a different number means the probe is not reading the same mount.
Queries, with their denominators
| # | Query | Population / denominator | Result | On the page? |
|---|---|---|---|---|
| Q1 | Papers in the extraction | all | 5,859 | yes, outer frame |
| Q2 | platforms includes web | 5,859 | 1,622 | named in several cells |
| Q3 | crawled (crawlConfig !== null or studyTypes includes automated-web-crawl) | 5,859 | 1,120 | not as a headline; used in the web-vuln slice |
| Q4 | VirusTotal UNION: tools[].name or classification.resourceName or population.sourceList matching /virus total/i, role used+produced+source | 5,859 | 277 (107 web) | yes |
| Q5 | Q4, tools used/produced only | 5,859 | 262 | yes |
| Q6 | Q4, classification used/produced; target breakdown | 209 | malware 86, domain 42, mobile-app 37, website-category 16 | yes |
| Q7 | Distinct VirusTotal raw strings (no fold) | Q4 any-role 279 | 64 | yes |
| Q8 | Tool category of used VT | 262 | classification-service 254 | yes |
| Q9 | Full-text VirusTotal or Virus Total | 5,855 with .cols (4 missing) | 332; schema-used ∩ ft = 277 / 0 / 55 | provenance only. 0 schema-only means the schema does not invent papers the full text never names; it does not mean every used label is true use |
| Q10 | population.unit == certificates | 5,859 | 50 (22 web) | yes |
| Q11 | TLS instruments used/produced: Censys, ZGrab, crt.sh, CT, sslyze, sslscan, testssl, Qualys SSL, SSL Labs, TLS-Scanner, Let's Encrypt, certbot. OpenSSL excluded. | 5,859 | 133 (45 web) | yes |
| Q12 | Full-text TLS-cert probe (TLS…certificat / X.509 / CT / Let's Encrypt) | 5,855 | 525 / 213 web | yes, labelled upper bound |
| Q13 | Full-text CSP/headers probe (Content-Security-Policy, CSP plus header/directive, HSTS, X-Frame-Options, SRI) | 5,855 | 317 / 176 web | yes, labelled upper bound |
| Q14 | Schema header-token match on detection.phenomenon, technique, or metric | 5,859 | 443, dominated by kernel/covert-channel | no — rejected, see folding |
| Q15 | Phishing schema union (detection/classification/population/slug /phish/) | 5,859 | 139 (92 web) | yes |
| Q16 | PhishTank as sourceList or resourceName among Q15 | 139 | 21 | yes |
| Q17 | Full-text phish / phishing / phishtank | 5,855 | 1,097 | yes, as the reason a keyword search is not a population |
| Q18 | Full-text Google Safe Browsing | 5,855 | 110 | yes, in the GSB-not-a-sixth-page sentence |
| Q19 | Full-text PhishTank/OpenPhish/APWG | 5,855 | 219 | provenance only |
| Q20 | classification.target == malware used/produced | 5,859 | 159 (51 web) | yes, in the rejected-malware row |
| Q21 | classification.target == vulnerability used/produced | 5,859 | 880; web 209; crawled 152; web+crawled 99; platform offline 433 (49.2%) | yes |
| Q22 | Full-text nuclei/nikto/w3af/openvas | 5,855 | 24 | yes, rejected-scanner row |
| Q23 | Full-text zmap/zgrab/masscan | 5,855 | 246 | provenance only |
| Q24 | Full-text censys | 5,855 | 145 | provenance only |
| Q25 | tools.category == network-scanner used/produced | 5,859 | 445 (ZMap 91, nmap 41, Censys 40, …) | provenance only; not a child, because traceroute and RIPE Atlas dominate the category |
Folding, and its residue
VirusTotal. No name-fold: 64 distinct raw strings, printed in full in the report (section B). They are products, feeds, thresholds, combinations and datasets, not merely orthographic spellings. The page publishes the count, not the list. The list is the residue of “not folding”.
TLS instruments. OpenSSL is excluded from Q11. A first regex that included openssl and lego returned 224 papers and a row for “LEGO Mindstorms NXT”. OpenSSL-the-library is a crypto paper's compiler, not a TLS measurement instrument. The 88-paper OpenSSL count is in the report (section C) and is not on the content page.
CSP / headers. Two probes, only one published:
- Rejected (schema, Q14).
hdrNameoriginally ended with anelalternative meant to catch Network Error Logging. It matches the last three letters of “channel”. The top “header” phenomena were then kernel code coverage and cross-VM covert channels. That 443 is not a CSP population. Do not resurrect it. - Published (full-text, Q13). The probe requires Content-Security-Policy, or CSP near “header”/“directive”/“policy”, or HSTS / X-Frame-Options / SRI. Sample contexts (report section H) are real web-header papers (X-Frame-Options 2011–2014, CSP headers, LastPass's CSP). 176 web is still an upper bound: SRI is a homograph and “CSP” still matches some policy-language papers. The child item requires a hand map.
Phishing. Schema union 139, not full-text 1,097. The 1,097 is published only as a warning. No attempt was made to fold “phishing” out of a security-venue corpus with a tighter regex — that is the child's job.
Quotes spot-checked
Not a results page; the load-bearing objects are population rules. The report prints the first eight CSP contexts and the first eight TLS contexts (section H). That is a sample of the head of the sweep, not a validation of 317 or 525. Reading them: several CSP hits are bibliography entries for the CSP spec (Chrome-extension security architecture, privilege-separation in HTML5), and several are body uses (X-Frame-Options, LastPass's CSP). The TLS sample is mixed similarly (Holz 2011 is the paper itself; some others are citations of X.509). The VirusTotal sample of eight includes a reference-list-only hit (PeerPress CCS 2012) that the extractor had labelled used — so Q4's used/produced/source filter does not guarantee true use, and 277 is an upper bound. The 0 schema-only vs full-text figure only shows the schema does not invent papers the full text never names.
External sources
None on the content page. Vendor status, API rate limits, and current CSP spec level belong on the children (VirusTotal's cap is already dated on website_classification). The namespace page does not repeat them.
What could not be established
- A precise CSP-measurement paper count. Q13 is an upper bound. The child page closes this. The namespace still proposes the child, because the topic is real (Weichselbaum, Roth, Steffens) and a hand map is the right next sitting, not a reason to omit the red link.
- Whether the 99 web+crawled vulnerability papers found those vulnerabilities by crawling. Paper-level conjunction only.
- Whether OVERVIEW.md's folded “VirusTotal 239 / 4th most-used tool” and this script's 262 used-or-produced are the same population under a different fold, or a real disagreement. The content page publishes 277 / 262 as an extractor upper bound and does not claim a rank. The drain item's “182, 5th” was the pre-extension corpus and is not on the page.
- Homographs inside the phishing full-text 1,097. Not investigated; not used as a denominator.
Report output (unedited)
- report_security_namespace-output.txt
======================================================================== A. CORPUS ======================================================================== corpus papers 5859 empirical 5118 crawled (crawlConfig or automated-web-crawl) 1120 web platform 1622 network-scan-or-probe 930 classified 4439 ======================================================================== B. VIRUSTOTAL ======================================================================== tools[].name matches VirusTotal (any role) 263 used or produced 262 classification.resourceName matches (any role) 211 used or produced 209 population.sourceList matches 62 UNION used/produced/source 277 UNION any role 279 --- VirusTotal classification targets (used/produced; papers, multi) --- Target Papers Share of 209 ------------------------------------------------------------ ------ ------------ malware 86 41.1% domain 42 20.1% mobile-app 37 17.7% website-category 16 7.7% ip-address 11 5.3% web-request 10 4.8% vulnerability 5 2.4% other:downloaded software 1 0.5% other:advertiser binaries and software families 1 0.5% other:exposed URLs 1 0.5% other:malicious URLs 1 0.5% other:VirusTotal engine detection labels 1 0.5% other:PDF documents as malicious 1 0.5% other:website blacklist status 1 0.5% other:STIX indicator values as malicious or non-malicious 1 0.5% other:threat type of files 1 0.5% other:shared files, proxy IPs, VPN configurations, and HTTP 1 0.5% other:malware behavior risk reports 1 0.5% user-generated-text 1 0.5% other:cryptomining processes 1 0.5% --- VirusTotal tool categories (used/produced) --- Category Papers ---------------------- ------ classification-service 254 other 5 infrastructure 3 http-client 1 blocklist 1 ml-model-or-algorithm 1 --- VirusTotal used-union by year --- Window Papers Share of 277 used-union ---------- ------ ----------------------- 2010–2014 23 8.3% 2015–2018 62 22.4% 2019–2021 71 25.6% 2022–2024 76 27.4% 2025–2026* 45 16.2% --- VirusTotal used-union by venue --- Venue Papers Share ------- ------ ----- USENIX 59 21.3% CCS 54 19.5% IEEE-SP 48 17.3% NDSS 46 16.6% WWW 33 11.9% IMC 31 11.2% PETS 6 2.2% distinct VirusTotal spellings: 64 30% VirusTotal detection threshold (custom) | AMD reports and VirusTotal reports | Bazaar and VirusTotal | Drebin dataset and VirusTotal | ForcePoint engine via VirusTotal | Google Play and third-party Android markets plus VirusTotal | Google Safe Browsing and VirusTotal | Hacking forums and VirusTotal | Malsign, Malcert, Symantec data set, Samples from WINE and VirusTotal | MalwareBazaar, Hybrid Analysis, VirusTotal, and direct botnet downloads | MalwareConfig, Shodan, VirusTotal, @Scum, and ReversingLabs | MalwareConfig, Shodan, VirusTotal, and ReversingLabs | Mozilla's PDF.js test suite and VirusTotal | PDF.js test suite, VirusTotal, and V8 regression test suite | SEISMIC; MineSweeper; Musch et al.; VirusTotal; VirusShare; NoCoin; MadeWithWasm | VX Heaven, VirusShare, and VirusTotal | Virus Total | VirusShare, VirusTotal, and the AMD dataset | VirusTotal | VirusTotal API | VirusTotal API and Google SafeBrowsing | VirusTotal AV engines | VirusTotal AV-engine threshold (t=4) | VirusTotal Balanced Dataset | VirusTotal CVE tags and AV-vendor scanner | VirusTotal Hunting | VirusTotal IP and graph APIs | VirusTotal Intelligence | VirusTotal Intelligence API | VirusTotal Intelligence Search | VirusTotal Premium API | VirusTotal Public API v2.0 | VirusTotal Relations | VirusTotal Retrohunt | VirusTotal URL Feed | VirusTotal URL feed | VirusTotal and Github | VirusTotal and MalwareBazaar | VirusTotal and VirusShare | VirusTotal and abuse.ch | VirusTotal and malware.lu | VirusTotal antivirus detections | VirusTotal antivirus reports | VirusTotal blacklists | VirusTotal distribute API | VirusTotal feed | VirusTotal intelligence API | VirusTotal internal data | VirusTotal malware configurations | VirusTotal private API v3.0 | VirusTotal public API | VirusTotal report APIs | VirusTotal score | VirusTotal vhash | VirusTotal's URL reputation service | VirusTotal, HybridAnalysis, and MetaDefender | VirusTotal, MetaDefender, and HybridAnalysis | VirusTotal, MetaMask, SEAL-ISAC, Google Safe Browsing, WalletGuard, Phishfort, and ChainPatrol | VirusTotal, Qihoo 360, and Baidu | VirusTotal, URLQuery, Malware Domain List, and VxVault | VirusTotal, certificate-matched potentially benign samples | VirusTotal-labeled malware subset | filtered VirusTotal APK corpus | six VirusTotal machine-learning engines ======================================================================== C. TLS / CERTIFICATES (schema signals) ======================================================================== population.unit == certificates 50 tools[].name TLS/CT/Censys-ish used/produced 224 top TLS-ish tool names (folded, used/produced): Name Papers Spellings ----------------------------- ------ --------- OpenSSL 88 2 Censys 65 1 ZGrab2 18 3 crt.sh 16 1 ZGrab 12 2 TLS-Scanner 5 1 Certbot 4 1 OpenSSL s_client 3 1 PyOpenSSL 3 2 Qualys SSL Server Test 2 1 Certificate Transparency Logs 2 2 Certificate Transparency 2 1 SSL Labs 2 1 OpenSSL 1.0.1e 2 1 sslscan 2 1 LEGO Mindstorms NXT 1 1 OpenSSL-AES 1 1 OpenSSL v1.0.1e 1 1 Gueron/Krasnov OpenSSL patch 1 1 OpenSSL 1.0.1f 1 1 detection phenomenon/technique TLS-ish 352 top detection.phenomenon (folded) among those: Phenomenon Papers n spellings ------------------------------------- ------ ----------- HTTPS adoption 7 1 TLS interception 6 1 Certificate revocation 4 2 Certificate pinning 3 2 incomplete certificate chains 2 1 Certificate key reuse 2 2 HSTS deployment 2 1 Invalid SSL certificates 2 1 TLS version adoption 2 2 HTTPS interception 2 1 Fraudulent certificate issuance 2 1 Certificate-chain validation failures 2 2 Certificate validation failures 2 1 TLS certificate validation failures 2 1 HTTPS webmail traffic 1 1 tools category network-scanner used/produced 445 top network-scanner names: Name Papers n spellings ---------------- ------ ----------- ZMap 91 3 NMAP 41 4 Censys 40 1 Traceroute 24 2 RIPE Atlas 24 1 Scamper 22 2 Shodan 19 1 ZGrab2 18 3 XMAP 15 2 ZDNS 12 2 ping 11 1 ZGrab 11 2 MIDAR 9 1 Paris traceroute 6 2 Snort 6 1 ======================================================================== D. SECURITY HEADERS / CSP (schema signals) ======================================================================== detection mentions a header/CSP token 443 classification resource/targetDetail does 91 tools[].name does (used/produced) 105 top detection.phenomenon among header papers: Phenomenon Papers n spellings ------------------------------ ------ ----------- Kernel code coverage 5 2 Cross-VM covert channel 3 1 memory bus covert channel 2 2 HSTS deployment 2 1 CSP adoption 2 1 Linux kernel vulnerabilities 2 1 Kernel vulnerability discovery 2 1 side-channel vulnerabilities 2 1 Cross-core covert channel 2 2 covert-channel capacity 2 1 Side-channel leakage 2 2 Previously unknown kernel bugs 2 2 kernel-module loading 2 1 Kernel branch coverage 2 2 kernel data races 2 1 ======================================================================== E. PHISHING / MALWARE (schema signals) ======================================================================== classification.target malware used/produced 159 (any role 160) classification.target vulnerability used/produced 880 (any role 883) detection phenomenon/technique /phish/ 111 classification resource/targetDetail /phish/ used 51 population.sourceList /phish/ 51 UNION schema+slug /phish/ 139 top phishing sourceList / resourceName: Name Papers n spellings ------------------------------------------------------- ------ ----------- PhishTank 21 2 OpenPhish 4 1 VisualPhishNet 2 1 PhishPedia 2 2 PhishIntention 2 1 Beyond Phish 2 2 five live feeds of phishing and malware-hosting sites 1 1 Google Safe Browsing, malware feeds, phishing feeds, sc 1 1 phishing and malware feeds 1 1 Google Safe Browsing API, Malware Patrol, PhishTank, AP 1 1 Spamhaus DBL; Google Safe Browsing; PhishTank; Wepawet; 1 1 PhishTrack (custom) 1 1 PhishNet 1 1 user-reported phishing emails 1 1 SafeBrowsing anti-phishing pipeline 1 1 malware classification methods (used/produced): Method Papers Share of 159 ------------------- ------ ------------ third-party-service 96 60.4% heuristic-rules 30 18.9% curated-database 29 18.2% supervised-ml 18 11.3% manual-labelling 15 9.4% regex-or-signature 12 7.5% other 11 6.9% unsupervised-ml 6 3.8% dynamic-analysis 5 3.1% static-analysis 3 1.9% graph-analysis 3 1.9% malware resourceName top: Resource Papers n spellings ---------------------- ------ ----------- VirusTotal 80 1 AV-Class 16 2 AVCLASS2 8 2 Random Forest (custom) 5 2 ClamAV 3 1 custom 3 1 Random Forest 3 2 JaSt 3 1 Cujo 3 1 YARA rules 2 2 Logistic Regression 2 1 PEiD 2 1 vulnerability classification methods: Method Papers Share of 880 ------------------- ------ ------------ heuristic-rules 303 34.4% manual-labelling 242 27.5% static-analysis 233 26.5% dynamic-analysis 159 18.1% curated-database 78 8.9% supervised-ml 27 3.1% other 22 2.5% graph-analysis 22 2.5% regex-or-signature 21 2.4% llm 14 1.6% third-party-service 11 1.3% unsupervised-ml 6 0.7% blocklist 1 0.1% vulnerability targetDetail / resourceName top (folded): Detail Papers n spellings ------------------------------------- ------ ----------- custom 22 1 CodeQL 11 1 custom manual analysis 9 1 custom static analysis 6 1 custom manual inspection 6 1 Common Weakness Enumeration (CWE) 6 1 Address Sanitizer (ASAN) 5 3 manual analysis (custom) 4 1 VirusTotal 4 1 NVD 4 1 National Vulnerability Database (NVD) 4 1 custom manual categorization 4 1 CVE database 4 1 National Vulnerability Database 3 1 random forest (custom) 3 1 ======================================================================== F. FULL-TEXT SWEEPS (paper.cols.txt) ======================================================================== tls-cert hits= 525 missing_cols=4 csp-header hits= 317 missing_cols=4 virustotal-ft hits= 332 missing_cols=4 phishing-ft hits=1097 missing_cols=4 malware-url hits= 539 missing_cols=4 zmap-zgrab hits= 246 missing_cols=4 censys hits= 145 missing_cols=4 nuclei-nikto hits= 24 missing_cols=4 gsb hits= 110 missing_cols=4 phishtank hits= 219 missing_cols=4 VirusTotal schema-used ∩ full-text: both=277 schema-only=0 ft-only=55 ======================================================================== G. WEB-PLATFORM NARROWING ======================================================================== TLS-cert full-text n= 525 web=213 (40.6%) crawled=113 (21.5%) CSP/headers full-text n= 317 web=176 (55.5%) crawled=112 (35.3%) VirusTotal used-union n= 277 web=107 (38.6%) crawled=107 (38.6%) phishing UNION schema+slug n= 139 web=92 (66.2%) crawled=64 (46.0%) malware class used n= 159 web=51 (32.1%) crawled=52 (32.7%) vulnerability class used n= 880 web=209 (23.8%) crawled=152 (17.3%) network-scanner used n= 445 web=103 (23.1%) crawled=52 (11.7%) zmap/zgrab/masscan ft n= 246 web=67 (27.2%) crawled=40 (16.3%) censys ft n= 145 web=42 (29.0%) crawled=16 (11.0%) --- web-platform TLS-cert ft by year --- Window Papers Share of 213 web TLS-ft ---------- ------ ----------------------- 2010–2014 21 9.9% 2015–2018 54 25.4% 2019–2021 54 25.4% 2022–2024 55 25.8% 2025–2026* 29 13.6% --- web-platform CSP ft by year --- Window Papers Share of 176 web CSP-ft ---------- ------ ----------------------- 2010–2014 20 11.4% 2015–2018 44 25.0% 2019–2021 43 24.4% 2022–2024 44 25.0% 2025–2026* 25 14.2% ======================================================================== H. SAMPLE CONTEXTS (8 papers each, first hits, for hand reading) ======================================================================== --- CSP full-text --- CCS/2011/fortifying-web-based-applications-automatically kies [2] that enable web developers to specify cookies that should be inaccessible from JavaScript, X-Frame-Options [21] to enable web developers to prevent their pages from being framed, and JSON.parse() [1] to ena USENIX/2011/toward-secure-embedded-web-interfaces e need to contact external web sites. Correspondingly our server is configured to offer restrictive CSP [14] directives to browsers, limiting the impact of any injected code in the page. S-CSP (Server-side Content Sec USENIX/2012/an-evaluation-of-the-google-chrome-extension-security-architecture function-a-bad-idea. [28] B. Sterne and A. Barth. Content security policy. https://dvcs.w3.org/hg/ content-security-policy/raw-file/tip/ csp-specification.dev.html. [29] Brandon Sterne and Adam Barth. Content security po USENIX/2012/clickjacking-attacks-and-defenses sure it is the top-level document [37], or with newly added browser support, using features called X-Frame-Options [21] and CSP's frame-ancestors [39]. A fundamental limitation of framebusting is its incompatibilit USENIX/2012/privilege-separation-in-html5-applications . Sterne and A. Barth, "Content security policy: W3c editor's draft," 2012. https://dvcs. w3.org/hg/content-security-policy/ raw-file/tip/csp-specification.dev. html. [35] diigo.com, "Awesome screenshot : Capture annotate CCS/2013/cross-origin-pixel-stealing-timing-attacks-using-css-filters ause stance we could use this data to quickly determine which they access cross-origin content when X-Frame-Options are web users are T-Mobile customers. not used. As a result, setting X-Frame-Options to Deny is the NDSS/2013/the-postman-always-rings-twice-attacking-and-defending-postmessage-in-html5-webs ). This defense is independent and complementary to the defenses described in Sections 6.1 and 6.2. CSP is an HTTP header string starting with X-Content-Security-Policy or X-WebKit-CSP [7]. It instructs Web browsers how USENIX/2014/the-emperor-s-new-password-manager-security-analysis-of-web-based-password-manag se red flags for users and reviewers. In the applications we studied, only Last-Pass shipped with a Content-Security-Policy header, albeit with an unsafe policy that allows eval and inline scripts. CSRF. The prevalence of C --- TLS-cert full-text --- IMC/2011/the-ssl-landscape-a-thorough-analysis-of-the-x-509-pki-using-active-and-passive The SSL Landscape - A Thorough Analysis of the X.509 PKI Using Active and Passive Measurements Ralph Holz, Lothar Braun, Nils Kammenhuber, Georg Carl CCS/2012/an-historical-examination-of-open-source-releases-and-their-vulnerabilities aracter in a Common Name (CN) field of an there were no CVE entries for 2004 and 2005 and then five X.509 certificate...". Although release 8.14.4 is not included in 2006 (8.13.5) (see appendix A table 4). CCS/2012/the-most-dangerous-code-in-the-world-validating-ssl-certificates-in-non-browser e focus on the client's validation of the server certificate. All SSL implementations we tested use X.509 certificates. The complete algorithm for validating X.509 certificates can be found in RFC 5280 [15 CCS/2012/why-eve-and-mallory-love-android-an-analysis-of-android-ssl-in-security ain access to the public key of the server. In most client/server setups, the server ob- tains an X.509 certificate that contains the server's public key and is signed by a Certificate Authority (CA). W IMC/2013/analysis-of-the-https-certificate-ecosystem tion]: [Public key cryptosystems, Standards] Keywords TLS; SSL; HTTPS; public-key infrastructure; X.509; certificates; security; measurement; Internet-wide scanning 1. INTRODUCTION Nearly all secure we WWW/2013/heres-my-cert-so-trust-me-maybe-understanding-tls-errors-on-the-web wsers. In summary, we make the following contributions: ullet We discuss how browsers validate TLS certificates and highlight the importance of relying on browser code for such measurement studies. We iden CCS/2013/rethinking-ssl-development-in-an-appified-world o strengthen the security of certificate validation, including Perspectives [18], Convergence [13], Certificate Transparency [11], Sovereign Keys [3], TACK [12], and DANE [9]. However, none of these systems has achieved wide CCS/2013/predictability-of-android-openssls-pseudo-random-number-generator d the vulnerability of weak public key pairs in network devices. They performed largescale scans of TLS certificates and SSH host keys. After analyzing the scanned data, they discovered that there were many vulnera --- VirusTotal used --- CCS/2010/blade-an-attack-agnostic-approach-for-preventing-drive-by-malware-infections tion rate of these binaries wherein BLADE and independent instrumentation ! tools are loaded. from virustotal.com was only 28.43%. These include procmon [3] to monitor Windows system call events Only about ha CCS/2011/bitshred-feature-hashing-malware-for-scalable-triage-and-semantic-analysis To create a reference clustering data set, we used 30â¼40 different anti-virus labels provided by VirusTotal [6]. First, we chose samples that were detected as malware by at least 20 anti-virus programs to ge CCS/2012/detecting-money-stealing-apps-in-alternative-android-markets he ground truth for SMSrelated money-stealing applications, we submitted all 56,000 applications to VirusTotal, which identified 1,278 android applications as being labeled malicious by at least one AV company CCS/2012/manufacturing-compromise-the-emergence-of-exploit-as-a-service ENTS We would like to thank the Arbor Networks ASERT Team for providing us with malware samples and VirusTotal for access to the thousands of virus scanner reports we used during classification. This material i CCS/2012/peerpress-utilizing-enemies-p2p-strength-against-them usiness/theme.jsp?themeid= threatreport. [9] Temu . http://bitblaze.cs.berkeley.edu/temu.html. [10] Virustotal. https://www.virustotal.com/. [11] Z3 EMT Solver . http://research.microsoft.com/en-us/ um/redmond/ CCS/2012/vanity-cracks-and-malware-insights-into-the-anti-copy-protection-ecosystem nalysis environment (Fig. 2). In order to conduct both static and dynamic analysis, we utilized the Virustotal [14] service and the Anubis [16] environment. Virustotal [14] is a publicly available service that USENIX/2012/b-bel-leveraging-email-delivery-for-spam-mitigation 1 groups bots according to the most frequent label assigned by the anti-virus products deployed by VirusTotal [44]. Our dataset contained 13 legitimate MUAs and MTAs, and 91 distinct malware samples5 . We pick WWW/2013/bitsquatting-exploiting-bit-flips-for-fun-or-profit all of which were pointing to the same executable. We downloaded the executable and submitted it to VirusTotal, an online service that scans user-submitted files against the signature databases of popular antiv ======================================================================== I. VULNERABILITY TARGET: is it web, or everything? ======================================================================== vulnerability used, web platform 209 / 880 studyTypes among vulnerability-used (multi): studyType Papers Share of 880 -------------------------- ------ ------------ system-or-defence-proposal 728 82.7% code-or-binary-analysis 580 65.9% existing-dataset-analysis 370 42.0% manual-audit 368 41.8% network-scan-or-probe 136 15.5% automated-web-crawl 115 13.1% mobile-app-analysis 112 12.7% user-study 57 6.5% interview-or-survey 40 4.5% simulation-or-theory-only 29 3.3% platforms among vulnerability-used (multi): platform Papers Share of 880 -------------------- ------ ------------ offline 433 49.2% other-online-service 261 29.7% web 209 23.8% mobile 179 20.3% iot 95 10.8% not-applicable 3 0.3% ======================================================================== J. ALL classification.target COUNTS (used/produced, papers) ======================================================================== target Papers --------------------- ------ other 2592 vulnerability 880 website-category 424 user-generated-text 419 network-traffic 382 domain 351 ip-address 295 mobile-app 280 web-request 258 malware 159 privacy-policy 102 sdk-or-library 77 email-message 54 cookie 53 javascript 44 consent-notice 39 fingerprinting-script 31 website-popularity 15 dark-pattern 13 ======================================================================== K. CHILD-PAGE POPULATIONS (the numbers the namespace page publishes) ======================================================================== TLS instruments used/produced (no OpenSSL) 133 of which web platform 45 population.unit == certificates 50 of which web platform 22 vulnerability used AND crawled 152 vulnerability used AND crawled AND web 99 vulnerability used AND web (any study type) 209 --- TLS instruments by year --- Window Papers Share of 133 TLS-instrument ---------- ------ --------------------------- 2010–2014 0 0.0% 2015–2018 26 19.5% 2019–2021 40 30.1% 2022–2024 45 33.8% 2025–2026* 22 16.5% --- phishing schema+slug by year --- Window Papers Share of 139 phishing-union ---------- ------ --------------------------- 2010–2014 16 11.5% 2015–2018 18 12.9% 2019–2021 29 20.9% 2022–2024 40 28.8% 2025–2026* 36 25.9% PUBLISHED_CORPUS 5859 PUBLISHED_EMPIRICAL 5118 PUBLISHED_CRAWLED 1120 PUBLISHED_WEB 1622 PUBLISHED_VT_USED 277 PUBLISHED_VT_WEB 107 PUBLISHED_VT_TOOL_USED 262 PUBLISHED_VT_CLASS_USED 209 PUBLISHED_VT_SPELLINGS 64 PUBLISHED_VT_MALWARE_TARGET 86 PUBLISHED_VT_FT 332 PUBLISHED_VT_FT_ONLY 55 PUBLISHED_CERT_UNIT 50 PUBLISHED_TLS_INSTR 133 PUBLISHED_TLS_FT 525 PUBLISHED_TLS_FT_WEB 213 PUBLISHED_CSP_FT 317 PUBLISHED_CSP_FT_WEB 176 PUBLISHED_PHISH_UNION 139 PUBLISHED_PHISH_WEB 92 PUBLISHED_PHISH_FT 1097 PUBLISHED_MALWARE_USED 159 PUBLISHED_MALWARE_WEB 51 PUBLISHED_VULN_USED 880 PUBLISHED_VULN_WEB 209 PUBLISHED_VULN_CRAWL 152 PUBLISHED_VULN_WEBCRAWL 99 PUBLISHED_SCANNER 445 PUBLISHED_ZMAP_FT 246 PUBLISHED_CENSYS_FT 145 PUBLISHED_GSB_FT 110 PUBLISHED_PHISHTANK_FT 219 PUBLISHED_NUCLEI_FT 24 PUBLISHED_MISSING_COLS 4 PUBLISHED_VT_CLASS_SERVICE 254 PUBLISHED_PHISHTANK_SCHEMA 21 PUBLISHED_TLS_INSTR_WEB 45 PUBLISHED_CERT_UNIT_WEB 22 PUBLISHED_N_CHILDREN 5 PUBLISHED_N_REJECTED 3 PUBLISHED_VENUES 7 PUBLISHED_CITE_YEARS 2011 2013 2015 2016 2018 2019 2020 2021 2024 PUBLISHED_VT_OVERVIEW_RANK 4 PUBLISHED_VT_OVERVIEW_N 239 PUBLISHED_OLD_TASK_HINT_VT 182 PUBLISHED_READABLE 5855 ======================================================================== Z. EXTERNAL / ARITHMETIC — none; all figures are corpus counts from this run ========================================================================
Review
Frozen drafts for the three focused passes: out/freeze_security/ (sha256 in that directory). The content page was not edited while those three ran. Findings below were applied before the generic pass. The generic pass saw the post-apply snapshot, not the freeze.
Models. User asked for luna medium on the focused three and Luna max on the generic pass. Neither slug is in this session's Task allow-list. Available GPT family: sol medium only. The three focused passes ran as sol medium; the generic pass uses the same model. This is not the requested tier split.
Focused A — figures vs script (sol medium)
| # | Finding | Verdict |
|---|---|---|
| 1 | The 24 is labelled nuclei/nikto/w3af on the content page; the script also includes OpenVAS. | Accepted. Content row and provenance scope-table now name all four. The published 24 was already the four-name count. |
| 2 | Q14 claimed only detection.phenomenon; the script searches phenomenon, technique, and metric. | Accepted. Q14 wording corrected. The 443 stays unpublished. |
| 3 | Q11 instrument list omitted sslscan. | Accepted. sslscan added to the Q11 list. The published 133 was already the script's number. |
Focused B — citations and quotes (sol medium)
| # | Finding | Verdict |
|---|---|---|
| 1 | CrawlPhish parenthetical said a “researcher-looking crawler does not see the phish”. The paper is about cloaking against anti-phishing crawlers, not every researcher crawler. | Accepted. Narrowed to “cloaking against anti-phishing crawlers”. |
All keys resolved or were in the new-keys file; reused keys were not re-added; author surnames matched.
Focused C — external currency (sol medium)
All-clear. Contributing still requires a namespace outline; live start still has a lone Security; twelve paper identifiers resolve; Censys/ZMap/VirusTotal/Let's Encrypt/PhishTank/GSB/crt.sh still exist under those names (Censys Legacy Search is transitioning to Censys Platform, name retained); six neighbour-page overlap claims hold on the live wiki.
Generic — no checklist (sol medium)
Saw the post-focused-apply snapshot (then the start page was re-exported and the findings below were applied).
| # | Finding | Verdict |
|---|---|---|
| 1 | The start-page draft was stale and would have deleted live Programming:Registration and Programming:Docker links. | Accepted. Re-exported live start after the generic pass and patched only the Security bullet. |
| 2 | PeerPress is a references-only VT hit inside the “used” sample, so “schema does not overcount” overclaims. | Accepted. 277 is now an extractor upper bound; Q9 wording narrowed. |
| 3 | “All eight CSP samples were web-header uses” is false (bibliography hits in the head of the sweep). | Accepted. Quote-check section rewritten as a head-of-sweep sample, not a validation. |
| 4 | Headers child is proposed on an unvalidated upper bound. | Accepted. Content page now says 176 is not a population; provenance records that the red link is still the right next sitting. |
| 5 | 99 web+crawled is a conjunction, not “vulnerabilities found by visiting pages”. | Accepted. Lead and table cell qualified. |
| 6 | The phishing drain item still said “you will not see the phish if you crawl like a researcher”. | Accepted. Drain item context corrected in the same sitting. |
| 7 | The TLS drain item still omitted sslscan. | Accepted. Drain item context corrected. |
| 8 | “64 distinct spellings” overstates: the strings are products, feeds, thresholds. | Accepted. Now “64 distinct raw strings”. |
| 9 | “security and privacy venues” vs corpus page's “security, privacy and measurement venues”. | Accepted. |
| 10 | “Three rejected” vs GSB discussed as a sixth. | Accepted. Three not given a page; GSB folded into phishing. |
Finding 11 was an acceptance that the page answers its assignment.
Bibliography keys added this sitting (collision-checked against a fresh export of the live bibliography; no collisions): holz2011_landscape, durumeric2013_https, durumeric2015_search, weichselbaum2016_dead, peng2019_opening, zhang2021_crawlphish, aas2019_encrypt, durumeric2024_years.
Keys reused, not re-added: kotzias2018_coming, roth2020complex, steffens2021_blockparty, steffens2019_dont.
