User Tools

Site Tools


provenance:security

Provenance: security

Working log behind security. Corpus-wide caveats are on corpus. Citations use the shared bibliography; this page adds no keys of its own.

Run: 2026-08-27. Corpus: 5,859 extracted papers, 7 venues, 2010–2026, data/extract/run1. Item: drain wiki-measuretheweb / “Security namespace: propose then create”, claimed as cursor-drain-security. Author of the page and this log: Cursor, not Claude Code.

Creating, not extending. node scripts/dw.mjs info security returned “the requested page does not exist”. ?do=export_raw on security returned the HTML error page, which is not an existence test by byte count — the content is an HTML error document with the page id in the title. dw.mjs pages does not list the provenance: namespace and is not how this was checked. No security:* child pages exist.

Scope decision

The item said: do not write a namespace page that outlines nothing; first propose 4–6 child pages the corpus can support; record them as new drain items; then write the namespace page.

Decision Why What a reasonable person might have done instead
Five children, not a sixth Five distinct measurement jobs (PKI, headers, VT-as-oracle, phishing feeds, in-the-wild web vulns). A sixth would have been GSB, which is a feed and belongs on the phishing page. Six pages by splitting GSB out, or four by merging phishing+VT.
Short namespace page (outline + support counts + rejections) contributing: a namespace page outlines the pages inside it rather than carrying its own content. The children own the deep figures. Write TLS/CSP/VT as sections of this page and never create children. Rejected: that is the “outlines nothing” failure mode in reverse.
Rejected scanning-ethics as a child ethics already has “Server-side scanning: the one place this field has a real checklist”. Duplicate it under security: so the namespace looks fuller.
Rejected malware-datasets as a child 159 malware-classification papers, 51 web. Two thirds are not web measurement. A page anyway, padded with mobile malware that mobile_and_app_measurement already covers.
Rejected nuclei/nikto/w3af/openvas as a child Full-text 24 papers. A “scanner tools” page with four names.
VirusTotal under security: not programming: The student question is “what does a VT label mean”, not “how do I call the API”. The API-cap / vendor-taxonomy material for topic labels is already on website_classification. programming:virustotal as a sibling of tranco. Defensible; the split is recorded here so the next run does not undo it.
No ~~DISCUSSION~~ on this provenance page Established default: comments belong on the content page.

Drain items added (source=agent, effort high), to be written as their own sittings:

  • security:tls_certificates (new)
  • security:headers (new)
  • security:virustotal (new)
  • security:phishing (new)
  • security:web_vulnerabilities (new)

Report script

scripts/report_security_namespace.mjs. Deterministic. Re-run:

node scripts/report_security_namespace.mjs > out/report_security_namespace.txt

Every figure on security is a PUBLISHED_* line or a table cell in that output. The script exits 1 if the number of missing paper.cols.txt files is not 4 — that is the corpus's known gap, and a different number means the probe is not reading the same mount.

Queries, with their denominators

# Query Population / denominator Result On the page?
Q1 Papers in the extraction all 5,859 yes, outer frame
Q2 platforms includes web 5,859 1,622 named in several cells
Q3 crawled (crawlConfig !== null or studyTypes includes automated-web-crawl) 5,859 1,120 not as a headline; used in the web-vuln slice
Q4 VirusTotal UNION: tools[].name or classification.resourceName or population.sourceList matching /virus total/i, role used+produced+source 5,859 277 (107 web) yes
Q5 Q4, tools used/produced only 5,859 262 yes
Q6 Q4, classification used/produced; target breakdown 209 malware 86, domain 42, mobile-app 37, website-category 16 yes
Q7 Distinct VirusTotal raw strings (no fold) Q4 any-role 279 64 yes
Q8 Tool category of used VT 262 classification-service 254 yes
Q9 Full-text VirusTotal or Virus Total 5,855 with .cols (4 missing) 332; schema-used ∩ ft = 277 / 0 / 55 provenance only. 0 schema-only means the schema does not invent papers the full text never names; it does not mean every used label is true use
Q10 population.unit == certificates 5,859 50 (22 web) yes
Q11 TLS instruments used/produced: Censys, ZGrab, crt.sh, CT, sslyze, sslscan, testssl, Qualys SSL, SSL Labs, TLS-Scanner, Let's Encrypt, certbot. OpenSSL excluded. 5,859 133 (45 web) yes
Q12 Full-text TLS-cert probe (TLS…certificat / X.509 / CT / Let's Encrypt) 5,855 525 / 213 web yes, labelled upper bound
Q13 Full-text CSP/headers probe (Content-Security-Policy, CSP plus header/directive, HSTS, X-Frame-Options, SRI) 5,855 317 / 176 web yes, labelled upper bound
Q14 Schema header-token match on detection.phenomenon, technique, or metric 5,859 443, dominated by kernel/covert-channel no — rejected, see folding
Q15 Phishing schema union (detection/classification/population/slug /phish/) 5,859 139 (92 web) yes
Q16 PhishTank as sourceList or resourceName among Q15 139 21 yes
Q17 Full-text phish / phishing / phishtank 5,855 1,097 yes, as the reason a keyword search is not a population
Q18 Full-text Google Safe Browsing 5,855 110 yes, in the GSB-not-a-sixth-page sentence
Q19 Full-text PhishTank/OpenPhish/APWG 5,855 219 provenance only
Q20 classification.target == malware used/produced 5,859 159 (51 web) yes, in the rejected-malware row
Q21 classification.target == vulnerability used/produced 5,859 880; web 209; crawled 152; web+crawled 99; platform offline 433 (49.2%) yes
Q22 Full-text nuclei/nikto/w3af/openvas 5,855 24 yes, rejected-scanner row
Q23 Full-text zmap/zgrab/masscan 5,855 246 provenance only
Q24 Full-text censys 5,855 145 provenance only
Q25 tools.category == network-scanner used/produced 5,859 445 (ZMap 91, nmap 41, Censys 40, …) provenance only; not a child, because traceroute and RIPE Atlas dominate the category

Folding, and its residue

VirusTotal. No name-fold: 64 distinct raw strings, printed in full in the report (section B). They are products, feeds, thresholds, combinations and datasets, not merely orthographic spellings. The page publishes the count, not the list. The list is the residue of “not folding”.

TLS instruments. OpenSSL is excluded from Q11. A first regex that included openssl and lego returned 224 papers and a row for “LEGO Mindstorms NXT”. OpenSSL-the-library is a crypto paper's compiler, not a TLS measurement instrument. The 88-paper OpenSSL count is in the report (section C) and is not on the content page.

CSP / headers. Two probes, only one published:

  • Rejected (schema, Q14). hdrName originally ended with a nel alternative meant to catch Network Error Logging. It matches the last three letters of “channel”. The top “header” phenomena were then kernel code coverage and cross-VM covert channels. That 443 is not a CSP population. Do not resurrect it.
  • Published (full-text, Q13). The probe requires Content-Security-Policy, or CSP near “header”/“directive”/“policy”, or HSTS / X-Frame-Options / SRI. Sample contexts (report section H) are real web-header papers (X-Frame-Options 2011–2014, CSP headers, LastPass's CSP). 176 web is still an upper bound: SRI is a homograph and “CSP” still matches some policy-language papers. The child item requires a hand map.

Phishing. Schema union 139, not full-text 1,097. The 1,097 is published only as a warning. No attempt was made to fold “phishing” out of a security-venue corpus with a tighter regex — that is the child's job.

Quotes spot-checked

Not a results page; the load-bearing objects are population rules. The report prints the first eight CSP contexts and the first eight TLS contexts (section H). That is a sample of the head of the sweep, not a validation of 317 or 525. Reading them: several CSP hits are bibliography entries for the CSP spec (Chrome-extension security architecture, privilege-separation in HTML5), and several are body uses (X-Frame-Options, LastPass's CSP). The TLS sample is mixed similarly (Holz 2011 is the paper itself; some others are citations of X.509). The VirusTotal sample of eight includes a reference-list-only hit (PeerPress CCS 2012) that the extractor had labelled used — so Q4's used/produced/source filter does not guarantee true use, and 277 is an upper bound. The 0 schema-only vs full-text figure only shows the schema does not invent papers the full text never names.

External sources

None on the content page. Vendor status, API rate limits, and current CSP spec level belong on the children (VirusTotal's cap is already dated on website_classification). The namespace page does not repeat them.

What could not be established

  • A precise CSP-measurement paper count. Q13 is an upper bound. The child page closes this. The namespace still proposes the child, because the topic is real (Weichselbaum, Roth, Steffens) and a hand map is the right next sitting, not a reason to omit the red link.
  • Whether the 99 web+crawled vulnerability papers found those vulnerabilities by crawling. Paper-level conjunction only.
  • Whether OVERVIEW.md's folded “VirusTotal 239 / 4th most-used tool” and this script's 262 used-or-produced are the same population under a different fold, or a real disagreement. The content page publishes 277 / 262 as an extractor upper bound and does not claim a rank. The drain item's “182, 5th” was the pre-extension corpus and is not on the page.
  • Homographs inside the phishing full-text 1,097. Not investigated; not used as a denominator.

Report output (unedited)

report_security_namespace-output.txt
========================================================================
A. CORPUS
========================================================================
corpus papers                                       5859
empirical                                           5118
crawled (crawlConfig or automated-web-crawl)        1120
web platform                                        1622
network-scan-or-probe                               930
classified                                          4439
 
========================================================================
B. VIRUSTOTAL
========================================================================
tools[].name matches VirusTotal (any role)          263
  used or produced                                  262
classification.resourceName matches (any role)      211
  used or produced                                  209
population.sourceList matches                       62
UNION used/produced/source                          277
UNION any role                                      279
 
--- VirusTotal classification targets (used/produced; papers, multi) ---
Target                                                        Papers  Share of 209
------------------------------------------------------------  ------  ------------
malware                                                       86      41.1%
domain                                                        42      20.1%
mobile-app                                                    37      17.7%
website-category                                              16      7.7%
ip-address                                                    11      5.3%
web-request                                                   10      4.8%
vulnerability                                                 5       2.4%
other:downloaded software                                     1       0.5%
other:advertiser binaries and software families               1       0.5%
other:exposed URLs                                            1       0.5%
other:malicious URLs                                          1       0.5%
other:VirusTotal engine detection labels                      1       0.5%
other:PDF documents as malicious                              1       0.5%
other:website blacklist status                                1       0.5%
other:STIX indicator values as malicious or non-malicious     1       0.5%
other:threat type of files                                    1       0.5%
other:shared files, proxy IPs, VPN configurations, and HTTP   1       0.5%
other:malware behavior risk reports                           1       0.5%
user-generated-text                                           1       0.5%
other:cryptomining processes                                  1       0.5%
 
--- VirusTotal tool categories (used/produced) ---
Category                Papers
----------------------  ------
classification-service  254
other                   5
infrastructure          3
http-client             1
blocklist               1
ml-model-or-algorithm   1
 
--- VirusTotal used-union by year ---
Window      Papers  Share of 277 used-union
----------  ------  -----------------------
2010–2014   23      8.3%
2015–2018   62      22.4%
2019–2021   71      25.6%
2022–2024   76      27.4%
2025–2026*  45      16.2%
--- VirusTotal used-union by venue ---
Venue    Papers  Share
-------  ------  -----
USENIX   59      21.3%
CCS      54      19.5%
IEEE-SP  48      17.3%
NDSS     46      16.6%
WWW      33      11.9%
IMC      31      11.2%
PETS     6       2.2%
 
distinct VirusTotal spellings: 64
30% VirusTotal detection threshold (custom) | AMD reports and VirusTotal reports | Bazaar and VirusTotal | Drebin dataset and VirusTotal | ForcePoint engine via VirusTotal | Google Play and third-party Android markets plus VirusTotal | Google Safe Browsing and VirusTotal | Hacking forums and VirusTotal | Malsign, Malcert, Symantec data set, Samples from WINE and VirusTotal | MalwareBazaar, Hybrid Analysis, VirusTotal, and direct botnet downloads | MalwareConfig, Shodan, VirusTotal, @Scum, and ReversingLabs | MalwareConfig, Shodan, VirusTotal, and ReversingLabs | Mozilla's PDF.js test suite and VirusTotal | PDF.js test suite, VirusTotal, and V8 regression test suite | SEISMIC; MineSweeper; Musch et al.; VirusTotal; VirusShare; NoCoin; MadeWithWasm | VX Heaven, VirusShare, and VirusTotal | Virus Total | VirusShare, VirusTotal, and the AMD dataset | VirusTotal | VirusTotal API | VirusTotal API and Google SafeBrowsing | VirusTotal AV engines | VirusTotal AV-engine threshold (t=4) | VirusTotal Balanced Dataset | VirusTotal CVE tags and AV-vendor scanner | VirusTotal Hunting | VirusTotal IP and graph APIs | VirusTotal Intelligence | VirusTotal Intelligence API | VirusTotal Intelligence Search | VirusTotal Premium API | VirusTotal Public API v2.0 | VirusTotal Relations | VirusTotal Retrohunt | VirusTotal URL Feed | VirusTotal URL feed | VirusTotal and Github | VirusTotal and MalwareBazaar | VirusTotal and VirusShare | VirusTotal and abuse.ch | VirusTotal and malware.lu | VirusTotal antivirus detections | VirusTotal antivirus reports | VirusTotal blacklists | VirusTotal distribute API | VirusTotal feed | VirusTotal intelligence API | VirusTotal internal data | VirusTotal malware configurations | VirusTotal private API v3.0 | VirusTotal public API | VirusTotal report APIs | VirusTotal score | VirusTotal vhash | VirusTotal's URL reputation service | VirusTotal, HybridAnalysis, and MetaDefender | VirusTotal, MetaDefender, and HybridAnalysis | VirusTotal, MetaMask, SEAL-ISAC, Google Safe Browsing, WalletGuard, Phishfort, and ChainPatrol | VirusTotal, Qihoo 360, and Baidu | VirusTotal, URLQuery, Malware Domain List, and VxVault | VirusTotal, certificate-matched potentially benign samples | VirusTotal-labeled malware subset | filtered VirusTotal APK corpus | six VirusTotal machine-learning engines
 
========================================================================
C. TLS / CERTIFICATES (schema signals)
========================================================================
population.unit == certificates                     50
tools[].name TLS/CT/Censys-ish used/produced        224
top TLS-ish tool names (folded, used/produced):
Name                           Papers  Spellings
-----------------------------  ------  ---------
OpenSSL                        88      2
Censys                         65      1
ZGrab2                         18      3
crt.sh                         16      1
ZGrab                          12      2
TLS-Scanner                    5       1
Certbot                        4       1
OpenSSL s_client               3       1
PyOpenSSL                      3       2
Qualys SSL Server Test         2       1
Certificate Transparency Logs  2       2
Certificate Transparency       2       1
SSL Labs                       2       1
OpenSSL 1.0.1e                 2       1
sslscan                        2       1
LEGO Mindstorms NXT            1       1
OpenSSL-AES                    1       1
OpenSSL v1.0.1e                1       1
Gueron/Krasnov OpenSSL patch   1       1
OpenSSL 1.0.1f                 1       1
 
detection phenomenon/technique TLS-ish              352
top detection.phenomenon (folded) among those:
Phenomenon                             Papers  n spellings
-------------------------------------  ------  -----------
HTTPS adoption                         7       1
TLS interception                       6       1
Certificate revocation                 4       2
Certificate pinning                    3       2
incomplete certificate chains          2       1
Certificate key reuse                  2       2
HSTS deployment                        2       1
Invalid SSL certificates               2       1
TLS version adoption                   2       2
HTTPS interception                     2       1
Fraudulent certificate issuance        2       1
Certificate-chain validation failures  2       2
Certificate validation failures        2       1
TLS certificate validation failures    2       1
HTTPS webmail traffic                  1       1
 
tools category network-scanner used/produced        445
top network-scanner names:
Name              Papers  n spellings
----------------  ------  -----------
ZMap              91      3
NMAP              41      4
Censys            40      1
Traceroute        24      2
RIPE Atlas        24      1
Scamper           22      2
Shodan            19      1
ZGrab2            18      3
XMAP              15      2
ZDNS              12      2
ping              11      1
ZGrab             11      2
MIDAR             9       1
Paris traceroute  6       2
Snort             6       1
 
========================================================================
D. SECURITY HEADERS / CSP (schema signals)
========================================================================
detection mentions a header/CSP token               443
classification resource/targetDetail does           91
tools[].name does (used/produced)                   105
top detection.phenomenon among header papers:
Phenomenon                      Papers  n spellings
------------------------------  ------  -----------
Kernel code coverage            5       2
Cross-VM covert channel         3       1
memory bus covert channel       2       2
HSTS deployment                 2       1
CSP adoption                    2       1
Linux kernel vulnerabilities    2       1
Kernel vulnerability discovery  2       1
side-channel vulnerabilities    2       1
Cross-core covert channel       2       2
covert-channel capacity         2       1
Side-channel leakage            2       2
Previously unknown kernel bugs  2       2
kernel-module loading           2       1
Kernel branch coverage          2       2
kernel data races               2       1
 
========================================================================
E. PHISHING / MALWARE (schema signals)
========================================================================
classification.target malware used/produced         159  (any role 160)
classification.target vulnerability used/produced   880  (any role 883)
detection phenomenon/technique /phish/              111
classification resource/targetDetail /phish/ used   51
population.sourceList /phish/                       51
UNION schema+slug /phish/                           139
 
top phishing sourceList / resourceName:
Name                                                     Papers  n spellings
-------------------------------------------------------  ------  -----------
PhishTank                                                21      2
OpenPhish                                                4       1
VisualPhishNet                                           2       1
PhishPedia                                               2       2
PhishIntention                                           2       1
Beyond Phish                                             2       2
five live feeds of phishing and malware-hosting sites    1       1
Google Safe Browsing, malware feeds, phishing feeds, sc  1       1
phishing and malware feeds                               1       1
Google Safe Browsing API, Malware Patrol, PhishTank, AP  1       1
Spamhaus DBL; Google Safe Browsing; PhishTank; Wepawet;  1       1
PhishTrack (custom)                                      1       1
PhishNet                                                 1       1
user-reported phishing emails                            1       1
SafeBrowsing anti-phishing pipeline                      1       1
 
malware classification methods (used/produced):
Method               Papers  Share of 159
-------------------  ------  ------------
third-party-service  96      60.4%
heuristic-rules      30      18.9%
curated-database     29      18.2%
supervised-ml        18      11.3%
manual-labelling     15      9.4%
regex-or-signature   12      7.5%
other                11      6.9%
unsupervised-ml      6       3.8%
dynamic-analysis     5       3.1%
static-analysis      3       1.9%
graph-analysis       3       1.9%
 
malware resourceName top:
Resource                Papers  n spellings
----------------------  ------  -----------
VirusTotal              80      1
AV-Class                16      2
AVCLASS2                8       2
Random Forest (custom)  5       2
ClamAV                  3       1
custom                  3       1
Random Forest           3       2
JaSt                    3       1
Cujo                    3       1
YARA rules              2       2
Logistic Regression     2       1
PEiD                    2       1
 
vulnerability classification methods:
Method               Papers  Share of 880
-------------------  ------  ------------
heuristic-rules      303     34.4%
manual-labelling     242     27.5%
static-analysis      233     26.5%
dynamic-analysis     159     18.1%
curated-database     78      8.9%
supervised-ml        27      3.1%
other                22      2.5%
graph-analysis       22      2.5%
regex-or-signature   21      2.4%
llm                  14      1.6%
third-party-service  11      1.3%
unsupervised-ml      6       0.7%
blocklist            1       0.1%
 
vulnerability targetDetail / resourceName top (folded):
Detail                                 Papers  n spellings
-------------------------------------  ------  -----------
custom                                 22      1
CodeQL                                 11      1
custom manual analysis                 9       1
custom static analysis                 6       1
custom manual inspection               6       1
Common Weakness Enumeration (CWE)      6       1
Address Sanitizer (ASAN)               5       3
manual analysis (custom)               4       1
VirusTotal                             4       1
NVD                                    4       1
National Vulnerability Database (NVD)  4       1
custom manual categorization           4       1
CVE database                           4       1
National Vulnerability Database        3       1
random forest (custom)                 3       1
 
========================================================================
F. FULL-TEXT SWEEPS (paper.cols.txt)
========================================================================
tls-cert               hits= 525  missing_cols=4
csp-header             hits= 317  missing_cols=4
virustotal-ft          hits= 332  missing_cols=4
phishing-ft            hits=1097  missing_cols=4
malware-url            hits= 539  missing_cols=4
zmap-zgrab             hits= 246  missing_cols=4
censys                 hits= 145  missing_cols=4
nuclei-nikto           hits=  24  missing_cols=4
gsb                    hits= 110  missing_cols=4
phishtank              hits= 219  missing_cols=4
 
VirusTotal schema-used ∩ full-text: both=277 schema-only=0 ft-only=55
 
========================================================================
G. WEB-PLATFORM NARROWING
========================================================================
TLS-cert full-text           n= 525  web=213 (40.6%)  crawled=113 (21.5%)
CSP/headers full-text        n= 317  web=176 (55.5%)  crawled=112 (35.3%)
VirusTotal used-union        n= 277  web=107 (38.6%)  crawled=107 (38.6%)
phishing UNION schema+slug   n= 139  web=92 (66.2%)  crawled=64 (46.0%)
malware class used           n= 159  web=51 (32.1%)  crawled=52 (32.7%)
vulnerability class used     n= 880  web=209 (23.8%)  crawled=152 (17.3%)
network-scanner used         n= 445  web=103 (23.1%)  crawled=52 (11.7%)
zmap/zgrab/masscan ft        n= 246  web=67 (27.2%)  crawled=40 (16.3%)
censys ft                    n= 145  web=42 (29.0%)  crawled=16 (11.0%)
 
--- web-platform TLS-cert ft by year ---
Window      Papers  Share of 213 web TLS-ft
----------  ------  -----------------------
2010–2014   21      9.9%
2015–2018   54      25.4%
2019–2021   54      25.4%
2022–2024   55      25.8%
2025–2026*  29      13.6%
--- web-platform CSP ft by year ---
Window      Papers  Share of 176 web CSP-ft
----------  ------  -----------------------
2010–2014   20      11.4%
2015–2018   44      25.0%
2019–2021   43      24.4%
2022–2024   44      25.0%
2025–2026*  25      14.2%
 
========================================================================
H. SAMPLE CONTEXTS (8 papers each, first hits, for hand reading)
========================================================================
--- CSP full-text ---
  CCS/2011/fortifying-web-based-applications-automatically
    kies [2] that enable web developers to specify cookies that should be inaccessible from JavaScript, X-Frame-Options [21] to enable web developers to prevent their pages from being framed, and JSON.parse() [1] to ena
  USENIX/2011/toward-secure-embedded-web-interfaces
    e need to contact external web sites. Correspondingly our server is configured to offer restrictive CSP [14] directives to browsers, limiting the impact of any injected code in the page. S-CSP (Server-side Content Sec
  USENIX/2012/an-evaluation-of-the-google-chrome-extension-security-architecture
    function-a-bad-idea. [28] B. Sterne and A. Barth. Content security policy. https://dvcs.w3.org/hg/ content-security-policy/raw-file/tip/ csp-specification.dev.html. [29] Brandon Sterne and Adam Barth. Content security po
  USENIX/2012/clickjacking-attacks-and-defenses
     sure it is the top-level document [37], or with newly added browser support, using features called X-Frame-Options [21] and CSP's frame-ancestors [39]. A fundamental limitation of framebusting is its incompatibilit
  USENIX/2012/privilege-separation-in-html5-applications
    . Sterne and A. Barth, "Content security policy: W3c editor's draft," 2012. https://dvcs. w3.org/hg/content-security-policy/ raw-file/tip/csp-specification.dev. html. [35] diigo.com, "Awesome screenshot : Capture annotate 
  CCS/2013/cross-origin-pixel-stealing-timing-attacks-using-css-filters
    ause stance we could use this data to quickly determine which they access cross-origin content when X-Frame-Options are web users are T-Mobile customers. not used. As a result, setting X-Frame-Options to Deny is the
  NDSS/2013/the-postman-always-rings-twice-attacking-and-defending-postmessage-in-html5-webs
    ). This defense is independent and complementary to the defenses described in Sections 6.1 and 6.2. CSP is an HTTP header string starting with X-Content-Security-Policy or X-WebKit-CSP [7]. It instructs Web browsers how 
  USENIX/2014/the-emperor-s-new-password-manager-security-analysis-of-web-based-password-manag
    se red flags for users and reviewers. In the applications we studied, only Last-Pass shipped with a Content-Security-Policy header, albeit with an unsafe policy that allows eval and inline scripts. CSRF. The prevalence of C
 
--- TLS-cert full-text ---
  IMC/2011/the-ssl-landscape-a-thorough-analysis-of-the-x-509-pki-using-active-and-passive
    The SSL Landscape - A Thorough Analysis of the X.509 PKI Using Active and Passive Measurements Ralph Holz, Lothar Braun, Nils Kammenhuber, Georg Carl
  CCS/2012/an-historical-examination-of-open-source-releases-and-their-vulnerabilities
    aracter in a Common Name (CN) field of an there were no CVE entries for 2004 and 2005 and then five X.509 certificate...". Although release 8.14.4 is not included in 2006 (8.13.5) (see appendix A table 4).
  CCS/2012/the-most-dangerous-code-in-the-world-validating-ssl-certificates-in-non-browser
    e focus on the client's validation of the server certificate. All SSL implementations we tested use X.509 certificates. The complete algorithm for validating X.509 certificates can be found in RFC 5280 [15
  CCS/2012/why-eve-and-mallory-love-android-an-analysis-of-android-ssl-in-security
    ain access to the public key of the server. In most client/server setups, the server ob- tains an X.509 certificate that contains the server's public key and is signed by a Certificate Authority (CA). W
  IMC/2013/analysis-of-the-https-certificate-ecosystem
    tion]: [Public key cryptosystems, Standards] Keywords TLS; SSL; HTTPS; public-key infrastructure; X.509; certificates; security; measurement; Internet-wide scanning 1. INTRODUCTION Nearly all secure we
  WWW/2013/heres-my-cert-so-trust-me-maybe-understanding-tls-errors-on-the-web
    wsers. In summary, we make the following contributions: ullet We discuss how browsers validate TLS certificates and highlight the importance of relying on browser code for such measurement studies. We iden
  CCS/2013/rethinking-ssl-development-in-an-appified-world
    o strengthen the security of certificate validation, including Perspectives [18], Convergence [13], Certificate Transparency [11], Sovereign Keys [3], TACK [12], and DANE [9]. However, none of these systems has achieved wide
  CCS/2013/predictability-of-android-openssls-pseudo-random-number-generator
    d the vulnerability of weak public key pairs in network devices. They performed largescale scans of TLS certificates and SSH host keys. After analyzing the scanned data, they discovered that there were many vulnera
 
--- VirusTotal used ---
  CCS/2010/blade-an-attack-agnostic-approach-for-preventing-drive-by-malware-infections
    tion rate of these binaries wherein BLADE and independent instrumentation ! tools are loaded. from virustotal.com was only 28.43%. These include procmon [3] to monitor Windows system call events Only about ha
  CCS/2011/bitshred-feature-hashing-malware-for-scalable-triage-and-semantic-analysis
     To create a reference clustering data set, we used 30∼40 different anti-virus labels provided by VirusTotal [6]. First, we chose samples that were detected as malware by at least 20 anti-virus programs to ge
  CCS/2012/detecting-money-stealing-apps-in-alternative-android-markets
    he ground truth for SMSrelated money-stealing applications, we submitted all 56,000 applications to VirusTotal, which identified 1,278 android applications as being labeled malicious by at least one AV company
  CCS/2012/manufacturing-compromise-the-emergence-of-exploit-as-a-service
    ENTS We would like to thank the Arbor Networks ASERT Team for providing us with malware samples and VirusTotal for access to the thousands of virus scanner reports we used during classification. This material i
  CCS/2012/peerpress-utilizing-enemies-p2p-strength-against-them
    usiness/theme.jsp?themeid= threatreport. [9] Temu . http://bitblaze.cs.berkeley.edu/temu.html. [10] Virustotal. https://www.virustotal.com/. [11] Z3 EMT Solver . http://research.microsoft.com/en-us/ um/redmond/
  CCS/2012/vanity-cracks-and-malware-insights-into-the-anti-copy-protection-ecosystem
    nalysis environment (Fig. 2). In order to conduct both static and dynamic analysis, we utilized the Virustotal [14] service and the Anubis [16] environment. Virustotal [14] is a publicly available service that
  USENIX/2012/b-bel-leveraging-email-delivery-for-spam-mitigation
     1 groups bots according to the most frequent label assigned by the anti-virus products deployed by VirusTotal [44]. Our dataset contained 13 legitimate MUAs and MTAs, and 91 distinct malware samples5 . We pick
  WWW/2013/bitsquatting-exploiting-bit-flips-for-fun-or-profit
    all of which were pointing to the same executable. We downloaded the executable and submitted it to VirusTotal, an online service that scans user-submitted files against the signature databases of popular antiv
 
========================================================================
I. VULNERABILITY TARGET: is it web, or everything?
========================================================================
vulnerability used, web platform                    209 / 880
studyTypes among vulnerability-used (multi):
studyType                   Papers  Share of 880
--------------------------  ------  ------------
system-or-defence-proposal  728     82.7%
code-or-binary-analysis     580     65.9%
existing-dataset-analysis   370     42.0%
manual-audit                368     41.8%
network-scan-or-probe       136     15.5%
automated-web-crawl         115     13.1%
mobile-app-analysis         112     12.7%
user-study                  57      6.5%
interview-or-survey         40      4.5%
simulation-or-theory-only   29      3.3%
 
platforms among vulnerability-used (multi):
platform              Papers  Share of 880
--------------------  ------  ------------
offline               433     49.2%
other-online-service  261     29.7%
web                   209     23.8%
mobile                179     20.3%
iot                   95      10.8%
not-applicable        3       0.3%
 
========================================================================
J. ALL classification.target COUNTS (used/produced, papers)
========================================================================
target                 Papers
---------------------  ------
other                  2592
vulnerability          880
website-category       424
user-generated-text    419
network-traffic        382
domain                 351
ip-address             295
mobile-app             280
web-request            258
malware                159
privacy-policy         102
sdk-or-library         77
email-message          54
cookie                 53
javascript             44
consent-notice         39
fingerprinting-script  31
website-popularity     15
dark-pattern           13
 
========================================================================
K. CHILD-PAGE POPULATIONS (the numbers the namespace page publishes)
========================================================================
TLS instruments used/produced (no OpenSSL)          133
  of which web platform                             45
population.unit == certificates                     50
  of which web platform                             22
 
vulnerability used AND crawled                      152
vulnerability used AND crawled AND web              99
vulnerability used AND web (any study type)         209
 
--- TLS instruments by year ---
Window      Papers  Share of 133 TLS-instrument
----------  ------  ---------------------------
2010–2014   0       0.0%
2015–2018   26      19.5%
2019–2021   40      30.1%
2022–2024   45      33.8%
2025–2026*  22      16.5%
--- phishing schema+slug by year ---
Window      Papers  Share of 139 phishing-union
----------  ------  ---------------------------
2010–2014   16      11.5%
2015–2018   18      12.9%
2019–2021   29      20.9%
2022–2024   40      28.8%
2025–2026*  36      25.9%
 
PUBLISHED_CORPUS 5859
PUBLISHED_EMPIRICAL 5118
PUBLISHED_CRAWLED 1120
PUBLISHED_WEB 1622
PUBLISHED_VT_USED 277
PUBLISHED_VT_WEB 107
PUBLISHED_VT_TOOL_USED 262
PUBLISHED_VT_CLASS_USED 209
PUBLISHED_VT_SPELLINGS 64
PUBLISHED_VT_MALWARE_TARGET 86
PUBLISHED_VT_FT 332
PUBLISHED_VT_FT_ONLY 55
PUBLISHED_CERT_UNIT 50
PUBLISHED_TLS_INSTR 133
PUBLISHED_TLS_FT 525
PUBLISHED_TLS_FT_WEB 213
PUBLISHED_CSP_FT 317
PUBLISHED_CSP_FT_WEB 176
PUBLISHED_PHISH_UNION 139
PUBLISHED_PHISH_WEB 92
PUBLISHED_PHISH_FT 1097
PUBLISHED_MALWARE_USED 159
PUBLISHED_MALWARE_WEB 51
PUBLISHED_VULN_USED 880
PUBLISHED_VULN_WEB 209
PUBLISHED_VULN_CRAWL 152
PUBLISHED_VULN_WEBCRAWL 99
PUBLISHED_SCANNER 445
PUBLISHED_ZMAP_FT 246
PUBLISHED_CENSYS_FT 145
PUBLISHED_GSB_FT 110
PUBLISHED_PHISHTANK_FT 219
PUBLISHED_NUCLEI_FT 24
PUBLISHED_MISSING_COLS 4
PUBLISHED_VT_CLASS_SERVICE 254
PUBLISHED_PHISHTANK_SCHEMA 21
PUBLISHED_TLS_INSTR_WEB 45
PUBLISHED_CERT_UNIT_WEB 22
PUBLISHED_N_CHILDREN 5
PUBLISHED_N_REJECTED 3
PUBLISHED_VENUES 7
PUBLISHED_CITE_YEARS 2011 2013 2015 2016 2018 2019 2020 2021 2024
PUBLISHED_VT_OVERVIEW_RANK 4
PUBLISHED_VT_OVERVIEW_N 239
PUBLISHED_OLD_TASK_HINT_VT 182
PUBLISHED_READABLE 5855
 
========================================================================
Z. EXTERNAL / ARITHMETIC — none; all figures are corpus counts from this run
========================================================================

Review

Frozen drafts for the three focused passes: out/freeze_security/ (sha256 in that directory). The content page was not edited while those three ran. Findings below were applied before the generic pass. The generic pass saw the post-apply snapshot, not the freeze.

Models. User asked for luna medium on the focused three and Luna max on the generic pass. Neither slug is in this session's Task allow-list. Available GPT family: sol medium only. The three focused passes ran as sol medium; the generic pass uses the same model. This is not the requested tier split.

Focused A — figures vs script (sol medium)

# Finding Verdict
1 The 24 is labelled nuclei/nikto/w3af on the content page; the script also includes OpenVAS. Accepted. Content row and provenance scope-table now name all four. The published 24 was already the four-name count.
2 Q14 claimed only detection.phenomenon; the script searches phenomenon, technique, and metric. Accepted. Q14 wording corrected. The 443 stays unpublished.
3 Q11 instrument list omitted sslscan. Accepted. sslscan added to the Q11 list. The published 133 was already the script's number.

Focused B — citations and quotes (sol medium)

# Finding Verdict
1 CrawlPhish parenthetical said a “researcher-looking crawler does not see the phish”. The paper is about cloaking against anti-phishing crawlers, not every researcher crawler. Accepted. Narrowed to “cloaking against anti-phishing crawlers”.

All keys resolved or were in the new-keys file; reused keys were not re-added; author surnames matched.

Focused C — external currency (sol medium)

All-clear. Contributing still requires a namespace outline; live start still has a lone Security; twelve paper identifiers resolve; Censys/ZMap/VirusTotal/Let's Encrypt/PhishTank/GSB/crt.sh still exist under those names (Censys Legacy Search is transitioning to Censys Platform, name retained); six neighbour-page overlap claims hold on the live wiki.

Generic — no checklist (sol medium)

Saw the post-focused-apply snapshot (then the start page was re-exported and the findings below were applied).

# Finding Verdict
1 The start-page draft was stale and would have deleted live Programming:Registration and Programming:Docker links. Accepted. Re-exported live start after the generic pass and patched only the Security bullet.
2 PeerPress is a references-only VT hit inside the “used” sample, so “schema does not overcount” overclaims. Accepted. 277 is now an extractor upper bound; Q9 wording narrowed.
3 “All eight CSP samples were web-header uses” is false (bibliography hits in the head of the sweep). Accepted. Quote-check section rewritten as a head-of-sweep sample, not a validation.
4 Headers child is proposed on an unvalidated upper bound. Accepted. Content page now says 176 is not a population; provenance records that the red link is still the right next sitting.
5 99 web+crawled is a conjunction, not “vulnerabilities found by visiting pages”. Accepted. Lead and table cell qualified.
6 The phishing drain item still said “you will not see the phish if you crawl like a researcher”. Accepted. Drain item context corrected in the same sitting.
7 The TLS drain item still omitted sslscan. Accepted. Drain item context corrected.
8 “64 distinct spellings” overstates: the strings are products, feeds, thresholds. Accepted. Now “64 distinct raw strings”.
9 “security and privacy venues” vs corpus page's “security, privacy and measurement venues”. Accepted.
10 “Three rejected” vs GSB discussed as a sixth. Accepted. Three not given a page; GSB folded into phishing.

Finding 11 was an acceptance that the page answers its assignment.

Bibliography keys added this sitting (collision-checked against a fresh export of the live bibliography; no collisions): holz2011_landscape, durumeric2013_https, durumeric2015_search, weichselbaum2016_dead, peng2019_opening, zhang2021_crawlphish, aas2019_encrypt, durumeric2024_years.

Keys reused, not re-added: kotzias2018_coming, roth2020complex, steffens2021_blockparty, steffens2019_dont.

provenance/security.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki