User Tools

Site Tools


provenance:security:web_vulnerabilities

Provenance: Security:Web vulnerabilities

Back to Web vulnerabilities. Corpus-wide selection and extraction notes are on corpus. Citations use the shared bibliography; this page adds no keys of its own. No discussion block on this page — comments belong on the content page.

Run record

  • Run date: 2026-08-27 (UTC).
  • Drain item: security:web_vulnerabilities (new), store id 224, run 62, claimed by cursor-drain-webvuln. Executed as Cursor, not via claude -p / drain-sandbox.sh.
  • Review models requested by this sitting: GPT 5.6 Luna medium (gpt-5.6-luna-medium) for the three focused passes and the generic pass, instead of the spec's sonnet/fable split. Logged here so a later reader does not assume the default mix.
  • Corpus at run time: 5,859 extracted papers, 2010–2026, CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P. Read-only inputs under /workspace/publications_dataset/data/.
  • Read first: the wiki-measuretheweb spec (audience, currency, 5,859-paper corpus, provenance page, start link, bibtex), corpus, security, ethics, notifying_websites, javascript, foxhound, interaction, live start (already promised Web vulnerabilities).
  • Target pages had no revision (core.getPageInfo “does not exist”). dw.mjs pages does not list provenance:. ?do=export_raw on a missing page returns an HTML error document — that is not an existence test. This is a creation, not an extension.
  • Live start (rev 1787836787) already links the child; this sitting does not edit start. Live security (rev 1787835259) still says all five children are red links — WRAP updated to name this child as written. Live VirusTotal does not exist; the local draft of that page is not this item and is not claimed live.
  • No write to the publication mount. Wiki saves through scripts/dw.mjs (JSON-RPC only). Live literature:bibliography was re-exported immediately before append (rev 1787836782), not taken from a stale local copy. The live bibliography already has a stray extra } after [1Stock, Ben; Johns, Martin; Steffens, Marius; Backes, Michael (2017): "How the Web Tangled Itself: Uncovering the History of Client-Side Web (In) Security", in: 26th USENIX Security Symposium (USENIX Security 17), pp. 971-987. USENIX Association, Vancouver, BC. (Link)]; this sitting does not “fix” it.
  • Hosted code: pages/count_vuln_sites.py (--demo / --selftest both run). Report: scripts/report_web_vulnerabilities.mjs. Role map: scripts/vuln_fold.mjs. Figure check: scripts/verify_web_vuln_figures.mjs. External re-fetch: scripts/external_checks_web_vuln.sh.

Why this page, not an overlap

Neighbour What it already answers What it does not
ethics scanning-ethics checklist, Hantke et al. [2Hantke, Florian; Roth, Sebastian; Mrowczynski, Rafael; Utz, Christine; Stock, Ben (2024): "Where Are the Red Lines? Towards Ethical Server-Side Scans in Security and Privacy Research", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] methods and denominators for XSS/CSRF/SOP on live sites
notifying_websites how to tell an operator how you counted the finding you are about to tell them
javascript script behaviour without an exploitability claim “is this live origin exploitable”
headers CSP/HSTS as crawlable artefacts (still unwritten) payload-executed XSS
security schema populations for five children, including the 880 / 209 / 99 the hand map of the 99

A reasonable person might have treated the 99 as the population. That is the mistake the parent cell already forbids. Another might have written a nuclei/ZAP tool page; full-text union of nuclei/nikto/w3af/openvas is 24 papers — residue here, not a page (already rejected on security).

Population and queries

All counts are papers unless labelled tuples. Sentinels are not answers. Four corpus papers have no paper.cols.txt; none of the 99 are among those four.

Membership of the 99 is mechanical: a used or produced classification.target == vulnerability tuple, platforms includes web, and the paper is in POPULATIONS.crawled. scripts/report_web_vulnerabilities.mjs exits 1 if that conjunction and the ROLE map disagree.

Query Denominator Result
classification.target == vulnerability, any role 5,859 883 (15.1%)
… used or produced 5,859 880 (15.0%) — published lead
… used + web 880 209 (23.8%)
… used + crawled 880 152 (17.3%)
… used + web + crawled 880 99 (11.3%); 8.8% of 1,120 crawled
… used + offline (multi-label platform) 880 433 (49.2%)
… used + offline-only 880 260 (29.5%)
ROLE = wild among the 99 99; 1,120 crawled 30 (30.3% of 99; 2.7% of crawled) — page population for in-the-wild figures
… of which archive-surface (Lerner rewriting-history; Stock Wayback [1Stock, Ben; Johns, Martin; Steffens, Marius; Backes, Michael (2017): "How the Web Tangled Itself: Uncovering the History of Client-Side Web (In) Security", in: 26th USENIX Security Symposium (USENIX Security 17), pp. 971-987. USENIX Association, Vancouver, BC. (Link)]) 30 2
… of which crawled then-live sites 30 28
wild classification method: dynamic-analysis / heuristic-rules / manual-labelling 30, multi 14 (46.7%) / 12 (40.0%) / 8 (26.7%)
ROLE = lab 99 27 (27.3%)
ROLE = cve 99 8 (8.1%)
ROLE = offtopic 99 34 (34.3%)
kind fold xss / csrf / clickjacking / sop / sqli / client-other / server-other 99, multi-label 25 / 7 / 1 / 4 / 4 / 7 / 6
kind-fold residue (matched no family) 99 62 — printed in the report, not a failure
wild ∩ xss 30 14
posters in the 99 99 1 (CCS 2024 login-interface ZAP poster, ROLE=wild)
corpus posters 5,859 138 (from corpus, not re-derived)
full-text clickjack among web+crawled 857 39 — upper bound
full-text nuclei|nikto|w3af|openvas 5,859 24 — upper bound
wild interactionDepth deep-crawl 30 10 (33.3%)
wild landing-plus-subpages 30 9 (30.0%)
wild single-target-page 30 5 (16.7%)
wild landing-page-only 30 2 (6.7%)
wild not-stated 30 2 (6.7%)
wild no crawlConfig 30 2 (6.7%)
wild shallow (landing-only + single-target) 30 7 (23.3%)
wild with a stated subpagesPerSite 30 10
wild venues USENIX / CCS / IEEE-SP / NDSS / PETS / WWW / IMC 30 10 / 7 / 6 / 6 / 1 / 0 / 0
wild 2022–2024 30 12 (40.0%)
wild 2025–2026* 30 6 (20.0%)
heuristic-rules / manual-labelling / dynamic-analysis among the 99 99, multi 33 (33.3%) / 28 (28.3%) / 19 (19.2%)
full-text “false positive” among wild 30 25 — upper bound on the words
full-text manually confirm/verify/inspect/analyse among wild 30 20 — same caveat
interaction landing-page share on the site-depth axis 417 of 857 web crawls 155 (37.2%) — quoted from that page, not re-derived here

Do not treat 880 or 99 as a method count. The 30 is the in-the-wild population.

Folding

  • ROLE (scripts/vuln_fold.mjs): one label per paper of the 99. Deciding sentence inline, quote-checked against paper.cols.txt. The report throws if a key of the 99 is missing from ROLE or ROLE contains a key that is not in the 99.
  • Kind fold: ordered regex families over title, slug and vulnerability tuples; multi-label. Residue 62 of 99 is printed in full in the report block below. Clickjacking as a primary study is essentially absent from this schema slice (1 of 99, 0 wild).
  • Archive-surface split (added after generic review): two keys, not a regex — CCS/2017/rewriting-history and USENIX/2017/how-the-web-tangled-itself. White Rabbit also reports archived redirects as a second experiment; it stays in the 28 because the crawl was then-live. The report throws if the two sets do not partition the 30.
    • USENIX/2010/searching-the-searchers-with-searchaudit → wild (SQL error oracle on live search engines), not lab.
    • PETS/2023/comparing-large-scale-privacy-and-security-notifications → wild (they notified live operators about findings from a crawl), not offtopic.
    • CCS/2024 login-interface posterwild (ZAP on live login pages) and counted in the 30; a methods-compression caveat sits on the page.
    • CCS/2017/rewriting-history → wild (archived web as the measurement surface).
    • Lauinger et al. (NDSS 2017 outdated JS) → cve, not wild: library version ≠ origin exploitable.
    • Spider-Scents and Black Widow → lab.

Wild keys (30):

  • CCS/2013/25-million-flows-later-large-scale-detection-of-dom-based-xss
  • CCS/2015/from-facepalm-to-brain-bender-exploring-client-side-cross-site-scripting
  • CCS/2017/rewriting-history-changing-the-archived-web-from-the-present
  • CCS/2021/out-of-sight-out-of-mind-detecting-orphaned-web-pages-at-internet-scale
  • CCS/2022/distinct-identity-theft-using-in-browser-communications-in-dual-window-single-si
  • CCS/2024/poster-security-of-login-interfaces-in-modern-organizations
  • CCS/2025/in-the-dom-we-trust-exploring-the-hidden-dangers-of-reading-from-the-dom-on-the
  • IEEE-SP/2019/empoweb-empowering-web-applications-with-browser-extensions
  • IEEE-SP/2022/the-state-of-the-samesite-studying-the-usage-effectiveness-and-adequacy-of-sames
  • IEEE-SP/2022/towards-automated-auditing-for-account-and-session-management-flaws-in-single-si
  • IEEE-SP/2023/its-dom-clobbering-time-attack-techniques-prevalence-and-defenses
  • IEEE-SP/2024/the-great-request-robbery-an-empirical-study-of-client-side-request-hijacking-vu
  • IEEE-SP/2025/403-forbidden-ethically-evaluating-broken-access-control-in-the-wild
  • NDSS/2013/the-postman-always-rings-twice-attacking-and-defending-postmessage-in-html5-webs
  • NDSS/2019/dont-trust-the-locals-investigating-the-prevalence-of-persistent-client-side-cross-site-scripting-in-the-wild
  • NDSS/2022/auto-draft-207 (Kang, Li, Cao — Probe the Proto; DOI 10.14722/ndss.2022.24308, not the auto-draft URL)
  • NDSS/2025/do-not-follow-the-white-rabbit-challenging-the-myth-of-harmless-open-redirection
  • NDSS/2025/misdirection-of-trust-demystifying-the-abuse-of-dedicated-url-shortening-service
  • NDSS/2026/dom-xss-detection-via-webpage-interaction-fuzzing-and-url-component-synthesis
  • PETS/2023/comparing-large-scale-privacy-and-security-notifications
  • USENIX/2010/searching-the-searchers-with-searchaudit
  • USENIX/2014/precise-client-side-protection-against-dom-based-cross-site-scripting
  • USENIX/2014/ssoscan-automated-testing-of-web-applications-for-single-sign-on-vulnerabilities
  • USENIX/2017/how-the-web-tangled-itself-uncovering-the-history-of-client-side-web-in-security
  • USENIX/2020/cached-and-confused-web-cache-deception-in-the-wild
  • USENIX/2022/web-cache-deception-escalates
  • USENIX/2023/a-large-scale-measurement-of-website-login-policies
  • USENIX/2023/extending-a-hand-to-attackers-browser-privilege-escalation-attacks-via-extension
  • USENIX/2024/dancer-in-the-dark-synthesizing-and-evaluating-polyglots-for-blind-cross-site-sc
  • USENIX/2025/the-domino-effect-detecting-and-exploiting-dom-clobbering-gadgets-via-concolic-e

Lab / cve / offtopic keys are the rest of ROLE in scripts/vuln_fold.mjs (27 / 8 / 34).

Quotes checked against paper.cols.txt

ROLE deciding quotes: 68 exact / 31 partial (60% five-word-window threshold, hyphen-break normalised) / 0 below, of 99. Below-threshold would have failed the report. Partial is the usual column-splice, not an unsupported claim.

Cover-paper figures on the content page were additionally grepped in paper.cols.txt by scripts/verify_web_vuln_figures.mjs: 45 literals, 0 missing (Lekies, Steffens, Bau, Son, Chehade, Kang, clobbering, request robbery, Dancer, SameSite, Black Widow). Presence is not pairing; the pairings were read in the detection.prevalence sentences in report section L.

out/authors.json for Cached and Confused was hand-fixed: the HTML parser had emitted KU Leuven as an author. Live authors used: Seyed Ali Mirheidari, Sajjad Arshad, Kaan Onarlioglu, Bruno Crispo, Engin Kirda, William Robertson.

External and industry sources

Fetched 2026-08-27 by scripts/external_checks_web_vuln.sh (exits 1 on any miss):

Claim on the page Primary source Verdict
OWASP Top 10:2025 A01 Broken Access Control; XSS inside A05 Injection as CWE-79 https://owasp.org/Top10/2025/0x00_2025-Introduction/ and https://owasp.org/Top10/2025/A05_2025-Injection/ kept
Chrome XSS Auditor removed in Chrome 78 https://developer.chrome.com/blog/chrome-78-deps-rems and https://www.chromium.org/developers/design-documents/xss-auditor/ kept
Chrome SameSite Lax-by-default is Chrome 80 (2020-02) https://blog.chromium.org/2020/02/samesite-cookie-changes-in-february.html kept
ZAP rebranded ZAP by Checkmarx on 2024-09-24, still Apache-2.0 https://www.zaproxy.org/blog/2024-09-24-zap-has-joined-forces-with-checkmarx/ kept
projectdiscovery/nuclei exists and is not archived; templates live GitHub API /repos/projectdiscovery/nuclei and nuclei-templates kept as a residue, not a method

Rejected (so the next sitting does not re-add them):

  • OWASP Top 10 2017 A7 as the current XSS letter — superseded 2021 (A03) and 2025 (A05).
  • Chromium XSS Auditor as a 2026 defence to evaluate against.
  • A nuclei / nikto / w3af / OpenVAS scanner-tool page — 24 full-text hits, already rejected on security.
  • MDN / textbook explanations of X-Frame-Options to pad the clickjacking hole.
  • SEO “top web vulnerabilities 2026” listicles.
  • Treating 880, 209 or 99 as “papers that measured XSS in the wild”.

What could not be established

  • A clickjacking-in-the-wild prevalence from this schema slice. Kind fold: 0 wild, 1 of 99. Full-text 39/857 is an upper bound on mentions.
  • A field-wide confirmation rate. Full-text “false positive” / “manually confirm” hits are word counts.
  • Whether every wild paper's crawl is how the vulnerability was found — ROLE is a hand reading of the deciding sentence, not a second extraction pass over PDFs.
  • Kang et al.'s slug in the index is still auto-draft-207. The DOI used is 10.14722/ndss.2022.24308.
  • Spider-Scents stored-XSS counts on 12 known apps are a lab result; they are cited as a detector, not a prevalence.

Judgement calls

  • Create the promised child rather than widen ethics or javascript.
  • Single-label ROLE (a paper cannot be wild and lab). Kind fold is the multi-label one.
  • Kind-fold residue of 62 is printed, not collapsed into “other XSS”.
  • Quote 37.2% from interaction rather than re-derive a slightly different landing-page share on a different axis.
  • Leave the stray } after [1Stock, Ben; Johns, Martin; Steffens, Marius; Backes, Michael (2017): "How the Web Tangled Itself: Uncovering the History of Client-Side Web (In) Security", in: 26th USENIX Security Symposium (USENIX Security 17), pp. 971-987. USENIX Association, Vancouver, BC. (Link)] in the live bibliography.
  • Add eight bibliography keys (seven cover papers + Spider-Scents). Collision-checked against the live export: 0 hits.
  • count_vuln_sites.py refuses a missing confirmed column (KeyError) rather than defaulting unconfirmed rows to vulnerable.
  • The report prints FAILURE to stderr and process.exit(1) — a script that prints FAILURE and exits 0 is not failing loudly.

Bibliography keys added this sitting

bau2010_state, son2013_postman, eriksson2021_black, khodayari2022_state, kang2022_probe, khodayari2023_clobbering, mirheidari2020_cached, olsson2024_spider.

Already live and reused: steffens2019_dont, lekies2013_million, chehade2025_forbidden, stock2017web, khodayari2024_great, kirchner2024_dancer, mirheidari2022_cache, hantke2024_redlines, sabino2026_detection, drescher2025_trust, khodayari2025_follow.

Published code

pages/count_vuln_sites.py is byte-identical to the <file python count_vuln_sites.py> block. uv run python pages/count_vuln_sites.py –selftest OK. --demo output on 2026-08-27:

findings (rows)                         5
unique URLs                             5
unique hosts                            3
confirmed rows                          3
confirmed unique URLs                   3
confirmed unique hosts                  2
unconfirmed rows                        2
sinks:
  innerHTML	3
  eval	1
  document.write	1
Do not publish unique URLs as sites, or unconfirmed rows as vulnerabilities.

Reviews

Drafts frozen at out/freeze_web_vuln/ after first publish (content rev 1787838181, provenance rev 1787838183). Three focused GPT 5.6 Luna medium passes ran in parallel against that freeze. The content page was not edited while they ran. No GENERIC_REVIEW placeholder.

Pass Model Findings Disposition
1. Figures vs script GPT 5.6 Luna medium none Accepted as empty. Re-ran report, verify_web_vuln_figures.mjs, page-number checks (windowed, whole-page, –code), table/wrap guards, count_vuln_sites.py –demo/–selftest, freeze byte-compare.
2. Citations and quotes GPT 5.6 Luna medium none Accepted as empty. All 17 content-page keys resolve; eight additions unique; cover-paper figures present in paper.cols.txt.
3. External currency GPT 5.6 Luna medium none Accepted as empty. Re-fetched OWASP 2025, Chrome 78/80, ZAP Checkmarx blog, zaproxy/zaproxy licence Apache-2.0, nuclei not archived.

Generic pass (no checklist) ran after the focused log was on this page. Content was not edited while it ran.

Pass Model Finding Disposition
4. Generic GPT 5.6 Luna medium 1. Wild 30 mixes live and archived; no live-only count. Accepted. Hand-split 2 archive-surface / 28 then-live; named in WRAP and the role table. Not a regex.
4. Generic GPT 5.6 Luna medium 2. Method distribution is for the 99, not the 30. Accepted. Report now prints H2 on the 30: dynamic-analysis 14 (46.7%), heuristic-rules 12 (40.0%), manual-labelling 8 (26.7%). Page leads with that; 99 kept as the flipped conjunction summary.
4. Generic GPT 5.6 Luna medium 3. “2024–2026 specialised browsers still do” overgroups Foxhound / PanoptiChrome / Sabino. Accepted. Scoped: Foxhound is the pipeline; PanoptiChrome is the same family; Sabino adds interaction fuzzing on top of a taint pass.
4. Generic GPT 5.6 Luna medium 4. Provenance No DISCUSSION token rendered as broken italics. Accepted. Reworded without nesting the discussion token.
4. Generic (re-run) GPT 5.6 Luna medium Called the 28 a count of sites, not papers. Accepted. “28 papers” / “2 papers” in WRAP and the role table.
4. Generic (re-run) GPT 5.6 Luna medium Report H2 labelled “live-measurement” while including the 2 archive-surface papers. Accepted. Relabelled “wild-role method distribution”.

No GENERIC_REVIEW placeholder.

Report output (unedited)

Command: node scripts/report_web_vulnerabilities.mjs. This is the run the content page was written from.

report_web_vulnerabilities-output.txt
==========================================================================
A. CORPUS
==========================================================================
corpus papers                                       5859
crawled                                             1120
web platform                                        1622
web AND crawled                                     857
missing paper.cols.txt                              4
venues: CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P
years: 2010–2026 (2025–2026 provisional)
 
==========================================================================
B. SCHEMA CONJUNCTION (the 99, not a method count)
==========================================================================
Slice                                         Papers  Share
--------------------------------------------  ------  -----
classification.target=vulnerability any role  883     15.1%
… used/produced                               880     15.0%
… used + web                                  209     23.8%
… used + crawled                              152     17.3%
… used + web + crawled  ← page membership     99      11.3%
… used + offline (multi-label)                433     49.2%
… used + offline-only                         260     29.5%
PUBLISHED_VULN_USED 880
PUBLISHED_VULN_WEB 209
PUBLISHED_VULN_CRAWL 152
PUBLISHED_VULN_WEBCRAWL 99
PUBLISHED_OFFLINE_MULTI 433
PUBLISHED_OFFLINE_ONLY 260
PUBLISHED_OFFLINE_SHARE 49.2%
 
==========================================================================
C. ROLE HAND MAP (scripts/vuln_fold.mjs ROLE)
==========================================================================
Role      Papers of 99  Share of 99  Share of 1120 crawled
--------  ------------  -----------  ---------------------
wild      30            30.3%        2.7%
lab       27            27.3%        2.4%
cve       8             8.1%         0.7%
offtopic  34            34.3%        3.0%
PUBLISHED_WILD 30
PUBLISHED_LAB 27
PUBLISHED_CVE 8
PUBLISHED_OFFTOPIC 34
role sum 99 (must equal 99)
 
==========================================================================
D. KIND FOLD (multi-label; residue printed)
==========================================================================
Kind          Papers of 99  of which wild  of which lab
------------  ------------  -------------  ------------
xss           25            14             11
csrf          7             7              0
clickjacking  1             0              0
sop           4             4              0
sqli          4             0              4
client-other  7             7              0
server-other  6             1              5
kind-fold residue (matched no family): 62
  RESIDUE USENIX/2010/searching-the-searchers-with-searchaudit role=wild title=Searching the Searchers with SearchAudit
  RESIDUE USENIX/2010/securing-script-based-extensibility-in-web-browsers role=offtopic title=Securing Script-Based Extensibility in Web Browsers
  RESIDUE WWW/2010/detection-and-analysis-of-drive-by-download-attacks-and-malicious-javascript-cod role=offtopic title=Detection and analysis of drive-by-download attacks and malicious JavaScript code.
  RESIDUE CCS/2011/fashion-crimes-trending-term-exploitation-on-the-web role=offtopic title=Fashion crimes: trending-term exploitation on the web.
  RESIDUE WWW/2011/heat-seeking-honeypots-design-and-experience role=offtopic title=Heat-seeking honeypots: design and experience.
  RESIDUE USENIX/2014/ssoscan-automated-testing-of-web-applications-for-single-sign-on-vulnerabilities role=wild title=SSOScan: Automated Testing of Web Applications for Single Sign-On Vulnerabilities
  RESIDUE CCS/2015/an-empirical-study-of-web-vulnerability-discovery-ecosystems role=cve title=An Empirical Study of Web Vulnerability Discovery Ecosystems.
  RESIDUE USENIX/2014/automatically-detecting-vulnerable-websites-before-they-turn-malicious role=offtopic title=Automatically Detecting Vulnerable Websites Before They Turn Malicious
  RESIDUE IMC/2016/a-view-from-the-other-side-understanding-mobile-phone-characteristics-in-the-dev role=offtopic title=A View from the Other Side: Understanding Mobile Phone Characteristics in the Developing World.
  RESIDUE IMC/2016/browser-feature-usage-on-the-modern-web role=cve title=Browser Feature Usage on the Modern Web.
  RESIDUE CCS/2017/a-large-scale-empirical-study-of-security-patches role=cve title=A Large-Scale Empirical Study of Security Patches.
  RESIDUE CCS/2017/rewriting-history-changing-the-archived-web-from-the-present role=wild title=Rewriting History: Changing the Archived Web from the Present.
  RESIDUE IMC/2017/measuring-and-mitigating-oauth-access-token-abuse-by-collusion-networks role=offtopic title=Measuring and mitigating oauth access token abuse by collusion networks.
  RESIDUE NDSS/2017/thou-shalt-not-depend-on-me-analysing-the-use-of-outdated-javascript-libraries-o role=cve title=Thou Shalt Not Depend on Me: Analysing the Use of Outdated JavaScript Libraries on the Web
  RESIDUE USENIX/2018/acquisitional-rule-based-engine-for-discovering-internet-of-things-devices role=offtopic title=Acquisitional Rule-based Engine for Discovering Internet-of-Things Devices
  RESIDUE USENIX/2018/from-patching-delays-to-infection-symptoms-using-risk-profiles-for-an-early-disc role=cve title=From Patching Delays to Infection Symptoms: Using Risk Profiles for an Early Discovery of Vulnerabilities Exploited in the Wild
  RESIDUE USENIX/2018/understanding-the-reproducibility-of-crowd-reported-security-vulnerabilities role=offtopic title=Understanding the Reproducibility of Crowd-reported Security Vulnerabilities
  RESIDUE USENIX/2019/devils-in-the-guidance-predicting-logic-vulnerabilities-in-payment-syndication-s role=lab title=Devils in the Guidance: Predicting Logic Vulnerabilities in Payment Syndication Services through Automated Documentation Analysis
  RESIDUE USENIX/2019/less-is-more-quantifying-the-security-benefits-of-debloating-web-applications role=lab title=Less is More: Quantifying the Security Benefits of Debloating Web Applications
  RESIDUE WWW/2019/an-investigation-of-cyber-autonomy-on-government-websites role=offtopic title=An Investigation of Cyber Autonomy on Government Websites.
  RESIDUE USENIX/2020/firmscope-automatic-uncovering-of-privilege-escalation-vulnerabilities-in-pre-in role=offtopic title=FIRMSCOPE: Automatic Uncovering of Privilege-Escalation Vulnerabilities in Pre-Installed Apps in Android Firmware
  RESIDUE CCS/2021/spinner-automated-dynamic-command-subsystem-perturbation role=lab title=Spinner: Automated Dynamic Command Subsystem Perturbation.
  RESIDUE USENIX/2021/blind-in-on-path-attacks-and-applications-to-vpns role=offtopic title=Blind In/On-Path Attacks and Applications to VPNs
  RESIDUE USENIX/2021/messy-states-of-wiring-vulnerabilities-in-emerging-personal-payment-systems role=lab title=Messy States of Wiring: Vulnerabilities in Emerging Personal Payment Systems
  RESIDUE WWW/2021/an-empirical-study-of-real-world-webassembly-binaries-security-languages-use-cas role=offtopic title=An Empirical Study of Real-World WebAssembly Binaries: Security, Languages, Use Cases.
  RESIDUE WWW/2021/tls-1-3-in-practice-how-tls-1-3-contributes-to-the-internet role=offtopic title=TLS 1.3 in Practice: How TLS 1.3 Contributes to the Internet.
  RESIDUE CCS/2022/cart-ology-intercepting-targeted-advertising-via-ad-network-identity-entanglemen role=offtopic title=Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement.
  RESIDUE CCS/2022/distinct-identity-theft-using-in-browser-communications-in-dual-window-single-si role=wild title=DISTINCT: Identity Theft using In-Browser Communications in Dual-Window Single Sign-On.
  RESIDUE IMC/2022/exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services role=offtopic title=Exploring the security and privacy risks of chatbots in messaging services.
  RESIDUE PETS/2022/how-not-to-handle-keys-timing-attacks-on-fido-authenticator-privacy role=offtopic title=How Not to Handle Keys: Timing Attacks on FIDO Authenticator Privacy
  RESIDUE USENIX/2022/phish-in-sheeps-clothing-exploring-the-authentication-pitfalls-of-browser-finger role=offtopic title=Phish in Sheep's Clothing: Exploring the Authentication Pitfalls of Browser Fingerprinting
  RESIDUE IMC/2023/a-longitudinal-study-of-vulnerable-client-side-resources-and-web-developers-upda role=cve title=A Longitudinal Study of Vulnerable Client-side Resources and Web Developers' Updating Behaviors.
  RESIDUE CCS/2023/jack-in-the-box-an-empirical-study-of-javascript-bundling-on-the-web-and-its-sec role=cve title=Jack-in-the-box: An Empirical Study of JavaScript Bundling on the Web and its Security Implications.
  RESIDUE USENIX/2023/a-large-scale-measurement-of-website-login-policies role=wild title=A Large-Scale Measurement of Website Login Policies
  RESIDUE PETS/2023/comparing-large-scale-privacy-and-security-notifications role=wild title=Comparing Large-Scale Privacy and Security Notifications
  RESIDUE WWW/2023/bad-apples-understanding-the-centralized-security-risks-in-decentralized-ecosyst role=offtopic title=Bad Apples: Understanding the Centralized Security Risks in Decentralized Ecosystems.
  RESIDUE USENIX/2023/animatedead-debloating-web-applications-using-concolic-execution role=lab title=AnimateDead: Debloating Web Applications Using Concolic Execution
  RESIDUE USENIX/2023/extending-a-hand-to-attackers-browser-privilege-escalation-attacks-via-extension role=wild title=Extending a Hand to Attackers: Browser Privilege Escalation Attacks via Extensions
  RESIDUE USENIX/2023/minimalist-semi-automated-debloating-of-php-web-applications-through-static-anal role=lab title=Minimalist: Semi-automated Debloating of PHP Web Applications through Static Analysis
  RESIDUE NDSS/2024/quack-hindering-deserialization-attacks-via-static-duck-typing role=lab title=QUACK: Hindering Deserialization Attacks via Static Duck Typing
  RESIDUE CCS/2024/a-first-look-at-security-and-privacy-risks-in-the-rapidapi-ecosystem role=offtopic title=A First Look at Security and Privacy Risks in the RapidAPI Ecosystem.
  RESIDUE PETS/2024/a-black-box-privacy-analysis-of-messaging-service-providers-chat-message-process role=offtopic title=A Black-Box Privacy Analysis of Messaging Service Providers' Chat Message Processing
  RESIDUE IEEE-SP/2024/mawseo-adversarial-wiki-search-poisoning-for-illicit-online-promotion role=offtopic title=MAWSEO: Adversarial Wiki Search Poisoning for Illicit Online Promotion.
  RESIDUE IMC/2024/analyzing-the-impact-of-copying-and-pasting-vulnerable-solidity-code-snippets-fr role=offtopic title=Analyzing the Impact of Copying-and-Pasting Vulnerable Solidity Code Snippets from Question-and-Answer Websites.
  RESIDUE USENIX/2024/smudged-fingerprints-characterizing-and-improving-the-performance-of-web-applica role=cve title=Smudged Fingerprints: Characterizing and Improving the Performance of Web Application Fingerprinting
  RESIDUE IEEE-SP/2010/a-symbolic-execution-framework-for-javascript role=lab title=A Symbolic Execution Framework for JavaScript.
  RESIDUE IEEE-SP/2016/seeking-nonsense-looking-for-trouble-efficient-promotional-infection-detection-t role=offtopic title=Seeking Nonsense, Looking for Trouble: Efficient Promotional-Infection Detection through Semantic Inconsistency Search.
  RESIDUE CCS/2025/bacscan-automatic-black-box-detection-of-broken-access-control-vulnerabilities-i role=lab title=BACScan: Automatic Black-Box Detection of Broken-Access-Control Vulnerabilities in Web Applications.
  RESIDUE CCS/2025/in-the-dom-we-trust-exploring-the-hidden-dangers-of-reading-from-the-dom-on-the role=wild title=In the DOM We Trust: Exploring the Hidden Dangers of Reading from the DOM on the Web.
  RESIDUE NDSS/2025/attributing-open-source-contributions-is-critical-but-difficult-a-systematic-analysis-of-github-practices-and-their-impact-on-software-supply-chain-security role=offtopic title=Attributing Open-Source Contributions is Critical but Difficult: A Systematic Analysis of GitHub Practices and Their Impact on Software Supply Chain Security
  RESIDUE WWW/2025/whats-in-phishers-a-longitudinal-study-of-security-configurations-in-phishing-we role=offtopic title=What's in Phishers: A Longitudinal Study of Security Configurations in Phishing Websites and Kits.
  RESIDUE NDSS/2026/transparent-taint-style-vulnerability-detection-in-generic-single-page-applications-through-automated-framework-abstraction role=lab title=TranSPArent: Taint-style Vulnerability Detection in Generic Single Page Applications through Automated Framework Abstraction
  RESIDUE PETS/2026/the-masks-we-think-we-wear-privacy-threats-of-browser-extension-wallets-in-the-w role=offtopic title=The Masks We (Think We) Wear: Privacy Threats of Browser-Extension Wallets in the Web3 Ecosystem
  RESIDUE USENIX/2026/abuse-risks-are-often-inherent-to-product-features-exploring-ai-vendors-bug-boun role=offtopic title="Abuse Risks are Often Inherent to Product Features": Exploring AI Vendors' Bug Bounty and Responsible Disclosure Policies
  RESIDUE NDSS/2025/mens-sana-in-corpore-sano-sound-firmware-corpora-for-vulnerability-research role=offtopic title=Mens Sana In Corpore Sano: Sound Firmware Corpora for Vulnerability Research
  RESIDUE NDSS/2025/the-midas-touch-triggering-the-capability-of-llms-for-rm-api-misuse-detection role=offtopic title=The Midas Touch: Triggering the Capability of LLMs for RM-API Misuse Detection
  RESIDUE IEEE-SP/2012/lastor-a-low-latency-as-aware-tor-client role=offtopic title=LASTor: A Low-Latency AS-Aware Tor Client.
  RESIDUE NDSS/2025/misdirection-of-trust-demystifying-the-abuse-of-dedicated-url-shortening-service role=wild title=Misdirection of Trust: Demystifying the Abuse of Dedicated URL Shortening Service
  RESIDUE IEEE-SP/2012/evilseed-a-guided-approach-to-finding-malicious-web-pages role=offtopic title=EvilSeed: A Guided Approach to Finding Malicious Web Pages.
  RESIDUE IEEE-SP/2022/towards-automated-auditing-for-account-and-session-management-flaws-in-single-si role=wild title=Towards Automated Auditing for Account and Session Management Flaws in Single Sign-On Deployments.
  RESIDUE IEEE-SP/2023/devious-device-driven-side-channel-attacks-on-the-iommu role=offtopic title=DevIOus: Device-Driven Side-Channel Attacks on the IOMMU.
  RESIDUE IEEE-SP/2023/toss-a-fault-to-your-witcher-applying-grey-box-coverage-guided-mutational-fuzzin role=lab title=Toss a Fault to Your Witcher: Applying Grey-box Coverage-Guided Mutational Fuzzing to Detect SQL and Command Injection Vulnerabilities.
PUBLISHED_KIND_XSS 25
PUBLISHED_KIND_CSRF 7
PUBLISHED_KIND_CLICKJACK 1
PUBLISHED_KIND_SOP 4
PUBLISHED_KIND_RESIDUE 62
PUBLISHED_WILD_XSS 14
 
==========================================================================
E. YEAR BUCKETS
==========================================================================
 
── E1. conjunction of 99 ──
Window      Papers  Share of 99
----------  ------  -----------
2010–2013   11      11.1%
2014–2017   13      13.1%
2018–2021   18      18.2%
2022–2024   37      37.4%
2025–2026*  20      20.2%
 
── E2. wild role ──
Window      Wild papers  Share of 30
----------  -----------  -----------
2010–2013   3            10.0%
2014–2017   5            16.7%
2018–2021   4            13.3%
2022–2024   12           40.0%
2025–2026*  6            20.0%
 
── E3. per-year wild, corpus denominator, 2025–2026 starred ──
Year | Corpus | Wild | Share | Provisional
2010 | 119 | 1 | 0.8% | no
2011 | 116 | 0 | 0.0% | no
2012 | 151 | 0 | 0.0% | no
2013 | 125 | 2 | 1.6% | no
2014 | 166 | 2 | 1.2% | no
2015 | 190 | 1 | 0.5% | no
2016 | 182 | 0 | 0.0% | no
2017 | 231 | 2 | 0.9% | no
2018 | 254 | 0 | 0.0% | no
2019 | 402 | 2 | 0.5% | no
2020 | 404 | 1 | 0.2% | no
2021 | 379 | 1 | 0.3% | no
2022 | 546 | 5 | 0.9% | no
2023 | 719 | 4 | 0.6% | no
2024 | 690 | 3 | 0.4% | no
2025 | 770 | 5 | 0.6% | yes
2026 | 415 | 1 | 0.2% | yes
 
==========================================================================
F. VENUE
==========================================================================
Venue    Of 99  Wild
-------  -----  ----
USENIX   31     10
IEEE-SP  18     6
CCS      16     7
NDSS     15     6
WWW      9      0
IMC      6      0
PETS     4      1
 
==========================================================================
G. INTERACTION DEPTH (landing-page undercount)
==========================================================================
 
── G1. all 99 ──
interactionDepth       Papers  Share of 99 conjunction
---------------------  ------  -----------------------
landing-page-only      8       8.1%
single-target-page     23      23.2%
landing-plus-subpages  14      14.1%
deep-crawl             31      31.3%
not-stated             19      19.2%
no-crawlConfig         4       4.0%
 
── G2. wild only ──
interactionDepth       Papers  Share of 30 wild
---------------------  ------  ----------------
landing-page-only      2       6.7%
single-target-page     5       16.7%
landing-plus-subpages  9       30.0%
deep-crawl             10      33.3%
not-stated             2       6.7%
no-crawlConfig         2       6.7%
PUBLISHED_WILD_LANDING 2
PUBLISHED_WILD_SINGLE 5
PUBLISHED_WILD_DEEP 10
PUBLISHED_WILD_DEPTH_NOT_STATED 2
PUBLISHED_WILD_SUBPAGES_STATED 10
shallow (landing-page-only + single-target-page) among wild: 7 / 30 = 23.3%
PUBLISHED_WILD_ARCHIVE_SURFACE 2
PUBLISHED_WILD_THEN_LIVE 28
  archive-surface CCS/2017/rewriting-history-changing-the-archived-web-from-the-present Rewriting History: Changing the Archived Web from the Present.
  archive-surface USENIX/2017/how-the-web-tangled-itself-uncovering-the-history-of-client-side-web-in-security How the Web Tangled Itself: Uncovering the History of Client-Side Web (In)Security
 
==========================================================================
H. CLASSIFICATION METHOD (vuln tuples, sentinels skipped)
==========================================================================
 
── H1. conjunction of 99 ──
method               Papers of 99  Share
-------------------  ------------  -----
heuristic-rules      33            33.3%
manual-labelling     28            28.3%
dynamic-analysis     19            19.2%
static-analysis      13            13.1%
curated-database     12            12.1%
supervised-ml        6             6.1%
regex-or-signature   6             6.1%
third-party-service  3             3.0%
graph-analysis       2             2.0%
other                1             1.0%
blocklist            1             1.0%
 
── H2. wild only (the wild-role method distribution; includes the 2 archive-surface papers) ──
method              Papers of 30 wild  Share
------------------  -----------------  -----
dynamic-analysis    14                 46.7%
heuristic-rules     12                 40.0%
manual-labelling    8                  26.7%
static-analysis     4                  13.3%
graph-analysis      2                  6.7%
supervised-ml       1                  3.3%
regex-or-signature  1                  3.3%
 
==========================================================================
I. FULL-TEXT PROBES (upper bounds; missing .cols counted as negatives)
==========================================================================
xss: corpus 409 (miss 4) | web+crawled 158 (miss 0) | wild 25 of 30
csrf: corpus 182 (miss 4) | web+crawled 81 (miss 0) | wild 14 of 30
clickjack: corpus 95 (miss 4) | web+crawled 39 (miss 0) | wild 7 of 30
sop: corpus 195 (miss 4) | web+crawled 101 (miss 0) | wild 12 of 30
domxss: corpus 77 (miss 4) | web+crawled 50 (miss 0) | wild 14 of 30
stored-xss: corpus 50 (miss 4) | web+crawled 25 (miss 0) | wild 5 of 30
reflected-xss: corpus 41 (miss 4) | web+crawled 22 (miss 0) | wild 5 of 30
scanner-classic: corpus 24 (miss 4) | web+crawled 12 (miss 0) | wild 0 of 30
scanner-plus: corpus 81 (miss 4) | web+crawled 37 (miss 0) | wild 7 of 30
false-positive: corpus 2574 (miss 4) | web+crawled 493 (miss 0) | wild 25 of 30
manually-confirm: corpus 1279 (miss 4) | web+crawled 319 (miss 0) | wild 20 of 30
PUBLISHED_SCANNER_FT 24
PUBLISHED_CLICKJACK_WEBCRAWL_FT 39
clickjacking full-text among web+crawled is an UPPER BOUND, not a population.
 
==========================================================================
J. POSTERS in the 99
==========================================================================
posters in the 99: 1
  CCS/2024/poster-security-of-login-interfaces-in-modern-organizations role=wild title=Poster: Security of Login Interfaces in Modern Organizations.
PUBLISHED_POSTERS 1
 
==========================================================================
K. ROLE QUOTE CHECK against paper.cols.txt
==========================================================================
ROLE quotes: exact 68 / partial(>=60% 5-word windows) 31 / below 0 / of 99
PUBLISHED_QUOTE_EXACT 68
PUBLISHED_QUOTE_PARTIAL 31
PUBLISHED_QUOTE_BELOW 0
 
==========================================================================
L. PER-PAPER DETECTION FIGURES (wild XSS and named cover papers)
==========================================================================
--- CCS/2013/25-million-flows-later-large-scale-detection-of-dom-based-xss role=wild in99=true
  phen="DOM-based XSS vulnerabilities"
  metric="unique vulnerabilities and affected domains"
  prev="6,167 unique vulnerabilities on 480 domains; 9.6% of the top 5000 sites"
  phen="Potentially unsafe data flows"
  metric="number of captured flows"
  prev="24,474,306 data flows"
  phen="Exploitable DOM-based XSS flows"
  metric="validated exploit success rate"
  prev="69,987 of 181,238 generated payloads successfully executed injected JavaScript"
  phen="Chromium XSS Filter bypasses"
  metric="susceptible domains"
  prev="300 of 701 domains remained susceptible"
--- NDSS/2019/dont-trust-the-locals-investigating-the-prevalence-of-persistent-client-side-cross-site-scripting-in-the-wild role=wild in99=true
  phen="persistent client-side XSS"
  metric="share of Alexa Top 5,000 domains"
  prev="more than 8% exhibit exploitable flows from client-side storage to a dangerous sink"
  phen="persistent client-side XSS"
  metric="share among domains using persisted data in sinks"
  prev="21% of sites are vulnerable"
  phen="persistent client-side XSS"
  metric="number of exploitable domains"
  prev="418 of 1,324 domains"
  phen="Network Attacker exploitability"
  metric="share of theoretically exploitable domains"
  prev="293 of 418 domains"
  phen="Web Attacker exploitability"
  metric="number of exploitable domains"
  prev="65 of 418 domains"
  phen="reflected client-side XSS"
  metric="number of susceptible domains"
  prev="468 of the top 5,000 domains"
--- IEEE-SP/2010/state-of-the-art-automated-black-box-web-application-vulnerability-testing role=lab in99=true
  phen="scanner vulnerability detection"
  metric="detection rate"
  prev="Reflected XSS exceeded 60% average detection; second-order SQL injection was detected by no scanner."
  phen="link traversal coverage"
  metric="percentage of successful links crawled"
  prev="Coverage was low for Java applets, SilverLight, and Flash."
  phen="scanner network footprint"
  metric="network bytes sent and received"
  prev="Traffic ranged from 80 MB to nearly 1 GB."
  phen="scanner execution time"
  metric="elapsed scanning time"
  prev="Execution time ranged from 66 to 473 minutes."
  phen="false positives"
  metric="false-positive count"
  prev="Two scanners reported false positives for the benign script region."
  phen="stored XSS detection"
  metric="detection rate"
  prev="Stored XSS detection was 15%."
  phen="second-order SQL injection detection"
  metric="detection rate"
  prev="No scanner detected even one second-order SQL injection vulnerability."
--- IEEE-SP/2021/black-widow-blackbox-data-driven-web-scanning role=lab in99=true
  phen="server-side code coverage"
  metric="number of unique executed lines"
  prev="Black Widow had the highest coverage on 9 out of 10 applications"
  phen="reflected and stored XSS"
  metric="unique correctly executing XSS injections"
  prev="25 unique vulnerabilities, including 6 previously unknown"
  phen="false-positive XSS reports"
  metric="false-positive count"
  prev="No false positives reported by Black Widow on tested applications"
--- IEEE-SP/2022/the-state-of-the-samesite-studying-the-usage-effectiveness-and-adequacy-of-sames role=wild in99=true
  phen="SameSite cookie adoption"
  metric="share of sites using each policy"
  prev="18.94% of sites adopted one of the three valid policies by March 2021"
  phen="Cross-site functionality breakage"
  metric="share of sampled requests broken"
  prev="19% of affected cross-site requests were broken; 77.5% involved advertising networks"
  phen="State-changing GET CSRF"
  metric="vulnerable sampled requests"
  prev="7 of 264 GET requests were forgeable, affecting four websites"
  phen="Window-properties and postMessage XS-Leaks"
  metric="vulnerable URLs and websites"
  prev="1,302 vulnerable URLs across 40 distinct websites"
  phen="SameSite bypass via POST-to-GET"
  metric="vulnerable sampled POST requests"
  prev="9 of 602 requests were forgeable, affecting six websites"
  phen="SSO redirect bypass"
  metric="affected websites"
  prev="Six IdPs enabled bypass across 4,935 sites, over 49% of Alexa top 10K"
  phen="User-agent SameSite inconsistency"
  metric="vulnerable websites"
  prev="9,951 websites allowed a policy downgrade in April 2021"
  phen="Browser and framework divergence"
  metric="distinct browser behaviors and framework share"
  prev="Seven browser behaviors; 24% of frameworks set None by default"
--- IEEE-SP/2025/403-forbidden-ethically-evaluating-broken-access-control-in-the-wild role=wild in99=true
  phen="AC-sensitive HTTP endpoints"
  metric="unique probing URL templates"
  prev="584 unique URL templates"
  phen="Improper access control"
  metric="share of tested sites with improper responses"
  prev="30 endpoints across 15 of 100 sites"
  phen="Broken access-control vulnerabilities"
  metric="confirmed vulnerabilities"
  prev="19 vulnerabilities across 7 sites"
--- NDSS/2013/the-postman-always-rings-twice-attacking-and-defending-postmessage-in-html5-webs role=wild in99=true
  phen="postMessage receiver usage"
  metric="share of visited hosts"
  prev="2,245 hosts (22% of the visited hosts)"
  phen="missing origin checks"
  metric="distinct receivers and hosts"
  prev="65 receivers used by 1,585 hosts"
  phen="incorrect origin checks"
  metric="distinct receivers and hosts"
  prev="14 receivers used by 261 hosts"
  phen="missing or incorrect origin checks"
  metric="affected hosts"
  prev="1,712 hosts"
  phen="exploitable postMessage vulnerabilities"
  metric="distinct receivers and hosts"
  prev="13 receivers compromising 84 hosts"
  phen="incorrect-origin-check bypass domains"
  metric="existing domains passing checks"
  prev="Examples include 2,291 and 2,276 existing domains"
--- USENIX/2017/how-the-web-tangled-itself-uncovering-the-history-of-client-side-web-in-security role=wild in99=true
  phen="client-side XSS"
  metric="share of domains with a verified exploitable flaw"
  prev="about 8% of the 2016 sites exploitable"
  phen="insecure postMessage receivers"
  metric="share of receiving domains without origin checks"
  prev="48.0% in 2016"
  phen="wildcard postMessage targets"
  metric="share of domains sending wildcard-target messages"
  prev="50.3% in 2016"
  phen="dangerous Flash crossdomain policies"
  metric="share of domains"
  prev="about 7% had dangerous wildcards in 2008; at most 3% appeared vulnerable"
  phen="outdated vulnerable JavaScript libraries"
  metric="share of domains using vulnerable versions"
  prev="85% of sites running YUI used a vulnerable version in 2016"
  phen="security-header adoption"
  metric="share of domains deploying CSP"
  prev="less than 10% in 2016"
  phen="remote script inclusions"
  metric="average distinct remote origins per domain"
  prev="almost 12 distinct remote origins per domain in 2016"
  phen="JSONP usage"
  metric="share of sites using JSONP"
  prev="at most about 17% of all sites"
--- NDSS/2022/auto-draft-207 role=wild in99=true
  phen="client-side prototype pollution"
  metric="number of exploitable vulnerabilities and domains"
  prev="2,738 domains with 2,917 exploitable prototype pollution vulnerabilities among one million websites"
  phen="XSS consequences"
  metric="number of vulnerabilities"
  prev="48 vulnerabilities"
  phen="cookie manipulation"
  metric="number of vulnerabilities"
  prev="736 vulnerabilities"
  phen="URL manipulation"
  metric="number of vulnerabilities"
  prev="830 vulnerabilities"
  phen="real-world prototype-pollution defenses"
  metric="joint flows and domains by defense type"
  prev="Object sanitization: 22,235 joint flows across 1,489 domains"
--- IEEE-SP/2023/its-dom-clobbering-time-attack-techniques-prevalence-and-defenses role=wild in99=true
  phen="DOM-clobbering markups"
  metric="number of distinct markups working in at least one browser"
  prev="31,432 distinct DOM Clobbering markups"
  phen="Browser-specific clobbering behavior"
  metric="behavioral groups"
  prev="10 distinct groups of browser behaviours"
  phen="Native API clobbering"
  metric="APIs clobbered in at least one browser"
  prev="347 DOM APIs, including 114 window APIs"
  phen="DOM-clobbering vulnerabilities"
  metric="confirmed data flows and affected websites"
  prev="9,467 flows across 491 sites"
  phen="Website vulnerability prevalence"
  metric="share of tested websites"
  prev="9.8% (491 of 5,000)"
  phen="Exploitability"
  metric="websites with proof-of-concept exploits"
  prev="44 websites"
  phen="HTML sanitizer robustness"
  metric="sanitizers vulnerable by default"
  prev="16 of 29 sanitizers"
  phen="CSP mitigation coverage"
  metric="XSS vulnerabilities mitigated by CSP"
  prev="1,385 of 3,677 XSS vulnerabilities (37.7%)"
--- IEEE-SP/2024/the-great-request-robbery-an-empirical-study-of-client-side-request-hijacking-vu role=wild in99=true
  phen="client-side request hijacking data flows"
  metric="verified vulnerable data flows, affected webpages and sites"
  prev="202,834 verified flows affecting 17,805 webpages and 961 sites; 9.6% of the top 10K sites"
  phen="request-sending API usage"
  metric="API calls, webpages and domains"
  prev="Approximately 7.9M API calls across 1,032,795 webpages and 9,901 domains"
  phen="request hijacking exploitability"
  metric="proof-of-concept exploits and affected websites"
  prev="67 proof-of-concept exploits across 49 websites"
  phen="defense adoption and coverage"
  metric="pages and sites adopting defenses; mitigated flow share"
  prev="CSP mitigated information leakage and XSS in 58.7% of flows; 7.6% of webpages deployed the relevant CSP directive"
--- USENIX/2024/dancer-in-the-dark-synthesizing-and-evaluating-polyglots-for-blind-cross-site-sc role=wild in99=true
  phen="XSS polyglot coverage"
  metric="number of injection contexts solved"
  prev="Seven polyglots covered all 111 selected GFR test cases."
  phen="client-side XSS"
  metric="validated vulnerabilities"
  prev="147 vulnerabilities triggered by synthesized polyglots versus 145 by precise payload generation."
  phen="blind XSS"
  metric="vulnerabilities and affected websites"
  prev="20 vulnerabilities on 18 websites."
  phen="blind XSS"
  metric="share of backends by submission type"
  prev="Headers triggered 10 vulnerabilities, URLs 9, and forms 1."
  phen="crawler page failures"
  metric="share of visited pages failing"
  prev="Approximately 7.4% of 1,676,812 visited pages failed to load."
--- CCS/2025/in-the-dom-we-trust-exploring-the-hidden-dangers-of-reading-from-the-dom-on-the role=wild in99=true
  phen="DOM-to-sink data flows"
  metric="verified flow count and affected sites"
  prev="357,982 verified gadgets across 14,345 webpages and 2,259 sites"
  phen="Markup injection vulnerabilities"
  metric="verified dataflow count"
  prev="4,722 verified dataflows across 34,223 webpages"
  phen="DOM-gadget exploitability"
  metric="end-to-end verified flows and sites"
  prev="657 flows across 37 sites"
  phen="Missing sanitization or validation"
  metric="share of static flows without relevant patterns"
  prev="10.38% contained no sanitization or validation patterns"
  phen="DOM selector complexity"
  metric="mean and median complexity"
  prev="average 1.80 and median 2"
  phen="Element-order exploitation requirement"
  metric="share requiring reordering techniques"
  prev="34% of 253K combinations had injected markup after the selected element"
  phen="Detection false negatives"
  metric="false negative rate"
  prev="38.5%"
--- NDSS/2025/do-not-follow-the-white-rabbit-challenging-the-myth-of-harmless-open-redirection role=wild in99=true
  phen="open redirect vulnerabilities"
  metric="number of confirmed vulnerabilities and affected websites"
  prev="20,898 confirmed open redirections across 623 websites"
  phen="open redirect prevalence"
  metric="share of top-10K websites"
  prev="approximately 8.7% of the top 10K websites"
  phen="archived open redirects"
  metric="confirmed vulnerabilities and affected websites"
  prev="375 vulnerabilities across 326 websites"
  phen="DOM-based XSS escalation"
  metric="share of vulnerabilities and affected sites"
  prev="about 9% of vulnerabilities across 33.2% of affected sites"
  phen="client-side CSRF escalation"
  metric="number and share of open redirects"
  prev="42 vulnerabilities, over 2.4% of open redirects"
  phen="information leakage escalation"
  metric="number and share of open redirects"
  prev="3 vulnerabilities, about 0.2% of open redirects"
  phen="redirect mitigations"
  metric="share of audited sites"
  prev="six mitigation types; redirect notices used by 54.4%"
  phen="indicator false negatives"
  metric="false-negative rate"
  prev="76% for indicators compared with static analysis"
--- NDSS/2026/dom-xss-detection-via-webpage-interaction-fuzzing-and-url-component-synthesis role=wild in99=true
  phen="DOM-XSS vulnerabilities"
  metric="unique confirmed vulnerable flows and pages"
  prev="114 unique DOM-XSS vulnerable flows in 146 pages"
  phen="Interaction-triggered DOM-XSS"
  metric="confirmed-flow increase over passive analysis"
  prev="15% more confirmed flows than Passive"
  phen="URL-parameter and fragment-triggered DOM-XSS"
  metric="new confirmed vulnerabilities"
  prev="20 new vulnerabilities"
  phen="Synthesized GET parameters"
  metric="overlap with ffuf/wfuzz wordlists"
  prev="95.6% of DSE-synthesized keys absent from those wordlists"
--- USENIX/2020/cached-and-confused-web-cache-deception-in-the-wild role=wild in99=true
  phen="web cache deception"
  metric="share of sites"
  prev="16 of 295 sites (5.4%)"
  phen="web cache deception variants"
  metric="share of sites"
  prev="25 of 340 sites; encoded techniques exploited 23 of 25 sites"
  phen="private-information leakage"
  metric="share of vulnerable sites"
  prev="14 of 16 vulnerable sites leaked PII"
  phen="security-token leakage"
  metric="share of vulnerable sites"
  prev="6 of 16 sites leaked CSRF tokens; 6 leaked session identifiers or API tokens"
  phen="unauthenticated exploitation"
  metric="share of vulnerable sites"
  prev="All remaining vulnerabilities were manually found exploitable without authentication"
  phen="cache expiration"
  metric="number of exploitable sites after delay"
  prev="16 sites after 1 hour, 10 after 6 hours, and 9 after 1 day"
  phen="CDN caching behavior"
  metric="default caching behavior and honored headers"
  prev="Akamai, Cloudflare, CloudFront, and Fastly showed distinct defaults"
--- CCS/2015/from-facepalm-to-brain-bender-exploring-client-side-cross-site-scripting role=wild in99=true
  phen="client-side XSS vulnerabilities"
  metric="count of exploitable flows"
  prev="1,273 actual vulnerabilities from 1,146 URLs"
  phen="vulnerability complexity"
  metric="low/medium/high complexity classification"
  prev="63.9% low, 20.5% medium, and 15.6% high combined complexity"
  phen="non-linear data and control flows"
  metric="count of vulnerable flows"
  prev="98 non-linear data flows; 59 flows with both non-linear data and control flow"
  phen="third-party code involvement"
  metric="count of vulnerabilities"
  prev="273 exclusively third-party; 165 mixed self-hosted and third-party"
  phen="multiflows"
  metric="count of exploited Web pages"
  prev="344 multiflow vulnerabilities"
  phen="vulnerable sink types"
  metric="count of exploitable flows"
  prev="732 document.write, 495 innerHTML, and 46 eval or derivatives"
  phen="cross-browser exploitability"
  metric="URLs triggering payload"
  prev="109 URLs still triggered the payload in Firefox"
 
==========================================================================
M. STUDY TYPES among wild
==========================================================================
studyType                   Wild papers  Share of 30
--------------------------  -----------  -----------
automated-web-crawl         30           100.0%
manual-audit                23           76.7%
system-or-defence-proposal  23           76.7%
code-or-binary-analysis     16           53.3%
existing-dataset-analysis   12           40.0%
network-scan-or-probe       4            13.3%
user-study                  1            3.3%
interview-or-survey         1            3.3%
mobile-app-analysis         1            3.3%
 
==========================================================================
Y. ARITHMETIC
==========================================================================
880 used / 5859 = 15.0%
209 web / 880 = 23.8%
99 web+crawled / 880 = 11.3%
99 / 1120 crawled = 8.8%
30 wild / 99 = 30.3%
30 wild / 1120 crawled = 2.7%
27 lab / 99 = 27.3%
8 cve / 99 = 8.1%
34 offtopic / 99 = 34.3%
433 offline-multi / 880 = 49.2%
260 offline-only / 880 = 29.5%
kind xss 25/99 = 25.3%
wild xss 14/30 wild = 46.7%
clickjack kind in 99: 1
csrf kind in 99: 7
sop kind in 99: 4
 
==========================================================================
Z. NON-CORPUS FIGURES (primary sources; re-fetched by external_checks_web_vuln.sh)
==========================================================================
Chrome XSS Auditor removed: Chrome 78, 2019-10. Chromium bug 709804 / release blog.
Chrome SameSite Lax-by-default: Chrome 80, 2020-02. Chromium SameSite updates.
OWASP Top 10 2021: A03 Injection; XSS folded into Injection. owasp.org/Top10/2021/.
OWASP Top 10:2025: do not cite a 2017 XSS-as-A7 ranking as current.
OWASP ZAP is now ZAP by Checkmarx (rebrand 2024-09-24); zaproxy.org.
Nuclei is projectdiscovery/nuclei; templates in nuclei-templates. Too thin in this corpus to carry a page.
count_vuln_sites.py --demo: 5 rows, 5 URLs, 3 hosts, 3 confirmed rows, 3 confirmed URLs, 2 confirmed hosts
Lekies et al. CCS 2013: 6,167 unique vulnerabilities on 480 domains; 9.6% of Alexa top 5000; 24,474,306 flows; 69,987 of 181,238 payloads executed
Steffens et al. NDSS 2019: Alexa Top 5,000; more than 8% unfiltered storage-to-sink flows; 21% of sites that use stored data; 418 of 1,324 exploitable; 468 of 5,000 reflected Client-Side XSS
Bau et al. IEEE S&P 2010: reflected XSS detection over 60%; stored XSS 15%; second-order SQLi detected by no scanner
Chehade, Hantke and Stock, IEEE S&P 2025: 100 sites; 30 improper endpoints on 15 sites; 19 confirmed vulns across 7 sites
programming:interaction: 155 of 417 (37.2%) landing-page-only on the site-depth axis; 857 crawled-web papers. Quoted, not re-derived.
Foxhound lineage: programming:crawler:foxhound. Taint tracking is the current client-side XSS instrument in this corpus.
Hantke et al. Red Lines: practices:ethics. Notification: practices:notifying_websites.
Corpus posters: 138 of 5,859 (literature:corpus). One poster is in this 99.
ZAP rebrand 2024-09-24: ZAP by Checkmarx, Apache v2; zaproxy.org/blog/2024-09-24-zap-has-joined-forces-with-checkmarx/
Chrome SameSite Lax-by-default: Chrome 80, 2020-02
OWASP Top 10:2025 A01 Broken Access Control; A05 Injection includes XSS as CWE-79; XSS is not a standalone Top-10 letter
OWASP Top 10 2017 had XSS as A7; 2021 folded XSS into A03 Injection; 2025 is A05
review tokens: FIG-01 FIG-02 CIT-01 CUR-01 GENERIC_PLACEHOLDER_FORBIDDEN
provenance/security/web_vulnerabilities.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki