User Tools

Site Tools


provenance:writing:literature_review

This is an old revision of the document!


Provenance: Writing:Literature review

Back to Literature review. Corpus-wide selection and extraction notes are on corpus. Citations use the shared bibliography; this page adds no keys of its own. No ~~DISCUSSION~~ — comments belong on the content page.

Run record

  • Run date: 2026-08-27 (UTC).
  • Drain item: writing:literature_review (new), claimed as cursor-drain-litrev (item 178, run 55). Executed as Cursor, not via claude -p / drain-sandbox.sh.
  • Corpus at run time: 5,859 extracted papers, 2010–2026, CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P. Read-only inputs under /workspace/publications_dataset/data/.
  • Read first: data/extract/OVERVIEW.md, the wiki-measuretheweb spec (audience, currency, 5,859-paper corpus, provenance page, start link, bibtex), corpus, fingerprinting, study_preregistration, requests, crawler, live start (already promised Literature review).
  • Target pages had no revision (core.getPageInfo “does not exist”; ?do=export_raw on a missing page returns an HTML error document — that is not an existence test by byte count). dw.mjs pages does not list provenance:. This is a creation, not an extension.
  • Overlap judgement: create the promised child. Do not broaden corpus (funnel) or conferences (where to send the paper). The Writing namespace page itself does not exist; not this item.
  • No write to the publication mount. Wiki saves through scripts/dw.mjs (JSON-RPC). Live start and literature:bibliography were re-exported immediately before append, not taken from a stale local copy.

Why this page

The item asked for the SoK papers in the corpus plus this project as a worked example of structured extraction vs keyword search, and why a keyword-derived denominator is biased along the axis being measured. That is a writing/methods page, not a survey of every SoK. A reasonable person might have written a catalogue of the 43; that would miss the trap the item named.

Population and queries

All counts are papers unless labelled tuples. Sentinels are not answers.

Membership is mechanical: title starts SoK / SOK: or slug starts sok-. On extract/run1 those two tests agree: 43 / 43, intersection 43, residue 0. The report throws if they diverge.

Query Denominator Result
title-start SoK / SOK: 5,859 extracted 43 (0.7%)
slug-start sok- 5,859 43 (0.7%)
union (page population) 5,859 43
title-only, slug-only 43 0, 0
IEEE S&P among extracted SoKs 43 30 (69.8%)
CCS / IMC / WWW extracted SoKs 43 0 / 0 / 0
same three venues in the 16,864-record index 16,864 0 / 0 / 0
TOPIC = web_measurement 43 4 (9.3%)
TOPIC = web_adjacent 43 13 (30.2%)
TOPIC = not_web 43 26 (60.5%)
SEARCH = venue_then_filter 43 7 (16.3%)
SEARCH = database_keyword 43 3 (7.0%)
SEARCH = scholar_snowball 43 8 (18.6%)
SEARCH = not_stated 43 6 (14.0%)
SEARCH = not_a_census 43 19 (44.2%)
claims a literature sample 43 24 (55.8%)
…and states a protocol 24 18 (75.0%)
venue-then-filter among stated protocols 18 7 (38.9%)
index SoKs (title-start or slug-start) 16,864 index records 167
…with an abstract 167 163 (lost 4)
…labelled 163 163 (lost 0)
…selected by the over-inclusive –any rule 163 47 (lost 116)
…extracted 47 43 (lost 4)
PETS index SoKs vs extracted 43 index 2 extracted (41 not extracted)
crawled; crawlConfig.headless stated (not sentinel) 1,120 crawled 140 (12.5%)
crawled; full-text /\bheadless\b/i 1,120 137 (12.2%)
crawled; schema never stated headless 1,120 980 (87.5%)
detection.phenomenon names a fingerprint 5,859 280; browser-device 83 (29.6%); traffic-analysis 105 (37.5%)
PREREG_RE full-text hits 5,859 (4 missing .cols) 62; study sense 15; other-sense 47 (75.8%)
title contains “survey” 5,859 11; none are SoKs; 7/11 tagged user-study or interview-or-survey
full-text /systematization of knowledge/i 5,859 (4 missing .cols) 26; 15 not in the title/slug union; 32/43 extracted SoKs lack the phrase
crawled 2010–2022 (Krawlers-year recall) 1,120 crawled 687
those matching tranco or alexa 687 375 (54.6%)
those matching neither 687 312 (45.4%)
generous three-keyword union miss 687 148 (21.5%)not the published lead; “top AND site” as two independent words is looser than Krawlers' query
generous three-keyword hits 2010–2022 that are not crawled 1,493 hits 954 (63.9%)
keyword_denominator.py '\btranco\b' on all crawled years 1,120 194 (17.3%) match, 926 (82.7%) miss
keyword_denominator.py '\bheadless\b' 1,120 137 (12.2%) match, 983 (87.8%) miss

The 21.5% three-keyword miss is in the report because it was computed. It is not on the content page as a headline, because treating “combination of top and site” as two independent word hits is a different experiment from theirs.

Folding and residue

  • Membership is the union of two mechanical tests. Residue of title-vs-slug: 0. The report throws if they diverge, and throws if TOPIC or SEARCH misses a key of the 43 or contains a key that is not in the 43.
  • TOPIC and SEARCH are hand maps in scripts/litrev_fold.mjs, one primary label per paper, deciding sentence inline. Residue 0 by construction (the throw). A different reader might move Birrell (privacy-regulation impact) or Mathews (website-fingerprinting defenses) across the web / adjacent line; both deciding sentences are in the fold file. Mathews stays adjacent on purpose — it is the fingerprinting homograph.
  • The four web-measurement SoKs: Stafeev & Pellegrino (krawlers), Birrell et al. (privacy regulations), Blessing et al. (web auth), Alam et al. (PHILTER). Alam's key alam2026_philter was already in the live bibliography; it was not added again.

Quotes checked against paper.cols.txt

Needles the report requires (flatten whitespace). All OK on the run that the page was written from:

  • Krawlers: 7,840, 1,057, 654, 403, 32.3%, tranco, alexa, and the combination of top and site
  • Warford: 6,534 papers, 127 potentially relevant papers, reviewed 95 papers
  • Blessing: 245 papers, top-300 (the sentence top 300 of the Alexa is column-spliced in .cols; the page uses the hyphenated form that survives)
  • Usman: more effective than keyword searching in a database; the six-venue list including IEEE Security & Privacy, and PETS
  • Alam: In total, 55 papers
  • Birrell: ten selected venues

The Usman quote on the content page is shortened with an ellipsis. The two fragments it uses are both in .cols.

External and industry verification

Source Load-bearing fact Verification 2026-08-27 Decision
IEEE S&P 2026 CFP https://sp2026.ieee-security.org/cfpapers.html SoK: prefix; surveys without insights may be rejected fetched HTML Used
https://oaklandsok.github.io/ SoK at S&P since 2010; EuroS&P 2017; PETS 2019; USENIX 2024; NDSS 2026; not CCS/IMC/WWW fetched HTML Used. First SoK-titled paper in this index is 2013; the page says so rather than pretending to have 2010–2012.
arXiv:2506.14057 abs still v1 preprint (16 June 2025); living repo named fetched 2026-08-27 Used as orientation, dated as preprint, not as a denominator
github.com/privacysandstorm/sok-advances-open-problems-web-tracking living version of Vekaria named on the abs page Named, not counted
IEEE-2026.json Rieder slug in index, no abstract, never screened, not in 5,859 report prints the four no-abstract slugs Used as the mechanical-miss example
USENIX Security 2024 landing /presentation/stafeev authors Aleksei Stafeev, Giancarlo Pellegrino fetch_authors.py + landing Used in BibTeX
PETS 2025 landing popets-2025-0113 authors Jenny Blessing, Daniel Hugenroth, Ross Anderson, Alastair Beresford; DOI 10.56553/popets-2025-0113 fetch_authors.py failed on parse_popets; authors filled by hand from citation_author meta tags into out/authors.json Used. The failed parser is recorded so the next run does not treat the cache as automatic.
literature:corpus funnel 16,864 → 15,800 → 6,103 → 5,873 → 5,859 quoted from that page / report_corpus.mjs, not re-derived this sitting Used, labelled as quoted
privacy:requests S1 ∪ S2 = 254; both = 107 quoted from that page's report Used, labelled as quoted

Rejected:

  • Leading with the 21.5% three-keyword miss as “what Krawlers dropped” — the third keyword was not reproduced as they wrote it.
  • Counting title-“survey” papers as literature surveys — 7/11 are user/interview studies; Surveylance is survey scams.
  • Using “SoK” as a quality filter or as a web-measurement name — 30/43 are IEEE S&P; 26/43 are not this field.
  • Treating Vekaria “200+” as a reproducible denominator from this extraction — wrong venue (arXiv).
  • Adding a duplicate alam2026_philter.
  • SEO “how to write a literature review” listicles. None were consulted.
  • PRISMA / Scholar tutorials. Textbook; out of scope.
  • Re-deriving the 16,864-index funnel in this sitting — that is corpus's job.

What the corpus and sources do not establish

  • How many of the 116 screened-out SoKs a web-measurement PhD student would still want. Screening did its job on PETS crypto/DP; a handful of edge cases were not re-read.
  • Whether IEEE S&P 2010–2012 SoKs exist under a different title pattern in this index. Zero title-start hits; oaklandsok dates the track to 2010.
  • A venue version of Vekaria et al. As of 2026-08-27, arXiv v1 only.
  • Rieder et al. in the 5,859. No abstract → never screened.
  • Whether CCS / IMC / TheWebConf will add an SoK track.
  • A current, independent accuracy comparison of keyword search vs venue-complete on a question other than the ones this page measures.

Judgement calls

  • Create the promised page rather than a catalogue of 43 SoKs or a Scholar tutorial.
  • Hand-fold topic and search rather than regex-fold: the questions are not spelling variants. Residue 0 because the report throws, not because a regex covered everything.
  • Put Mathews in adjacent, not web_measurement.
  • Quote Krawlers' own 7,840 / 1,057 / 654 / 403 / 32.3% from their PDF (needles in .cols), and reproduce only the named-list half (tranco / alexa) as a recall experiment on this extraction's crawled 2010–2022.
  • Demo script scans crawled papers only and crashes if a crawled .cols is missing. An earlier draft scanned all papers and crashed on a 2010 USENIX paper with no full text — that paper is not crawled, so it is not in the denominator.
  • PETS authors for Blessing filled by hand after fetch_authors.py failed. Recorded here; not silently patched at the read site.
  • Start already linked the page. A one-line blurb was added on live start so the outline matches the Conferences line next to it.
  • check_attributions.mjs matches table-row “Name et al., VENUE YEAR [key]” only. This page attributes in prose, so that guard reports 0 checked and exits 1. That is not a pass and was not treated as one.

Embedded script

pages/keyword_denominator.py is byte-identical to the <file python keyword_denominator.py> block on the content page (diffed before publish). Quoted output is the \bheadless\b run, 2026-08-27, against extract/run1: 137 / 983 of 1,120. The \btranco\b run is 194 / 926 and is in the report; the page leads with headless because that is the reporting-rate trap.

Report script output (unedited)

Command: node scripts/report_literature_review.mjs. Output of the run the content page was written from:

corpus: 5859 papers, 7 venues, 2010–2026
crawled: 1120  empirical: 5118

==========================================================================
A. SOK MEMBERSHIP (title-start vs slug-start)
==========================================================================
Test                     Papers  Share of 5859
-----------------------  ------  -------------
title starts SoK / SOK:  43      0.7%
slug starts sok-         43      0.7%
union (the population)   43      0.7%
title-only (not slug)    0       0.0%
slug-only (not title)    0       0.0%
title contains word SoK anywhere: 43; of which outside the union: 0

==========================================================================
B. SHAPE OF THE 43 EXTRACTED SOKS
==========================================================================
Venue    Extracted SoKs  Extracted papers  SoKs / venue
-------  --------------  ----------------  ------------
CCS      0               990               0.0%
IEEE-SP  30              767               3.9%
IMC      0               638               0.0%
NDSS     4               701               0.6%
PETS     2               510               0.4%
USENIX   7               1410              0.5%
WWW      0               843               0.0%
venues with zero extracted SoKs: CCS, IMC, WWW

── by year (2025–2026 starred as provisional) ──
Year   Extracted SoKs  Extracted papers that year
-----  --------------  --------------------------
2010   0               119
2011   0               116
2012   0               151
2013   2               125
2014   0               166
2015   1               190
2016   2               182
2017   2               231
2018   1               254
2019   1               402
2020   2               404
2021   3               379
2022   4               546
2023   4               719
2024   8               690
2025*  7               770
2026*  6               415

── IEEE-SP share of extracted SoKs ──
IEEE-SP SoKs: 30 of 43 = 69.8%

── platforms (a paper can have several; shares of 43) ──
platform              SoKs  Share of 43
--------------------  ----  -----------
offline               26    60.5%
other-online-service  11    25.6%
web                   6     14.0%
not-applicable        5     11.6%
mobile                3     7.0%
iot                   1     2.3%
platforms includes web: 6 of 43 = 14.0%
  (platform=web is not the topic fold — printers and HPC counters sit here)

==========================================================================
C. TOPIC HAND FOLD (43 papers, 0 residue by construction)
==========================================================================
Topic                                                                                                      Papers  Share of 43
---------------------------------------------------------------------------------------------------------  ------  -----------
Web measurement (crawls, live-site audits, phishing-site detectors, privacy-regulation impact on the web)  4       9.3%
Adjacent online privacy / abuse / traffic analysis                                                         13      30.2%
Not this field (TEE, binaries, fuzzing, crypto, hardware, DeFi, …)                                         26      60.5%

── web_measurement ──
  2024 IEEE-SP  SoK: Technical Implementation and Human Impact of Internet Privacy Regulations.
  2024 USENIX   SoK: State of the Krawlers – Evaluating the Effectiveness of Crawling Algorithms for Web Security Measurements
  2025 PETS     SoK: Web Authentication and Recovery in the Age of End-to-End Encryption
  2026 USENIX   SoK: PHILTER: Uncovering Security and Functional Gaps in AI-based Phishing Website Detection Literature via an LLM-based Reasoning Framework

── web_adjacent ──
  2016 IEEE-SP  SoK: Towards Grounding Censorship Circumvention in Empiricism.
  2020 PETS     SoK: Anatomy of Data Breaches
  2021 IEEE-SP  SoK: Hate, Harassment, and the Changing Landscape of Online Abuse.
  2022 IEEE-SP  SoK: A Framework for Unifying At-Risk User Research.
  2022 IEEE-SP  SoK: Social Cybersecurity.
  2022 IEEE-SP  SoK: The Dual Nature of Technology in Sexual Abuse.
  2023 IEEE-SP  SoK: A Critical Evaluation of Efficient Website Fingerprinting Defenses.
  2024 IEEE-SP  SoK: Safer Digital-Safety Research Involving At-Risk Users.
  2024 USENIX   SoK (or SoLK?): On the Quantitative Study of Sociodemographic Factors and Computer Security Behaviors
  2025 IEEE-SP  SoK: A Privacy Framework for Security Research Using Social Media Data.
  2025 IEEE-SP  SoK: A Framework and Guide for Human-Centered Threat Modeling in Security and Privacy Research.
  2025 IEEE-SP  SoK: Self-Generated Nudes over Private Chats: How can Technology Contribute to a Safer Sexting?
  2025 IEEE-SP  SoK: Decoding the Enigma of Encrypted Network Traffic Classifiers.

── not_web ──
  2013 IEEE-SP  SoK: Secure Data Deletion.
  2013 IEEE-SP  SoK: P2PWNED - Modeling and Evaluating the Resilience of Peer-to-Peer Botnets.
  2015 IEEE-SP  SoK: Deep Packer Inspection: A Longitudinal Study of the Complexity of Run-Time Packers.
  2016 IEEE-SP  SOK: (State of) The Art of War: Offensive Techniques in Binary Analysis.
  2017 IEEE-SP  SoK: Cryptographically Protected Database Search.
  2017 IEEE-SP  SoK: Exploiting Network Printers.
  2018 IEEE-SP  SoK: Keylogging Side Channels.
  2019 IEEE-SP  SoK: The Challenges, Pitfalls, and Perils of Using Hardware Performance Counters for Security.
  2020 IEEE-SP  SoK: Understanding the Prevailing Security Vulnerabilities in TrustZone-assisted TEE Systems.
  2021 IEEE-SP  SoK: All You Ever Wanted to Know About x86/x64 Binary Disassembly But Were Afraid to Ask.
  2021 IEEE-SP  SoK: Quantifying Cyber Risk.
  2022 IEEE-SP  SoK: How Robust is Image Classification Deep Neural Network Watermarking?
  2023 IEEE-SP  SoK: Decentralized Finance (DeFi) Attacks.
  2023 IEEE-SP  SoK: History is a Vast Early Warning System: Auditing the Provenance of System Intrusions.
  2023 IEEE-SP  SoK: Taxonomy of Attacks on Open-Source Software Supply Chains.
  2024 IEEE-SP  SoK: Prudent Evaluation Practices for Fuzzing.
  2024 USENIX   SoK: All You Need to Know About On-Device ML Model Extraction - The Gap Between Research and Practice
  2024 USENIX   SoK: What Don't We Know? Understanding Security Vulnerabilities in SNARKs
  2024 IEEE-SP  SoK: SGX.Fail: How Stuff Gets eXposed.
  2025 USENIX   SoK: An Introspective Analysis of RPKI Security
  2025 IEEE-SP  SoK: Software Compartmentalization.
  2026 NDSS     SoK: Take a Deep Step into Linux Kernel Hardening Effectiveness from the Offensive-Defensive Perspective
  2026 NDSS     SoK: Understanding the Fundamentals and Implications of Sensor Out-of-band Vulnerabilities
  2026 NDSS     SoK: Cryptographic Authenticated Dictionaries
  2026 USENIX   SoK: Security of Cyber-physical Systems Under Intentional Electromagnetic Interference Attacks
  2026 NDSS     SoK: Analysis of Accelerator TEE Designs

web_measurement SoKs: 4 of 43 = 9.3%

==========================================================================
D. SEARCH-METHOD HAND FOLD (primary label, 43 papers)
==========================================================================
How the paper list was built                                   Papers  Share of 43
-------------------------------------------------------------  ------  -----------
Venue-complete (or venue-list) then filter                     7       16.3%
Digital-library / DBLP keyword query                           3       7.0%
Google Scholar and/or citation snowball                        8       18.6%
Not a paper-census (system, evaluation, vulnerability corpus)  19      44.2%
Literature sample claimed; search protocol not stated          6       14.0%
claims a literature sample: 24 of 43 = 55.8%
…and states a search protocol: 18 of 24 = 75.0%
venue-complete then filter: 7 of 43 = 16.3% (of the 18 with a stated protocol: 38.9%)

── web_measurement SoKs × search method ──
Paper                                                                                                                                    Year  Venue    Search
---------------------------------------------------------------------------------------------------------------------------------------  ----  -------  -----------------
Technical Implementation and Human Impact of Internet Privacy Regulations.                                                               2024  IEEE-SP  scholar_snowball
State of the Krawlers – Evaluating the Effectiveness of Crawling Algorithms for Web Security Measurements                                2024  USENIX   venue_then_filter
Web Authentication and Recovery in the Age of End-to-End Encryption                                                                      2025  PETS     venue_then_filter
PHILTER: Uncovering Security and Functional Gaps in AI-based Phishing Website Detection Literature via an LLM-based Reasoning Framework  2026  USENIX   database_keyword

==========================================================================
E. INDEX FUNNEL — SOKS IN THE SEVEN VENUES, NOT JUST THE 43
==========================================================================
index records: 16864
index SoKs (title-start or slug-start): 167
Venue    Index SoKs  Extracted SoKs
-------  ----------  --------------
IEEE-SP  84          30
PETS     43          2
USENIX   36          7
NDSS     4           4
index SoKs in CCS, IMC, WWW: 0, 0, 0
Stage                        SoKs  Lost here
---------------------------  ----  ---------
Index (title or slug)        167   —
…with an abstract            163   4
…labelled (screened)         163   0
…selected by the --any rule  47    116
…extracted                   43    4
no-abstract index SoKs (4):
  IEEE-SP/2026/sok-all-you-ever-wanted-to-know-about-bootloader-security-but-were-afraid-to-ask
  IEEE-SP/2026/sok-critical-evaluation-of-quantum-machine-learning-for-adversarial-robustness
  IEEE-SP/2026/sok-robustness-in-large-language-models-against-jailbreak-attacks
  IEEE-SP/2026/sok-after-decades-of-web-tracker-detection-whats-next
selected but not extracted (4):
  IEEE-SP/2024/sok-a-comprehensive-analysis-and-evaluation-of-docker-container-attack-and-defen
  USENIX/2026/sok-history-doesnt-repeat-itself-but-android-design-level-vulnerabilities-rhyme
  USENIX/2026/sok-darpas-ai-cyber-challenge-aixcc-competition-design-architectures-and-lessons
  USENIX/2026/sok-capability-operating-systems-is-the-future-finally-here
Rieder index record present: IEEE-SP/2026/sok-after-decades-of-web-tracker-detection-whats-next
  abstract present: false
  in labels: false
  in extractions: false
  fulltext dir exists: false

── PETS index SoKs vs extracted (most dropped SoKs are crypto / DP / sensing) ──
PETS index SoKs 43; extracted 2

==========================================================================
F. HOMOGRAPHS AND FULL-TEXT SOK MENTIONS
==========================================================================
title contains "survey": 11 of 5859
…of which are not SoKs: 11
…of those, studyTypes includes user-study or interview-or-survey: 7
title matches systematic (literature) review: 1

── non-SoK "survey" titles (these are mostly user surveys, not literature surveys) ──
  2016 CCS      user/interview=true  How I Learned to be Secure: a Census-Representative Survey of Security Advice Sources and Behavior.
  2018 IEEE-SP  user/interview=false  Surveylance: Automatically Detecting Online Survey Scams.
  2019 IEEE-SP  user/interview=true  How Well Do My Results Generalize? Comparing Security and Privacy Survey Results from MTurk, Web, and Telephone Samples.
  2020 WWW      user/interview=false  Attention Please: Your Attention Check Questions in Survey Studies Can Be Automatically Answered.
  2022 PETS     user/interview=false  A Global Survey of Android Dual-Use Applications Used in Intimate Partner Surveillance
  2022 WWW      user/interview=true  Beyond Bot Detection: Combating Fraudulent Online Survey Takers✱.
  2023 PETS     user/interview=true  Privacy Concerns and Acceptance Factors of OSINT for Cybersecurity: A Representative Survey
  2024 IEEE-SP  user/interview=true  Digital Security - A Question of Perspective A Large-Scale Telephone Survey with Four At-Risk User Groups.
  2024 USENIX   user/interview=true  Bridging Barriers: A Survey of Challenges and Priorities in the Censorship Circumvention Landscape
  2024 USENIX   user/interview=true  Engaging Company Developers in Security Research Studies: A Comprehensive Literature Review and Quantitative Survey
  2026 USENIX   user/interview=false  Inconsistent, Incomplete, and Insecure: A Survey of Account Security Interfaces

full-text /systematization of knowledge/i: 26 papers; missing .cols: 4
…not in the title/slug union (citation / passing mention): 15
extracted SoKs whose .cols does NOT contain the phrase: 32 of 43

==========================================================================
G. KRAWLERS KEYWORD CUT REPRODUCED ON THIS EXTRACTION
==========================================================================
Krawlers (USENIX 2024) quoted figures, checked against paper.cols.txt:
  OK    "7,840"
  OK    "1,057"
  OK    "654"
  OK    "403"
  OK    "32.3%"
  OK    "tranco, alexa, and the combination of top and site"
  OK    Warford "6,534 papers"
  OK    Warford "127 potentially relevant papers"
  OK    Warford "reviewed 95 papers"
  OK    Blessing "245 papers"
  OK    Blessing "top-300"
  OK    Usman "more effective than keyword searching in a database"
  OK    Usman "SOUPS, CHI, CSCW, USENIX Security, IEEE Security & Privacy, and PETS"
  OK    Alam "In total, 55 papers"
  OK    Birrell "ten selected venues"
extracted papers 2010–2022: 3265
crawled papers 2010–2022: 687 (denominator for recall)
Krawlers keyword on crawled 2010–2022                               Papers  Share of 687
------------------------------------------------------------------  ------  ------------
tranco                                                              56      8.2%
alexa                                                               351     51.1%
tranco OR alexa (the two named lists)                               375     54.6%
neither tranco nor alexa                                            312     45.4%
top AND site (both words; generous reading of their third keyword)  478     69.6%
any of the three (union under that generous reading)                539     78.5%
NONE of the three                                                   148     21.5%
missing .cols                                                       0       0.0%
stricter /top[\s-]*site/i on crawled 2010–2022: 162 (not used as the published recall; the paper said "combination of top and site")
keyword union on all extracted 2010–2022: 1493 hits; of which NOT crawled: 954 = 63.9% (Krawlers discarded 654 of 1,057 = 61.9% as not employing automated crawling; our extraction is already screened, so this is a lower bound on the false-positive rate)

── crawled 2010–2022 missed by all three keywords (first 20 of the miss list) ──
  CCS/2010/detecting-and-characterizing-social-spam-campaigns  Detecting and characterizing social spam campaigns.
  CCS/2010/fingerprinting-websites-using-remote-traffic-analysis  Fingerprinting websites using remote traffic analysis.
  IMC/2010/detecting-and-characterizing-social-spam-campaigns  Detecting and characterizing social spam campaigns.
  WWW/2010/analyzing-content-level-properties-of-the-web-adversphere  Analyzing content-level properties of the web adversphere.
  WWW/2010/diversifying-landmark-image-search-results-by-learning-interested-views-from-com  Diversifying landmark image search results by learning interested views from community pho
  WWW/2010/the-social-honeypot-project-protecting-online-communities-from-spammers  The social honeypot project: protecting online communities from spammers.
  WWW/2010/restler-crawling-restful-services  RESTler: crawling RESTful services.
  CCS/2011/automated-black-box-detection-of-side-channel-vulnerabilities-in-web-application  Automated black-box detection of side-channel vulnerabilities in web applications.
  CCS/2011/android-permissions-demystified  Android permissions demystified.
  IMC/2011/analyzing-facebook-privacy-settings-user-expectations-vs-reality  Analyzing facebook privacy settings: user expectations vs. reality.
  WWW/2011/arrow-generating-signatures-to-detect-drive-by-downloads  ARROW: GenerAting SignatuRes to Detect DRive-By DOWnloads.
  CCS/2012/detecting-money-stealing-apps-in-alternative-android-markets  Detecting money-stealing apps in alternative Android markets.
  IMC/2012/evolution-of-a-location-based-online-social-network-analysis-and-models  Evolution of a location-based online social network: analysis and models.
  IMC/2012/evolution-of-social-attribute-networks-measurements-modeling-and-implications-us  Evolution of social-attribute networks: measurements, modeling, and implications using goo
  NDSS/2012/you-are-what-you-like-information-leakage-through-users-interests  You are what you like! Information leakage through users’ Interests
  USENIX/2012/enemy-of-the-state-a-state-aware-black-box-web-vulnerability-scanner  Enemy of the State: A State-Aware Black-Box Web Vulnerability Scanner
  CCS/2013/delta-automatic-identification-of-unknown-web-based-infection-campaigns  Delta: automatic identification of unknown web-based infection campaigns.
  WWW/2012/understanding-and-combating-link-farming-in-the-twitter-social-network  Understanding and combating link farming in the twitter social network.
  WWW/2012/economics-of-bittorrent-communities  Economics of BitTorrent communities.
  WWW/2012/analyzing-spammers-social-networks-for-fun-and-profit-a-case-study-of-cyber-crim  Analyzing spammers' social networks for fun and profit: a case study of cyber criminal eco
  … 128 more
crawled 2010–2022 whose .cols matches /\bcrux\b/i: 7 (Krawlers dropped the keyword because CrUX rankings started in 2022)

==========================================================================
H. REPORTING-RATE TRAP — SEARCHING FOR THE THING UNDERCOUNTS THE SILENCE
==========================================================================
crawled papers: 1120
crawlConfig.headless stated (not a sentinel): 140 = 12.5%
full-text /\bheadless\b/i over crawled: 137 = 12.2%; missing .cols 0
word "headless" but schema sentinel/absent: 28 (mentions in related work, or a different sense)
schema-stated but word not in .cols: 31
trap: a Scholar search for "headless crawler" returns papers that used the word; it cannot estimate the 87.5% of crawls that never said (980 of 1120).

==========================================================================
I. SIBLING HOMOGRAPHS RE-DERIVED
==========================================================================
detection.phenomenon names a fingerprint: 280 of 5859
…browser-device family: 83 = 29.6% of the fingerprint papers
…traffic-analysis family: 105 = 37.5%
keyword precision for browser fingerprinting if you search this corpus for "fingerprint": 29.6%

── preregistration homograph (PREREG_RE over full text, sense from prereg_fold.mjs) ──
PREREG_RE hits: 62 (missing .cols 4)
…study preregistration sense: 15
…other sense (pre-registered domain / OAuth / FIDO / …): 47 = 75.8% of the keyword hits

==========================================================================
J. EXTERNAL FIGURES (non-corpus; each with its primary source)
==========================================================================
IEEE S&P SoK track: "As in past years, we solicit systematization of knowledge (SoK) papers"
IEEE S&P 2026 CFP: https://sp2026.ieee-security.org/cfpapers.html  (fetched 2026-08-27)
  "Submissions will be distinguished by the prefix “SoK:” in the title"
  "Survey papers without such insights are not appropriate and may be rejected without full review."
oaklandsok.github.io: SoK track at IEEE S&P since 2010; EuroS&P since 2017; PETS since 2019;
  USENIX Security since 2024; NDSS since 2026; SaTML since 2023. CCS, IMC, WWW: not listed.
  Fetched 2026-08-27. First SoK-titled paper in THIS index: 2013 (0 in 2010–2012).
Vekaria et al. arXiv:2506.14057 "SoK: Advances and Open Problems in Web Tracking"
  abs page fetched 2026-08-27: still a preprint; living version at
  github.com/privacysandstorm/sok-advances-open-problems-web-tracking
  Their method (from the HTML): "seven top web security and privacy venues" over 20 years
  (IEEE S&P, USENIX Security, ACM CCS, NDSS, ACM IMC, PETS, WWW) — the same seven.
  "A total of 200+ research papers were identified." Not in this extraction (wrong venue: arXiv).
Rieder et al. IEEE S&P 2026 / arXiv:2605.02982 listed on oaklandsok.github.io.
  In IEEE-2026.json; no abstract, so never screened, so not in the 5,859.
Krawlers discarded 654 of 1,057 = 61.9% of keyword hits as not automated crawls.
  130 of 403 = 32.3% navigate beyond a single page.
USENIX Security 2024 Krawlers landing: /conference/usenixsecurity24/presentation/stafeev
Blessing DOI 10.56553/popets-2025-0113 (fetch_authors.py failed on parse_popets; authors from citation_author meta)
drain item writing:literature_review (new) id 178 run 55 claimed as cursor-drain-litrev

==========================================================================
Z. ARITHMETIC-DERIVED (so a page figure is grep-able here even when it is a quotient)
==========================================================================
43 extracted SoKs of 5859 = 0.7%
IEEE-SP 30 of 43 = 69.8%
web_measurement 4 of 43 = 9.3%
index 167 SoKs → extracted 43
Krawlers 654/1057 = 61.9%
Krawlers 130/403 = 32.3%
keyword miss among crawled 2010-2022: 148/687 = 21.5%
neither tranco nor alexa: 312/687 = 45.4%
tranco OR alexa: 375/687 = 54.6%
crawled 2010-2022: 687
crawled any year matching /\btranco\b/i: 194 of 1120 = 17.3% (keyword_denominator.py demo)
crawled any year NOT matching /\btranco\b/i: 926 = 82.7%
headless never stated 980 of 1120 = 87.5%
keyword_denominator.py \bheadless\b: 137 match, 983 miss = 12.2% / 87.8%
Warford 6534 / 127 / 95; Blessing 245; Alam 55 papers from DBLP; Birrell ten selected venues; Alexa top-300
PETS 41 of 43 index SoKs not extracted (43 index, 2 extracted)
literature:corpus funnel (quoted, not re-derived this sitting): 16,864 index → 15,800 abstracts → 6,103 selected → 5,873 PDFs → 5,859 extracted
privacy:requests population S1 ∪ S2 = 254; overlap both = 107 (quoted from that page's report, 2026-08 corpus)
keyword_denominator.py pct formula uses 100 * n / d
survey titles 11 of 5859; not SoK 11; user/interview-tagged 7
index SoKs 167; selected 47; extracted 43; no-abstract 4; selected-not-extracted 4
PETS 43 index SoKs, 2 extracted
search: venue_then_filter 7, database_keyword 3, scholar_snowball 8, not_a_census 19, not_stated 6
census 24 of 43; stated protocol 18 of 24
full-text SoK phrase 26; citation-only 15; extracted SoKs without the phrase 32 of 43
web platform tag 6 of 43
year 2024 has 8 extracted SoKs, the peak complete year
headless stated 140/1120 = 12.5%
browser FP precision 83/280 = 29.6%
prereg homograph 47/62 = 75.8%

OK: no failures.

Review log

Focused-pass and generic-pass findings, with accept/reject, are appended below after the reviews return. This heading is the log's slot, filled in the same sitting — not an empty review stub.

provenance/writing/literature_review.1787835146.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki