User Tools

Site Tools


provenance:writing:literature_review

Provenance: Writing:Literature review

Back to Literature review. Corpus-wide selection and extraction notes are on corpus. Citations use the shared bibliography; this page adds no keys of its own. No ~~DISCUSSION~~ — comments belong on the content page.

Run record

  • Run date: 2026-08-27 (UTC).
  • Drain item: writing:literature_review (new), claimed as cursor-drain-litrev (item 178, run 55). Executed as Cursor, not via claude -p / drain-sandbox.sh.
  • Corpus at run time: 5,859 extracted papers, 2010–2026, CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P. Read-only inputs under /workspace/publications_dataset/data/.
  • Read first: data/extract/OVERVIEW.md, the wiki-measuretheweb spec (audience, currency, 5,859-paper corpus, provenance page, start link, bibtex), corpus, fingerprinting, study_preregistration, requests, crawler, live start (already promised Literature review).
  • Target pages had no revision (core.getPageInfo “does not exist”; ?do=export_raw on a missing page returns an HTML error document — that is not an existence test by byte count). dw.mjs pages does not list provenance:. This is a creation, not an extension.
  • Overlap judgement: create the promised child. Do not broaden corpus (funnel) or conferences (where to send the paper). The Writing namespace page itself does not exist; not this item.
  • No write to the publication mount. Wiki saves through scripts/dw.mjs (JSON-RPC). Live start and literature:bibliography were re-exported immediately before append, not taken from a stale local copy.

Why this page

The item asked for the SoK papers in the corpus plus this project as a worked example of structured extraction vs keyword search, and why a keyword-derived denominator is biased along the axis being measured. That is a writing/methods page, not a survey of every SoK. A reasonable person might have written a catalogue of the 43; that would miss the trap the item named.

Population and queries

All counts are papers unless labelled tuples. Sentinels are not answers.

Membership is mechanical: title starts SoK / SOK: or slug starts sok-. On extract/run1 those two tests agree: 43 / 43, intersection 43, residue 0. The report throws if they diverge.

Query Denominator Result
title-start SoK / SOK: 5,859 extracted 43 (0.7%)
slug-start sok- 5,859 43 (0.7%)
union (page population) 5,859 43
title-only, slug-only 43 0, 0
IEEE S&P among extracted SoKs 43 30 (69.8%)
CCS / IMC / WWW extracted SoKs 43 0 / 0 / 0
same three venues in the 16,864-record index 16,864 0 / 0 / 0
TOPIC = web_measurement 43 4 (9.3%)
TOPIC = web_adjacent 43 13 (30.2%)
TOPIC = not_web 43 26 (60.5%)
SEARCH = venue_then_filter 43 7 (16.3%)
SEARCH = database_keyword 43 3 (7.0%)
SEARCH = scholar_snowball 43 8 (18.6%)
SEARCH = not_stated 43 6 (14.0%)
SEARCH = not_a_census 43 19 (44.2%)
claims a literature sample 43 24 (55.8%)
…and states a protocol 24 18 (75.0%)
venue-then-filter among stated protocols 18 7 (38.9%)
index SoKs (title-start or slug-start) 16,864 index records 167
…with an abstract 167 163 (lost 4)
…labelled 163 163 (lost 0)
…selected by the over-inclusive –any rule 163 47 (lost 116)
…extracted 47 43 (lost 4)
PETS index SoKs vs extracted 43 index 2 extracted (41 not extracted)
crawled; crawlConfig.headless stated (not sentinel) 1,120 crawled 140 (12.5%)
crawled; full-text /\bheadless\b/i 1,120 137 (12.2%)
crawled; schema never stated headless 1,120 980 (87.5%)
detection.phenomenon names a fingerprint 5,859 280; browser-device 83 (29.6%); traffic-analysis 105 (37.5%)
PREREG_RE full-text hits 5,859 (4 missing .cols) 62; study sense 15; other-sense 47 (75.8%)
title contains “survey” 5,859 11; none are SoKs; 7/11 tagged user-study or interview-or-survey
full-text /systematization of knowledge/i 5,859 (4 missing .cols) 26; 15 not in the title/slug union; 32/43 extracted SoKs lack the phrase
crawled 2010–2022 (Krawlers-year recall) 1,120 crawled 687
those matching tranco or alexa 687 375 (54.6%)
those matching neither 687 312 (45.4%)
generous three-keyword union miss 687 148 (21.5%)not the published lead; “top AND site” as two independent words is looser than Krawlers' query
generous three-keyword hits 2010–2022 that are not crawled 1,493 hits 954 (63.9%)
keyword_denominator.py '\btranco\b' on all crawled years 1,120 194 (17.3%) match, 926 (82.7%) miss
keyword_denominator.py '\bheadless\b' 1,120 137 (12.2%) match, 983 (87.8%) miss

The 21.5% three-keyword miss is in the report because it was computed. It is not on the content page as a headline, because treating “combination of top and site” as two independent word hits is a different experiment from theirs.

Folding and residue

  • Membership is the union of two mechanical tests. Residue of title-vs-slug: 0. The report throws if they diverge, and throws if TOPIC or SEARCH misses a key of the 43 or contains a key that is not in the 43.
  • TOPIC and SEARCH are hand maps in scripts/litrev_fold.mjs, one primary label per paper, deciding sentence inline. Residue 0 by construction (the throw). A different reader might move Birrell (privacy-regulation impact) or Mathews (website-fingerprinting defenses) across the web / adjacent line; both deciding sentences are in the fold file. Mathews stays adjacent on purpose — it is the fingerprinting homograph.
  • The four web-measurement SoKs: Stafeev & Pellegrino (krawlers), Birrell et al. (privacy regulations), Blessing et al. (web auth), Alam et al. (PHILTER). Alam's key alam2026_philter was already in the live bibliography; it was not added again.

Quotes checked against paper.cols.txt

Needles the report requires (flatten whitespace). All OK on the run that the page was written from:

  • Krawlers: 7,840, 1,057, 654, 403, 32.3%, tranco, alexa, and the combination of top and site
  • Warford: 6,534 papers, 127 potentially relevant papers, reviewed 95 papers
  • Blessing: 245 papers, top-300 (the sentence top 300 of the Alexa is column-spliced in .cols; the page uses the hyphenated form that survives)
  • Usman: more effective than keyword searching in a database; the six-venue list including IEEE Security & Privacy, and PETS
  • Alam: In total, 55 papers
  • Birrell: ten selected venues

The Usman quote on the content page is shortened with an ellipsis. The two fragments it uses are both in .cols.

External and industry verification

Source Load-bearing fact Verification 2026-08-27 Decision
IEEE S&P 2026 CFP https://sp2026.ieee-security.org/cfpapers.html SoK: prefix; surveys without insights may be rejected fetched HTML Used
https://oaklandsok.github.io/ SoK at S&P since 2010; EuroS&P 2017; PETS 2019; USENIX 2024; NDSS 2026; not CCS/IMC/WWW fetched HTML Used. First SoK-titled paper in this index is 2013; the page says so rather than pretending to have 2010–2012.
arXiv:2506.14057 abs still v1 preprint (16 June 2025); living repo named fetched 2026-08-27 Used as orientation, dated as preprint, not as a denominator
github.com/privacysandstorm/sok-advances-open-problems-web-tracking living version of Vekaria named on the abs page Named, not counted
IEEE-2026.json Rieder slug in index, no abstract, never screened, not in 5,859 report prints the four no-abstract slugs Used as the mechanical-miss example
USENIX Security 2024 landing /presentation/stafeev authors Aleksei Stafeev, Giancarlo Pellegrino fetch_authors.py + landing Used in BibTeX
PETS 2025 landing popets-2025-0113 authors Jenny Blessing, Daniel Hugenroth, Ross Anderson, Alastair Beresford; DOI 10.56553/popets-2025-0113 fetch_authors.py failed on parse_popets; authors filled by hand from citation_author meta tags into out/authors.json Used. The failed parser is recorded so the next run does not treat the cache as automatic.
literature:corpus funnel 16,864 → 15,800 → 6,103 → 5,873 → 5,859 quoted from that page / report_corpus.mjs, not re-derived this sitting Used, labelled as quoted
privacy:requests S1 ∪ S2 = 254; both = 107 quoted from that page's report Used, labelled as quoted

Rejected:

  • Leading with the 21.5% three-keyword miss as “what Krawlers dropped” — the third keyword was not reproduced as they wrote it.
  • Counting title-“survey” papers as literature surveys — 7/11 are user/interview studies; Surveylance is survey scams.
  • Using “SoK” as a quality filter or as a web-measurement name — 30/43 are IEEE S&P; 26/43 are not this field.
  • Treating Vekaria “200+” as a reproducible denominator from this extraction — wrong venue (arXiv).
  • Adding a duplicate alam2026_philter.
  • SEO “how to write a literature review” listicles. None were consulted.
  • PRISMA / Scholar tutorials. Textbook; out of scope.
  • Re-deriving the 16,864-index funnel in this sitting — that is corpus's job.

What the corpus and sources do not establish

  • How many of the 116 screened-out SoKs a web-measurement PhD student would still want. Screening did its job on PETS crypto/DP; a handful of edge cases were not re-read.
  • Whether IEEE S&P 2010–2012 SoKs exist under a different title pattern in this index. Zero title-start hits; oaklandsok dates the track to 2010.
  • A venue version of Vekaria et al. As of 2026-08-27, arXiv v1 only.
  • Rieder et al. in the 5,859. No abstract → never screened.
  • Whether CCS / IMC / TheWebConf will add an SoK track.
  • A current, independent accuracy comparison of keyword search vs venue-complete on a question other than the ones this page measures.

Judgement calls

  • Create the promised page rather than a catalogue of 43 SoKs or a Scholar tutorial.
  • Hand-fold topic and search rather than regex-fold: the questions are not spelling variants. Residue 0 because the report throws, not because a regex covered everything.
  • Put Mathews in adjacent, not web_measurement.
  • Quote Krawlers' own 7,840 / 1,057 / 654 / 403 / 32.3% from their PDF (needles in .cols), and reproduce only the named-list half (tranco / alexa) as a recall experiment on this extraction's crawled 2010–2022.
  • Demo script scans crawled papers only and crashes if a crawled .cols is missing. An earlier draft scanned all papers and crashed on a 2010 USENIX paper with no full text — that paper is not crawled, so it is not in the denominator.
  • PETS authors for Blessing filled by hand after fetch_authors.py failed. Recorded here; not silently patched at the read site.
  • Start already linked the page. A one-line blurb was added on live start so the outline matches the Conferences line next to it.
  • check_attributions.mjs matches table-row “Name et al., VENUE YEAR [key]” only. This page attributes in prose, so that guard reports 0 checked and exits 1. That is not a pass and was not treated as one.

Embedded script

pages/keyword_denominator.py is byte-identical to the <file python keyword_denominator.py> block on the content page (diffed before publish). Quoted output is the \bheadless\b run, 2026-08-27, against extract/run1: 137 / 983 of 1,120. The \btranco\b run is 194 / 926 and is in the report; the page leads with headless because that is the reporting-rate trap.

Report script output (unedited)

Command: node scripts/report_literature_review.mjs. Output of the run the content page was written from:

corpus: 5859 papers, 7 venues, 2010–2026
crawled: 1120  empirical: 5118

==========================================================================
A. SOK MEMBERSHIP (title-start vs slug-start)
==========================================================================
Test                     Papers  Share of 5859
-----------------------  ------  -------------
title starts SoK / SOK:  43      0.7%
slug starts sok-         43      0.7%
union (the population)   43      0.7%
title-only (not slug)    0       0.0%
slug-only (not title)    0       0.0%
title contains word SoK anywhere: 43; of which outside the union: 0

==========================================================================
B. SHAPE OF THE 43 EXTRACTED SOKS
==========================================================================
Venue    Extracted SoKs  Extracted papers  SoKs / venue
-------  --------------  ----------------  ------------
CCS      0               990               0.0%
IEEE-SP  30              767               3.9%
IMC      0               638               0.0%
NDSS     4               701               0.6%
PETS     2               510               0.4%
USENIX   7               1410              0.5%
WWW      0               843               0.0%
venues with zero extracted SoKs: CCS, IMC, WWW

── by year (2025–2026 starred as provisional) ──
Year   Extracted SoKs  Extracted papers that year
-----  --------------  --------------------------
2010   0               119
2011   0               116
2012   0               151
2013   2               125
2014   0               166
2015   1               190
2016   2               182
2017   2               231
2018   1               254
2019   1               402
2020   2               404
2021   3               379
2022   4               546
2023   4               719
2024   8               690
2025*  7               770
2026*  6               415

── IEEE-SP share of extracted SoKs ──
IEEE-SP SoKs: 30 of 43 = 69.8%

── platforms (a paper can have several; shares of 43) ──
platform              SoKs  Share of 43
--------------------  ----  -----------
offline               26    60.5%
other-online-service  11    25.6%
web                   6     14.0%
not-applicable        5     11.6%
mobile                3     7.0%
iot                   1     2.3%
platforms includes web: 6 of 43 = 14.0%
  (platform=web is not the topic fold — printers and HPC counters sit here)

==========================================================================
C. TOPIC HAND FOLD (43 papers, 0 residue by construction)
==========================================================================
Topic                                                                                                      Papers  Share of 43
---------------------------------------------------------------------------------------------------------  ------  -----------
Web measurement (crawls, live-site audits, phishing-site detectors, privacy-regulation impact on the web)  4       9.3%
Adjacent online privacy / abuse / traffic analysis                                                         13      30.2%
Not this field (TEE, binaries, fuzzing, crypto, hardware, DeFi, …)                                         26      60.5%

── web_measurement ──
  2024 IEEE-SP  SoK: Technical Implementation and Human Impact of Internet Privacy Regulations.
  2024 USENIX   SoK: State of the Krawlers – Evaluating the Effectiveness of Crawling Algorithms for Web Security Measurements
  2025 PETS     SoK: Web Authentication and Recovery in the Age of End-to-End Encryption
  2026 USENIX   SoK: PHILTER: Uncovering Security and Functional Gaps in AI-based Phishing Website Detection Literature via an LLM-based Reasoning Framework

── web_adjacent ──
  2016 IEEE-SP  SoK: Towards Grounding Censorship Circumvention in Empiricism.
  2020 PETS     SoK: Anatomy of Data Breaches
  2021 IEEE-SP  SoK: Hate, Harassment, and the Changing Landscape of Online Abuse.
  2022 IEEE-SP  SoK: A Framework for Unifying At-Risk User Research.
  2022 IEEE-SP  SoK: Social Cybersecurity.
  2022 IEEE-SP  SoK: The Dual Nature of Technology in Sexual Abuse.
  2023 IEEE-SP  SoK: A Critical Evaluation of Efficient Website Fingerprinting Defenses.
  2024 IEEE-SP  SoK: Safer Digital-Safety Research Involving At-Risk Users.
  2024 USENIX   SoK (or SoLK?): On the Quantitative Study of Sociodemographic Factors and Computer Security Behaviors
  2025 IEEE-SP  SoK: A Privacy Framework for Security Research Using Social Media Data.
  2025 IEEE-SP  SoK: A Framework and Guide for Human-Centered Threat Modeling in Security and Privacy Research.
  2025 IEEE-SP  SoK: Self-Generated Nudes over Private Chats: How can Technology Contribute to a Safer Sexting?
  2025 IEEE-SP  SoK: Decoding the Enigma of Encrypted Network Traffic Classifiers.

── not_web ──
  2013 IEEE-SP  SoK: Secure Data Deletion.
  2013 IEEE-SP  SoK: P2PWNED - Modeling and Evaluating the Resilience of Peer-to-Peer Botnets.
  2015 IEEE-SP  SoK: Deep Packer Inspection: A Longitudinal Study of the Complexity of Run-Time Packers.
  2016 IEEE-SP  SOK: (State of) The Art of War: Offensive Techniques in Binary Analysis.
  2017 IEEE-SP  SoK: Cryptographically Protected Database Search.
  2017 IEEE-SP  SoK: Exploiting Network Printers.
  2018 IEEE-SP  SoK: Keylogging Side Channels.
  2019 IEEE-SP  SoK: The Challenges, Pitfalls, and Perils of Using Hardware Performance Counters for Security.
  2020 IEEE-SP  SoK: Understanding the Prevailing Security Vulnerabilities in TrustZone-assisted TEE Systems.
  2021 IEEE-SP  SoK: All You Ever Wanted to Know About x86/x64 Binary Disassembly But Were Afraid to Ask.
  2021 IEEE-SP  SoK: Quantifying Cyber Risk.
  2022 IEEE-SP  SoK: How Robust is Image Classification Deep Neural Network Watermarking?
  2023 IEEE-SP  SoK: Decentralized Finance (DeFi) Attacks.
  2023 IEEE-SP  SoK: History is a Vast Early Warning System: Auditing the Provenance of System Intrusions.
  2023 IEEE-SP  SoK: Taxonomy of Attacks on Open-Source Software Supply Chains.
  2024 IEEE-SP  SoK: Prudent Evaluation Practices for Fuzzing.
  2024 USENIX   SoK: All You Need to Know About On-Device ML Model Extraction - The Gap Between Research and Practice
  2024 USENIX   SoK: What Don't We Know? Understanding Security Vulnerabilities in SNARKs
  2024 IEEE-SP  SoK: SGX.Fail: How Stuff Gets eXposed.
  2025 USENIX   SoK: An Introspective Analysis of RPKI Security
  2025 IEEE-SP  SoK: Software Compartmentalization.
  2026 NDSS     SoK: Take a Deep Step into Linux Kernel Hardening Effectiveness from the Offensive-Defensive Perspective
  2026 NDSS     SoK: Understanding the Fundamentals and Implications of Sensor Out-of-band Vulnerabilities
  2026 NDSS     SoK: Cryptographic Authenticated Dictionaries
  2026 USENIX   SoK: Security of Cyber-physical Systems Under Intentional Electromagnetic Interference Attacks
  2026 NDSS     SoK: Analysis of Accelerator TEE Designs

web_measurement SoKs: 4 of 43 = 9.3%

==========================================================================
D. SEARCH-METHOD HAND FOLD (primary label, 43 papers)
==========================================================================
How the paper list was built                                   Papers  Share of 43
-------------------------------------------------------------  ------  -----------
Venue-complete (or venue-list) then filter                     7       16.3%
Digital-library / DBLP keyword query                           3       7.0%
Google Scholar and/or citation snowball                        8       18.6%
Not a paper-census (system, evaluation, vulnerability corpus)  19      44.2%
Literature sample claimed; search protocol not stated          6       14.0%
claims a literature sample: 24 of 43 = 55.8%
…and states a search protocol: 18 of 24 = 75.0%
venue-complete then filter: 7 of 43 = 16.3% (of the 18 with a stated protocol: 38.9%)

── web_measurement SoKs × search method ──
Paper                                                                                                                                    Year  Venue    Search
---------------------------------------------------------------------------------------------------------------------------------------  ----  -------  -----------------
Technical Implementation and Human Impact of Internet Privacy Regulations.                                                               2024  IEEE-SP  scholar_snowball
State of the Krawlers – Evaluating the Effectiveness of Crawling Algorithms for Web Security Measurements                                2024  USENIX   venue_then_filter
Web Authentication and Recovery in the Age of End-to-End Encryption                                                                      2025  PETS     venue_then_filter
PHILTER: Uncovering Security and Functional Gaps in AI-based Phishing Website Detection Literature via an LLM-based Reasoning Framework  2026  USENIX   database_keyword

==========================================================================
E. INDEX FUNNEL — SOKS IN THE SEVEN VENUES, NOT JUST THE 43
==========================================================================
index records: 16864
index SoKs (title-start or slug-start): 167
Venue    Index SoKs  Extracted SoKs
-------  ----------  --------------
IEEE-SP  84          30
PETS     43          2
USENIX   36          7
NDSS     4           4
index SoKs in CCS, IMC, WWW: 0, 0, 0
Stage                        SoKs  Lost here
---------------------------  ----  ---------
Index (title or slug)        167   —
…with an abstract            163   4
…labelled (screened)         163   0
…selected by the --any rule  47    116
…extracted                   43    4
no-abstract index SoKs (4):
  IEEE-SP/2026/sok-all-you-ever-wanted-to-know-about-bootloader-security-but-were-afraid-to-ask
  IEEE-SP/2026/sok-critical-evaluation-of-quantum-machine-learning-for-adversarial-robustness
  IEEE-SP/2026/sok-robustness-in-large-language-models-against-jailbreak-attacks
  IEEE-SP/2026/sok-after-decades-of-web-tracker-detection-whats-next
selected but not extracted (4):
  IEEE-SP/2024/sok-a-comprehensive-analysis-and-evaluation-of-docker-container-attack-and-defen
  USENIX/2026/sok-history-doesnt-repeat-itself-but-android-design-level-vulnerabilities-rhyme
  USENIX/2026/sok-darpas-ai-cyber-challenge-aixcc-competition-design-architectures-and-lessons
  USENIX/2026/sok-capability-operating-systems-is-the-future-finally-here
Rieder index record present: IEEE-SP/2026/sok-after-decades-of-web-tracker-detection-whats-next
  abstract present: false
  in labels: false
  in extractions: false
  fulltext dir exists: false

── PETS index SoKs vs extracted (most dropped SoKs are crypto / DP / sensing) ──
PETS index SoKs 43; extracted 2

==========================================================================
F. HOMOGRAPHS AND FULL-TEXT SOK MENTIONS
==========================================================================
title contains "survey": 11 of 5859
…of which are not SoKs: 11
…of those, studyTypes includes user-study or interview-or-survey: 7
title matches systematic (literature) review: 1

── non-SoK "survey" titles (these are mostly user surveys, not literature surveys) ──
  2016 CCS      user/interview=true  How I Learned to be Secure: a Census-Representative Survey of Security Advice Sources and Behavior.
  2018 IEEE-SP  user/interview=false  Surveylance: Automatically Detecting Online Survey Scams.
  2019 IEEE-SP  user/interview=true  How Well Do My Results Generalize? Comparing Security and Privacy Survey Results from MTurk, Web, and Telephone Samples.
  2020 WWW      user/interview=false  Attention Please: Your Attention Check Questions in Survey Studies Can Be Automatically Answered.
  2022 PETS     user/interview=false  A Global Survey of Android Dual-Use Applications Used in Intimate Partner Surveillance
  2022 WWW      user/interview=true  Beyond Bot Detection: Combating Fraudulent Online Survey Takers✱.
  2023 PETS     user/interview=true  Privacy Concerns and Acceptance Factors of OSINT for Cybersecurity: A Representative Survey
  2024 IEEE-SP  user/interview=true  Digital Security - A Question of Perspective A Large-Scale Telephone Survey with Four At-Risk User Groups.
  2024 USENIX   user/interview=true  Bridging Barriers: A Survey of Challenges and Priorities in the Censorship Circumvention Landscape
  2024 USENIX   user/interview=true  Engaging Company Developers in Security Research Studies: A Comprehensive Literature Review and Quantitative Survey
  2026 USENIX   user/interview=false  Inconsistent, Incomplete, and Insecure: A Survey of Account Security Interfaces

full-text /systematization of knowledge/i: 26 papers; missing .cols: 4
…not in the title/slug union (citation / passing mention): 15
extracted SoKs whose .cols does NOT contain the phrase: 32 of 43

==========================================================================
G. KRAWLERS KEYWORD CUT REPRODUCED ON THIS EXTRACTION
==========================================================================
Krawlers (USENIX 2024) quoted figures, checked against paper.cols.txt:
  OK    "7,840"
  OK    "1,057"
  OK    "654"
  OK    "403"
  OK    "32.3%"
  OK    "tranco, alexa, and the combination of top and site"
  OK    Warford "6,534 papers"
  OK    Warford "127 potentially relevant papers"
  OK    Warford "reviewed 95 papers"
  OK    Warford "139 potentially relevant papers"
  OK    Blessing "245 papers"
  OK    Blessing "top-300"
  OK    Usman "more effective than keyword searching in a database"
  OK    Usman "SOUPS, CHI, CSCW, USENIX Security, IEEE Security & Privacy, and PETS"
  OK    Alam "In total, 55 papers"
  OK    Alam "a total of 38 papers"
  OK    Alam "17 additional papers"
  OK    Birrell "ten selected venues"
extracted papers 2010–2022: 3265
crawled papers 2010–2022: 687 (denominator for recall)
Krawlers keyword on crawled 2010–2022                               Papers  Share of 687
------------------------------------------------------------------  ------  ------------
tranco                                                              56      8.2%
alexa                                                               351     51.1%
tranco OR alexa (the two named lists)                               375     54.6%
neither tranco nor alexa                                            312     45.4%
top AND site (both words; generous reading of their third keyword)  478     69.6%
any of the three (union under that generous reading)                539     78.5%
NONE of the three                                                   148     21.5%
missing .cols                                                       0       0.0%
stricter /top[\s-]*site/i on crawled 2010–2022: 162 (not used as the published recall; the paper said "combination of top and site")
keyword union on all extracted 2010–2022: 1493 hits; of which NOT crawled: 954 = 63.9% (Krawlers discarded 654 of 1,057 = 61.9% as not employing automated crawling; our extraction is already screened, so this is a lower bound on the false-positive rate)

── crawled 2010–2022 missed by all three keywords (first 20 of the miss list) ──
  CCS/2010/detecting-and-characterizing-social-spam-campaigns  Detecting and characterizing social spam campaigns.
  CCS/2010/fingerprinting-websites-using-remote-traffic-analysis  Fingerprinting websites using remote traffic analysis.
  IMC/2010/detecting-and-characterizing-social-spam-campaigns  Detecting and characterizing social spam campaigns.
  WWW/2010/analyzing-content-level-properties-of-the-web-adversphere  Analyzing content-level properties of the web adversphere.
  WWW/2010/diversifying-landmark-image-search-results-by-learning-interested-views-from-com  Diversifying landmark image search results by learning interested views from community pho
  WWW/2010/the-social-honeypot-project-protecting-online-communities-from-spammers  The social honeypot project: protecting online communities from spammers.
  WWW/2010/restler-crawling-restful-services  RESTler: crawling RESTful services.
  CCS/2011/automated-black-box-detection-of-side-channel-vulnerabilities-in-web-application  Automated black-box detection of side-channel vulnerabilities in web applications.
  CCS/2011/android-permissions-demystified  Android permissions demystified.
  IMC/2011/analyzing-facebook-privacy-settings-user-expectations-vs-reality  Analyzing facebook privacy settings: user expectations vs. reality.
  WWW/2011/arrow-generating-signatures-to-detect-drive-by-downloads  ARROW: GenerAting SignatuRes to Detect DRive-By DOWnloads.
  CCS/2012/detecting-money-stealing-apps-in-alternative-android-markets  Detecting money-stealing apps in alternative Android markets.
  IMC/2012/evolution-of-a-location-based-online-social-network-analysis-and-models  Evolution of a location-based online social network: analysis and models.
  IMC/2012/evolution-of-social-attribute-networks-measurements-modeling-and-implications-us  Evolution of social-attribute networks: measurements, modeling, and implications using goo
  NDSS/2012/you-are-what-you-like-information-leakage-through-users-interests  You are what you like! Information leakage through users’ Interests
  USENIX/2012/enemy-of-the-state-a-state-aware-black-box-web-vulnerability-scanner  Enemy of the State: A State-Aware Black-Box Web Vulnerability Scanner
  CCS/2013/delta-automatic-identification-of-unknown-web-based-infection-campaigns  Delta: automatic identification of unknown web-based infection campaigns.
  WWW/2012/understanding-and-combating-link-farming-in-the-twitter-social-network  Understanding and combating link farming in the twitter social network.
  WWW/2012/economics-of-bittorrent-communities  Economics of BitTorrent communities.
  WWW/2012/analyzing-spammers-social-networks-for-fun-and-profit-a-case-study-of-cyber-crim  Analyzing spammers' social networks for fun and profit: a case study of cyber criminal eco
  … 128 more
crawled 2010–2022 whose .cols matches /\bcrux\b/i: 7 (Krawlers dropped the keyword because CrUX rankings started in 2022)

==========================================================================
H. REPORTING-RATE TRAP — SEARCHING FOR THE THING UNDERCOUNTS THE SILENCE
==========================================================================
crawled papers: 1120
crawlConfig.headless stated (not a sentinel): 140 = 12.5%
full-text /\bheadless\b/i over crawled: 137 = 12.2%; missing .cols 0
word "headless" but schema sentinel/absent: 28 (mentions in related work, or a different sense)
schema-stated but word not in .cols: 31
headless membership disagreement 59 papers (28 word-only, 31 schema-only); totals differ by 3
trap: a Scholar search for "headless crawler" returns papers that used the word; it cannot estimate the 87.5% of crawls that never said (980 of 1120).

==========================================================================
I. SIBLING HOMOGRAPHS RE-DERIVED
==========================================================================
detection.phenomenon names a fingerprint: 280 of 5859
…browser-device family: 83 = 29.6% of the fingerprint papers
…traffic-analysis family: 105 = 37.5%
keyword precision for browser fingerprinting if you search this corpus for "fingerprint": 29.6%

── preregistration homograph (PREREG_RE over full text, sense from prereg_fold.mjs) ──
PREREG_RE hits: 62 (missing .cols 4)
…study preregistration sense: 15
…other sense (pre-registered domain / OAuth / FIDO / …): 47 = 75.8% of the keyword hits

==========================================================================
J. EXTERNAL FIGURES (non-corpus; each with its primary source)
==========================================================================
IEEE S&P SoK track: "As in past years, we solicit systematization of knowledge (SoK) papers"
IEEE S&P 2026 CFP: https://sp2026.ieee-security.org/cfpapers.html  (fetched 2026-08-27)
  "Submissions will be distinguished by the prefix “SoK:” in the title"
  "Survey papers without such insights are not appropriate and may be rejected without full review."
oaklandsok.github.io: SoK track at IEEE S&P since 2010; EuroS&P since 2017; PETS since 2019;
  USENIX Security since 2024; NDSS since 2026; SaTML since 2023. CCS, IMC, WWW: not listed.
  Fetched 2026-08-27. First SoK-titled paper in THIS index: 2013 (0 in 2010–2012).
Vekaria et al. arXiv:2506.14057 "SoK: Advances and Open Problems in Web Tracking"
  abs page fetched 2026-08-27: still a preprint; living version at
  github.com/privacysandstorm/sok-advances-open-problems-web-tracking
  Their method (from the HTML): "seven top web security and privacy venues" over 20 years
  (IEEE S&P, USENIX Security, ACM CCS, NDSS, ACM IMC, PETS, WWW) — the same seven.
  "A total of 200+ research papers were identified." Not in this extraction (wrong venue: arXiv).
Rieder et al. IEEE S&P 2026 / arXiv:2605.02982 listed on oaklandsok.github.io.
  In IEEE-2026.json; no abstract, so never screened, so not in the 5,859.
Krawlers discarded 654 of 1,057 = 61.9% of keyword hits as not automated crawls.
  130 of 403 = 32.3% navigate beyond a single page.
USENIX Security 2024 Krawlers landing: /conference/usenixsecurity24/presentation/stafeev
Blessing DOI 10.56553/popets-2025-0113 (fetch_authors.py failed on parse_popets; authors from citation_author meta)
drain item writing:literature_review (new) id 178 run 55 claimed as cursor-drain-litrev
CCS 2026 and IMC 2026 have not been held; IEEE-SP 2026 and WWW 2026 abstracts are not in OpenAlex (two more 2026 venue-years incompletely selected). Source: literature:corpus / extract README, quoted.
review-log tokens: FIG-01 FIG-02 FIG-03 FIG-04 FIG-05 CIT-01 CIT-02 CIT-03 CIT-04 VEKARIA-01 CFP-01 TRACKS-01 TRACKS-02 GITHUB-01 KRAWLERS-01 BLESSING-01 LINKS-01 GEN-01 GEN-02 GEN-03 GEN-04 GEN-05 GEN-06 GEN-07
gpt-5.6-luna-max gpt-5.6-sol-medium frozen 12:53:21Z
reviewer ids 254169ad-89ee-4fdd-8d5c-5cf7a2bd990f bd1ebd8e-9534-4acf-9a89-4ed058490491 29c63fa2-4842-43fb-a227-161e420f6d2b 5f83de07-74a2-499b-9a8f-16cc7f52affc f08c070e-bc3a-4726-9acf-e8a8991bdbfa 5e30b90a-106c-469c-93c3-378e3738bda2

==========================================================================
Z. ARITHMETIC-DERIVED (so a page figure is grep-able here even when it is a quotient)
==========================================================================
43 extracted SoKs of 5859 = 0.7%
IEEE-SP 30 of 43 = 69.8%
web_measurement 4 of 43 = 9.3%
index 167 SoKs → extracted 43
Krawlers 654/1057 = 61.9%
Krawlers 130/403 = 32.3%
keyword miss among crawled 2010-2022: 148/687 = 21.5%
neither tranco nor alexa: 312/687 = 45.4%
tranco OR alexa: 375/687 = 54.6%
crawled 2010-2022: 687
crawled any year matching /\btranco\b/i: 194 of 1120 = 17.3% (keyword_denominator.py demo)
crawled any year NOT matching /\btranco\b/i: 926 = 82.7%
headless never stated 980 of 1120 = 87.5%
keyword_denominator.py \bheadless\b: 137 match, 983 miss = 12.2% / 87.8%
Warford 6534 / 127 / 12 / 139 / 95; Blessing 245; Alam 38 DBLP + 17 snowball = 55; Birrell ten selected venues then citation expansion; Alexa top-300
PETS 41 of 43 index SoKs not extracted (43 index, 2 extracted)
literature:corpus funnel (quoted, not re-derived this sitting): 16,864 index → 15,800 abstracts → 6,103 selected → 5,873 PDFs → 5,859 extracted
privacy:requests population S1 ∪ S2 = 254; overlap both = 107 (quoted from that page's report, 2026-08 corpus)
keyword_denominator.py pct formula uses 100 * n / d
survey titles 11 of 5859; not SoK 11; user/interview-tagged 7
index SoKs 167; selected 47; extracted 43; no-abstract 4; selected-not-extracted 4
PETS 43 index SoKs, 2 extracted
search: venue_then_filter 7, database_keyword 3, scholar_snowball 8, not_a_census 19, not_stated 6
census 24 of 43; stated protocol 18 of 24
full-text SoK phrase 26; citation-only 15; extracted SoKs without the phrase 32 of 43
web platform tag 6 of 43
year 2024 has 8 extracted SoKs, the peak complete year
headless stated 140/1120 = 12.5%
browser FP precision 83/280 = 29.6%
prereg homograph 47/62 = 75.8%

OK: no failures.

Review log

Focused drafts frozen at 2026-08-27T12:53:21Z under out/frozen_litrev/ (sha256 in SHA256SUMS). Three focused passes were requested as luna@medium; the Task tool rejected gpt-5.6-luna-max and there is no luna-medium slug in this session. Fallback: gpt-5.6-sol-medium (the only GPT 5.6 slug the tool accepts). Recorded so this is not a silent substitute.

Reviewer agent ids (Task tool): figures 254169ad-89ee-4fdd-8d5c-5cf7a2bd990f; citations bd1ebd8e-9534-4acf-9a89-4ed058490491; external 29c63fa2-4842-43fb-a227-161e420f6d2b.

Id Pass Finding Decision
FIG-01 figures vs script “disagree by three papers” — totals differ by 3; membership differs on 59 (28 word-only, 31 schema-only) Accepted. Prose now says totals vs membership.
FIG-02 figures vs script “Among those 18” headed a table of all 43 Accepted. Now “Across all 43”.
FIG-03 figures vs script Krawlers' 403 called “crawls that named a popular list” — third keyword is “top”+“site” Accepted.
FIG-04 figures vs script missing .cols counted and computation continued Accepted. Report now fail()s if crawled 2010–2022 or the headless sweep has missing full text. Current missing counts are 0.
FIG-05 figures vs script “two more 2026 venue-years are incompletely selected” not printed by the report Accepted. Printed in section J, quoted from literature:corpus (IEEE-SP 2026 and WWW 2026 abstracts not in OpenAlex).
CIT-01 citations Alam “55 papers from DBLP” — DBLP yielded 38, snowball +17 Accepted. Needles in .cols.
CIT-02 citations Warford 6,534→127→95 omitted the +12→139 pool Accepted.
CIT-03 citations Birrell “snowball inside ten venues” — seeds were ten venues; snowball admitted outside-venue papers Accepted.
CIT-04 citations same as FIG-03 Accepted.
VEKARIA-01 external still arXiv v1; optional note of IEEE S&P 2026 poster / USENIX under-review from author CV Rejected the USENIX-under-review claim (CV is not a primary venue source). Preprint hedge already on the page. Poster not added without a primary URL fetched this sitting.
CFP-01, TRACKS-01/02, GITHUB-01, KRAWLERS-01, BLESSING-01, LINKS-01 external live facts match; hedge CCS/IMC/WWW as “not listed on oaklandsok” Accepted the hedge on the venue table sentence and the TODO. Other live checks: no change.

Generic pass (requested gpt-5.6-luna-max; same Task-tool fallback gpt-5.6-sol-medium), no checklist. Reviewer id 5e30b90a-106c-469c-93c3-378e3738bda2. Re-review of figures/citations after focused fixes: CLEAR. Generic re-check after applying GEN-01..05: CLEAR.

Id Finding Decision
GEN-01 980 “never said headless” conflates schema absence with textual silence (28 word-only) Accepted. Schema wording is now “did not state whether their own crawl was headless”.
GEN-02 29.6% presented as keyword-search precision; 280 is detection.phenomenon Accepted. WRAP and the homograph closer now name the schema field.
GEN-03 “correctly dropped” / “the label did its job” overruns the un-reread 116 Accepted. Spot-check language; heading no longer says “correctly dropped”.
GEN-04 Snowballing collapsed backward and forward Accepted. One-sentence distinction; not a methods-textbook on snowballing.
GEN-05 “do not run” / “will not give you one” vs oaklandsok listing Accepted in WRAP and the SoK-filter bullet.
GEN-06 Actionable guidance too late; move script to provenance Partial. Added a <WRAP tip> with the three sentences a PhD student needs first. Rejected moving the demo script off the page — it is the worked example, not an appendix.
GEN-07 <WRAP todo> looks like a review stub; use lowercase wrap Rejected. Uppercase <WRAP todo> is the site convention (check_wrap.mjs forbids lowercase). Other pages ship open items that way; this is not an empty review placeholder.
provenance/writing/literature_review.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki