This is an old revision of the document!
Table of Contents
Provenance: Writing:Literature review
Back to Literature review. Corpus-wide selection and extraction notes are on corpus. Citations use the shared bibliography; this page adds no keys of its own. No ~~DISCUSSION~~ — comments belong on the content page.
Run record
- Run date: 2026-08-27 (UTC).
- Drain item:
writing:literature_review (new), claimed ascursor-drain-litrev(item 178, run 55). Executed as Cursor, not viaclaude -p/drain-sandbox.sh. - Corpus at run time: 5,859 extracted papers, 2010–2026, CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P. Read-only inputs under
/workspace/publications_dataset/data/. - Read first:
data/extract/OVERVIEW.md, the wiki-measuretheweb spec (audience, currency, 5,859-paper corpus, provenance page, start link, bibtex), corpus, fingerprinting, study_preregistration, requests, crawler, livestart(already promised Literature review). - Target pages had no revision (
core.getPageInfo“does not exist”;?do=export_rawon a missing page returns an HTML error document — that is not an existence test by byte count).dw.mjs pagesdoes not listprovenance:. This is a creation, not an extension. - Overlap judgement: create the promised child. Do not broaden corpus (funnel) or conferences (where to send the paper). The
Writingnamespace page itself does not exist; not this item. - No write to the publication mount. Wiki saves through
scripts/dw.mjs(JSON-RPC). Livestartandliterature:bibliographywere re-exported immediately before append, not taken from a stale local copy.
Why this page
The item asked for the SoK papers in the corpus plus this project as a worked example of structured extraction vs keyword search, and why a keyword-derived denominator is biased along the axis being measured. That is a writing/methods page, not a survey of every SoK. A reasonable person might have written a catalogue of the 43; that would miss the trap the item named.
Population and queries
All counts are papers unless labelled tuples. Sentinels are not answers.
Membership is mechanical: title starts SoK / SOK: or slug starts sok-. On extract/run1 those two tests agree: 43 / 43, intersection 43, residue 0. The report throws if they diverge.
| Query | Denominator | Result |
|---|---|---|
| title-start SoK / SOK: | 5,859 extracted | 43 (0.7%) |
| slug-start sok- | 5,859 | 43 (0.7%) |
| union (page population) | 5,859 | 43 |
| title-only, slug-only | 43 | 0, 0 |
| IEEE S&P among extracted SoKs | 43 | 30 (69.8%) |
| CCS / IMC / WWW extracted SoKs | 43 | 0 / 0 / 0 |
| same three venues in the 16,864-record index | 16,864 | 0 / 0 / 0 |
| TOPIC = web_measurement | 43 | 4 (9.3%) |
| TOPIC = web_adjacent | 43 | 13 (30.2%) |
| TOPIC = not_web | 43 | 26 (60.5%) |
| SEARCH = venue_then_filter | 43 | 7 (16.3%) |
| SEARCH = database_keyword | 43 | 3 (7.0%) |
| SEARCH = scholar_snowball | 43 | 8 (18.6%) |
| SEARCH = not_stated | 43 | 6 (14.0%) |
| SEARCH = not_a_census | 43 | 19 (44.2%) |
| claims a literature sample | 43 | 24 (55.8%) |
| …and states a protocol | 24 | 18 (75.0%) |
| venue-then-filter among stated protocols | 18 | 7 (38.9%) |
| index SoKs (title-start or slug-start) | 16,864 index records | 167 |
| …with an abstract | 167 | 163 (lost 4) |
| …labelled | 163 | 163 (lost 0) |
| …selected by the over-inclusive –any rule | 163 | 47 (lost 116) |
| …extracted | 47 | 43 (lost 4) |
| PETS index SoKs vs extracted | 43 index | 2 extracted (41 not extracted) |
crawled; crawlConfig.headless stated (not sentinel) | 1,120 crawled | 140 (12.5%) |
crawled; full-text /\bheadless\b/i | 1,120 | 137 (12.2%) |
| crawled; schema never stated headless | 1,120 | 980 (87.5%) |
detection.phenomenon names a fingerprint | 5,859 | 280; browser-device 83 (29.6%); traffic-analysis 105 (37.5%) |
PREREG_RE full-text hits | 5,859 (4 missing .cols) | 62; study sense 15; other-sense 47 (75.8%) |
| title contains “survey” | 5,859 | 11; none are SoKs; 7/11 tagged user-study or interview-or-survey |
full-text /systematization of knowledge/i | 5,859 (4 missing .cols) | 26; 15 not in the title/slug union; 32/43 extracted SoKs lack the phrase |
| crawled 2010–2022 (Krawlers-year recall) | 1,120 crawled | 687 |
| those matching tranco or alexa | 687 | 375 (54.6%) |
| those matching neither | 687 | 312 (45.4%) |
| generous three-keyword union miss | 687 | 148 (21.5%) — not the published lead; “top AND site” as two independent words is looser than Krawlers' query |
| generous three-keyword hits 2010–2022 that are not crawled | 1,493 hits | 954 (63.9%) |
keyword_denominator.py '\btranco\b' on all crawled years | 1,120 | 194 (17.3%) match, 926 (82.7%) miss |
keyword_denominator.py '\bheadless\b' | 1,120 | 137 (12.2%) match, 983 (87.8%) miss |
The 21.5% three-keyword miss is in the report because it was computed. It is not on the content page as a headline, because treating “combination of top and site” as two independent word hits is a different experiment from theirs.
Folding and residue
- Membership is the union of two mechanical tests. Residue of title-vs-slug: 0. The report throws if they diverge, and throws if
TOPICorSEARCHmisses a key of the 43 or contains a key that is not in the 43. - TOPIC and SEARCH are hand maps in
scripts/litrev_fold.mjs, one primary label per paper, deciding sentence inline. Residue 0 by construction (the throw). A different reader might move Birrell (privacy-regulation impact) or Mathews (website-fingerprinting defenses) across the web / adjacent line; both deciding sentences are in the fold file. Mathews stays adjacent on purpose — it is the fingerprinting homograph. - The four web-measurement SoKs: Stafeev & Pellegrino (krawlers), Birrell et al. (privacy regulations), Blessing et al. (web auth), Alam et al. (PHILTER). Alam's key
alam2026_philterwas already in the live bibliography; it was not added again.
Quotes checked against paper.cols.txt
Needles the report requires (flatten whitespace). All OK on the run that the page was written from:
- Krawlers:
7,840,1,057,654,403,32.3%,tranco, alexa, and the combination of top and site - Warford:
6,534 papers,127 potentially relevant papers,reviewed 95 papers - Blessing:
245 papers,top-300(the sentencetop 300 of the Alexais column-spliced in.cols; the page uses the hyphenated form that survives) - Usman:
more effective than keyword searching in a database; the six-venue list includingIEEE Security & Privacy, and PETS - Alam:
In total, 55 papers - Birrell:
ten selected venues
The Usman quote on the content page is shortened with an ellipsis. The two fragments it uses are both in .cols.
External and industry verification
| Source | Load-bearing fact | Verification 2026-08-27 | Decision |
|---|---|---|---|
| IEEE S&P 2026 CFP https://sp2026.ieee-security.org/cfpapers.html | SoK: prefix; surveys without insights may be rejected | fetched HTML | Used |
| https://oaklandsok.github.io/ | SoK at S&P since 2010; EuroS&P 2017; PETS 2019; USENIX 2024; NDSS 2026; not CCS/IMC/WWW | fetched HTML | Used. First SoK-titled paper in this index is 2013; the page says so rather than pretending to have 2010–2012. |
| arXiv:2506.14057 abs | still v1 preprint (16 June 2025); living repo named | fetched 2026-08-27 | Used as orientation, dated as preprint, not as a denominator |
| github.com/privacysandstorm/sok-advances-open-problems-web-tracking | living version of Vekaria | named on the abs page | Named, not counted |
| IEEE-2026.json Rieder slug | in index, no abstract, never screened, not in 5,859 | report prints the four no-abstract slugs | Used as the mechanical-miss example |
| USENIX Security 2024 landing /presentation/stafeev | authors Aleksei Stafeev, Giancarlo Pellegrino | fetch_authors.py + landing | Used in BibTeX |
| PETS 2025 landing popets-2025-0113 | authors Jenny Blessing, Daniel Hugenroth, Ross Anderson, Alastair Beresford; DOI 10.56553/popets-2025-0113 | fetch_authors.py failed on parse_popets; authors filled by hand from citation_author meta tags into out/authors.json | Used. The failed parser is recorded so the next run does not treat the cache as automatic. |
| literature:corpus funnel | 16,864 → 15,800 → 6,103 → 5,873 → 5,859 | quoted from that page / report_corpus.mjs, not re-derived this sitting | Used, labelled as quoted |
| privacy:requests | S1 ∪ S2 = 254; both = 107 | quoted from that page's report | Used, labelled as quoted |
Rejected:
- Leading with the 21.5% three-keyword miss as “what Krawlers dropped” — the third keyword was not reproduced as they wrote it.
- Counting title-“survey” papers as literature surveys — 7/11 are user/interview studies; Surveylance is survey scams.
- Using “SoK” as a quality filter or as a web-measurement name — 30/43 are IEEE S&P; 26/43 are not this field.
- Treating Vekaria “200+” as a reproducible denominator from this extraction — wrong venue (arXiv).
- Adding a duplicate
alam2026_philter. - SEO “how to write a literature review” listicles. None were consulted.
- PRISMA / Scholar tutorials. Textbook; out of scope.
- Re-deriving the 16,864-index funnel in this sitting — that is corpus's job.
What the corpus and sources do not establish
- How many of the 116 screened-out SoKs a web-measurement PhD student would still want. Screening did its job on PETS crypto/DP; a handful of edge cases were not re-read.
- Whether IEEE S&P 2010–2012 SoKs exist under a different title pattern in this index. Zero title-start hits; oaklandsok dates the track to 2010.
- A venue version of Vekaria et al. As of 2026-08-27, arXiv v1 only.
- Rieder et al. in the 5,859. No abstract → never screened.
- Whether CCS / IMC / TheWebConf will add an SoK track.
- A current, independent accuracy comparison of keyword search vs venue-complete on a question other than the ones this page measures.
Judgement calls
- Create the promised page rather than a catalogue of 43 SoKs or a Scholar tutorial.
- Hand-fold topic and search rather than regex-fold: the questions are not spelling variants. Residue 0 because the report throws, not because a regex covered everything.
- Put Mathews in adjacent, not web_measurement.
- Quote Krawlers' own 7,840 / 1,057 / 654 / 403 / 32.3% from their PDF (needles in
.cols), and reproduce only the named-list half (tranco / alexa) as a recall experiment on this extraction's crawled 2010–2022. - Demo script scans crawled papers only and crashes if a crawled
.colsis missing. An earlier draft scanned all papers and crashed on a 2010 USENIX paper with no full text — that paper is not crawled, so it is not in the denominator. - PETS authors for Blessing filled by hand after
fetch_authors.pyfailed. Recorded here; not silently patched at the read site. - Start already linked the page. A one-line blurb was added on live
startso the outline matches the Conferences line next to it. check_attributions.mjsmatches table-row “Name et al., VENUE YEAR [key]” only. This page attributes in prose, so that guard reports 0 checked and exits 1. That is not a pass and was not treated as one.
Embedded script
pages/keyword_denominator.py is byte-identical to the <file python keyword_denominator.py> block on the content page (diffed before publish). Quoted output is the \bheadless\b run, 2026-08-27, against extract/run1: 137 / 983 of 1,120. The \btranco\b run is 194 / 926 and is in the report; the page leads with headless because that is the reporting-rate trap.
Report script output (unedited)
Command: node scripts/report_literature_review.mjs. Output of the run the content page was written from:
corpus: 5859 papers, 7 venues, 2010–2026 crawled: 1120 empirical: 5118 ========================================================================== A. SOK MEMBERSHIP (title-start vs slug-start) ========================================================================== Test Papers Share of 5859 ----------------------- ------ ------------- title starts SoK / SOK: 43 0.7% slug starts sok- 43 0.7% union (the population) 43 0.7% title-only (not slug) 0 0.0% slug-only (not title) 0 0.0% title contains word SoK anywhere: 43; of which outside the union: 0 ========================================================================== B. SHAPE OF THE 43 EXTRACTED SOKS ========================================================================== Venue Extracted SoKs Extracted papers SoKs / venue ------- -------------- ---------------- ------------ CCS 0 990 0.0% IEEE-SP 30 767 3.9% IMC 0 638 0.0% NDSS 4 701 0.6% PETS 2 510 0.4% USENIX 7 1410 0.5% WWW 0 843 0.0% venues with zero extracted SoKs: CCS, IMC, WWW ── by year (2025–2026 starred as provisional) ── Year Extracted SoKs Extracted papers that year ----- -------------- -------------------------- 2010 0 119 2011 0 116 2012 0 151 2013 2 125 2014 0 166 2015 1 190 2016 2 182 2017 2 231 2018 1 254 2019 1 402 2020 2 404 2021 3 379 2022 4 546 2023 4 719 2024 8 690 2025* 7 770 2026* 6 415 ── IEEE-SP share of extracted SoKs ── IEEE-SP SoKs: 30 of 43 = 69.8% ── platforms (a paper can have several; shares of 43) ── platform SoKs Share of 43 -------------------- ---- ----------- offline 26 60.5% other-online-service 11 25.6% web 6 14.0% not-applicable 5 11.6% mobile 3 7.0% iot 1 2.3% platforms includes web: 6 of 43 = 14.0% (platform=web is not the topic fold — printers and HPC counters sit here) ========================================================================== C. TOPIC HAND FOLD (43 papers, 0 residue by construction) ========================================================================== Topic Papers Share of 43 --------------------------------------------------------------------------------------------------------- ------ ----------- Web measurement (crawls, live-site audits, phishing-site detectors, privacy-regulation impact on the web) 4 9.3% Adjacent online privacy / abuse / traffic analysis 13 30.2% Not this field (TEE, binaries, fuzzing, crypto, hardware, DeFi, …) 26 60.5% ── web_measurement ── 2024 IEEE-SP SoK: Technical Implementation and Human Impact of Internet Privacy Regulations. 2024 USENIX SoK: State of the Krawlers – Evaluating the Effectiveness of Crawling Algorithms for Web Security Measurements 2025 PETS SoK: Web Authentication and Recovery in the Age of End-to-End Encryption 2026 USENIX SoK: PHILTER: Uncovering Security and Functional Gaps in AI-based Phishing Website Detection Literature via an LLM-based Reasoning Framework ── web_adjacent ── 2016 IEEE-SP SoK: Towards Grounding Censorship Circumvention in Empiricism. 2020 PETS SoK: Anatomy of Data Breaches 2021 IEEE-SP SoK: Hate, Harassment, and the Changing Landscape of Online Abuse. 2022 IEEE-SP SoK: A Framework for Unifying At-Risk User Research. 2022 IEEE-SP SoK: Social Cybersecurity. 2022 IEEE-SP SoK: The Dual Nature of Technology in Sexual Abuse. 2023 IEEE-SP SoK: A Critical Evaluation of Efficient Website Fingerprinting Defenses. 2024 IEEE-SP SoK: Safer Digital-Safety Research Involving At-Risk Users. 2024 USENIX SoK (or SoLK?): On the Quantitative Study of Sociodemographic Factors and Computer Security Behaviors 2025 IEEE-SP SoK: A Privacy Framework for Security Research Using Social Media Data. 2025 IEEE-SP SoK: A Framework and Guide for Human-Centered Threat Modeling in Security and Privacy Research. 2025 IEEE-SP SoK: Self-Generated Nudes over Private Chats: How can Technology Contribute to a Safer Sexting? 2025 IEEE-SP SoK: Decoding the Enigma of Encrypted Network Traffic Classifiers. ── not_web ── 2013 IEEE-SP SoK: Secure Data Deletion. 2013 IEEE-SP SoK: P2PWNED - Modeling and Evaluating the Resilience of Peer-to-Peer Botnets. 2015 IEEE-SP SoK: Deep Packer Inspection: A Longitudinal Study of the Complexity of Run-Time Packers. 2016 IEEE-SP SOK: (State of) The Art of War: Offensive Techniques in Binary Analysis. 2017 IEEE-SP SoK: Cryptographically Protected Database Search. 2017 IEEE-SP SoK: Exploiting Network Printers. 2018 IEEE-SP SoK: Keylogging Side Channels. 2019 IEEE-SP SoK: The Challenges, Pitfalls, and Perils of Using Hardware Performance Counters for Security. 2020 IEEE-SP SoK: Understanding the Prevailing Security Vulnerabilities in TrustZone-assisted TEE Systems. 2021 IEEE-SP SoK: All You Ever Wanted to Know About x86/x64 Binary Disassembly But Were Afraid to Ask. 2021 IEEE-SP SoK: Quantifying Cyber Risk. 2022 IEEE-SP SoK: How Robust is Image Classification Deep Neural Network Watermarking? 2023 IEEE-SP SoK: Decentralized Finance (DeFi) Attacks. 2023 IEEE-SP SoK: History is a Vast Early Warning System: Auditing the Provenance of System Intrusions. 2023 IEEE-SP SoK: Taxonomy of Attacks on Open-Source Software Supply Chains. 2024 IEEE-SP SoK: Prudent Evaluation Practices for Fuzzing. 2024 USENIX SoK: All You Need to Know About On-Device ML Model Extraction - The Gap Between Research and Practice 2024 USENIX SoK: What Don't We Know? Understanding Security Vulnerabilities in SNARKs 2024 IEEE-SP SoK: SGX.Fail: How Stuff Gets eXposed. 2025 USENIX SoK: An Introspective Analysis of RPKI Security 2025 IEEE-SP SoK: Software Compartmentalization. 2026 NDSS SoK: Take a Deep Step into Linux Kernel Hardening Effectiveness from the Offensive-Defensive Perspective 2026 NDSS SoK: Understanding the Fundamentals and Implications of Sensor Out-of-band Vulnerabilities 2026 NDSS SoK: Cryptographic Authenticated Dictionaries 2026 USENIX SoK: Security of Cyber-physical Systems Under Intentional Electromagnetic Interference Attacks 2026 NDSS SoK: Analysis of Accelerator TEE Designs web_measurement SoKs: 4 of 43 = 9.3% ========================================================================== D. SEARCH-METHOD HAND FOLD (primary label, 43 papers) ========================================================================== How the paper list was built Papers Share of 43 ------------------------------------------------------------- ------ ----------- Venue-complete (or venue-list) then filter 7 16.3% Digital-library / DBLP keyword query 3 7.0% Google Scholar and/or citation snowball 8 18.6% Not a paper-census (system, evaluation, vulnerability corpus) 19 44.2% Literature sample claimed; search protocol not stated 6 14.0% claims a literature sample: 24 of 43 = 55.8% …and states a search protocol: 18 of 24 = 75.0% venue-complete then filter: 7 of 43 = 16.3% (of the 18 with a stated protocol: 38.9%) ── web_measurement SoKs × search method ── Paper Year Venue Search --------------------------------------------------------------------------------------------------------------------------------------- ---- ------- ----------------- Technical Implementation and Human Impact of Internet Privacy Regulations. 2024 IEEE-SP scholar_snowball State of the Krawlers – Evaluating the Effectiveness of Crawling Algorithms for Web Security Measurements 2024 USENIX venue_then_filter Web Authentication and Recovery in the Age of End-to-End Encryption 2025 PETS venue_then_filter PHILTER: Uncovering Security and Functional Gaps in AI-based Phishing Website Detection Literature via an LLM-based Reasoning Framework 2026 USENIX database_keyword ========================================================================== E. INDEX FUNNEL — SOKS IN THE SEVEN VENUES, NOT JUST THE 43 ========================================================================== index records: 16864 index SoKs (title-start or slug-start): 167 Venue Index SoKs Extracted SoKs ------- ---------- -------------- IEEE-SP 84 30 PETS 43 2 USENIX 36 7 NDSS 4 4 index SoKs in CCS, IMC, WWW: 0, 0, 0 Stage SoKs Lost here --------------------------- ---- --------- Index (title or slug) 167 — …with an abstract 163 4 …labelled (screened) 163 0 …selected by the --any rule 47 116 …extracted 43 4 no-abstract index SoKs (4): IEEE-SP/2026/sok-all-you-ever-wanted-to-know-about-bootloader-security-but-were-afraid-to-ask IEEE-SP/2026/sok-critical-evaluation-of-quantum-machine-learning-for-adversarial-robustness IEEE-SP/2026/sok-robustness-in-large-language-models-against-jailbreak-attacks IEEE-SP/2026/sok-after-decades-of-web-tracker-detection-whats-next selected but not extracted (4): IEEE-SP/2024/sok-a-comprehensive-analysis-and-evaluation-of-docker-container-attack-and-defen USENIX/2026/sok-history-doesnt-repeat-itself-but-android-design-level-vulnerabilities-rhyme USENIX/2026/sok-darpas-ai-cyber-challenge-aixcc-competition-design-architectures-and-lessons USENIX/2026/sok-capability-operating-systems-is-the-future-finally-here Rieder index record present: IEEE-SP/2026/sok-after-decades-of-web-tracker-detection-whats-next abstract present: false in labels: false in extractions: false fulltext dir exists: false ── PETS index SoKs vs extracted (most dropped SoKs are crypto / DP / sensing) ── PETS index SoKs 43; extracted 2 ========================================================================== F. HOMOGRAPHS AND FULL-TEXT SOK MENTIONS ========================================================================== title contains "survey": 11 of 5859 …of which are not SoKs: 11 …of those, studyTypes includes user-study or interview-or-survey: 7 title matches systematic (literature) review: 1 ── non-SoK "survey" titles (these are mostly user surveys, not literature surveys) ── 2016 CCS user/interview=true How I Learned to be Secure: a Census-Representative Survey of Security Advice Sources and Behavior. 2018 IEEE-SP user/interview=false Surveylance: Automatically Detecting Online Survey Scams. 2019 IEEE-SP user/interview=true How Well Do My Results Generalize? Comparing Security and Privacy Survey Results from MTurk, Web, and Telephone Samples. 2020 WWW user/interview=false Attention Please: Your Attention Check Questions in Survey Studies Can Be Automatically Answered. 2022 PETS user/interview=false A Global Survey of Android Dual-Use Applications Used in Intimate Partner Surveillance 2022 WWW user/interview=true Beyond Bot Detection: Combating Fraudulent Online Survey Takers✱. 2023 PETS user/interview=true Privacy Concerns and Acceptance Factors of OSINT for Cybersecurity: A Representative Survey 2024 IEEE-SP user/interview=true Digital Security - A Question of Perspective A Large-Scale Telephone Survey with Four At-Risk User Groups. 2024 USENIX user/interview=true Bridging Barriers: A Survey of Challenges and Priorities in the Censorship Circumvention Landscape 2024 USENIX user/interview=true Engaging Company Developers in Security Research Studies: A Comprehensive Literature Review and Quantitative Survey 2026 USENIX user/interview=false Inconsistent, Incomplete, and Insecure: A Survey of Account Security Interfaces full-text /systematization of knowledge/i: 26 papers; missing .cols: 4 …not in the title/slug union (citation / passing mention): 15 extracted SoKs whose .cols does NOT contain the phrase: 32 of 43 ========================================================================== G. KRAWLERS KEYWORD CUT REPRODUCED ON THIS EXTRACTION ========================================================================== Krawlers (USENIX 2024) quoted figures, checked against paper.cols.txt: OK "7,840" OK "1,057" OK "654" OK "403" OK "32.3%" OK "tranco, alexa, and the combination of top and site" OK Warford "6,534 papers" OK Warford "127 potentially relevant papers" OK Warford "reviewed 95 papers" OK Warford "139 potentially relevant papers" OK Blessing "245 papers" OK Blessing "top-300" OK Usman "more effective than keyword searching in a database" OK Usman "SOUPS, CHI, CSCW, USENIX Security, IEEE Security & Privacy, and PETS" OK Alam "In total, 55 papers" OK Alam "a total of 38 papers" OK Alam "17 additional papers" OK Birrell "ten selected venues" extracted papers 2010–2022: 3265 crawled papers 2010–2022: 687 (denominator for recall) Krawlers keyword on crawled 2010–2022 Papers Share of 687 ------------------------------------------------------------------ ------ ------------ tranco 56 8.2% alexa 351 51.1% tranco OR alexa (the two named lists) 375 54.6% neither tranco nor alexa 312 45.4% top AND site (both words; generous reading of their third keyword) 478 69.6% any of the three (union under that generous reading) 539 78.5% NONE of the three 148 21.5% missing .cols 0 0.0% stricter /top[\s-]*site/i on crawled 2010–2022: 162 (not used as the published recall; the paper said "combination of top and site") keyword union on all extracted 2010–2022: 1493 hits; of which NOT crawled: 954 = 63.9% (Krawlers discarded 654 of 1,057 = 61.9% as not employing automated crawling; our extraction is already screened, so this is a lower bound on the false-positive rate) ── crawled 2010–2022 missed by all three keywords (first 20 of the miss list) ── CCS/2010/detecting-and-characterizing-social-spam-campaigns Detecting and characterizing social spam campaigns. CCS/2010/fingerprinting-websites-using-remote-traffic-analysis Fingerprinting websites using remote traffic analysis. IMC/2010/detecting-and-characterizing-social-spam-campaigns Detecting and characterizing social spam campaigns. WWW/2010/analyzing-content-level-properties-of-the-web-adversphere Analyzing content-level properties of the web adversphere. WWW/2010/diversifying-landmark-image-search-results-by-learning-interested-views-from-com Diversifying landmark image search results by learning interested views from community pho WWW/2010/the-social-honeypot-project-protecting-online-communities-from-spammers The social honeypot project: protecting online communities from spammers. WWW/2010/restler-crawling-restful-services RESTler: crawling RESTful services. CCS/2011/automated-black-box-detection-of-side-channel-vulnerabilities-in-web-application Automated black-box detection of side-channel vulnerabilities in web applications. CCS/2011/android-permissions-demystified Android permissions demystified. IMC/2011/analyzing-facebook-privacy-settings-user-expectations-vs-reality Analyzing facebook privacy settings: user expectations vs. reality. WWW/2011/arrow-generating-signatures-to-detect-drive-by-downloads ARROW: GenerAting SignatuRes to Detect DRive-By DOWnloads. CCS/2012/detecting-money-stealing-apps-in-alternative-android-markets Detecting money-stealing apps in alternative Android markets. IMC/2012/evolution-of-a-location-based-online-social-network-analysis-and-models Evolution of a location-based online social network: analysis and models. IMC/2012/evolution-of-social-attribute-networks-measurements-modeling-and-implications-us Evolution of social-attribute networks: measurements, modeling, and implications using goo NDSS/2012/you-are-what-you-like-information-leakage-through-users-interests You are what you like! Information leakage through users’ Interests USENIX/2012/enemy-of-the-state-a-state-aware-black-box-web-vulnerability-scanner Enemy of the State: A State-Aware Black-Box Web Vulnerability Scanner CCS/2013/delta-automatic-identification-of-unknown-web-based-infection-campaigns Delta: automatic identification of unknown web-based infection campaigns. WWW/2012/understanding-and-combating-link-farming-in-the-twitter-social-network Understanding and combating link farming in the twitter social network. WWW/2012/economics-of-bittorrent-communities Economics of BitTorrent communities. WWW/2012/analyzing-spammers-social-networks-for-fun-and-profit-a-case-study-of-cyber-crim Analyzing spammers' social networks for fun and profit: a case study of cyber criminal eco … 128 more crawled 2010–2022 whose .cols matches /\bcrux\b/i: 7 (Krawlers dropped the keyword because CrUX rankings started in 2022) ========================================================================== H. REPORTING-RATE TRAP — SEARCHING FOR THE THING UNDERCOUNTS THE SILENCE ========================================================================== crawled papers: 1120 crawlConfig.headless stated (not a sentinel): 140 = 12.5% full-text /\bheadless\b/i over crawled: 137 = 12.2%; missing .cols 0 word "headless" but schema sentinel/absent: 28 (mentions in related work, or a different sense) schema-stated but word not in .cols: 31 headless membership disagreement 59 papers (28 word-only, 31 schema-only); totals differ by 3 trap: a Scholar search for "headless crawler" returns papers that used the word; it cannot estimate the 87.5% of crawls that never said (980 of 1120). ========================================================================== I. SIBLING HOMOGRAPHS RE-DERIVED ========================================================================== detection.phenomenon names a fingerprint: 280 of 5859 …browser-device family: 83 = 29.6% of the fingerprint papers …traffic-analysis family: 105 = 37.5% keyword precision for browser fingerprinting if you search this corpus for "fingerprint": 29.6% ── preregistration homograph (PREREG_RE over full text, sense from prereg_fold.mjs) ── PREREG_RE hits: 62 (missing .cols 4) …study preregistration sense: 15 …other sense (pre-registered domain / OAuth / FIDO / …): 47 = 75.8% of the keyword hits ========================================================================== J. EXTERNAL FIGURES (non-corpus; each with its primary source) ========================================================================== IEEE S&P SoK track: "As in past years, we solicit systematization of knowledge (SoK) papers" IEEE S&P 2026 CFP: https://sp2026.ieee-security.org/cfpapers.html (fetched 2026-08-27) "Submissions will be distinguished by the prefix “SoK:” in the title" "Survey papers without such insights are not appropriate and may be rejected without full review." oaklandsok.github.io: SoK track at IEEE S&P since 2010; EuroS&P since 2017; PETS since 2019; USENIX Security since 2024; NDSS since 2026; SaTML since 2023. CCS, IMC, WWW: not listed. Fetched 2026-08-27. First SoK-titled paper in THIS index: 2013 (0 in 2010–2012). Vekaria et al. arXiv:2506.14057 "SoK: Advances and Open Problems in Web Tracking" abs page fetched 2026-08-27: still a preprint; living version at github.com/privacysandstorm/sok-advances-open-problems-web-tracking Their method (from the HTML): "seven top web security and privacy venues" over 20 years (IEEE S&P, USENIX Security, ACM CCS, NDSS, ACM IMC, PETS, WWW) — the same seven. "A total of 200+ research papers were identified." Not in this extraction (wrong venue: arXiv). Rieder et al. IEEE S&P 2026 / arXiv:2605.02982 listed on oaklandsok.github.io. In IEEE-2026.json; no abstract, so never screened, so not in the 5,859. Krawlers discarded 654 of 1,057 = 61.9% of keyword hits as not automated crawls. 130 of 403 = 32.3% navigate beyond a single page. USENIX Security 2024 Krawlers landing: /conference/usenixsecurity24/presentation/stafeev Blessing DOI 10.56553/popets-2025-0113 (fetch_authors.py failed on parse_popets; authors from citation_author meta) drain item writing:literature_review (new) id 178 run 55 claimed as cursor-drain-litrev CCS 2026 and IMC 2026 have not been held; IEEE-SP 2026 and WWW 2026 abstracts are not in OpenAlex (two more 2026 venue-years incompletely selected). Source: literature:corpus / extract README, quoted. review-log tokens: FIG-01 FIG-02 FIG-03 FIG-04 FIG-05 CIT-01 CIT-02 CIT-03 CIT-04 VEKARIA-01 CFP-01 TRACKS-01 TRACKS-02 GITHUB-01 KRAWLERS-01 BLESSING-01 LINKS-01 gpt-5.6-luna-max gpt-5.6-sol-medium frozen 12:53:21Z reviewer ids 254169ad-89ee-4fdd-8d5c-5cf7a2bd990f bd1ebd8e-9534-4acf-9a89-4ed058490491 29c63fa2-4842-43fb-a227-161e420f6d2b ========================================================================== Z. ARITHMETIC-DERIVED (so a page figure is grep-able here even when it is a quotient) ========================================================================== 43 extracted SoKs of 5859 = 0.7% IEEE-SP 30 of 43 = 69.8% web_measurement 4 of 43 = 9.3% index 167 SoKs → extracted 43 Krawlers 654/1057 = 61.9% Krawlers 130/403 = 32.3% keyword miss among crawled 2010-2022: 148/687 = 21.5% neither tranco nor alexa: 312/687 = 45.4% tranco OR alexa: 375/687 = 54.6% crawled 2010-2022: 687 crawled any year matching /\btranco\b/i: 194 of 1120 = 17.3% (keyword_denominator.py demo) crawled any year NOT matching /\btranco\b/i: 926 = 82.7% headless never stated 980 of 1120 = 87.5% keyword_denominator.py \bheadless\b: 137 match, 983 miss = 12.2% / 87.8% Warford 6534 / 127 / 12 / 139 / 95; Blessing 245; Alam 38 DBLP + 17 snowball = 55; Birrell ten selected venues then citation expansion; Alexa top-300 PETS 41 of 43 index SoKs not extracted (43 index, 2 extracted) literature:corpus funnel (quoted, not re-derived this sitting): 16,864 index → 15,800 abstracts → 6,103 selected → 5,873 PDFs → 5,859 extracted privacy:requests population S1 ∪ S2 = 254; overlap both = 107 (quoted from that page's report, 2026-08 corpus) keyword_denominator.py pct formula uses 100 * n / d survey titles 11 of 5859; not SoK 11; user/interview-tagged 7 index SoKs 167; selected 47; extracted 43; no-abstract 4; selected-not-extracted 4 PETS 43 index SoKs, 2 extracted search: venue_then_filter 7, database_keyword 3, scholar_snowball 8, not_a_census 19, not_stated 6 census 24 of 43; stated protocol 18 of 24 full-text SoK phrase 26; citation-only 15; extracted SoKs without the phrase 32 of 43 web platform tag 6 of 43 year 2024 has 8 extracted SoKs, the peak complete year headless stated 140/1120 = 12.5% browser FP precision 83/280 = 29.6% prereg homograph 47/62 = 75.8% OK: no failures.
Review log
Focused drafts frozen at 2026-08-27T12:53:21Z under out/frozen_litrev/ (sha256 in SHA256SUMS). Three focused passes were requested as luna@medium; the Task tool rejected gpt-5.6-luna-max and there is no luna-medium slug in this session. Fallback: gpt-5.6-sol-medium (the only GPT 5.6 slug the tool accepts). Recorded so this is not a silent substitute.
Reviewer agent ids (Task tool): figures 254169ad-89ee-4fdd-8d5c-5cf7a2bd990f; citations bd1ebd8e-9534-4acf-9a89-4ed058490491; external 29c63fa2-4842-43fb-a227-161e420f6d2b.
| Id | Pass | Finding | Decision |
|---|---|---|---|
| FIG-01 | figures vs script | “disagree by three papers” — totals differ by 3; membership differs on 59 (28 word-only, 31 schema-only) | Accepted. Prose now says totals vs membership. |
| FIG-02 | figures vs script | “Among those 18” headed a table of all 43 | Accepted. Now “Across all 43”. |
| FIG-03 | figures vs script | Krawlers' 403 called “crawls that named a popular list” — third keyword is “top”+“site” | Accepted. |
| FIG-04 | figures vs script | missing .cols counted and computation continued | Accepted. Report now fail()s if crawled 2010–2022 or the headless sweep has missing full text. Current missing counts are 0. |
| FIG-05 | figures vs script | “two more 2026 venue-years are incompletely selected” not printed by the report | Accepted. Printed in section J, quoted from literature:corpus (IEEE-SP 2026 and WWW 2026 abstracts not in OpenAlex). |
| CIT-01 | citations | Alam “55 papers from DBLP” — DBLP yielded 38, snowball +17 | Accepted. Needles in .cols. |
| CIT-02 | citations | Warford 6,534→127→95 omitted the +12→139 pool | Accepted. |
| CIT-03 | citations | Birrell “snowball inside ten venues” — seeds were ten venues; snowball admitted outside-venue papers | Accepted. |
| CIT-04 | citations | same as FIG-03 | Accepted. |
| VEKARIA-01 | external | still arXiv v1; optional note of IEEE S&P 2026 poster / USENIX under-review from author CV | Rejected the USENIX-under-review claim (CV is not a primary venue source). Preprint hedge already on the page. Poster not added without a primary URL fetched this sitting. |
| CFP-01, TRACKS-01/02, GITHUB-01, KRAWLERS-01, BLESSING-01, LINKS-01 | external | live facts match; hedge CCS/IMC/WWW as “not listed on oaklandsok” | Accepted the hedge on the venue table sentence and the TODO. Other live checks: no change. |
Generic pass (requested gpt-5.6-luna-max; same Task-tool fallback) is below, after the applied page, with no checklist.
