User Tools

Site Tools


provenance:privacy:policies

Table of Contents

Provenance: privacy:policies

Working log behind Measuring Privacy Policies and Terms. Corpus-wide caveats — which venues are in, how papers were selected, what the extraction gets wrong — are on corpus and are not restated here. Citations use the shared bibliography; this page adds no entries of its own. No ~~DISCUSSION~~ block: comments belong on the content page.

Voice here is a working log, not prose. It is read by somebody checking a number.

The run

Item Value
Dates Drafted 2026-09-09; verification, review and publication 2026-09-10. The 2026-09-09 sitting was cut off mid-draft by an API quota stop, so parts of this log were reconstructed on 2026-09-10 from the committed scripts, their outputs and the page draft rather than written as the decisions were made. Where a decision's reasoning could not be recovered, this page says so rather than inventing one
Process failures in this run Two, both mine. The pages were published before the generic review returned, so the first published revision carried the twelve defects that pass listed (including the PrivaSeer error) for about half an hour. And I edited the content page while that reviewer was reading it, which is why its report opens by saying the page moved under it. Neither is how this should go: publish after the last pass, and freeze the file while a reviewer holds it
Corpus data/extract/run1/extractions.jsonl, 5,859 papers with extracted full text; 7 venues (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P), 2010–2026
Page status New page. Before writing: a wiki search for “privacy policy”, “OPP-115”, “Polisis” and “privacy label” found no content page with a policy section. OPP-115 and Polisis occurred only inside provenance logs. consent covers the banner, mobile_and_app_measurement covers store listings, legal_enforcement covers what regulators did — none covers the document. So this is a creation, not an extension
Page id privacy:policies, not privacy:privacy_policies. Decided 2026-09-07 on roadmap (see roadmap §3) because privacy_policies sorts next to the site's own privacy_policy page in search. scripts/sitemap.mjs already gated on privacy:policies and the roadmap's Queued table already promised it, so any other id would have left a dangling promise
Models Page, scripts and this log: Claude (Opus 5, with the 2026-09-09 draft written by the same model in an earlier session). Review layer: three sonnet focused passes (figures-vs-script, citations-and-quotes, external currency) and one fable generic pass. Findings and verdicts below
Scripts added scripts/report_policies.mjs (+ -output.txt), scripts/policy_fold.mjs, scripts/policies_fulltext_probe.mjs (+ -output.txt), scripts/policies_quotecheck.mjs (+ -output.txt), scripts/policies_significance.py (+ -output.txt), scripts/policies_external_checks.sh (+ -output.txt), scripts/policies_gh_search.py, scripts/policies_w3c_p3p_check.mjs (+ -output.txt), scripts/policies_table_check.mjs (+ -output.txt), scripts/policies_fetch_pets_authors.py, scripts/bib_additions_policies.bib, scripts/build_provenance_policies.py
Write path node scripts/dw.mjs put (JSON-RPC) with –if-rev on every save
Accidental exposure None. Credentials stayed in .env and were never echoed. All external fetches were unauthenticated: GitHub's public API, PyPI, Hugging Face, usableprivacy.org, privaseer.ist.psu.edu, developer.apple.com, support.google.com, w3.org, secartifacts.github.io

Scope and judgement calls

Decision Why What a reasonable person might have done instead
New page rather than a section on consent A cookie banner is a UI measurement — you click it and watch what changes. A privacy policy is a document-retrieval and NLP measurement. The two literatures share almost no method and barely cite each other; the corpus enum separates them too (consent-notice 39 papers, privacy-policy 102, and the overlap is small) Fold policies into the consent page as “the long version of the notice”. Rejected: the reader who needs to retrieve and label 100,000 documents would find a page about clicking buttons
Scope the page to web and mobile, and say so in the first screen The platform split is mobile 60 / web 56 of the 123 UNION papers (43 mobile-only, 39 web-only, 17 both). A web-only page would silently drop the tool lineage, which was built for Android, and half the availability table Write a web-only page and send app policies to mobile_and_app_measurement. Rejected: PolicyLint→PoliCheck→PoliGraph is the spine of this literature and it is Android work
POLICY (102) is the population; UNION (123) is a candidate set classification[].target is a structured enum, so 102 needs no folding and is stable between extraction runs. The union depends on a regex I chose. Figures denominated on 123 do appear on the page — the UNION share columns of the year and venue tables, the full-text probe table, and the sentences quoting probe rows — and every one is either labelled UNION in its column header or says it is a candidate set. No rate about the field is over 123: the method, validation and significance tables are all over the 102. An earlier draft claimed the page carried no 123-denominated figures at all, which was false; the generic review caught it Publish rates over 123 because it is the larger, more intuitive number. Rejected — that is the “mention threshold is a candidate set” failure
Widened the title probe from four alternatives to nine The 2026-09-02 gap analysis used privacy polic|privacy notice|terms of service|privacy label and got a union of 121. Adding terms and conditions, data safety, nutrition label, policy text and privacy statement takes it to 123, and the papers it adds are the store-declaration half of the literature ([1Ali, Mir Masood; Balash, David G.; Kodwani, Monica; Kanich, Chris; Aviv, Adam J. (2024): "Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps' Privacy Policies", in: Proceedings on Privacy Enhancing Technologies. (DOI)]-adjacent Data-safety and nutrition-label work). Probe width decides the claim, so both widths are recorded below Keep the narrow probe for comparability with the gap analysis. Rejected: the narrow probe misses a family the page is about, and the gap analysis was a scoping exercise, not a published figure
Dated the method table by era, not by year 102 papers over 12 years cannot carry a per-year method trend; four era buckets (2014–2018 n=7, 2019–2021 n=21, 2022–2024 n=50, 2025–2026 n=24) each have enough papers to read Publish per-year percentages. Rejected: n=1 and n=3 years would produce 100% and 0% cells
Called supervised OPP-115 classifiers “declining, still the reproducible baseline” rather than superseded supervised-ml falls 42.9% → 12.5% across the eras, but PrivBERT is used by 3 papers all in 2024 and OPP-115 by 6 papers in 2024+. A method still in use is not superseded Call them historical, matching the LLM narrative. Rejected on the counts
Called readability-as-a-headline “historical” 17 of 123 papers measure readability, but the last paper whose contribution is a readability finding is [2Amos, Ryan; Acar, Gunes; Lucherini, Eli; Kshirsagar, Mihir; Narayanan, Arvind; Mayer, Jonathan (2021): "Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset", in: Proceedings of the Web Conference 2021, pp. 2165–2176. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] (2021) and [3Adhikari, Andrick; Das, Sanchari; Dewri, Rinku (2023): "Evolution of Composition, Readability, and Structure of Privacy Policies over Two Decades", in: Proceedings on Privacy Enhancing Technologies. (DOI)] (2023); after that it is a descriptive table inside a larger study. This is a judgement from reading, not a count — flagged as such Leave it undated. Rejected: a fresh student re-running Flesch–Kincaid on 10,000 policies as a paper is the exact mistake this page exists to prevent
Called P3P dead Verified against the primary W3C sources, not recalled — see External sources Omit P3P entirely as ancient history. Rejected: it is the obvious “why not just make it machine-readable?” question and a student deserves the answer
LLM extraction called “current” on 12 papers 11 of the 24 papers in the provisional 2025–2026 slice, against 1 in 2022–2024. The page labels the slice provisional in the table itself and says the claim rests on the corpus's thinnest years Wait for a complete 2026. Rejected: this is the one thing about this literature that a 2024-trained reader would get wrong
No page-level “what fraction of the web has a privacy policy” The availability table spans 4.35% to 94.2% across ecosystems and the page says the spread is the finding. There is no defensible single number Quote the [4Cui, Hao; Trimananda, Rahmadi; Markopoulou, Athina (2025): "Understanding Privacy Norms through Web Forms", in: Proceedings on Privacy Enhancing Technologies. (DOI)] 94.2% as “the web figure”. Rejected: its population is sites with a personal-information-collecting web form, not the web
Cited OPP-115 and PrivaSeer although both are outside the seven venues They are the substrate under most of the corpus's own supervised work — 14 corpus papers use OPP-115. The page says explicitly that they are outside the corpus and that PrivaSeer is named by zero corpus papers Restrict the page to corpus artefacts. Rejected: it would send a student to build a corpus that already exists
No ~~DISCUSSION~~ on this provenance page Established default across the provenance: namespace

Populations, and every query behind a figure

Three populations. They are not interchangeable and no table on the content page mixes them.

Name Definition n What it may be used for
corpus every extracted paper 5,859 denominators for corpus base rates only
classified classification[] non-empty 4,439 (75.8% of corpus) the denominator for “2.3% of papers that classified anything classified policy text”
POLICY classification[].target === “privacy-policy” — a structured enum 102 every method, validation and rate figure on the page
PROBE title+summary regex, recall-oriented 77 never published alone; reported so the page can state how much each side misses
UNION POLICYPROBE 123 rankings and “does this literature do X” only. The one exception is the full-text probe table, whose header says it is a candidate set

Overlap: POLICYPROBE = 56; POLICY only = 46; PROBE only = 21. The 46 enum-only papers are the reason a title probe alone is not enough — they are Alexa-skill, IoT-companion-app and VR audits whose policy analysis is one component of a wider study.

report_policies.mjs asserts three invariants at load and throws rather than drifting: POLICY ⊆ UNION, PROBE ⊆ UNION, and UNION has no duplicate venue/year/slug keys.

Figure on the page Population Query
102, 2.3% classified classification[].target === “privacy-policy”, papers
123, 56, 46, 21 set arithmetic printed by §1 of the report
179 tuples POLICY tuple count, printed so the page can say it counts papers not tuples
per-year table, 2014–2026, both share columns corpus per year §1 of the report. The UNION share column is a share of a candidate set and its header says so
per-venue table, both share columns that venue's corpus slice §1 of the report, which also prints the top-venue ratio (PETS is 4.8x the next venue on the POLICY share and 5.5x on the UNION share) so the page does not have to divide two percentages by eye
platform split 39 web / 43 mobile / 17 both UNION population[].platform, multi-valued
method-by-era table POLICY, split into four era buckets §2 of the report; era n printed in the column header, and the four buckets partition the 102
classification.validation rows privacy-policy tuples in POLICY §2; none-reported is printed as a real value, never subtracted into “states a value”
tool lineage counts whole corpus, not UNION — a tool used outside the union is still a use §4, regex per artefact against tools[].name, otherToolsMentioned[].name, classification[].resourceName/targetDetail, population[].sourceList, detection[].phenomenon/technique
every measured result (availability table, consistency table, text-description bullets) the paper's own denominator, quoted as the paper states it §5 of the report, then re-checked needle-by-needle against the paper's own text — see Quotes
five Fisher's exact rows UNION or POLICY vs the corpus base rate, both printed policies_significance.py
full-text probe table UNION (123), labelled a candidate set policies_fulltext_probe.mjs

The probes, at both widths

Title+summary probe (defines PROBE, and with it UNION). Nine alternatives, case-insensitive, tested against title • summary:

privacy polic|privacy notice|terms of service|terms and conditions|privacy label|data safety|nutrition label|policy text|privacy statement

The 2026-09-02 gap analysis that proposed this page used four alternatives (privacy polic|privacy notice|terms of service|privacy label) and reported a union of 121. The nine-alternative form gives 123. Both are recorded because probe width decides the claim; the page quotes the nine-alternative figure and says so.

Whole-corpus lineage scan (added 2026-09-10, same script). For each of the fifteen artefacts in the lineage, a case-insensitive scan of all 5,869 readable paper.cols.txt files, counting papers whose text names it. This is a different question from the extraction fold in §4 of the report and gives systematically larger answers (Polisis 13 by extraction, 68 in text; PolicyLint 18 and 63). The distinction is the fix for this run's worst error: “no corpus paper's extraction records PrivaSeer” is true, “nobody in these venues cites PrivaSeer” is false, and only the full-text scan can tell the two apart. One artefact needs a different pattern in full text than in the extraction — MAPS collides with Google Maps and the Play Store's Maps & Navigation category, so the full-text pattern is the paper's title. The override and its reason are printed by the script; the shared regex list lives in policy_fold.mjs so the two scripts cannot drift.

Full-text probes (policies_fulltext_probe.mjs): thirteen questions, each with a narrow and a wide pattern, run over the 123 UNION papers' paper.cols.txt. The page quotes the narrow form throughout and labels it an upper bound on a candidate set — a probe counts papers whose text contains a phrase, not papers that did the thing. The wide form is printed beside it in the output below so a reader can see how far the answer moves: “reports dead policy links” is 7 narrow and 64 wide, and the page's “only 7 of 123” claim would be a different claim at the wide width. Three of the thirteen are not narrow/wide pairs but different questions, and the output marks them so.

Folding, and the complete residue

scripts/policy_fold.mjs. Two folds, both ordered — first match wins — and both leaving anything unmatched as raw residue that report_policies.mjs prints in full.

  • population[].sourceList → 13 families (app stores, ranking lists, the four annotated corpora, archives, participant panels, other app ecosystems, and an honest custom / hand-built list bucket). Residue: 138 distinct strings, printed in full in §3 of the output below.
  • tools[].name + otherToolsMentioned[].name → 23 families (the policy-analysis lineage, the NLP substrate, browser automation, boilerplate strippers, language detection, readability metrics). Residue: 572 distinct names, printed in full in §4.

Known limits of the fold, recorded rather than hidden:

  • The tool fold is deliberately narrow. It maps the policy lineage and the NLP substrate and leaves general-purpose tools unmapped, because guessing at families for 572 names would produce a table nobody could audit. That is why the residue is large.
  • participant panel is folded as its own family and kept out of every “where policies come from” table — Prolific is where the participants came from, not where the policies came from.
  • Alexa list (the ranking) is separated from Alexa Skills Store (the voice-app marketplace) by a negative lookahead. Strings like Alexa skill marketplaces satisfy neither rule and land in the residue, where they are visible.
  • MAPS is matched with an anchored, case-sensitive /^maps$/ to avoid folding the word “maps”; a paper writing “MAPS pipeline” falls to the residue.
  • No folded free-text count is published as a percentage anywhere on the page. They appear as rankings only.

Quotes and figures checked against the papers

scripts/policies_quotecheck.mjs. Every measured figure the page prints from a paper is a needle taken from the paper's own words, not from the extraction's evidence.quote — the point is to catch an extraction error, so checking the extraction against itself would prove nothing. Each needle is searched in four renderings: paper.cols.txt, paper.norm.txt, paper.txt and a pypdf extraction of paper.pdf cached under cache/pypdf/. Whitespace, soft hyphens, curly quotes and the several Unicode dashes are normalised before matching.

Result: 70 needles, 70 located, 0 missing. 69 in paper.cols.txt; 1 only outside it — [2Amos, Ryan; Acar, Gunes; Lucherini, Eli; Kshirsagar, Mihir; Narayanan, Arvind; Mayer, Jonathan (2021): "Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset", in: Proceedings of the Web Conference 2021, pp. 2165–2176. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]'s “corpus are from 2009-2019” is spliced by the de-columning and was found only by pypdf. A .cols-only check would have reported a false MISS on it. Full output below.

Seventeen of the 70 needles were added on 2026-09-10. Ten came from the whole-page number guard, which flagged ten figures the page printed that no earlier needle covered: [5Pan, Shidong; Zhang, Dawen; Staples, Mark; Xing, Zhenchang; Chen, Jieshan; Xu, Xiwei; Hoang, Thong (2024): "Is It a Trap? A Large-scale Empirical Study And Comprehensive Assessment of Online Automated Privacy Policy Generators for Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)]'s 10,375/46,472, 9,523/46,472 and 15.7%; [6Andow, Benjamin; Mahmud, Samin Yaseer; Whitaker, Justin; Enck, William; Reaves, Bradley; Singh, Kapil; Egelman, Serge (2020): "Actions Speak Louder than Words: Entity-Sensitive Privacy Policy and Data Flow Analysis with PoliCheck", in: Proceedings of the USENIX Security Symposium. (Link)]'s 37.1% and 31.1% (14,409/45,603); [1Ali, Mir Masood; Balash, David G.; Kodwani, Monica; Kanich, Chris; Aviv, Adam J. (2024): "Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps' Privacy Policies", in: Proceedings on Privacy Enhancing Technologies. (DOI)]'s 228,539 and n=306,404; [7Xiang, Anhao; Pei, Weiping; Yue, Chuan (2023): "PolicyChecker: Analyzing the GDPR Completeness of Mobile Apps' Privacy Policies", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]'s 98.1%; [8Zimmeck, Sebastian; Wang, Ziqi; Zou, Lieyong; Iyengar, Roger; Liu, Bin; Schaub, Florian; Wilson, Shomir; Sadeh, Norman; Bellovin, Steven M.; Reidenberg, Joel (2017): "Automated Analysis of Privacy Requirements for Mobile Apps", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]'s mean of 1.83; and the two [2Amos, Ryan; Acar, Gunes; Lucherini, Eli; Kshirsagar, Mihir; Narayanan, Arvind; Mayer, Jonathan (2021): "Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset", in: Proceedings of the Web Conference 2021, pp. 2165–2176. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] year-span needles. Five more came from the citations review, which found two availability rows naming a different population from the paper's own (see Review). Two more came from the generic review's back-calculated-denominator finding. All seventeen check out. This is the guard doing its job: the figures were right, but nothing had verified them.

Needle specificity, and what this check does not prove. Locating a needle proves the string is in the right PDF. It does not prove the string is the sentence the page is quoting: a bare 26% is in nine of these papers, and the generic review showed that changing the Degeling needle from 84.5 % to 84.9 % still passed, because both appear in that paper's tables. The check now prints a specificity report — how many other check papers contain each needle, and how many needles are numeric-only or under twelve characters — and nine of the thirteen weak needles it found were rewritten to include the paper's surrounding words. Four remain weak and are printed in the output: 1,071,488, 84.7, 54.5% and 1,035,853 each also occur in one other check paper. Each was read in context; none is load-bearing on its own.

Two things the quote check specifically caught or settled:

  • [5Pan, Shidong; Zhang, Dawen; Staples, Mark; Xing, Zhenchang; Chen, Jieshan; Xu, Xiwei; Hoang, Thong (2024): "Is It a Trap? A Large-scale Empirical Study And Comprehensive Assessment of Online Automated Privacy Policy Generators for Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)] states both “37.5% (37,150/99,194) of privacy policy links lead to unavailable websites” and “15.7% (15,572/99,194) of apps do not provide a privacy policy” — 99,194 is used as a denominator of links in one sentence and of apps in another. The page uses the paper's own wording for each figure and does not reconcile them.
  • An earlier draft of the page claimed “two needles needed a pypdf rendering”. The committed output showed zero. The claim was removed on 2026-09-10; after the ten new needles it is genuinely one, and the page now says one. A sentence about a check is a figure like any other.

Bibliography

The page cites 38 keys. 14 were already in bibliography and were reused unchanged; 24 were added in two rounds — 22 with the first publication, and 2 more after the generic review asked for citekeys on four papers the page had described in prose. Two of those four turned out to be in the bibliography already under other keys ([9Khatun, Mst Eshita; Noureddine, Lamine; Bello, Sideeq; Ali-Gombe, Aisha (2026): "Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale", Proceedings on Privacy Enhancing Technologies 2026(4):213-231. (DOI)] and [10Nenadić, Luka; Rodriguez, David; Calandrino, Joseph A. (2026): "Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact", Proceedings on Privacy Enhancing Technologies 2026(4):703-723. (DOI)]); the key-string check passed them and bib_dedup_scan.py caught them on DOI and title. They are in scripts/bib_additions_policies.bib, appended before the closing </bibtex> of a freshly exported copy of the live page (not a local snapshot — the local copy in this workdir was already stale by one page's worth of additions on the morning of 2026-09-10).

Check Result
Keys used on the page that resolve after both appends 38 of 38
New keys colliding with an existing key string 0
Definite duplicates by DOI or squashed title (bib_dedup_scan.py over the merged file) 0 in the first round of 22 (909 entries). 2 in the second round of 4 — both caught and dropped, the page repointed at the existing keys (911 entries)
Candidate duplicate pairs (rule C/D) touching a new key 4, all judged distinct by hand: cui2025_odyssey/cui2025_privacy (Cui, Jian vs Cui, Hao), wu2025_appprivacyreport/wu2025_depth (Wu, Xiaoyuan vs Wu, Yuhao), wu2025_depth/wu2025_revealing (Wu, Yuhao vs Wu, Mengying), zimmeck2017_automated/zimmeck2017_privacy (same first author, two different 2017 papers)
Literal ASCII @ inside any field of a new entry 0 — one would silently drop the entry and every marker to it
Stray non-BibTeX lines in the additions file 0. Checked by stripping every @entry{…} block and asserting the remainder is blank — a generator's QA chatter has gone live on the bibliography page before
DOI or landing URL present 14 entries carry a DOI (CCS, IMC, PoPETs), 10 carry a USENIX or NDSS landing URL. The four PoPETs 2026 DOIs were each resolved through doi.org and each returned HTTP 200 at its petsymposium.org landing page

Authors for PETS records. The corpus index carries no authors and no DOI for any PETS or USENIX record (100% of both venues). scripts/fetch_authors.py fails on every PETS landing page as of 2026-09-09 — the pages return HTTP 200 but its parser expects a byline layout petsymposium.org no longer serves. scripts/policies_fetch_pets_authors.py was written for this run: it reads the citation_author meta tags from the landing page with curl and a browser User-Agent, refuses to overwrite a cached entry, and fails loudly rather than guessing when no meta tag is present. USENIX authors were taken from the paper PDF, not from usenix.org's own metadata or DBLP — both are known to drop authors from long author lists.

Guards run before publication

Guard On what Result
node scripts/check_wrap.mjs both pages OK
node scripts/check_tables.mjs both pages OK — every table has one width
python3 scripts/check_wrapped_lists.py both pages OK. This is the guard for the DokuWiki rule that an indented continuation line under a bullet renders as preformatted text and swallows the rest of the list
node scripts/check_attributions.mjs content page + merged bibliography 0 attributions checked, in both table and prose mode — which is NOT a pass. This page cites by citekey without naming authors in prose, so the guard has nothing to match. Recorded rather than reported as green
node scripts/check_page_numbers.mjs the whole content page against the concatenation of all five outputs OK — every figure traces. Run whole-page, not windowed: a windowed run cannot see figures in the introduction or the Related Pages section, which is how 29 stale figures once survived a refresh
node scripts/policies_table_check.mjs the page's four corpus tables, cell by cell, against the report and the significance output OK — 11 method rows, 12 year rows, 7 venue rows, 5 significance rows. Written on 2026-09-10 because the number guard is a membership test: a mutation changing the llm count from 12 to 77 passed it, since 77 occurs elsewhere in the output
bib_dedup_scan.py live bibliography + the 22 new entries (909 entries) 0 definite duplicates; 4 candidate pairs touching a new key, all judged distinct by hand
build_provenance_policies.py structural assertions this page, at generation time 16 opening file tags = 16 closing, an even number of inline nowiki delimiters in the prose, no unescaped discussion macro, ≥10 level-2 headings

What the number guard does not prove: it checks that each numeral on the page appears somewhere in the script output, not that it appears in the right sentence, and it prints only the first occurrence's context. A stale figure that happens to collide with a live one passes it — demonstrated, not assumed: changing the llm row from 12 to 77 passes the number guard. policies_table_check.mjs was written to close that hole for the four tables it can re-derive, and it catches that mutation and six others. Everything outside those four tables still rests on reading; that is how the availability-denominator findings were caught, not by a guard.

The number guard's ALLOW map is shared across every page on this wiki, so whitelisting a value here would silently bless it elsewhere. Nothing was added to it for this page. The four figures that would have needed an entry were instead given real evidence: two became quote-check needles against the paper text, and two are now printed by policies_external_checks.sh.

External sources

scripts/policies_external_checks.sh re-fetches every external fact the content page states; its unedited output is at the foot of this page. Two rules it encodes, both from earlier mistakes on this wiki: print %{http_code} and the effective URL before any byte count, because a 302 recorded as an empty body has been published here as “empty 200”; and never date a repository from /releases/latest, /tags or pushed_at — ask the commits API on the repository's own default branch.

Source How it was verified Verdict
OPP-115 and APP-350 (usableprivacy.org/data) Fetched 2026-09-10, HTTP 200. Both entries present; OPP-115 = 115 policies with a documented commercial-licence path via CMU Flintbox; APP-350 = 350 app policies with the same research wording but no commercial path Accepted, with the licence asymmetry stated on the page rather than glossed as “same licence”
PrivaSeer corpus size and licence Fetched privaseer.ist.psu.edu/data, HTTP 200; the script prints the site's own sentence, “The PrivaSeer corpus is a collection of 3,967,487 privacy policies”, and “the corpus is available under a CC BY-NC-SA license” Accepted. Note the site advertises a different, smaller figure for its live search index; that is a different object and the page quotes the corpus
mukund/privbert on Hugging Face Fetched, HTTP 200 Accepted
Princeton–Leuven repositories GitHub commits API on each default branch: citp/privacy-policy-historical master, last commit 2023-10-12; citp/PrivacyPoliciesOverTime master, last commit 2022-06-06; no licence file on either Accepted. The earlier draft named only one of the two repositories; both are now named
PolicyLint / PoliCheck repository benandow/PrivacyPolicyAnalysis, master, last commit 2022-10-05. GitHub reports the licence as “Other”; LICENSE.txt fetched directly is a three-clause BSD naming NC State Accepted, with “GitHub says Other, the file says BSD-3” stated rather than picking one
PoliGraph UCI-Networking-Group/PoliGraph, master, MIT, last commit 2023-06-21. PyPI queried for both poligraph and poligraph-er: HTTP 404 for each Accepted
PoliGraph “needs a GPU” README fetched: “A GPU is required to enable hardware acceleration… Note that PoliGraph-er can run without a GPU, but the performance would be significantly lower” Corrected. The draft said “needs a GPU”; the page now states what the README states
Lalaine, Calpric, PolicyChecker repositories xiaoyue10131748/Lalaine (MIT, 2023-09-25), dlgroupuoft/Calpric (no licence, 2023-06-21), AndyXiang945/PolicyChecker (no licence, 2023-11-20) Accepted
PurPliance repository ducalpha/PurPlianceOpenSource, main, last commit 2024-03-11 — but that commit is “Update README.rst”; the last code commit is 2022-12-25. LICENSE.txt is BSD-style, U. Michigan 2022 Accepted, and it corrected the page. The draft said every tool “stopped being maintained within a year of its paper” and called PoliGraph the newest; PurPliance is newer, and the page now says so and distinguishes a README edit from a code commit
“Polisis has no official repo; the maintained lineage is a community reproduction, last commit 2023-02-02” GitHub name search returns quanmou/polisis (master last commit 2020-07-27) and SmartDataAnalytics/Polisis_Benchmark (pushed_at 2023-02-02, but master last commit 2020-02-13). Every branch of the latter was enumerated: the 2023-02-02 activity is on dependabot/pip/werkzeug-0.15.5 Rejected as written. pushed_at counts activity on any branch, and a Dependabot security bump is not maintenance. The page now gives the default-branch dates
“Calpric, Lalaine and PolicyChecker are effectively undiscoverable; only reachable through the USENIX artifact index” Tested. secartifacts.github.io/usenixsec2023/results (HTTP 200) lists Calpric and Lalaine but not PolicyChecker, which is a CCS paper. GitHub name search returns the right repository first for PolicyChecker, Calpric, Lalaine and PoliGraph — but for PolicyLint and PoliCheck it returns only unrelated projects, because the repository is named after the paper series Rejected and replaced. The claim was true of the wrong tools. PolicyChecker was removed from it; PolicyLint/PoliCheck were added, and the page now says which search finds what
w3.org/P3P has not changed since 2007” w3.org answers curl with a Cloudflare HTTP 403, so this was checked with Playwright's chromium (scripts/policies_w3c_p3p_check.mjs). w3.org/P3P/ returns 200 with Last-Modified: Fri, 02 Feb 2018; its newest news item is “3 October 2007: The Policy Languages Interest Group (PLING) was created” and its footer reads “Last updated $Date: 2018/02/02” Corrected. “Unchanged since 2007” was wrong; the page now states both dates and what each one is
W3C P3P 1.1 retirement Same Playwright fetch of w3.org/TR/P3P11/ (HTTP 200): the visible text contains “Retired 30 August 2018” and “should not be referenced in this form or implemented as-is”, both asserted by the script Accepted
Apple privacy manifests developer.apple.com/news/?id=pvszzano fetched, HTTP 200, dated 26 April 2024, effective 1 May 2024, containing “will expand to include the entire app binary”. The reviewer additionally fetched Apple's live upcoming requirements page and found no newer deadline superseding it Accepted, still current
Google Play Data safety and the User Data policy support.google.com/…/answer/10144311 fetched, HTTP 200, still containing “Apps that do not access any personal and sensitive user data must still submit a privacy policy.” verbatim Accepted for the quote. The “mandatory since 20 July 2022” date is not on the current Google page — it rests on contemporaneous reporting, and is flagged below as something the primary source no longer states
ADPC status Search only, no primary fetch: still a proposal, not in production Accepted as a weak claim, and it carries no figure
Any SEO listicle, vendor blog or “top 10 privacy policy tools” page Rejected on sight. None was consulted or cited

What could not be established

  • Google's own current pages do not state when Data safety became mandatory. The 20 July 2022 date rests on contemporaneous third-party reporting. The quote the page uses is on the live primary page; the date is not. Closing this needs an archived snapshot of the Play Console announcement.
  • Whether the visible text of w3.org/P3P/ has changed since 2007. Only the server's Last-Modified (2018-02-02) and the page's own newest news item (2007-10-03) could be established; the Wayback Machine rate-limited both attempts to diff them.
  • How much of the 16%–94.2% availability spread is ecosystem and how much is method. Nobody has run five link-detection heuristics over one site sample. This is on the page as an open question, not resolved.
  • Whether LLM extraction actually beats PoliGraph or a PrivBERT baseline. No head-to-head evaluation on the same documents with the same ground truth was found in these seven venues. The page says the switch is currently a preference, not a finding.
  • Anything about venues outside the seven. OPP-115 (ACL 2016), APP-350 and PrivaSeer (ACL 2021) are cited as artefacts because corpus papers use them, but ACL, EMNLP, CHI, SOUPS and the law reviews are not in the corpus and no count here covers them.
  • The ~20% free-text run-to-run stability and 0.9% unlocatable-quote rates quoted in this project's documentation were measured on the previous 4,322-paper extraction run and have not been re-measured on this one. They are used here only as an order of magnitude, and are the reason nothing on the content page publishes a percentage over a folded free-text field.
  • Figures deliberately not published: no single “fraction of the web with a privacy policy”; no per-year method trend (the era buckets are used instead); no percentage over any folded free-text field; no rate over the 123-paper UNION except the full-text probe table and the two prose sentences that quote its largest row, all three labelled.

Review

Four reviewers, all told explicitly that the author's context might not be exhaustive, and all handed the page text, every script and its unedited output. The three focused passes ran in parallel on the first revision; the generic pass ran afterwards on the corrected page and on this log.

Pass 1 — figures against the script (''sonnet'')

Re-ran all four scripts and diffed them against the committed outputs: byte-identical, no drift. Then checked every numeral on the page.

# Finding Verdict
1 The Fisher's-exact table states rates over the 123-paper UNION, which the page's own methodology section says never happens outside the probe table Accepted. The comparisons were moved to POLICY (102). While fixing it, two further improvements: the counts are no longer hand-keyed — report_policies.mjs now emits SIG| lines that policies_significance.py parses — and the base rate now excludes the subgroup, since comparing a set against a population containing it shrinks the difference. The validation row's p moved 0.545 → 0.613 and is still not significant
2 “65% of the papers in this literature run a policy-versus-behaviour comparison” is a full-text probe result stated twice as fact, outside the table that carries the caveat Accepted. Both sentences now say “80 of the 123 candidate papers” and name it as an upper bound, and the methodology bullet now lists all three places a 123-denominated figure appears
3 The EU availability row prints 6,579 while the paper's own availability table totals 6,357 Accepted, and the fix went further than the finding: the paper states three numbers (6,759 January domains, 6,579 in its abstract, 6,357 in Table II). The page's cell now says which is which, and two quote-check needles were added
4 “99,194 policy links” is really 99,194 apps Accepted. Both the page and the script's hand-keyed availability map now say “usable apps”, and the page points out that the paper divides link failures by its app count
5 “Every measured figure above was checked… 51 needles” overstates the coverage: several published figures had no needle Accepted. Ten needles were added (pan2024_trap ×3, andow2020_actions ×3, ali2024_honesty ×2, xiang2023_policychecker, zimmeck2017_automated), then five more from pass 2. The check now runs 68 needles, 68 located
6 The method table silently drops the other (8) and curated-database (1) rows Accepted. Both restored. A dropped enum row is exactly where something hides
7 The per-year table starts at 2019, silently dropping 8 UNION papers Accepted. 2014, 2016, 2017 and 2018 restored

The pass also independently re-derived a long list of figures it found correct, including the whole tool-lineage table, the per-venue table, the platform split and the probe table.

Pass 2 — citations and quotes (''sonnet'')

Independently re-extracted the page's citekeys and quoted spans rather than working from a supplied list, and verified all 22 new BibTeX entries against the corpus index, the paper PDFs and the PETS landing pages.

# Finding Verdict
1 The page says two PETS 2026 papers compare policy text against store-declared labels; [11Cory, Thomas; Rieder, Wolf; Krämer, Julia; Raschke, Philip; Herbke, Patrick; Küpper, Axel (2026): "Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Language Models", Proceedings on Privacy Enhancing Technologies 2026(1):509-528. (DOI)] does not — it is word-level GDPR-transparency annotation, and privacy label, data safety and nutrition label appear zero times in it Accepted. Re-checked directly (0 hits for each phrase against 108 for GDPR as a positive control). The sentence now names only the papers that do this, and cory2026_wordlevel was moved to the LLM row, where it belongs
2 MAPS's 50.5% is over the 1,035,853 analysed apps, not the 1,049,790 retrieved Accepted. Verified against the paper (“our analysis reveals that only 50.5% of apps have links”, in a section operating on the analysed set). Page and script both corrected, plus two needles
3 “across CS, NLP and law” is not the interview paper's own discipline breakdown Accepted. The page now gives the paper's own breakdown
4 The Degeling denominator is internally inconsistent in the source paper; flagged for awareness, not as a page defect Accepted as information, and merged with pass 1's finding 3

It also confirmed: 34 of 34 citekeys resolve, 0 key collisions, 0 DOI or title duplicates against the live bibliography, and all 22 new entries correct in author order, venue, year and identifier — including three where the PDF's front matter splits the byline across columns and a careless reader would drop authors.

Pass 3 — external currency (''sonnet'')

Fetched rather than recalled. Its findings and their verdicts are in the External sources table above rather than duplicated here; in summary it corrected the P3P currency claim, the PoliGraph GPU claim, the APP-350 licence symmetry and the artifact-discoverability footnote, confirmed every dataset host, every repository date and licence, and both platform-policy quotes, and could not verify two things now listed under What could not be established.

One of its corrections was itself incomplete and was corrected in turn: it proposed replacing the discoverability footnote with a claim about Calpric and Lalaine, but a direct GitHub name search showed the genuinely unfindable pair is PolicyLint and PoliCheck, whose repository is named after neither tool. A reviewer finding is a lead, not a verdict.

Pass 4 — generic (''fable'')

No checklist. It ran on the corrected page and on this log, and it was the most productive of the four. Its own disclosures, recorded because they matter: one of its mutation-test commands was denied, so a later sed ran against the committed scripts/policies_quotecheck.mjs instead of a copy; it reverted the change and said so. Verified independently afterwards — the script reproduces its committed output byte for byte. It also noted that the content page was edited while it read, which is true and is a process failure of mine, recorded below.

# Finding Verdict
1 “PrivaSeer is named by zero corpus papers” and “Nobody in these seven venues cites it” are false: 12 corpus papers name it in their full text, including four the page itself cites Accepted, and it is the worst error the four reviews found. Re-derived independently: a scan of all 5,869 paper.cols.txt finds exactly 12. The zero was a fold over tools/sourceList — a fact about the extraction, published as a fact about the literature. Fixed by adding a second count column to the lineage table (extraction vs full text) and a whole-corpus full-text scan to the probe script, and by rewording all three places. The scan also found a homonym: /\bMAPS\b/ matches “Google MAPS” and the Play category “Maps & Navigation”, so the full-text pattern for MAPS is its title, and the override is printed
2 The page publishes UNION-denominated shares (the “share of corpus” and “share of venue” columns, and 70/123 in an open question) while claiming it never does Accepted. The report now prints a POLICY share beside every UNION share, both are on the page with UNION named in the column header, the open question uses 64/102, and the methodology bullet now enumerates the 123-denominated figures instead of denying they exist. The earlier claim was written when the Fisher table was the only offender and was not revisited after pass 1 fixed that one
3 Probe counts stated as facts, several with “at all” Accepted. Six sentences rewritten to name the probe and its width. “Only 7 papers say anything at all” became “a narrow probe matches 7 of 123, the wide form 64, and the gap is how much probe width decides the answer”
4 The Discord denominator 15,528 is arithmetic on the paper's own non-reconciling numbers; the paper states 15,525 Accepted. The page now gives 15,525 as the paper states it and says the paper's own 14,852 does not reconcile. This is a back-calculated denominator, which is a named failure mode here
5 heuristic-rules peaked at 40% in 2022–2024” contradicts the table two screens above (42.9% in 2014–2018) Accepted, reworded
6 “three of the four corpora are still downloadable” while every row says live Accepted, “all four were still reachable on 2026-09-10”
7 This log was stale after pass 2 (63 needles, “per-year 2019–2026”) Accepted, corrected
8 [[roadmap]] on this page resolves to provenance:privacy:roadmap and renders red Accepted, and confirmed in the rendered DOM of the first published revision: exactly one wikilink2 span. Now [[:roadmap]]
9 The Apple December 2020 date is not primary-sourced while the parallel Google date is flagged Accepted, flagged in the same way
10 Generator guards: deleting two whole sections passed (floor of ≥10 headings), emptying an output passed, and the file-block count assertion was tautological Accepted, all three. The heading check now names the thirteen expected sections; each output must be over 200 bytes, contain a known terminal string, and be newer than the script that produced it; and OUTPUTS is compared against glob('scripts/policies_*-output.txt') rather than against itself. All five mutations are now caught
11 Quote-check needles are weak: 26 of 68 are bare numbers, and a wrong digit passes Accepted. See Quotes above: a specificity report was added and nine needles rewritten. Four remain weak and are printed
12 The 46 enum-only papers are characterised from memory as “mostly Alexa-skill, IoT and VR” Accepted. The report now lists all 46 with their platform enum (other-online-service 21, mobile 20, web 18, iot 4, offline 1) and the page says what is actually there
13 “Half of this literature is about Android apps” — the store-label half is iOS-heavy Accepted, “mobile apps”
14 “Both are current and both are enforced” — no source for enforcement Accepted. The page now says both requirements are current, that neither store publishes enforcement figures, and to treat “required” as a rule rather than a fact about the population
15 “every tool in the lineage is abandoned code” overstates last-commit dates Accepted, softened to unmaintained code, with the distinction stated
16 Four corpus papers referred to by description with no citekey, so a reader cannot find them Accepted. Keys added for all four. Two of them turned out to be already in the bibliography under other keys ([9Khatun, Mst Eshita; Noureddine, Lamine; Bello, Sideeq; Ali-Gombe, Aisha (2026): "Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale", Proceedings on Privacy Enhancing Technologies 2026(4):213-231. (DOI)], [10Nenadić, Luka; Rodriguez, David; Calandrino, Joseph A. (2026): "Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact", Proceedings on Privacy Enhancing Technologies 2026(4):703-723. (DOI)]) — caught by bib_dedup_scan.py, not by the key-string check, which is exactly what that scan exists for
17 “four in 2024–2026” store-label comparison papers is a reading, not a count Accepted, the number removed and the reason given
18 “the two most-starred” reproductions — the search is by relevance, not stars Accepted, “the top two hits”
19 Licence wordings are tighter than the sources: OPP-115 says “in the spirit of” CC BY-NC and its commercial licence covers the annotation files Accepted for OPP-115, APP-350 and PrivaSeer. The PolicyLint and PurPliance licence descriptions were left as they are: the external-check output prints the first lines of both LICENSE.txt files and the page says what they are
20 [12Chanenson, Jake; Pickering, Madison; Apthorpe, Noah (2025): "Automating Governing Knowledge Commons and Contextual Integrity (GKC-CI) Privacy Policy Annotations with Large Language Models", in: Proceedings on Privacy Enhancing Technologies. (DOI)]'s comparison arm is a custom RNN, not an OPP-115-trained classifier, so “the two to copy” misdescribes it Accepted, split into one model and one companion read
21 A footnote published this page's own draft history Accepted, moved here
22 P3P did define a well-known location, which the page's “no fixed address” bullet invites as an objection Accepted, one clause added
Two gaps it named but did not call defects: no reusable retrieval code is pointed at, and nothing on the cost of LLM extraction at scale Not fixed. Both are real. The retrieval gap is already the page's first open question; the cost question has no corpus source and would be a vendor-price claim with a shelf life of months. Recorded here rather than guessed at

Mutation tests of the guards this page publishes

Reading a guard does not tell you whether it asserts anything. Both the generic reviewer and I broke things on purpose in /tmp copies and checked that the guard failed.

Mutation Guard Result
A needle that is nowhere in the paper policies_quotecheck.mjs caught
A needle attributed to the wrong paper policies_quotecheck.mjs caught
A needle with the wrong digit (84.584.9) policies_quotecheck.mjs not caught — both strings are in that paper. Hence the specificity report
An empty needle, or % policies_quotecheck.mjs not caught. Left as it is: the specificity report now makes short needles visible
llm count 12 → 77 (collides with a live value) check_page_numbers.mjs not caughtpolicies_table_check.mjs written, which catches it
An era percentage changed to another value from the same table policies_table_check.mjs caught
A method row silently deleted policies_table_check.mjs caught
A venue share changed to a value used elsewhere policies_table_check.mjs caught
A significance base swapped for the naive base rate policies_table_check.mjs caught
A p-value exponent moved from 10⁻⁸ to 10⁻⁵ policies_table_check.mjs not caught at first — the check compared three digits and ignored the exponent. Now compares the numbers, and catches it
Odd inline nowiki count in the prose build_provenance_policies.py caught
An unescaped discussion macro build_provenance_policies.py caught
A published script containing a closing file tag build_provenance_policies.py caught
Two whole sections deleted from the prose build_provenance_policies.py not caught (floor of ≥10 headings) → now names all thirteen, and catches it
An output file emptied, or truncated build_provenance_policies.py not caught → now caught
An output dropped from the published list build_provenance_policies.py not caught (tautological) → now compared against the filesystem, and caught
A script edited without re-running it build_provenance_policies.py not caught → now caught by an mtime comparison

Seven of seventeen mutations survived the guards as first written. That ratio is the argument for mutation-testing every published check rather than reading it.

The scripts, as committed

Every block below is the file itself, inserted by scripts/build_provenance_policies.py at build time — not a sample, not an abridgement. Re-running the generator re-inserts whatever is on disk.

The report script — ''report_policies.mjs''

report_policies.mjs
// Every figure on privacy:policies, with its denominator.
//
//   node scripts/report_policies.mjs > scripts/report_policies-output.txt
//
// THREE populations are used and they are NOT interchangeable. Every table
// says which one it is on, and no table mixes them:
//
//   POLICY   classification[].target === 'privacy-policy'   — a STRUCTURED ENUM.
//            "the extractor recorded this paper as classifying privacy-policy
//            text". No folding needed, stable between runs. This is the page's
//            primary population.
//   PROBE    /privacy polic|.../ over title + summary only  — a CANDIDATE SET,
//            recall-oriented, NOT a population. Reported so the page can say how
//            much the enum misses and vice versa. Never published as a rate.
//   UNION    POLICY ∪ PROBE — used only for "which tools/sources appear in this
//            literature at all" rankings, never for a percentage.
//
// Counts are of PAPERS, never tuples. Sentinels are printed but never counted
// as an answer. Free-text sourceList and tool names are folded through
// policy_fold.mjs and the unmapped residue is printed in full at the end.
 
import { loadExtractions, pct, table, isSentinel } from './lib.mjs';
import { foldSource, foldTool, LINEAGE } from './policy_fold.mjs';
 
const P = loadExtractions();
const key = (p) => `${p.venue}/${p.year}/${p.slug}`;
 
// The probe. Widened from the 2026-09-02 brainstorm's four alternatives to nine:
// 'privacy label', 'data safety', 'nutrition label' and 'privacy statement' are
// the store-declaration half of this literature and the narrower probe missed
// them. Recorded here because probe width decides the claim.
const PROBE_RE =
  /privacy polic|privacy notice|terms of service|terms and conditions|privacy label|data safety|nutrition label|policy text|privacy statement/i;
const titleSummary = (p) => `${p.title ?? ''} • ${p.summary ?? ''}`;
 
const POLICY = P.filter((p) => p.classification.some((c) => c.target === 'privacy-policy'));
const PROBE = P.filter((p) => PROBE_RE.test(titleSummary(p)));
const UNION = [...new Set([...POLICY, ...PROBE])];
const CLASSIFIED = P.filter((p) => p.classification.length > 0);
 
const h = (s) => console.log(`\n${'='.repeat(78)}\n${s}\n${'='.repeat(78)}`);
const sub = (s) => console.log(`\n--- ${s}`);
 
// Sanity invariant: every published population must be reproducible from the
// definitions above. If the corpus moves, this throws rather than drifting.
if (POLICY.some((p) => !UNION.includes(p))) throw new Error('POLICY not a subset of UNION');
if (PROBE.some((p) => !UNION.includes(p))) throw new Error('PROBE not a subset of UNION');
if (UNION.length !== new Set(UNION.map(key)).size) throw new Error('UNION has duplicate keys');
 
// ------------------------------------------------------------ 1. populations
h('1. POPULATIONS');
console.log(`corpus                                                   ${P.length}`);
console.log(`classified   classification[] non-empty                  ${CLASSIFIED.length}   ${pct(CLASSIFIED.length, P.length)} of corpus`);
console.log(`POLICY       classification[].target=='privacy-policy'   ${POLICY.length}   ${pct(POLICY.length, CLASSIFIED.length)} of classified`);
console.log(`PROBE        title+summary regex (candidate set)         ${PROBE.length}`);
console.log(`UNION        POLICY u PROBE                              ${UNION.length}`);
console.log(`  POLICY n PROBE                                         ${POLICY.filter((p) => PROBE.includes(p)).length}`);
console.log(`  POLICY only (enum fires, title silent)                 ${POLICY.filter((p) => !PROBE.includes(p)).length}`);
console.log(`  PROBE only (title fires, enum silent)                  ${PROBE.filter((p) => !POLICY.includes(p)).length}`);
console.log(`\nprivacy-policy tuples in POLICY                          ${POLICY.reduce((n, p) => n + p.classification.filter((c) => c.target === 'privacy-policy').length, 0)}`);
 
// The enum-only papers are the argument for not using a title probe alone, so
// print them rather than characterising them from memory. A 2026-09-10 review
// found the page describing this set as "mostly Alexa-skill, IoT and VR studies"
// on no evidence; the list below is the evidence.
sub('POLICY-only: the enum fires and the title probe is silent (why a title probe is not enough)');
{
  const only = POLICY.filter((p) => !PROBE.includes(p))
    .sort((a, b) => a.year - b.year || a.venue.localeCompare(b.venue));
  console.log(`${only.length} papers. Platform measured (enum, multi-valued):`);
  const plat = {};
  for (const p of only) for (const x of new Set(p.platforms)) plat[x] = (plat[x] ?? 0) + 1;
  console.log(table(['platform', 'papers of the enum-only set'],
    Object.entries(plat).sort((a, b) => b[1] - a[1])));
  if (!Object.keys(plat).length) throw new Error('enum-only platform table came out empty');
  for (const p of only) console.log(`  ${p.year} ${p.venue.padEnd(8)} ${p.title}`);
}
 
sub('classification[].target, whole corpus, papers (enum — publishable)');
{
  const m = {};
  for (const p of CLASSIFIED) for (const t of new Set(p.classification.map((c) => c.target))) m[t] = (m[t] ?? 0) + 1;
  console.log(table(['target', 'papers', 'share of 4,439 classified'],
    Object.entries(m).sort((a, b) => b[1] - a[1]).map(([k, v]) => [k, v, pct(v, CLASSIFIED.length)])));
}
 
sub('POLICY and UNION per year (2026 PROVISIONAL: CCS/IMC 2026 not held, IEEE S&P/WWW 2026 under-selected)');
{
  const y = {};
  for (const p of UNION) {
    y[p.year] ??= { u: 0, e: 0, t: 0 };
    y[p.year].u += 1;
    if (POLICY.includes(p)) y[p.year].e += 1;
    if (PROBE.includes(p)) y[p.year].t += 1;
  }
  const allY = {};
  for (const p of P) allY[p.year] = (allY[p.year] ?? 0) + 1;
  // Both shares are printed. UNION/corpus is a share of a regex-widened
  // candidate set and must be labelled as such wherever it is published;
  // POLICY/corpus is the share of the stable enum.
  console.log(table(['Year', 'corpus', 'POLICY', 'PROBE', 'UNION', 'POLICY share of corpus', 'UNION share of corpus'],
    Object.keys(y).sort().map((k) => [k + (k >= '2025' ? '*' : ''), allY[k], y[k].e, y[k].t, y[k].u,
      pct(y[k].e, allY[k]), pct(y[k].u, allY[k])])));
}
 
sub('UNION per venue (denominator: that venue\'s whole corpus slice)');
{
  const v = {}, vAll = {};
  for (const p of P) vAll[p.venue] = (vAll[p.venue] ?? 0) + 1;
  for (const p of UNION) v[p.venue] = (v[p.venue] ?? 0) + 1;
  console.log(table(['Venue', 'UNION', 'POLICY', 'venue papers', 'POLICY share of venue', 'UNION share of venue'],
    Object.entries(v).sort((a, b) => b[1] - a[1]).map(([k, n]) => {
      const pol = POLICY.filter((p) => p.venue === k).length;
      return [k, n, pol, vAll[k], pct(pol, vAll[k]), pct(n, vAll[k])];
    })));
  // The page says "PETS by a factor of five over the next venue"; print the ratio
  // rather than leaving the reader to divide two percentages by eye.
  {
    const byPolicy = Object.keys(vAll)
      .map((k) => [k, POLICY.filter((p) => p.venue === k).length / vAll[k]])
      .sort((a, b) => b[1] - a[1]);
    const byUnion = Object.entries(v).map(([k, n]) => [k, n / vAll[k]]).sort((a, b) => b[1] - a[1]);
    console.log(`top venue over the next, POLICY share: ${byPolicy[0][0]} / ${byPolicy[1][0]} = ${(byPolicy[0][1] / byPolicy[1][1]).toFixed(1)}x`);
    console.log(`top venue over the next, UNION share : ${byUnion[0][0]} / ${byUnion[1][0]} = ${(byUnion[0][1] / byUnion[1][1]).toFixed(1)}x`);
  }
}
 
sub('UNION by platform measured (multi-valued: an app+web paper is in two rows)');
{
  const m = {};
  for (const p of UNION) for (const x of new Set(p.platforms)) m[x] = (m[x] ?? 0) + 1;
  console.log(table(['platform', 'papers', 'share of UNION'],
    Object.entries(m).sort((a, b) => b[1] - a[1]).map(([k, v]) => [k, v, pct(v, UNION.length)])));
  const web = UNION.filter((p) => p.platforms.includes('web'));
  const mob = UNION.filter((p) => p.platforms.includes('mobile'));
  console.log(`\nweb only   ${web.filter((p) => !p.platforms.includes('mobile')).length}`);
  console.log(`mobile only ${mob.filter((p) => !p.platforms.includes('web')).length}`);
  console.log(`both        ${web.filter((p) => p.platforms.includes('mobile')).length}`);
  console.log(`neither     ${UNION.filter((p) => !p.platforms.includes('web') && !p.platforms.includes('mobile')).length}`);
}
 
// ------------------------------------------------- 2. how policies are labelled
h('2. HOW THE FIELD LABELS POLICY TEXT  (population: POLICY, n=' + POLICY.length + ')');
console.log('classification[].method is an ENUM. Multi-valued: a paper with two');
console.log('privacy-policy tuples using two methods appears in two rows.');
 
const polTuples = (p) => p.classification.filter((c) => c.target === 'privacy-policy');
const methodSet = (p) => new Set(polTuples(p).map((c) => c.method));
 
sub('method, all years');
{
  const m = {};
  for (const p of POLICY) for (const x of methodSet(p)) m[x] = (m[x] ?? 0) + 1;
  console.log(table(['method', 'papers', 'share of POLICY'],
    Object.entries(m).sort((a, b) => b[1] - a[1]).map(([k, v]) => [k, v, pct(v, POLICY.length)])));
}
 
// ERAS. 2025-2026 is starred everywhere it appears.
const ERAS = [
  ['2014-2018', (y) => y <= 2018],
  ['2019-2021', (y) => y >= 2019 && y <= 2021],
  ['2022-2024', (y) => y >= 2022 && y <= 2024],
  ['2025-2026*', (y) => y >= 2025],
];
 
sub('method by era — this is the CURRENCY table the page leans on');
{
  const methods = [...new Set(POLICY.flatMap((p) => [...methodSet(p)]))];
  const rows = methods.map((mm) => {
    const cells = ERAS.map(([, f]) => {
      const sub2 = POLICY.filter((p) => f(p.year));
      const n = sub2.filter((p) => methodSet(p).has(mm)).length;
      return `${n} (${pct(n, sub2.length)})`;
    });
    const total = POLICY.filter((p) => methodSet(p).has(mm)).length;
    return [mm, total, ...cells];
  }).sort((a, b) => b[1] - a[1]);
  console.log(table(['method', 'all', ...ERAS.map(([l]) => `${l} n=${POLICY.filter((p) => ERAS.find(([ll]) => ll === l)[1](p.year)).length}`)], rows));
}
 
sub('classification[].validation on privacy-policy tuples (enum; none-reported is a REAL value, not a gap in the data)');
{
  const m = {};
  for (const p of POLICY) for (const x of new Set(polTuples(p).map((c) => c.validation))) m[x] = (m[x] ?? 0) + 1;
  console.log(table(['validation', 'papers', 'share of POLICY'],
    Object.entries(m).sort((a, b) => b[1] - a[1]).map(([k, v]) => [k, v, pct(v, POLICY.length)])));
  // Same row for the whole classified corpus, so the page can print the base rate
  // beside the subgroup share instead of implying the subgroup is unusual.
  const base = {};
  for (const p of CLASSIFIED) for (const x of new Set(p.classification.map((c) => c.validation))) base[x] = (base[x] ?? 0) + 1;
  console.log('\nBASE RATE, all 4,439 papers that classified anything:');
  console.log(table(['validation', 'papers', 'share of classified'],
    Object.entries(base).sort((a, b) => b[1] - a[1]).map(([k, v]) => [k, v, pct(v, CLASSIFIED.length)])));
}
 
sub('LLM as the labelling method — POLICY papers whose privacy-policy tuple has method=="llm", listed in full');
{
  const llm = POLICY.filter((p) => methodSet(p).has('llm')).sort((a, b) => a.year - b.year);
  console.log(`n=${llm.length} of ${POLICY.length} POLICY papers`);
  for (const p of llm) console.log(`  ${p.year} ${p.venue.padEnd(8)} ${p.title}`);
}
 
// --------------------------------------------------- 3. where policies come from
h('3. WHERE THE POLICIES COME FROM  (population: UNION, n=' + UNION.length + ')');
 
sub('population[].unit (enum)');
{
  const m = {};
  for (const p of UNION) for (const x of new Set(p.population.map((x2) => x2.unit))) if (!isSentinel(x)) m[x] = (m[x] ?? 0) + 1;
  console.log(table(['unit', 'papers'], Object.entries(m).sort((a, b) => b[1] - a[1])));
}
 
sub('population[].sourceList, FOLDED through policy_fold.mjs (free text — a ranking, not percentages)');
{
  const m = {}, residue = {};
  for (const p of UNION) {
    const seen = new Set();
    for (const s of p.population) {
      if (isSentinel(s.sourceList) || !s.sourceList) continue;
      const f = foldSource(s.sourceList);
      if (f === null) { residue[s.sourceList] = (residue[s.sourceList] ?? 0) + 1; continue; }
      seen.add(f);
    }
    for (const f of seen) m[f] = (m[f] ?? 0) + 1;
  }
  console.log(table(['folded source', 'papers'], Object.entries(m).sort((a, b) => b[1] - a[1])));
  console.log(`\nRESIDUE — sourceList strings matching no fold rule (${Object.keys(residue).length} distinct, printed in full):`);
  for (const [k, v] of Object.entries(residue).sort((a, b) => b[1] - a[1])) console.log(`  ${String(v).padStart(2)}  ${k}`);
}
 
// ------------------------------------------------------------- 4. the lineage
h('4. THE TOOL LINEAGE  (population: whole corpus, so a tool used outside UNION is visible)');
console.log('Matched against tools[].name, otherToolsMentioned[].name,');
console.log('classification[].resourceName/targetDetail, population[].sourceList,');
console.log('detection[].phenomenon/technique. Counts are PAPERS.');
{
  const NAMED = LINEAGE;   // defined in policy_fold.mjs, shared with the full-text probe
  const hay = (p) => {
    const a = [];
    for (const t of [...p.tools, ...p.otherToolsMentioned]) a.push(t.name ?? '', t.purpose ?? '');
    for (const c of p.classification) a.push(c.resourceName ?? '', c.targetDetail ?? '');
    for (const s of p.population) a.push(s.sourceList ?? '');
    for (const d of p.detection) a.push(d.phenomenon ?? '', d.technique ?? '');
    return a.join(' | ');
  };
  const rows = NAMED.map(([label, re]) => {
    const all = P.filter((p) => re.test(hay(p)));
    const yrs = all.map((p) => p.year).sort();
    return [label, all.length, all.filter((p) => UNION.includes(p)).length,
      yrs.length ? `${yrs[0]}-${yrs[yrs.length - 1]}` : '—',
      all.filter((p) => p.year >= 2024).length];
  });
  console.log(table(['artefact (first paper)', 'papers, corpus', 'of those in UNION', 'year range of use', 'used 2024+'], rows));
}
 
sub('tools[] + otherToolsMentioned in UNION, FOLDED (free text — ranking only)');
{
  const m = {}, residue = {};
  for (const p of UNION) {
    const seen = new Set();
    for (const t of [...p.tools, ...p.otherToolsMentioned]) {
      if (!t.name) continue;
      const f = foldTool(t.name);
      if (f === null) { residue[t.name] = (residue[t.name] ?? 0) + 1; continue; }
      seen.add(f);
    }
    for (const f of seen) m[f] = (m[f] ?? 0) + 1;
  }
  console.log(table(['folded tool family', 'papers in UNION'], Object.entries(m).sort((a, b) => b[1] - a[1])));
  const res = Object.entries(residue).sort((a, b) => b[1] - a[1]);
  console.log(`\nRESIDUE — ${res.length} distinct tool names matching no fold rule. Top 40 by frequency:`);
  for (const [k, v] of res.slice(0, 40)) console.log(`  ${String(v).padStart(2)}  ${k}`);
  console.log(`  … and ${Math.max(0, res.length - 40)} more, each in 1-2 papers.`);
}
 
// --------------------------------------------------- 5. measured results
h('5. MEASURED RESULTS: POLICY AVAILABILITY BY ECOSYSTEM');
console.log('Hand-keyed from detection[].prevalence tuples in UNION papers. The map');
console.log('below is INSIDE this script on purpose: the page must not carry a');
console.log('per-paper figure the script cannot print. Every row was read back');
console.log('against the paper\'s own quote; see report_policies_quotecheck.mjs.');
{
  // [citekey, venue/year/slug, ecosystem, denominator as the paper states it, figure]
  const AVAIL = [
    ['degeling2019_value', 'NDSS/2019/we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy', 'web, EU', '6,357 (the total of the paper\'s own availability table); the paper also states 6,759 domains in its January lists and 6,579 in its abstract', '84.5% had a policy after 25 May 2018, up from 79.6% in January'],
    ['vallina2019_porn', 'IMC/2019/tales-from-the-porn-a-comprehensive-privacy-analysis-of-the-web-porn-ecosystem', 'web, adult sites', '6,843 pornographic websites', 'only 16% had an accessible privacy policy'],
    ['cui2025_privacy', 'PETS/2025/understanding-privacy-norms-through-web-forms', 'web, sites with a PI-collecting form', '10,143 websites', '94.2% (9,559) had a privacy-policy link'],
    ['zimmeck2019_maps', 'PETS/2019/maps-scaling-privacy-compliance-analysis-to-a-million-apps', 'Android', '1,035,853 analysed apps, of 1,049,790 retrieved', 'only 50.5% had a policy link on the Play Store page'],
    ['pan2024_trap', 'USENIX/2024/is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin', 'Android', '99,194 usable apps — the paper divides link failures by its app count', '37.5% (37,150/99,194) of policy links led to an unavailable page; separately 15.7% (15,572/99,194) of apps have no link at all'],
    ['manandhar2022_smart', 'USENIX/2022/smart-home-privacy-policies-demystified-a-study-of-availability-content-and-cove', 'smart-home vendors', '596 vendors on 7 platforms', '48.99% had a device-applicable policy; 10.57% had none at all'],
    ['lentzsch2021_alexa', 'NDSS/2021/hey-alexa-is-this-skill-safe-taking-a-closer-look-at-the-alexa-skill-ecosystem', 'Alexa skills', '150,708 skills across 7 country stores', '36,475 (24.2%) provided a policy link'],
    ['yan2024_quality', 'PETS/2024/on-the-quality-of-privacy-policy-documents-of-virtual-personal-assistant-applica', 'Alexa skills', '65,195 skills', '21,063 of 65,195 provided a policy link'],
    ['zhan2024_vpvet', 'CCS/2024/vpvet-vetting-privacy-policies-of-virtual-reality-apps', 'VR apps', '11,923 apps on 10 VR platforms', '29.5% had a findable privacy policy'],
    ['edu2022_exploring', 'IMC/2022/exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services', 'Discord chatbots requesting permissions', "15,525 unique active chatbots (the paper's own Table 2 total); its 14,852-without figure does not reconcile with it", '676 (4.35%) had a policy; 14,852 (95.67%) did not'],
    ['wu2025_depth', 'IMC/2025/an-in-depth-investigation-of-data-collection-in-llm-app-ecosystems', 'GPT Actions', 'Actions declaring a legal_info_url', '93.96% of those policies were reachable'],
  ];
  const byKey = new Map(P.map((p) => [key(p), p]));
  const rows = AVAIL.map(([ck, k, eco, den, fig]) => {
    if (!byKey.has(k)) throw new Error(`AVAIL row names a paper not in the corpus: ${k}`);
    const p = byKey.get(k);
    return [ck, `${p.venue} ${p.year}`, eco, den, fig];
  });
  console.log(table(['citekey', 'venue', 'ecosystem', 'denominator (paper\'s own)', 'policy availability'], rows));
  console.log(`\nrows: ${rows.length}; every one resolves to a corpus paper.`);
}
 
h('6. MEASURED RESULTS: POLICY-VERSUS-BEHAVIOUR CONSISTENCY');
{
  const CONS = [
    ['zimmeck2017_automated', 'NDSS/2017/automated-analysis-of-privacy-requirements-for-mobile-apps', '9,050 apps with policies', 'mean 1.83 potential inconsistencies per app'],
    ['libert2018_automated', 'WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w', '1,807,491 identified third-party transmissions', 'only 14.80% were disclosed in the policy'],
    ['andow2019_policylint', 'USENIX/2019/policylint-investigating-internal-privacy-policy-contradictions-on-google-play', '11,430 policies', '14.2% (1,618) contained logical contradictions; 17.7% (2,028) contradictions or narrowing definitions'],
    ['andow2020_actions', 'USENIX/2020/actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an', '13,796 applications / 45,603 data flows', '42.4% of apps had an omitted or incorrect disclosure; 31.1% of flows were omitted; only 0.5% of flows were clearly disclosed'],
    ['trimananda2022_ovrseen', 'USENIX/2022/ovrseen-auditing-network-traffic-and-privacy-policies-in-oculus-vr', '1,135 data flows in Oculus VR apps', '68% (776) inconsistent disclosures'],
    ['cui2023_poligraph', 'USENIX/2023/poligraph-automated-privacy-policy-analysis-using-knowledge-graphs', '1,566 mapped statement pairs', '13.5% (211) conflicting; 25.5% (1,339/5,255) of policies define a term differently from the CCPA-based ontology'],
    ['xiang2023_policychecker', 'CCS/2023/policychecker-analyzing-the-gdpr-completeness-of-mobile-apps-privacy-policies', '163,068 analysable policies', '99.3% incomplete under GDPR; 98.1% had at least one mandatory-requirement violation'],
    ['xiao2023_lalaine', 'USENIX/2023/lalaine-measuring-and-characterizing-non-compliance-of-apple-privacy-labels', '5,102 fully tested iOS apps', '3,423 (67.1%) non-compliant privacy labels'],
    ['samarin2023_lessons', 'PETS/2023/lessons-in-vcr-repair-compliance-of-android-app-developers-with-the-california-c', '69 apps with CCPA disclosures', '80% (55) collected an identifier they did not disclose'],
    ['ali2024_honesty', 'PETS/2024/honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a', 'iOS apps labelled "Data Not Collected"', '97% had policy statements indicating data collection'],
    ['cui2025_privacy', 'PETS/2025/understanding-privacy-norms-through-web-forms', 'websites collecting PI via web forms', 'phi coefficient between observed collection and PoliGraph-er disclosure was < 0.20 for every PI type'],
  ];
  const byKey = new Map(P.map((p) => [key(p), p]));
  const rows = CONS.map(([ck, k, den, fig]) => {
    if (!byKey.has(k)) throw new Error(`CONS row names a paper not in the corpus: ${k}`);
    const p = byKey.get(k);
    return [ck, `${p.venue} ${p.year}`, den, fig];
  });
  console.log(table(['citekey', 'venue', 'denominator (paper\'s own)', 'finding'], rows));
  console.log(`\nrows: ${rows.length}; every one resolves to a corpus paper.`);
}
 
// ------------------------------------------------------- 7. reading / language
h('7. WHAT ELSE THE UNION MEASURES');
sub('UNION papers that are also in the `legal` population (legal[] non-empty)');
{
  const L = UNION.filter((p) => p.legal.length > 0);
  const LC = P.filter((p) => p.legal.length > 0);
  console.log(`${L.length} of ${UNION.length} UNION papers assess a law; base rate ${LC.length} of ${P.length} (${pct(LC.length, P.length)}) corpus-wide.`);
  const laws = {};
  for (const p of L) for (const x of new Set(p.legal.map((l) => l.law).filter((x2) => x2 && !isSentinel(x2)))) laws[x] = (laws[x] ?? 0) + 1;
  console.log(table(['law (free text, unfolded — ranking only)', 'papers'],
    Object.entries(laws).sort((a, b) => b[1] - a[1]).slice(0, 15)));
}
 
sub('human annotation in UNION — policy labelling is a hand-coding literature');
{
  const ann = UNION.filter((p) => p.humanAnnotation.length > 0);
  const withMetric = ann.filter((p) => p.humanAnnotation.some((a) => !isSentinel(a.agreementMetric)));
  const withCount = ann.filter((p) => p.humanAnnotation.some((a) => !isSentinel(a.annotatorCount)));
  console.log(`UNION papers with humanAnnotation[]            ${ann.length} of ${UNION.length}  ${pct(ann.length, UNION.length)}`);
  console.log(`  …stating an agreement metric                 ${withMetric.length}  ${pct(withMetric.length, ann.length)} of annotated`);
  console.log(`  …stating an annotator count                  ${withCount.length}  ${pct(withCount.length, ann.length)} of annotated`);
  const ANN = P.filter((p) => p.humanAnnotation.length > 0);
  const bMetric = ANN.filter((p) => p.humanAnnotation.some((a) => !isSentinel(a.agreementMetric)));
  console.log(`\nBASE RATE, all ${ANN.length} papers that coded data by hand:`);
  console.log(`  …stating an agreement metric                 ${bMetric.length}  ${pct(bMetric.length, ANN.length)}`);
}
 
sub('temporal[].mode in UNION (enum) vs corpus base rate — is this a longitudinal literature?');
{
  const count = (set) => {
    const m = {};
    for (const p of set) for (const x of new Set(p.temporal.map((t) => t.mode))) if (!isSentinel(x)) m[x] = (m[x] ?? 0) + 1;
    return m;
  };
  const u = count(UNION), b = count(P);
  const uT = UNION.filter((p) => p.temporal.length > 0).length;
  const bT = P.filter((p) => p.temporal.length > 0).length;
  console.log(`UNION papers with a temporal[] tuple ${uT} of ${UNION.length}; corpus ${bT} of ${P.length}`);
  console.log(table(['temporal.mode', 'UNION', 'share of UNION w/ temporal', 'corpus', 'share of corpus w/ temporal'],
    [...new Set([...Object.keys(u), ...Object.keys(b)])].sort((x, y) => (u[y] ?? 0) - (u[x] ?? 0))
      .map((k) => [k, u[k] ?? 0, pct(u[k] ?? 0, uT), b[k] ?? 0, pct(b[k] ?? 0, bT)])));
}
 
sub('artifact availability in UNION vs corpus base rate (enum)');
{
  const a = {}, b = {};
  for (const p of UNION) if (p.artifacts) a[p.artifacts.availability] = (a[p.artifacts.availability] ?? 0) + 1;
  for (const p of P) if (p.artifacts) b[p.artifacts.availability] = (b[p.artifacts.availability] ?? 0) + 1;
  const withArt = P.filter((p) => p.artifacts);
  const uWithArt = UNION.filter((p) => p.artifacts);
  console.log(table(['availability', 'UNION', 'share', 'corpus', 'share'],
    [...new Set([...Object.keys(a), ...Object.keys(b)])].sort((x, y) => (b[y] ?? 0) - (b[x] ?? 0))
      .map((k) => [k, a[k] ?? 0, pct(a[k] ?? 0, uWithArt.length), b[k] ?? 0, pct(b[k] ?? 0, withArt.length)])));
}
 
// --------------------------------------- 7b. inputs for the significance test
//
// policies_significance.py PARSES these lines. They are emitted here, from the
// same population objects every other table uses, so the p-values can never be
// computed over a hand-keyed count that has drifted from the report.
//
// Each comparison is emitted TWICE: once over POLICY (the stable enum, 102) and
// once over UNION (the regex-widened candidate set, 123). The page publishes the
// POLICY rows; the UNION rows are published on the provenance page so a reader
// can see whether the choice of population changes any conclusion.
h('7b. SIGNIFICANCE INPUTS — parsed by policies_significance.py');
{
  const hasTemporalArchive = (p) => p.temporal.some((x) => x.mode === 'web-archive');
  const hasTemporal = (p) => p.temporal.length > 0;
  const annotated = (p) => p.humanAnnotation.length > 0;
  const statesMetric = (p) => p.humanAnnotation.some((a) => !isSentinel(a.agreementMetric));
  const hasArtifacts = (p) => Boolean(p.artifacts);
  const publicArtifacts = (p) => Boolean(p.artifacts) && p.artifacts.availability === 'public';
  const assessesLaw = (p) => p.legal.length > 0;
  const policyTuples = (p) => p.classification.filter((c) => c.target === 'privacy-policy');
  const noValidation = (p) => policyTuples(p).some((c) => c.validation === 'none-reported');
  const hasPolicyTuple = (p) => policyTuples(p).length > 0;
 
  // The validation comparison is the one asymmetric case: on the subgroup side
  // the question is about the paper's PRIVACY-POLICY tuples, on the corpus side
  // it is about ANY classification tuple. Giving each comparison its own pair of
  // predicates keeps that visible instead of hiding it in one shared predicate.
  const anyNoValidation = (p) => p.classification.some((c) => c.validation === 'none-reported');
  const classifiedAnything = (p) => p.classification.length > 0;
 
  // [label, subgroup restrict, subgroup property, base restrict, base property]
  const COMPARISONS = [
    ['temporal.mode == web-archive', hasTemporal, hasTemporalArchive, hasTemporal, hasTemporalArchive],
    ['humanAnnotation states an agreement metric', annotated, statesMetric, annotated, statesMetric],
    ['artifacts.availability == public', hasArtifacts, publicArtifacts, hasArtifacts, publicArtifacts],
    ['assesses a law (legal[] non-empty)', () => true, assessesLaw, () => true, assessesLaw],
    ['classification.validation == none-reported', hasPolicyTuple, noValidation, classifiedAnything, anyNoValidation],
  ];
  for (const [label, sRestrict, sProp, bRestrict, bProp] of COMPARISONS) {
    const base = P.filter(bRestrict);
    for (const [popName, pop] of [['POLICY', POLICY], ['UNION', UNION]]) {
      const sub2 = pop.filter(sRestrict);
      // The base rate must EXCLUDE the subgroup's own papers, otherwise the
      // subgroup is compared against a population that contains it.
      const rest = base.filter((p) => !pop.includes(p));
      console.log(`SIG|${label}|${popName}|${sub2.filter(sProp).length}|${sub2.length}|`
        + `${rest.filter(bProp).length}|${rest.length}|`
        + `${base.filter(bProp).length}|${base.length}`);
    }
  }
  console.log('\nColumns: SIG|comparison|population|sub hits|sub n|rest-of-corpus hits|rest n|whole-corpus hits|whole n');
  console.log('"rest" excludes the subgroup itself; "whole" is the figure a naive base rate would use.');
}
 
// --------------------------------------------------------------- 8. full list
h('8. THE UNION, IN FULL — the audit surface for every count above');
for (const p of [...UNION].sort((a, b) => a.year - b.year || a.venue.localeCompare(b.venue))) {
  const tag = POLICY.includes(p) ? (PROBE.includes(p) ? 'BOTH ' : 'ENUM ') : 'PROBE';
  console.log(`${tag} ${p.year} ${p.venue.padEnd(8)} ${p.title}`);
}

The fold — ''policy_fold.mjs''

policy_fold.mjs
// Fold the free-text names that privacy:policies aggregates, and nothing else.
//
// WHY: two fields on that page are free text and ~20% stable run-to-run by exact
// string (data/extract/README.md), so aggregating them raw undercounts:
//
//   population[].sourceList  — "Google Play Store" (15 papers) and "Google Play"
//     (15) are one source under two spellings; "OPP-115" (11) and "OPP-115
//     corpus" (2) are one dataset; five different Alexa spellings are one list.
//   tools[].name / otherToolsMentioned[].name — "Beautiful Soup" / "BeautifulSoup"
//     / "bs4", "PoliGraph" / "PoliGraph-er", "Selenium" / "Selenium WebDriver".
//
// Rules are ordered; the first match wins. Anything matching no rule keeps its
// raw name and is printed as residue by report_policies.mjs, so the part this
// file cannot classify stays visible instead of vanishing.
 
// [canonical family, matcher tested against the raw name]
export const SOURCE_FAMILIES = [
  // --- app stores and app-set sources
  ['Google Play', /google ?play|play ?store|androzoo|playdrone|google-play-scraper/i],
  ['Apple App Store', /app ?store(?! optimization)|apple ?app|ios app store|app annie|appfigures/i],
  ['Alexa list', /\balexa\b(?!.*skill)/i],
  ['Alexa Skills Store', /alexa (skill|store).*skill|skill store|alexa skills/i],
  ['Tranco', /\btranco\b/i],
  ['Majestic / Umbrella / Quantcast', /majestic|umbrella|quantcast|chrome ux|\bcrux\b/i],
 
  // --- annotated policy corpora
  ['OPP-115', /opp-?115|usable ?privacy|acl\/coling 2014|ramanath/i],
  ['APP-350', /app-?350/i],
  ['PrivaSeer', /privaseer/i],
  ['Princeton Policies-over-Time', /princeton.*polic|policies over time|million-document/i],
 
  // --- archives
  ['Wayback Machine', /wayback|internet archive|common ?crawl/i],
 
  // --- participant panels (these are NOT policy sources; kept separate so they
  // cannot be silently mixed into a "where policies come from" table)
  ['participant panel', /prolific|mechanical turk|\bmturk\b|qualtrics panel|clickworker|respondi/i],
 
  // --- other app / extension / package ecosystems
  ['other app ecosystem', /sidequest|oculus|meta quest|steam|chrome web store|cocoapods|rapidapi|wordpress|github|npm|pypi/i],
 
  // --- the honest bucket: the paper made its own list
  ['custom / hand-built list', /custom|hand-?(built|picked|curated)|manual(ly)? (selected|compiled)|own (list|corpus|selection)|seed list/i],
];
 
export const TOOL_FAMILIES = [
  ['Polisis / PriBot', /polisis|pribot/i],
  ['PolicyLint', /policylint/i],
  ['PoliCheck', /policheck/i],
  ['PoliGraph', /poligraph/i],
  ['PolicyChecker', /policychecker/i],
  ['PurPliance', /purpliance/i],
  ['PrivBERT', /privbert/i],
  ['MAPS', /^maps$/],
  ['spaCy', /^spacy/i],
  ['NLTK', /^nltk/i],
  ['Stanford CoreNLP / AllenNLP', /corenlp|allennlp|stanza/i],
  ['BERT family (non-privacy)', /^bert|roberta|distilbert|legal-?bert|^albert/i],
  ['LLM (commercial API)', /gpt-?[345]|chatgpt|openai|claude|gemini|\bbard\b/i],
  ['LLM (open weights)', /llama|mistral|qwen|vicuna|falcon|deepseek/i],
  ['boilerplate stripper', /boilerpipe|readability|trafilatura|html2text|justext|goose/i],
  ['BeautifulSoup', /beautiful ?soup|^bs4$/i],
  ['Selenium', /^selenium/i],
  ['Playwright', /^playwright/i],
  ['Puppeteer', /^puppeteer/i],
  ['OpenWPM', /openwpm/i],
  ['language detection', /langdetect|langid|\bcld[23]?\b|fasttext.*lang/i],
  ['readability metric', /flesch|kincaid|gunning|smog|coleman|dale-?chall|textstat/i],
];
 
function fold(raw, families) {
  const s = String(raw ?? '').trim();
  if (!s) return null;
  for (const [canon, re] of families) if (re.test(s)) return canon;
  return null; // residue: caller keeps the raw string and prints it
}
 
export const foldSource = (raw) => fold(raw, SOURCE_FAMILIES);
export const foldTool = (raw) => fold(raw, TOOL_FAMILIES);
 
// The policy-analysis lineage, one regex per artefact. Exported so that
// report_policies.mjs (which matches it against EXTRACTION fields) and
// policies_fulltext_probe.mjs (which matches it against the paper's FULL TEXT)
// cannot drift apart. The two ask different questions and give different
// answers: PrivaSeer is named as a tool or data source by no corpus paper and
// cited in the text of twelve.
export const LINEAGE = [
  ['Privee (USENIX 2014)', /\bPrivee\b/],
  ['OPP-115 corpus (ACL 2016)', /OPP-?115/i],
  ['Polisis / PriBot (USENIX 2018)', /polisis|pribot/i],
  ['PolicyLint (USENIX 2019)', /policylint/i],
  ['MAPS (PETS 2019)', /\bMAPS\b/],
  ['APP-350 corpus (2019)', /APP-?350/i],
  ['PoliCheck (USENIX 2020)', /policheck/i],
  ['PurPliance (2021)', /purpliance/i],
  ['PrivBERT (2021)', /privbert/i],
  ['Calpric (USENIX 2023)', /calpric/i],
  ['PoliGraph / PoliGraph-er (USENIX 2023)', /poligraph/i],
  ['PolicyChecker (CCS 2023)', /policychecker/i],
  ['Lalaine (USENIX 2023)', /lalaine/i],
  ['PolicyComp (USENIX 2023)', /policycomp/i],
  ['PrivaSeer', /privaseer/i],
];
 
// Matching a tool name against a paper's FULL TEXT is not the same problem as
// matching it against the extraction's tool fields. In full text an acronym
// collides with unrelated uses: /\bMAPS\b/ matches "Google MAPS abuse" and the
// Play Store's "MAPS & NAVIGATION" category. Where that happens, the full-text
// scan uses the artefact's own title instead of its acronym, and the override is
// recorded here rather than buried in the scanning script.
export const LINEAGE_FULLTEXT_OVERRIDE = new Map([
  ['MAPS (PETS 2019)', {
    re: /MAPS:\s*Scaling/i,
    why: '/\\bMAPS\\b/ also matches "Google MAPS" and the Play category "MAPS & NAVIGATION"; the paper is always cited by its title',
  }],
]);

The full-text probe — ''policies_fulltext_probe.mjs''

policies_fulltext_probe.mjs
// Full-text probes for the methodological questions the extraction schema has no
// field for: how a paper FINDS a privacy policy, what it does about language, and
// what it does about the fact that the policy is a moving target.
//
//   node scripts/policies_fulltext_probe.mjs > scripts/policies_fulltext_probe-output.txt
//
// Population: the 123 UNION papers from report_policies.mjs, re-derived here from
// the same two definitions so the two scripts cannot drift apart.
//
// A probe count is a CANDIDATE SET, not a measurement: it says the phrase is in
// the paper, not that the paper did the thing. Two forms are printed for every
// probe, because probe width decides the claim. Whitespace is collapsed and
// hyphenation rejoined first — a PDF line break inside a phrase otherwise
// silently undercounts.
//
// CAVEAT ON THE SECOND COLUMN: 'wide' is a looser phrasing of the same question
// for most rows, and there it is a superset. For three rows it is a DIFFERENT
// question (a tool-name probe rather than a phrase probe), so it can be smaller
// than 'narrow'. Those rows are flagged in the output; do not read them as a
// narrow-vs-wide comparison.
 
import fs from 'node:fs';
import path from 'node:path';
import { loadExtractions, dataRoot, pct, table } from './lib.mjs';
import { LINEAGE, LINEAGE_FULLTEXT_OVERRIDE } from './policy_fold.mjs';
 
const ROOT = path.join(dataRoot(), 'fulltext');
const P = loadExtractions();
const PROBE_RE =
  /privacy polic|privacy notice|terms of service|terms and conditions|privacy label|data safety|nutrition label|policy text|privacy statement/i;
const ts = (p) => `${p.title ?? ''} • ${p.summary ?? ''}`;
const UNION = [...new Set([
  ...P.filter((p) => p.classification.some((c) => c.target === 'privacy-policy')),
  ...P.filter((p) => PROBE_RE.test(ts(p))),
])];
 
const norm = (s) => s.replace(/­/g, '').replace(/-\n/g, '').replace(/\s+/g, ' ');
function text(p) {
  const f = path.join(ROOT, String(p.year), p.venue, p.slug, 'paper.cols.txt');
  if (!fs.existsSync(f)) return null;
  return norm(fs.readFileSync(f, 'utf8'));
}
 
const TEXTS = new Map();
let noText = 0;
for (const p of UNION) {
  const t = text(p);
  if (t === null) { noText += 1; continue; }
  TEXTS.set(p, t);
}
console.log(`UNION = ${UNION.length} papers; full text present for ${TEXTS.size}; missing ${noText}.`);
console.log('Every count below is over the ' + TEXTS.size + ' papers with full text.\n');
 
// [label, narrow regex, wide regex]. Narrow is what the page quotes.
const PROBES = [
  ['finds the policy by LINK TEXT / anchor keyword',
    /(link|anchor)s? (text|label)|matching the (link|anchor)|keyword(s)? (in|on) the (link|footer)|footer link/i,
    /link text|anchor text|keyword/i],
  ['names a link-detection SEED PHRASE list',
    /seed phrase|seed keyword|list of (candidate )?(keywords|phrases) (used )?to (find|locate|identify)/i,
    /seed (phrase|term|keyword)/i],
  ['follows the policy link and reports FAILURES (404 / dead / unreachable)',
    /polic(y|ies)[^.]{0,80}(404|dead link|broken link|unreachable|did not resolve|failed to (load|download|retrieve))/i,
    /(404|broken link|dead link|unreachable)/i],
  ['handles a policy served as a PDF',
    /polic(y|ies)[^.]{0,60}\bPDF\b|\bPDF\b[^.]{0,60}polic(y|ies)/i,
    /\bPDF\b/],
  ['states the LANGUAGE of the policies it analysed',
    /english[- ]language (privacy )?polic|polic(y|ies)[^.]{0,60}in english|non-?english polic|language of the polic/i,
    /langdetect|langid|\bcld[23]\b|language detection/i],
  ['analyses policies in more than one language',
    /multilingual|bilingual|(translat(e|ed|ing|ion))[^.]{0,50}polic|polic(y|ies)[^.]{0,50}translat/i,
    /multilingual|bilingual|translat/i],
  ['strips BOILERPLATE / extracts the main content of the policy page',
    /boilerpipe|trafilatura|readability(-lxml| library|\.js)|justext|main content extraction|html to text|html2text/i,
    /boilerplate|extract(ed|ing)? the (main )?(text|content)/i],
  ['measures READABILITY of the policy',
    /flesch|kincaid|gunning fog|\bSMOG\b|coleman-liau|dale-chall|reading (ease|grade|level|time)/i,
    /readab(le|ility)/i],
  ['compares the policy against OBSERVED BEHAVIOUR (traffic, code, or storage)',
    /(consisten|inconsisten|discrepan|mismatch|divergen|misalign)[a-z]*[^.]{0,90}(polic|label|disclos)/i,
    /consistency analysis|policy[- ]to[- ]?(code|flow|traffic)/i],
  ['handles the policy being VERSIONED / changing under it',
    /polic(y|ies)[^.]{0,80}(version|snapshot|revision|updated?|chang(e|ed|es))|wayback|internet archive/i,
    /longitudinal|over time|snapshot/i],
  ['reports the sentence/segment SEGMENTATION step',
    /segment(ation|ed|ing)? (the )?polic|polic[^.]{0,40}(into|by) (sentences|segments|paragraphs)|sentence[- ]level/i,
    /segment/i],
  ['says which policy applies (app vs developer vs platform vs layered)',
    /(which|applicable|relevant|correct) privacy polic|polic(y|ies) (that )?appl(y|ies)|layered polic|multiple privacy polic|generic (company|corporate) polic/i,
    /applicable polic|multiple polic/i],
  ['deduplicates identical / templated policies',
    /(duplicate|template|boilerplate|reuse[d]?|identical)[^.]{0,60}polic|polic[^.]{0,60}(duplicat|templat|reus)/i,
    /duplicat|templat/i],
];
 
const rows = [];
for (const [label, narrow, wide] of PROBES) {
  let n = 0, w = 0;
  for (const [, t] of TEXTS) {
    if (narrow.test(t)) n += 1;
    if (wide.test(t)) w += 1;
  }
  const flag = w < n ? '  <-- not a superset: different question' : '';
  rows.push([label + flag, n, pct(n, TEXTS.size), w, pct(w, TEXTS.size)]);
}
console.log(table(['probe (narrow form is what the page quotes)', 'narrow', 'share', 'wide', 'share'], rows));
 
console.log('\n--- The two probes the page leans on hardest, with the matching papers named');
for (const [label, narrow] of [
  ['follows the policy link and reports FAILURES (404 / dead / unreachable)', PROBES[2][1]],
  ['analyses policies in more than one language', PROBES[5][1]],
]) {
  const hits = [...TEXTS.entries()].filter(([, t]) => narrow.test(t)).map(([p]) => p)
    .sort((a, b) => a.year - b.year);
  console.log(`\n${label}  —  ${hits.length} papers`);
  for (const p of hits) console.log(`  ${p.year} ${p.venue.padEnd(8)} ${p.title.slice(0, 92)}`);
}
 
// ---------------------------------------------------------------------------
// The lineage, counted a second way: WHOLE CORPUS full text rather than the
// extraction's tool/source fields.
//
// These are different questions and they give different answers. "No corpus
// paper's extraction names PrivaSeer" is a fact about the extraction. "Nobody
// in these seven venues cites PrivaSeer" would be a fact about the literature,
// and it is false. A fold over a structured field cannot support an "at all"
// claim; only a full-text scan can even try.
//
// Scans every paper.cols.txt in the corpus, not just the UNION.
{
  const ROOTDIR = ROOT;
  const all = [];
  for (const y of fs.readdirSync(ROOTDIR)) {
    const yd = path.join(ROOTDIR, y);
    if (!fs.statSync(yd).isDirectory()) continue;
    for (const v of fs.readdirSync(yd)) {
      const vd = path.join(yd, v);
      if (!fs.statSync(vd).isDirectory()) continue;
      for (const s of fs.readdirSync(vd)) {
        const f = path.join(vd, s, 'paper.cols.txt');
        if (fs.existsSync(f)) all.push([`${y} ${v.padEnd(8)} ${s}`, f]);
      }
    }
  }
  console.log(`\n\n--- LINEAGE ARTEFACTS: extraction fold vs FULL-TEXT mention, whole corpus`);
  console.log(`Full-text denominator: ${all.length} papers with a readable paper.cols.txt.`);
  console.log(`"extraction" is the column report_policies.mjs section 4 prints.`);
  const PATTERNS = LINEAGE.map(([label, re]) => {
    const o = LINEAGE_FULLTEXT_OVERRIDE.get(label);
    return [label, o ? o.re : re, o ? o.why : ''];
  });
  const seen = new Map(PATTERNS.map(([label]) => [label, []]));
  for (const [id, f] of all) {
    const t2 = norm(fs.readFileSync(f, 'utf8').replace(/\0/g, ''));
    for (const [label, re] of PATTERNS) if (re.test(t2)) seen.get(label).push(id);
  }
  console.log(table(['artefact', 'full-text pattern', 'papers whose FULL TEXT names it'],
    PATTERNS.map(([label, re]) => [label, String(re), seen.get(label).length])));
  const overridden = PATTERNS.filter(([, , why]) => why);
  console.log('\nFull-text patterns that differ from the extraction pattern, and why:');
  for (const [label, re, why] of overridden) console.log(`  ${label}: ${re} — ${why}`);
  console.log('\nThe two artefacts where the gap changes what may be said:');
  for (const label of ['PrivaSeer', 'Calpric (USENIX 2023)']) {
    console.log(`\n${label} — ${seen.get(label).length} papers:`);
    for (const id of seen.get(label)) console.log('  ' + id);
  }
}

The quote check — ''policies_quotecheck.mjs''

policies_quotecheck.mjs
// Verify every measured figure privacy:policies publishes against the paper's
// own text — not against the extraction's evidence.quote, which can be a
// different sentence from the one carrying the number.
//
//   node scripts/policies_quotecheck.mjs > scripts/policies_quotecheck-output.txt
//
// Checked against FOUR renderings of each PDF, because they fail on different
// sentences: paper.cols.txt (column order repaired — the one to quote from),
// paper.norm.txt, paper.txt, and a pypdf extraction cached under cache/pypdf/.
// A needle matching any of them is LOCATED, and the rendering that matched is
// named, so a .cols-only miss is visible. The pypdf pass is not decoration:
// PETS 2024 'Honesty is the Best Policy' states its 65% template figure in a
// sentence that .cols splices with an unrelated one, so .cols alone reports a
// false MISS on a figure the paper does state.
//
// Whitespace is collapsed and hyphenation rejoined before matching: a PDF line
// break inside a phrase otherwise produces a false MISS.
//
// Exit status is non-zero if any needle is unlocated, so this cannot pass by
// being ignored.
 
import fs from 'node:fs';
import { spawnSync } from 'node:child_process';
import path from 'node:path';
import { dataRoot } from './lib.mjs';
 
const ROOT = path.join(dataRoot(), 'fulltext');
 
const norm = (s) =>
  s
    .replace(/­/g, '')
    .replace(/-\n/g, '')
    .replace(/[‘’ʼ]/g, "'")
    .replace(/[“”]/g, '"')
    .replace(/[‐-―−]/g, '-')
    .replace(/\s+/g, ' ')
    .trim();
 
const RENDERINGS = ['paper.cols.txt', 'paper.norm.txt', 'paper.txt'];
const PYCACHE = 'cache/pypdf'; // the dataset mount is read-only, so cache locally
 
function pypdfText(dir, venue, year, slug) {
  const out = path.join(PYCACHE, `${year}_${venue}_${slug}.txt`);
  if (!fs.existsSync(out)) {
    const pdf = path.join(dir, 'paper.pdf');
    if (!fs.existsSync(pdf)) return null;
    fs.mkdirSync(PYCACHE, { recursive: true });
    const r = spawnSync('python3', ['-c',
      'import sys,pypdf;print("".join(p.extract_text() or "" for p in pypdf.PdfReader(sys.argv[1]).pages))',
      pdf], { encoding: 'utf8', maxBuffer: 64 * 1024 * 1024 });
    if (r.status !== 0) return null;
    fs.writeFileSync(out, r.stdout);
  }
  return fs.readFileSync(out, 'utf8');
}
 
const cache = new Map();
function renderings(venue, year, slug) {
  const k = `${venue}/${year}/${slug}`;
  if (!cache.has(k)) {
    const dir = path.join(ROOT, String(year), venue, slug);
    const out = {};
    for (const r of RENDERINGS) {
      const f = path.join(dir, r);
      if (fs.existsSync(f)) out[r] = norm(fs.readFileSync(f, 'utf8'));
    }
    const py = pypdfText(dir, venue, year, slug);
    if (py !== null) out['pypdf'] = norm(py);
    if (Object.keys(out).length === 0) throw new Error(`no full text at all: ${dir}`);
    cache.set(k, out);
  }
  return cache.get(k);
}
 
// [venue, year, slug, label as it appears on the page, needle from the PAPER]
// The needle is the paper's own words carrying the number the page prints.
const CHECKS = [
  // --- section: how many sites/apps have a policy at all
  ['NDSS', 2019, 'we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy',
    'web policy availability 84.5% / 79.6%', '84.5 %'],
  ['NDSS', 2019, 'we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy',
    'Degeling sample 6,579 sites, 500 per member state', '500 most popular websites'],
  ['IMC', 2019, 'tales-from-the-porn-a-comprehensive-privacy-analysis-of-the-web-porn-ecosystem',
    'adult sites: only 16% have an accessible policy, of 6,843', 'only 16% of the analyzed websites have an accessible privacy policy'],
  ['PETS', 2025, 'understanding-privacy-norms-through-web-forms',
    'web forms 94.2% policy link', '94.2% (9,559)'],
  ['PETS', 2019, 'maps-scaling-privacy-compliance-analysis-to-a-million-apps',
    'Play policy links 50.5%', '50.5%'],
  ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin',
    'APPG: 37.5% of links unavailable', 'we found that 37.5%'],
  ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin',
    'APPG: 20.5% non-English', '20.5%'],
  ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin',
    'APPG: 22.3% low quality under 2KB/200 words', 'we identify 22.3% (10,375/46,472)'],
  ['USENIX', 2022, 'smart-home-privacy-policies-demystified-a-study-of-availability-content-and-cove',
    'smart-home vendors 48.99% (292/596)', '48.99%'],
  ['NDSS', 2021, 'hey-alexa-is-this-skill-safe-taking-a-closer-look-at-the-alexa-skill-ecosystem',
    'Alexa skills 24.2% policy link', '36,475 (24.2 %)'],
  ['CCS', 2024, 'vpvet-vetting-privacy-policies-of-virtual-reality-apps',
    'VR apps 29.5% have a policy', '(i.e., 29.5%) of privacy policies were successfully found'],
  ['IMC', 2022, 'exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services',
    'chatbots 95.67% lack a policy', '95.67%'],
 
  ['IMC', 2019, 'tales-from-the-porn-a-comprehensive-privacy-analysis-of-the-web-porn-ecosystem',
    'adult sample 6,843 sites', '6,843 pornographic websites'],
  ['PETS', 2025, 'understanding-privacy-norms-through-web-forms',
    'web-forms denominator 10,143', 'leaving 10,143 websites for this analysis'],
  ['PETS', 2019, 'maps-scaling-privacy-compliance-analysis-to-a-million-apps',
    'MAPS denominator 1,049,790 retrieved / 1,035,853 analysed', '1,049,790 retrieved apps, 1,035,853'],
  ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin',
    'APPG link denominator 37,150/99,194', '(37,150/99,194)'],
  ['USENIX', 2022, 'smart-home-privacy-policies-demystified-a-study-of-availability-content-and-cove',
    'smart home 10.57% no policy at all', '10.57% not providing privacy policies at all'],
  ['NDSS', 2021, 'hey-alexa-is-this-skill-safe-taking-a-closer-look-at-the-alexa-skill-ecosystem',
    'Alexa denominator 150,708 skills, 36,475 with a link', 'Combined 150,708 36,475 (24.2 %)'],
  ['PETS', 2024, 'on-the-quality-of-privacy-policy-documents-of-virtual-personal-assistant-applica',
    'Alexa 21,063 of 65,195', '65,195 skills, of which 21,063'],
  ['CCS', 2024, 'vpvet-vetting-privacy-policies-of-virtual-reality-apps',
    'VR denominator 11,923 apps', '11,923 VR apps'],
  ['IMC', 2022, 'exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services',
    'Discord chatbots 676 (4.35%) have a policy', '676 (4.35%)'],
 
  // --- section: policy versus behaviour
  ['WWW', 2018, 'an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w',
    'only 14.80% of transmissions disclosed', '14.80%'],
  ['USENIX', 2019, 'policylint-investigating-internal-privacy-policy-contradictions-on-google-play',
    'PolicyLint 14.2% (1,618/11,430) contradictions', '1,618/11,430'],
  ['USENIX', 2020, 'actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an',
    'PoliCheck 42.4% of apps', '42.4%'],
  ['USENIX', 2020, 'actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an',
    'PoliCheck only 0.5% of flows clearly disclosed', 'Only 0.5% of data flows were explicitly discussed'],
  ['USENIX', 2022, 'ovrseen-auditing-network-traffic-and-privacy-policies-in-oculus-vr',
    'OVRseen 68% inconsistent', '68% (776/1,135)'],
  ['USENIX', 2023, 'poligraph-automated-privacy-policy-analysis-using-knowledge-graphs',
    'PoliGraph 70.6% recall / 96.9% precision', '70.6% recall'],
  ['USENIX', 2023, 'poligraph-automated-privacy-policy-analysis-using-knowledge-graphs',
    'PoliGraph 25.5% of policies redefine a term', '25.5%'],
  ['CCS', 2023, 'policychecker-analyzing-the-gdpr-completeness-of-mobile-apps-privacy-policies',
    'PolicyChecker 99.3% incomplete', '99.3%'],
  ['CCS', 2023, 'policychecker-analyzing-the-gdpr-completeness-of-mobile-apps-privacy-policies',
    'PolicyChecker 163,068 analysable of 205,973', '163,068'],
  ['USENIX', 2023, 'lalaine-measuring-and-characterizing-non-compliance-of-apple-privacy-labels',
    'Lalaine 3,423 of 5,102 apps', '3,423'],
  ['PETS', 2023, 'lessons-in-vcr-repair-compliance-of-android-app-developers-with-the-california-c',
    'CCPA VCR 80% undisclosed identifier', '(55 apps, 80%)'],
  ['PETS', 2024, 'honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a',
    'Apple labels: 97% of "Data Not Collected" contradicted', 'almost all (97%) apps that indicate in their privacy'],
  ['PETS', 2025, 'understanding-privacy-norms-through-web-forms',
    'phi < 0.20 policy-vs-form association', 'the association appears to be weak (< 0.20) for all PI types'],
 
  // --- section: what the policy text itself looks like
  ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset',
    'Princeton corpus 1,071,488 policies / 130,000 sites', '1,071,488'],
  ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset',
    'median length 876 -> 1,522 words', '1,522'],
  ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset',
    'FKGL 11.9 -> 13.2', 'to 2019B (13.2)'],
  ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset',
    'beacons: 25.8% of policies vs 94.6% of top-10K sites', '25.8%'],
  ['WWW', 2018, 'an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w',
    '84.7 minutes to read applicable policies', '84.7'],
  ['PETS', 2020, 'the-privacy-policy-landscape-after-the-gdpr',
    'EU policies gained 35% words / 33% sentences, Global 25% / 22%', '(+35%, +25%)'],
  ['CCS', 2024, 'vpvet-vetting-privacy-policies-of-virtual-reality-apps',
    'VR policy reuse 54.5% (1,919/3,521)', '54.5%'],
  ['PETS', 2024, 'honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a',
    'template reuse 65%', 'privacy policies of 65%'],
 
  // --- section: the annotated corpora
  ['USENIX', 2018, 'polisis-automated-analysis-and-presentation-of-privacy-policies-using-deep-learn',
    'Polisis trained on 65 of the OPP-115 policies, 50 held out', 'we used the data from 65 policies in the OPP-115 dataset, and we kept 50 policies as a testing set'],
  ['USENIX', 2018, 'polisis-automated-analysis-and-presentation-of-privacy-policies-using-deep-learn',
    'Polisis average F1 0.84', 'Average 0.87 0.83 0.84 0.84'],
  ['USENIX', 2023, 'calpric-inclusive-and-fine-grain-labeling-of-privacy-policies-with-crowdsourcing',
    'Calpric 16,856 labelled segments', '16856 labeled text segments'],
 
  ['PETS', 2023, 'researchers-experiences-in-analyzing-privacy-policies-challenges-and-opportuniti',
    'no best practices have emerged (26 interviews)', 'no clear best practices or methodologies for privacy policy analysis have emerged'],
  ['PETS', 2023, 'researchers-experiences-in-analyzing-privacy-policies-challenges-and-opportuniti',
    '26 researchers interviewed', 'semi-structured interviews with 26 researchers'],
  ['PETS', 2023, 'evolution-of-composition-readability-and-structure-of-privacy-policies-over-two',
    'User Choice/Control semantic change 26%', 'semantic modifications are made to 26% of the policies'],
 
  // --- section: LLM extraction (2024-2026)
  ['IMC', 2024, 'analyzing-corporate-privacy-policies-using-ai-chatbots',
    'GPT-4 chatbot annotation of corporate policies', 'GPT-4'],
  ['IMC', 2025, 'an-in-depth-investigation-of-data-collection-in-llm-app-ecosystems',
    'LLM disclosure classifier 87.44% accuracy', '87.44%'],
  ['IMC', 2025, 'an-in-depth-investigation-of-data-collection-in-llm-app-ecosystems',
    'only 5.8% of Actions clearly disclose', 'only 5.8% of Actions clearly disclosing'],
 
  // --- added 2026-09-10 after the whole-page number guard flagged these as
  //     unaccounted: figures the page prints that no earlier needle covered.
  ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin',
    'APPG low-quality count 10,375/46,472', '22.3% (10,375/46,472)'],
  ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin',
    'APPG non-English count 9,523/46,472', '(9,523/46,472)'],
  ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin',
    'APPG 15.7% of apps provide no policy', '15.7% (15,572/99,194)'],
  ['USENIX', 2020, 'actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an',
    'PoliCheck entity-insensitive false-consistency 37.1%', '37.1% of inconsistent third-party flows as consistent'],
  ['USENIX', 2020, 'actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an',
    'PoliCheck 31.1% (14,409/45,603) omitted flows', '31.1% (14,409/45,603)'],
  ['USENIX', 2020, 'actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an',
    'PoliCheck 14,409 omitted disclosures (the 31.1% numerator)', 'Of the 14,409 omitted'],
  ['PETS', 2024, 'honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a',
    'Apple labels: 228,539 apps policy-but-not-label', 'an additional 228,539 apps'],
  ['PETS', 2024, 'honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a',
    'template count n=306,404 behind the 65%', '306, 404'],
  ['CCS', 2023, 'policychecker-analyzing-the-gdpr-completeness-of-mobile-apps-privacy-policies',
    'PolicyChecker 98.1% mandatory-requirement violation', '98.1% of them had at least one'],
  ['NDSS', 2017, 'automated-analysis-of-privacy-requirements-for-mobile-apps',
    'Zimmeck mean 1.83 inconsistencies per app', 'mean of 1.83'],
  ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset',
    'Princeton-Leuven corpus reaches back to 1997', 'policies from as early as 1997'],
  ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset',
    'the length/readability series is 2009-2019', 'corpus are from 2009-2019'],
 
  // --- added 2026-09-10 after the citations reviewer found two availability
  //     denominators that named a different population from the paper's own.
  ['NDSS', 2019, 'we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy',
    'Degeling: 6357 is the total of the availability table', 'Total 6357 79.6 % 84.5 %'],
  ['NDSS', 2019, 'we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy',
    'Degeling: the same paper also says 6,759 domains', 'contained 6,759 different domains'],
  ['PETS', 2019, 'maps-scaling-privacy-compliance-analysis-to-a-million-apps',
    'MAPS 50.5% is over the analysed set, not the retrieved set', 'our analysis reveals that only 50.5% of apps have links to privacy policies'],
  ['PETS', 2019, 'maps-scaling-privacy-compliance-analysis-to-a-million-apps',
    'MAPS analysed 1,035,853 of 1,049,790 retrieved', '1,035,853'],
  ['PETS', 2026, 'word-level-annotation-of-gdpr-transparency-compliance-in-privacy-policies-using',
    'cory2026 is word-level GDPR transparency annotation by LLM', 'word-level'],
 
  // --- added 2026-09-10 after the generic review found a back-calculated
  //     denominator: 15,528 was 14,852 + 676, not a number the paper states.
  ['IMC', 2022, 'exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services',
    'Discord: the paper states 15,525 unique active chatbots', 'Unique active chatbots 15,525 100%'],
  ['IMC', 2022, 'exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services',
    'Discord: and separately 14,852 (95.67%) without a policy', 'the remaining 14,852 (95.67%)'],
];
 
let miss = 0;
let colsOnlyMiss = 0;
for (const [venue, year, slug, label, needle] of CHECKS) {
  const rs = renderings(venue, year, slug);
  const n = norm(needle).toLowerCase();
  const hits = Object.entries(rs).filter(([, t]) => t.toLowerCase().includes(n)).map(([r]) => r);
  if (hits.length === 0) {
    miss += 1;
    console.log(`MISS      ${venue} ${year}  ${label}\n            needle: ${needle}`);
  } else {
    if (!hits.includes('paper.cols.txt')) colsOnlyMiss += 1;
    const tag = hits.includes('paper.cols.txt') ? 'OK  cols' : `OK  ${hits[0]} (NOT in .cols)`;
    console.log(`${tag.padEnd(28)} ${venue} ${year}  ${label}`);
  }
}
// SPECIFICITY. Locating a needle proves the string is in the right PDF; it does
// not prove the string is the sentence the page is quoting. A bare "26%" is in
// nine of these papers. A 2026-09-10 review showed that changing the Degeling
// needle from "84.5 %" to "84.9 %" still passed, because both appear in that
// paper's tables. So: report, for every needle, how many OTHER check papers also
// contain it, and how long it is. A needle that is short, numeric and shared is
// weak evidence and the page should not lean on it alone.
{
  const papers = [...new Set(CHECKS.map(([v, y, s]) => `${v}/${y}/${s}`))];
  const texts = new Map(papers.map((k) => {
    const [v, y, s] = k.split('/');
    return [k, Object.values(renderings(v, Number(y), s)).join(' \n ').toLowerCase()];
  }));
  const rows = [];
  for (const [venue, year, slug, label, needle] of CHECKS) {
    const n = norm(needle).toLowerCase();
    const own = `${venue}/${year}/${slug}`;
    const elsewhere = papers.filter((k) => k !== own && texts.get(k).includes(n));
    const numericOnly = !/[a-z]/i.test(needle);
    rows.push([needle.length, numericOnly, elsewhere.length, label, needle]);
  }
  const weak = rows.filter(([len, num, el]) => num && el > 0);
  console.log(`\nSPECIFICITY of the ${CHECKS.length} needles`);
  console.log(`  needles with no letters (pure number/punctuation): ${rows.filter((r) => r[1]).length}`);
  console.log(`  of those, also present in another check paper     : ${weak.length}`);
  console.log(`  needles under 12 characters                       : ${rows.filter((r) => r[0] < 12).length}`);
  if (weak.length) {
    console.log('\n  WEAK — numeric-only and not unique to the paper they are attributed to:');
    for (const [len, , el, label, needle] of weak.sort((a, b) => b[2] - a[2])) {
      console.log(`    ${String(el).padStart(2)} other check papers also contain ${JSON.stringify(needle)}  (${label})`);
    }
    console.log('\n  These are not wrong — each was read in context — but they are the needles');
    console.log('  a future edit could break without this check noticing.');
  }
}
 
console.log(`\n${CHECKS.length} needles, ${CHECKS.length - miss} located, ${miss} MISSING.`);
console.log(`${CHECKS.length - miss - colsOnlyMiss} located in paper.cols.txt; ${colsOnlyMiss} located only outside paper.cols.txt.`);
if (miss > 0) process.exitCode = 1;

The significance test — ''policies_significance.py''

policies_significance.py
"""Fisher's exact test for every subgroup-vs-base-rate claim privacy:policies makes.
 
    node scripts/report_policies.mjs > scripts/report_policies-output.txt
    python3 scripts/policies_significance.py > scripts/policies_significance-output.txt
 
WHY: five of the page's comparisons are of the form "policy papers do X more
often than the rest of the corpus does". Four survive; one does not, and the one
that does not (validation reporting) is the one that reads most like a finding.
 
Counts are NOT hand-keyed. report_policies.mjs emits them as `SIG|...` lines from
the same population objects every other table in that report uses, and this
script parses those lines — so a p-value here cannot drift away from the report
the way a re-typed count can.
 
Each comparison is tested twice, once over POLICY (the stable enum, 102 papers)
and once over UNION (the regex-widened candidate set, 123). The page publishes
the POLICY rows because a rate over UNION is a rate over a set a regex chose;
both are printed here so the reader can see whether that choice changes anything.
 
The base rate is the REST of the corpus — the subgroup's own papers removed —
because comparing a subgroup against a population that contains it shrinks the
difference. The "naive" column shows what the un-removed base rate would have
been, so the size of that effect stays visible.
 
No scipy in this container, so the two-sided Fisher p-value is computed directly
from the hypergeometric distribution (sum of all tables at most as probable as
the observed one).
"""
import sys
from math import comb
 
REPORT = 'scripts/report_policies-output.txt'
 
 
def fisher_two_sided(a, b, c, d):
    n = a + b + c + d
    def p(x):
        return comb(a + b, x) * comb(c + d, a + c - x) / comb(n, a + c)
    obs = p(a)
    lo, hi = max(0, a + c - (c + d)), min(a + b, a + c)
    return sum(p(x) for x in range(lo, hi + 1) if p(x) <= obs * (1 + 1e-9))
 
 
rows = []
for line in open(REPORT, encoding='utf-8'):
    if not line.startswith('SIG|'):
        continue
    _, label, pop, a, an, c, cn, wc, wn = line.rstrip('\n').split('|')
    rows.append((label, pop, int(a), int(an), int(c), int(cn), int(wc), int(wn)))
 
if not rows:
    sys.exit(f'no SIG| lines in {REPORT} — re-run report_policies.mjs first')
 
print(f'parsed {len(rows)} SIG rows from {REPORT}\n')
hdr = (f"{'comparison':46s} {'pop':6s} {'subgroup':>14s} {'rest of corpus':>16s} "
       f"{'p (Fisher, 2-sided)':>20s}  {'naive base':>14s}")
print(hdr)
print('-' * len(hdr))
for label, pop, a, an, c, cn, wc, wn in rows:
    p = fisher_two_sided(a, an - a, c, cn - c)
    verdict = 'significant' if p < 0.05 else 'NOT SIGNIFICANT — do not publish as a movement'
    print(f'{label:46s} {pop:6s} {a}/{an} ({100*a/an:4.1f}%) {c}/{cn} ({100*c/cn:4.1f}%) '
          f'{p:>13.3g}  {verdict}   [{wc}/{wn} = {100*wc/wn:4.1f}%]')
 
print('\n"rest of corpus" removes the subgroup\'s own papers from the base; the')
print('bracketed column is the naive base rate that leaves them in.')

The cell-by-cell table check — ''policies_table_check.mjs''

policies_table_check.mjs
// Assert that privacy:policies' four corpus tables match report_policies.mjs
// CELL BY CELL, not merely that each numeral occurs somewhere in the output.
//
//   node scripts/policies_table_check.mjs
//
// WHY this exists in addition to check_page_numbers.mjs: that guard tests
// whether every numeral on the page appears anywhere in the concatenated script
// output. A mutation test on 2026-09-10 changed the `llm` row's count from 12 to
// 77 and the guard still passed, because 77 occurs elsewhere in the output (the
// PROBE size, and the public-artifacts count). A wrong figure that collides with
// a live one is invisible to a membership test. This one re-derives each row
// from the report and compares positionally.
//
// Exits non-zero on the first mismatch, printing the row it disagrees with.
import fs from 'node:fs';
 
const PAGE = process.argv[2] ?? 'pages/privacy_policies.txt';
const REPORT = process.argv[3] ?? 'scripts/report_policies-output.txt';
const SIG = process.argv[4] ?? 'scripts/policies_significance-output.txt';
 
const page = fs.readFileSync(PAGE, 'utf8');
const report = fs.readFileSync(REPORT, 'utf8');
const sig = fs.readFileSync(SIG, 'utf8');
 
let failures = 0;
const fail = (what, expected, got) => {
  failures += 1;
  console.log(`MISMATCH  ${what}\n  report: ${expected}\n  page  : ${got}`);
};
const ok = (what) => console.log(`OK        ${what}`);
 
// A page table row -> array of trimmed cells, markup stripped.
const cells = (line) =>
  line.replace(/^\||\|$/g, '').split('|')
    .map((c) => c.replace(/\*\*|''|\/\//g, '').trim());
 
const pageRows = (headerMatch) => {
  const lines = page.split('\n');
  const i = lines.findIndex((l) => l.startsWith('^') && headerMatch.test(l));
  if (i < 0) throw new Error(`no page table whose header matches ${headerMatch}`);
  const out = [];
  for (let j = i + 1; j < lines.length && lines[j].startsWith('|'); j += 1) out.push(cells(lines[j]));
  if (!out.length) throw new Error(`page table ${headerMatch} has no rows`);
  return out;
};
 
const reportSection = (marker) => {
  const i = report.indexOf(marker);
  if (i < 0) throw new Error(`report has no section ${JSON.stringify(marker)}`);
  const rest = report.slice(i + marker.length);
  // A section ends at the next "--- " subsection header or "===" banner. The
  // column-rule line under each table is also dashes, so match the space.
  const end = rest.search(/\n(--- |=====)/);
  return (end < 0 ? rest : rest.slice(0, end)).split('\n').filter((l) => l.trim());
};
 
// ---------------------------------------------------------------- 1. by era
{
  const want = new Map();
  for (const l of reportSection('--- method by era')) {
    const m = l.match(/^(\S+)\s+(\d+)\s+\d+ \(\s*([\d.]+)%\)\s+\d+ \(\s*([\d.]+)%\)\s+\d+ \(\s*([\d.]+)%\)\s+\d+ \(\s*([\d.]+)%\)/);
    if (m) want.set(m[1], m.slice(2));
  }
  if (want.size < 5) throw new Error('parsed too few method-by-era rows from the report');
  const rows = pageRows(/classification\.method/);
  if (rows.length !== want.size) fail('method-by-era row count', want.size, rows.length);
  for (const r of rows) {
    const [name, all, ...eras] = r;
    if (!want.has(name)) { fail(`method row "${name}"`, '(no such method in the report)', r.join(' | ')); continue; }
    const w = want.get(name);
    const got = [all, ...eras.map((e) => e.replace('%', ''))];
    const exp = [w[0], ...w.slice(1).map((x) => String(parseFloat(x)))];
    const norm = got.map((x) => String(parseFloat(x)));
    if (norm.join(',') !== exp.join(',')) fail(`method row "${name}"`, exp.join(' | '), norm.join(' | '));
  }
  ok(`method-by-era table: ${rows.length} rows`);
}
 
// ---------------------------------------------------------------- 2. by year
{
  const want = new Map();
  for (const l of reportSection('--- POLICY and UNION per year')) {
    const m = l.match(/^(\d{4})\*?\s+(\d+)\s+(\d+)\s+(\d+)\s+(\d+)\s+([\d.]+)%/);
    if (m) want.set(m[1], [m[2], m[3], m[5], m[6]]); // corpus, POLICY, UNION, share
  }
  const rows = pageRows(/\^ Year \^/);
  for (const r of rows) {
    const year = r[0].replace('*', '');
    if (!want.has(year)) { fail(`year row ${year}`, '(not in the report)', r.join(' | ')); continue; }
    const w = want.get(year);
    const got = [r[1].replace(/,/g, ''), r[2], r[3], r[4].replace('%', '')];
    const exp = [w[0], w[1], w[2], String(parseFloat(w[3]))];
    if (got.map((x) => String(parseFloat(x))).join(',') !== exp.join(',')) {
      fail(`year row ${year}`, exp.join(' | '), got.join(' | '));
    }
  }
  if (rows.length !== want.size) fail('per-year row count', want.size, rows.length);
  ok(`per-year table: ${rows.length} rows`);
}
 
// --------------------------------------------------------------- 3. by venue
{
  const want = new Map();
  for (const l of reportSection('--- UNION per venue')) {
    const m = l.match(/^(\S+)\s+(\d+)\s+(\d+)\s+(\d+)\s+([\d.]+)%/);
    if (m) want.set(m[1], [m[2], m[3], m[4], m[5]]);
  }
  const alias = { 'USENIX Sec': 'USENIX', TheWebConf: 'WWW', 'IEEE S&P': 'IEEE-SP' };
  const rows = pageRows(/\^ Venue \^/);
  for (const r of rows) {
    const v = alias[r[0]] ?? r[0];
    if (!want.has(v)) { fail(`venue row ${r[0]}`, '(not in the report)', r.join(' | ')); continue; }
    const w = want.get(v);
    const got = [r[1], r[2], r[3].replace(/,/g, ''), r[4].replace('%', '')];
    const exp = [w[0], w[1], w[2], String(parseFloat(w[3]))];
    if (got.map((x) => String(parseFloat(x))).join(',') !== exp.join(',')) {
      fail(`venue row ${r[0]}`, exp.join(' | '), got.join(' | '));
    }
  }
  if (rows.length !== want.size) fail('per-venue row count', want.size, rows.length);
  ok(`per-venue table: ${rows.length} rows`);
}
 
// ---------------------------------------------------------- 4. Fisher's exact
{
  // Only the POLICY rows are published; the UNION rows stay on the provenance page.
  const want = [];
  for (const l of sig.split('\n')) {
    const m = l.match(/^(.+?)\s+POLICY\s+(\d+)\/(\d+) \(\s*[\d.]+%\)\s+(\d+)\/(\d+) \(\s*[\d.]+%\)\s+(\S+)/);
    if (m) want.push({ sub: `${m[2]}/${m[3]}`, base: `${m[4]}/${m[5]}`, p: m[6] });
  }
  if (want.length !== 5) throw new Error(`expected 5 POLICY significance rows, parsed ${want.length}`);
  const rows = pageRows(/\^ Property \^/);
  if (rows.length !== want.length) fail('significance row count', want.length, rows.length);
  // The page orders rows for reading; match on the subgroup fraction, not position.
  const bySub = new Map(want.map((w) => [w.sub, w]));
  for (const r of rows) {
    const sub = r[1].replace(/,/g, '').split(' ')[0];
    if (!bySub.has(sub)) { fail(`significance row "${r[0]}"`, '(no POLICY row with that numerator/denominator)', r[1]); continue; }
    const w = bySub.get(sub);
    const base = r[2].replace(/,/g, '').split(' ')[0];
    if (base !== w.base) fail(`significance base for "${r[0]}"`, w.base, base);
    // p is printed on the page in scientific form with superscript digits
    // ("2.65 x 10^-8"); the script prints "2.65e-08". Normalise BOTH to a number
    // and compare the value, exponent included. An earlier form of this check
    // compared only the first three digits and a mutation test on 2026-09-10
    // showed it could not see an exponent moved from -8 to -5.
    const SUP = { '⁰': '0', '¹': '1', '²': '2', '³': '3', '⁴': '4', '⁵': '5', '⁶': '6', '⁷': '7', '⁸': '8', '⁹': '9', '⁻': '-' };
    const pageP = (() => {
      const s = r[3].split('—')[0].trim();
      const sci = s.match(/^([\d.]+)\s*×\s*10([⁻⁰¹²³⁴⁵⁶⁷⁸⁹]+)$/);
      if (sci) {
        const exp = [...sci[2]].map((c) => SUP[c] ?? c).join('');
        return Number(sci[1]) * 10 ** Number(exp);
      }
      return Number(s);
    })();
    const repP = Number(w.p);
    if (!Number.isFinite(pageP)) {
      fail(`significance p for "${r[0]}" is unparseable`, w.p, r[3]);
    } else if (Math.abs(pageP - repP) > Math.abs(repP) * 0.02) {
      fail(`significance p for "${r[0]}"`, `${w.p} (${repP})`, `${r[3]} (${pageP})`);
    }
  }
  ok(`significance table: ${rows.length} rows`);
}
 
console.log(failures ? `\n${failures} MISMATCHES` : '\nAll four tables match the report cell by cell.');
process.exitCode = failures ? 1 : 0;

The external checks — ''policies_external_checks.sh''

policies_external_checks.sh
#!/usr/bin/env bash
# Every external, non-corpus fact privacy:policies states, re-fetched.
#
#   bash scripts/policies_external_checks.sh > scripts/policies_external_checks-output.txt 2>&1
#
# Unauthenticated GitHub API: 60 requests/hour. A re-run that prints <none>
# everywhere has been rate-limited, not found an empty repository — check
# https://api.github.com/rate_limit before believing a negative result.
#
# Rules this script exists to enforce:
#   - print %{http_code} and %{redirect_url} before any byte count, because a
#     302 recorded as an empty body has been published as "empty 200" here before;
#   - never ask GitHub for /releases/latest or /tags to date a repository — the
#     first 404s on tag-only repos and the second is unsorted. Ask the commits
#     API on the repository's own default branch;
#   - print the raw JSON field, not a summary of it.
set -u
API_ERRORS=0
UA='Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0 Safari/537.36'
GH='https://api.github.com/repos'
 
hdr () { printf '\n=== %s ===\n' "$1"; }
 
probe () { # probe <label> <url>
  local out; out=$(curl -sL -A "$UA" -o /tmp/_ext.$$ -w '%{http_code} %{num_redirects} %{url_effective}' "$2")
  printf '%-46s HTTP %s  redirects=%s  bytes=%s\n  final: %s\n' \
    "$1" "${out%% *}" "$(echo "$out" | cut -d' ' -f2)" "$(wc -c < /tmp/_ext.$$)" "$(echo "$out" | cut -d' ' -f3-)"
  rm -f /tmp/_ext.$$
}
 
# GitHub's unauthenticated API allows 60 requests/hour and this script makes
# more than that. Use `gh api` when it is available (it supplies its own
# credentials and never puts a token on a command line); fall back to curl.
# `gh auth status` can exit non-zero while gh is perfectly usable (a stale
# secondary account is enough), so probe the thing we actually need instead.
if command -v gh >/dev/null 2>&1 && gh api rate_limit >/dev/null 2>&1; then
  ghapi () { gh api "$1"; }
else
  ghapi () { curl -s -A "$UA" "https://api.github.com/$1"; }
fi
 
repo () { # repo <owner/name>
  local meta last
  meta=$(ghapi "repos/$1" | python3 -c 'import json,sys
d = json.load(sys.stdin)
if "full_name" not in d:
    print("<API error: %s>" % d.get("message", d))
else:
    print("%s | %s | archived=%s | pushed_at=%s"
          % (d["default_branch"], (d.get("license") or {}).get("spdx_id", "<no license file>"),
             d["archived"], d["pushed_at"]))' 2>&1)
  last=$(ghapi "repos/$1/commits?per_page=1" | python3 -c 'import json,sys
d = json.load(sys.stdin)
print(d[0]["commit"]["committer"]["date"] if isinstance(d, list) and d else "<none>")' 2>&1)
  printf '%-46s default|license|archived|pushed_at: %s\n%-46s last commit on default branch: %s\n' \
    "$1" "$meta" "" "$last"
  # A rate-limited run prints an error string where a date belongs. That has been
  # published as provenance on this wiki before; fail the script instead.
  case "$meta$last" in *"API error"*|*"rate limit"*) API_ERRORS=$((API_ERRORS + 1)) ;; esac
}
 
hdr 'GitHub repositories named on the page (commits API on the default branch)'
for r in citp/privacy-policy-historical citp/PrivacyPoliciesOverTime \
         benandow/PrivacyPolicyAnalysis UCI-Networking-Group/PoliGraph \
         AndyXiang945/PolicyChecker xiaoyue10131748/Lalaine \
         dlgroupuoft/Calpric ducalpha/PurPlianceOpenSource \
         SmartDataAnalytics/Polisis_Benchmark quanmou/polisis; do repo "$r"; done
 
hdr 'pushed_at is not the default branch: per-branch last commit where they differ'
# SmartDataAnalytics/Polisis_Benchmark reports pushed_at 2023-02-02 while its
# master last moved in 2020. Reading pushed_at as "last commit" would call a
# Dependabot security bump on a side branch evidence of maintenance.
for b in $(ghapi "repos/SmartDataAnalytics/Polisis_Benchmark/branches?per_page=100" |
           python3 -c 'import json,sys; print(" ".join(b["name"] for b in json.load(sys.stdin)))' 2>/dev/null); do
  d=$(ghapi "repos/SmartDataAnalytics/Polisis_Benchmark/commits?sha=$b&per_page=1" |
      python3 -c 'import json,sys; d=json.load(sys.stdin); print(d[0]["commit"]["committer"]["date"] if isinstance(d,list) and d else "<none>")' 2>/dev/null)
  printf '  Polisis_Benchmark %-40s %s\n' "$b" "$d"
done
 
hdr 'LICENSE.txt of PurPliance (GitHub reports NOASSERTION)'
curl -s -A "$UA" 'https://raw.githubusercontent.com/ducalpha/PurPlianceOpenSource/main/LICENSE.txt' | head -2 | sed 's/^/  /'
 
hdr 'PurPliance: last three commits on the default branch'
ghapi "repos/ducalpha/PurPlianceOpenSource/commits?per_page=3" |
  python3 -c 'import json,sys
for c in json.load(sys.stdin):
    print("  %s  %s" % (c["commit"]["committer"]["date"], c["commit"]["message"].split(chr(10))[0][:70]))' 2>/dev/null
 
hdr 'Is each lineage repo findable by GitHub search on the tool name?'
for q in PoliGraph PolicyChecker Lalaine Calpric Polisis PolicyLint; do
  printf '%-16s ' "$q"
  curl -s -A "$UA" "https://api.github.com/search/repositories?q=$q+in:name" |
    python3 -c 'import json,sys
d = json.load(sys.stdin)
items = d.get("items", [])
top = items[0]["full_name"] if items else "<none>"
print("total={}  top={}".format(d.get("total_count"), top))'
done
 
hdr 'LICENSE files GitHub cannot auto-detect (first line of each)'
for r in benandow/PrivacyPolicyAnalysis; do
  for f in LICENSE LICENSE.txt LICENSE.md; do
    u="https://raw.githubusercontent.com/$r/master/$f"
    code=$(curl -s -o /tmp/_lic.$$ -w '%{http_code}' -A "$UA" "$u")
    if [ "$code" = 200 ]; then
      printf '%-40s %s -> HTTP 200, first 2 non-empty lines:\n' "$r" "$f"
      grep -v '^[[:space:]]*$' /tmp/_lic.$$ | head -2 | sed 's/^/    /'
    else
      printf '%-40s %s -> HTTP %s\n' "$r" "$f" "$code"
    fi
    rm -f /tmp/_lic.$$
  done
done
 
hdr 'Last-Modified on the W3C P3P pages (the page claims a date here)'
for u in https://www.w3.org/P3P/ https://www.w3.org/TR/P3P11/; do
  printf '%-34s ' "$u"
  curl -sI -A "$UA" "$u" | tr -d '\r' | grep -i '^last-modified' || echo '<no last-modified header>'
done
 
hdr 'PyPI: is there an installable package for the lineage?'
for p in poligraph poligraph-er policylint policheck privbert polisis; do
  code=$(curl -s -o /dev/null -w '%{http_code}' "https://pypi.org/pypi/$p/json")
  printf '%-24s pypi.org/pypi/%s/json -> HTTP %s\n' "$p" "$p" "$code"
done
 
hdr 'Top GitHub name-search hits — is the FIRST hit the real artefact?'
python3 scripts/policies_gh_search.py
 
hdr 'Polisis: the paper has no code repository; the demo site is claimed alive'
probe 'pribot.org/polisis (the Polisis demo)' 'https://pribot.org/polisis'
printf '  <title>: '
curl -sL -A "$UA" 'https://pribot.org/polisis' | grep -o -i '<title>[^<]*</title>' | head -1
 
hdr 'Dataset and model hosts'
probe 'OPP-115 / APP-350 (usableprivacy.org/data)' 'https://usableprivacy.org/data'
probe 'PrivaSeer search engine'                    'https://privaseer.ist.psu.edu/'
probe 'PrivaSeer data + licence page'              'https://privaseer.ist.psu.edu/data'
# The page quotes the corpus size and licence from this sentence; print it, do
# not summarise it. The site's own homepage advertises a DIFFERENT, smaller
# number for its live search index — the two are not the same thing.
curl -sL -A "$UA" 'https://privaseer.ist.psu.edu/data' |
  python3 -c 'import html, re, sys
txt = re.sub(r"\s+", " ", html.unescape(re.sub(r"<[^>]+>", " ", sys.stdin.read())))
for pat in (r"The PrivaSeer corpus is a collection of [\d,]+ privacy policies",
            r"the corpus is available under a [A-Z][^.]{0,40}license"):
    m = re.search(pat, txt)
    print("  " + (m.group(0) if m else "<NOT FOUND: " + pat + ">"))'
probe 'PrivBERT model card (Hugging Face)'         'https://huggingface.co/mukund/privbert'
 
hdr 'Standards and platform policy sources'
probe 'W3C P3P home'                       'https://www.w3.org/P3P/'
probe 'W3C P3P 1.1 (obsoleted note)'       'https://www.w3.org/TR/P3P11/'
probe 'Apple: third-party SDK requirements' 'https://developer.apple.com/news/?id=pvszzano'
probe 'Google Play User Data policy'        'https://support.google.com/googleplay/android-developer/answer/10144311'
probe 'USENIX Sec 2023 artifact index'      'https://secartifacts.github.io/usenixsec2023/results'
 
if [ "$API_ERRORS" -gt 0 ]; then
  printf '\nFAILED: %s repository lookups returned an API error rather than a date.\n' "$API_ERRORS"
  printf 'This output is NOT evidence of anything. Re-run with gh authenticated,\n'
  printf 'or wait for the unauthenticated limit to reset (api.github.com/rate_limit).\n'
  exit 1
fi
printf '\nAll repository lookups returned a date, not an API error.\n'

The GitHub name-search check — ''policies_gh_search.py''

policies_gh_search.py
#!/usr/bin/env python3
"""Is the first GitHub name-search hit for each tool the real artefact?
 
The page claims some of the policy-analysis lineage is hard to find. That claim
is testable: search GitHub for the tool name and look at what comes back.
 
    python3 scripts/policies_gh_search.py
"""
import json
import shutil
import subprocess
import urllib.request
 
 
def api(path):
    """GitHub API. Prefer `gh` (it supplies its own credentials and raises the
    rate limit from 60/hour to 5,000); fall back to an unauthenticated fetch."""
    if shutil.which('gh'):
        r = subprocess.run(['gh', 'api', path], capture_output=True, text=True)
        if r.returncode == 0:
            return json.loads(r.stdout)
    req = urllib.request.Request('https://api.github.com/' + path,
                                 headers={'User-Agent': 'curl'})
    return json.load(urllib.request.urlopen(req))
 
TOOLS = ['Polisis', 'PolicyLint', 'PoliCheck', 'PoliGraph', 'PolicyChecker',
         'Lalaine', 'Calpric', 'PurPliance', 'PolicyComp']
 
for q in TOOLS:
    d = api(f'search/repositories?q={q}+in:name&per_page=3')
    print(f'=== {q}   total={d.get("total_count")}')
    if not d.get('items'):
        print('    <no repository has this name>')
    for i in d['items']:
        print('    %-46s stars=%4d pushed=%s  %s'
              % (i['full_name'], i['stargazers_count'], i['pushed_at'],
                 (i.get('description') or '')[:56]))

The W3C P3P check (Playwright: w3.org 403s curl) — ''policies_w3c_p3p_check.mjs''

policies_w3c_p3p_check.mjs
// w3.org sits behind a Cloudflare challenge that answers curl with HTTP 403,
// so the P3P claims on privacy:policies are checked with Playwright's own
// chromium (PLAYWRIGHT_BROWSERS_PATH=/workspace/.playwright).
//
//   node scripts/policies_w3c_p3p_check.mjs
import { chromium } from 'playwright';
 
const URLS = ['https://www.w3.org/P3P/', 'https://www.w3.org/TR/P3P11/'];
const NEEDLES = [
  'Retired 30 August 2018',
  'should not be referenced in this form or implemented as-is',
];
 
const browser = await chromium.launch();
const page = await browser.newPage();
for (const url of URLS) {
  const resp = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60000 });
  const headers = resp.headers();
  const body = await page.evaluate(() => document.body.innerText);
  console.log(`\n=== ${url}`);
  console.log(`  final URL      : ${page.url()}`);
  console.log(`  HTTP status    : ${resp.status()}`);
  console.log(`  last-modified  : ${headers['last-modified'] ?? '<none>'}`);
  console.log(`  body chars     : ${body.length}`);
  for (const n of NEEDLES) {
    console.log(`  contains ${JSON.stringify(n)}: ${body.includes(n)}`);
  }
  // Any four-digit year the page states about itself, in document order.
  const years = [...new Set((body.match(/\b(19|20)\d\d\b/g) ?? []))].sort();
  console.log(`  years appearing in the visible text: ${years.join(' ')}`);
  const copy = body.match(/.{0,90}(Copyright|©).{0,90}/);
  if (copy) console.log(`  copyright line : ${copy[0].replace(/\s+/g, ' ').trim()}`);
}
await browser.close();
 
// Second pass: print the sentences of w3.org/P3P/ that carry a year, so the
// page's claim about the home page's own currency rests on its words.
const b2 = await chromium.launch();
const pg2 = await b2.newPage();
await pg2.goto('https://www.w3.org/P3P/', { waitUntil: 'domcontentloaded', timeout: 60000 });
const txt = (await pg2.evaluate(() => document.body.innerText)).replace(/\s+/g, ' ');
console.log('\n=== w3.org/P3P/ — every sentence containing a year');
for (const s of txt.split(/(?<=[.!?]) /)) if (/\b(19|20)\d\d\b/.test(s)) console.log('  ' + s.trim());
await b2.close();

The PETS author fetch — ''policies_fetch_pets_authors.py''

policies_fetch_pets_authors.py
#!/usr/bin/env python3
"""Fill out/authors.json for the PETS papers privacy:policies cites.
 
scripts/fetch_authors.py fails on every PETS landing page as of 2026-09-09 (its
parser expects a byline layout petsymposium.org no longer serves) while the pages
themselves return HTTP 200 and carry the authors in `citation_author` meta tags.
This reads those tags instead. It is deliberately narrow: PETS only, and it
refuses to overwrite an entry that is already cached.
 
    python3 scripts/policies_fetch_pets_authors.py PETS/2020/the-privacy-... [...]
"""
import json
import os
import re
import subprocess
import sys
 
UA = ('Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 '
      '(KHTML, like Gecko) Chrome/131.0 Safari/537.36')
IDX = '/workspace/publications_dataset/data/corpus2/.meta'
OUT = 'out/authors.json'
 
 
def landing(venue, year, slug):
    f = os.path.join(IDX, f'{venue}-{year}.json')
    for rec in json.load(open(f))['papers']:
        if rec['slug'] == slug:
            return rec['landingUrl']
    raise KeyError(f'{venue}/{year}/{slug} not in {f}')
 
 
def main():
    cache = json.load(open(OUT)) if os.path.exists(OUT) else {}
    added = 0
    for k in sys.argv[1:]:
        venue, year, slug = k.split('/', 2)
        if venue != 'PETS':
            raise SystemExit(f'{k}: this script is PETS-only')
        if k in cache:
            print(f'CACHED  {k}  {len(cache[k])} authors')
            continue
        url = landing(venue, int(year), slug)
        html = subprocess.run(['curl', '-sL', '-A', UA, url],
                              capture_output=True, text=True, check=True).stdout
        authors = re.findall(r'name="citation_author"\s+content="([^"]+)"', html)
        if not authors:
            raise SystemExit(f'FAILED  {k}  {url}  (no citation_author meta tags in '
                             f'{len(html)} bytes — do not guess, read the page)')
        cache[k] = authors
        added += 1
        print(f'FETCHED {k}  {len(authors)} authors: {"; ".join(authors)}')
    json.dump(cache, open(OUT, 'w'), indent=1, ensure_ascii=False)
    print(f'\n{added} added, {len(cache)} cached total')
 
 
main()

This page's own generator — ''build_provenance_policies.py''

build_provenance_policies.py
#!/usr/bin/env python3
"""Assemble pages/provenance_privacy_policies.txt from the prose log + the real
script sources and their real, unedited outputs.
 
    python3 scripts/build_provenance_policies.py
 
Why a generator rather than a hand-written page: a published script must BE the
committed script, and a published output must be the output that script actually
produced. Pasting either by hand is how an abridged sample ends up reproducing
four of nine table rows. The structural assertions at the end exist because an
in-place generator can silently drop a section.
"""
import glob
import os
import re
import sys
 
PROSE = 'scripts/provenance_policies_prose.txt'
OUT = 'pages/provenance_privacy_policies.txt'
 
# (heading, path, dokuwiki <file> language)
SCRIPTS = [
    ('The report script', 'scripts/report_policies.mjs', 'javascript'),
    ('The fold', 'scripts/policy_fold.mjs', 'javascript'),
    ('The full-text probe', 'scripts/policies_fulltext_probe.mjs', 'javascript'),
    ('The quote check', 'scripts/policies_quotecheck.mjs', 'javascript'),
    ('The significance test', 'scripts/policies_significance.py', 'python'),
    ('The cell-by-cell table check', 'scripts/policies_table_check.mjs', 'javascript'),
    ('The external checks', 'scripts/policies_external_checks.sh', 'bash'),
    ('The GitHub name-search check', 'scripts/policies_gh_search.py', 'python'),
    ('The W3C P3P check (Playwright: w3.org 403s curl)',
     'scripts/policies_w3c_p3p_check.mjs', 'javascript'),
    ('The PETS author fetch', 'scripts/policies_fetch_pets_authors.py', 'python'),
    ('This page\'s own generator', 'scripts/build_provenance_policies.py', 'python'),
]
 
# (producing script, output path, a substring the output MUST end with).
# The third field exists because an empty or truncated output file is the exact
# way this generator can publish nothing and still pass: a 2026-09-10 mutation
# test emptied one of these files and the page shipped an empty <file> block.
OUTPUTS = [
    ('report_policies.mjs', 'scripts/report_policies-output.txt',
     'PROBE 2026 PETS'),
    ('policies_fulltext_probe.mjs', 'scripts/policies_fulltext_probe-output.txt',
     'Calpric (USENIX 2023) —'),
    ('policies_significance.py', 'scripts/policies_significance-output.txt',
     'bracketed column is the naive base rate that leaves them in.'),
    ('policies_table_check.mjs', 'scripts/policies_table_check-output.txt',
     'All four tables match the report cell by cell.'),
    ('policies_quotecheck.mjs', 'scripts/policies_quotecheck-output.txt',
     'located only outside paper.cols.txt.'),
    ('policies_external_checks.sh', 'scripts/policies_external_checks-output.txt',
     'All repository lookups returned a date, not an API error.'),
    ('policies_w3c_p3p_check.mjs', 'scripts/policies_w3c_p3p_check-output.txt',
     'Last updated'),
]
 
 
def read(path):
    with open(path, encoding='utf-8') as f:
        return f.read()
 
 
# A literal closing <file> tag inside a <file> block terminates it and renders
# the rest of the page as markup. The sentinel is assembled from pieces so that
# THIS file — which is itself published in a <file> block below — does not
# contain the literal it is guarding against.
CLOSE_TAG = '</' + 'file' + '>'
 
 
def block(lang, name, body):
    if CLOSE_TAG in body:
        sys.exit(f'{name}: contains a literal closing file tag; it cannot go in a <file> block')
    return f'<file {lang} {name}>\n{body.rstrip()}\n{CLOSE_TAG}\n'
 
 
parts = [read(PROSE).rstrip(), '']
 
parts.append('===== The scripts, as committed =====\n')
parts.append('Every block below is the file itself, inserted by '
             "''scripts/build_provenance_policies.py'' at build time — not a sample, not an "
             'abridgement. Re-running the generator re-inserts whatever is on disk.\n')
for heading, path, lang in SCRIPTS:
    parts.append(f"==== {heading} — ''{os.path.basename(path)}'' ====\n")
    parts.append(block(lang, os.path.basename(path), read(path)))
 
parts.append('===== The outputs, unedited =====\n')
for name, path, tail in OUTPUTS:
    parts.append(f"==== Output of ''{name}'' ====\n")
    parts.append(block('text', os.path.basename(path), read(path)))
 
parts.append('====== References ======\n')
parts.append('<bibtex bibliography></bibtex>\n')
 
# --- Is every output actually an output? An empty or truncated file publishes an
# empty <file> block and every other assertion still passes.
for _name, _path, _tail in OUTPUTS:
    _body = read(_path)
    # 100 bytes is a floor against an emptied file, not a quality bar — one of
    # these outputs is legitimately six lines long. The terminal-string check
    # below is what catches truncation.
    if len(_body.strip()) < 100:
        sys.exit(f'{_path}: {len(_body)} bytes — that is not an output, it is a stub')
    if _tail not in _body:
        sys.exit(f'{_path}: does not contain {_tail!r} — truncated, or the script it '
                 f'came from changed and this expectation was not updated')
    if os.path.getmtime(_path) < os.path.getmtime(f'scripts/{_name}'):
        sys.exit(f'{_path} is older than scripts/{_name} — re-run the script before publishing')
 
# --- Does OUTPUTS cover everything on disk? Comparing len(OUTPUTS) against a
# count derived from OUTPUTS is tautological; compare it against the filesystem.
_on_disk = set(glob.glob('scripts/policies_*-output.txt')) | {'scripts/report_policies-output.txt'}
_listed = {o[1] for o in OUTPUTS}
if _on_disk != _listed:
    sys.exit('OUTPUTS does not match the outputs on disk.\n'
             f'  not published: {sorted(_on_disk - _listed)}\n'
             f'  listed but absent: {sorted(_listed - _on_disk)}')
 
page = '\n'.join(parts).rstrip() + '\n'
 
# --- structural assertions: an in-place generator deletes silently.
h1 = page.count('\n====== ')
h2 = page.count('\n===== ')
# Count only tags at the start of a line: the generator's own source is
# published below and mentions the tag in prose several times.
files = len([l for l in page.split('\n') if l.startswith('<file ')])
closes = len([l for l in page.split('\n') if l == CLOSE_TAG])
tables = page.count('\n^ ')
assert files == len(SCRIPTS) + len(OUTPUTS), f'{files} <file> blocks, expected {len(SCRIPTS) + len(OUTPUTS)}'
assert files == closes, f'{files} opening vs {closes} closing file tags'
# %% inside a <file> block is literal and harmless; only the prose matters. An
# odd count in the prose silently kills every heading, table and citation below it.
prose_only = re.sub(r'(?ms)^<file [^>]*>.*?^' + re.escape(CLOSE_TAG) + r'$', '', page)
assert prose_only.count('%%') % 2 == 0, 'odd number of %% delimiters in the prose'
# The macro fires even inside ''monospace'' — only %%...%% escapes it — so the
# guard looks for an UNESCAPED occurrence, not for the string.
DISCUSSION = '~~' + 'DISCUSSION' + '~~'   # split so this file can be published
assert not re.search(r'(?<!%%)' + re.escape(DISCUSSION), prose_only), \
    'an unescaped discussion macro would put a comment box on a provenance page'
# `h2 >= 10` was the first form of this check and a mutation test on 2026-09-10
# showed it asserts almost nothing: deleting the whole second half of the prose
# still leaves ten headings. Name them instead.
EXPECTED_H2 = [
    'The run',
    'Scope and judgement calls',
    'Populations, and every query behind a figure',
    'The probes, at both widths',
    'Folding, and the complete residue',
    'Quotes and figures checked against the papers',
    'Bibliography',
    'Guards run before publication',
    'External sources',
    'What could not be established',
    'Review',
    'The scripts, as committed',
    'The outputs, unedited',
]
found = re.findall(r'(?m)^=====\s*(.+?)\s*=====$', page)
missing_h2 = [h for h in EXPECTED_H2 if h not in found]
assert not missing_h2, f'sections missing from the page: {missing_h2}'
assert h2 == len(EXPECTED_H2), f'{h2} level-2 headings, expected {len(EXPECTED_H2)}: {found}'
 
with open(OUT, 'w', encoding='utf-8') as f:
    f.write(page)
 
print(f'wrote {OUT}')
print(f'  bytes            {len(page):,}')
print(f'  level-1 headings {h1}')
print(f'  level-2 headings {h2}')
print(f'  <file> blocks    {files}')
print(f'  table rows       {tables}')

The outputs, unedited

Output of ''report_policies.mjs''

report_policies-output.txt
==============================================================================
1. POPULATIONS
==============================================================================
corpus                                                   5859
classified   classification[] non-empty                  4439   75.8% of corpus
POLICY       classification[].target=='privacy-policy'   102   2.3% of classified
PROBE        title+summary regex (candidate set)         77
UNION        POLICY u PROBE                              123
  POLICY n PROBE                                         56
  POLICY only (enum fires, title silent)                 46
  PROBE only (title fires, enum silent)                  21
 
privacy-policy tuples in POLICY                          179
 
--- POLICY-only: the enum fires and the title probe is silent (why a title probe is not enough)
46 papers. Platform measured (enum, multi-valued):
platform              papers of the enum-only set
--------------------  ---------------------------
other-online-service  21
mobile                20
web                   18
iot                   4
offline               1
  2017 NDSS     Automated Analysis of Privacy Requirements for Mobile Apps
  2017 PETS     Analyzing Remote Server Locations for Personal Data Transfers in Mobile Apps
  2019 IMC      Tales from the Porn: A Comprehensive Privacy Analysis of the Web Porn Ecosystem.
  2019 PETS     MAPS: Scaling Privacy Compliance Analysis to a Million Apps
  2019 WWW      Understanding the Evolution of Mobile App Ecosystems: A Longitudinal Measurement Study of Google Play.
  2020 PETS     Angel or Devil? A Privacy Study of Mobile Parental Control Apps
  2020 PETS     CanaryTrap: Detecting Data Misuse by Third-Party Apps on Online Social Networks
  2020 PETS     The Price is (Not) Right: Comparing Privacy in Free and Paid Apps
  2020 USENIX   SkillExplorer: Understanding the Behavior of Skills in Large Scale
  2021 NDSS     Hey Alexa, is this Skill Safe?: Taking a Closer Look at the Alexa Skill Ecosystem
  2022 IEEE-SP  Scraping Sticky Leftovers: App User Information Left on Servers After Account Deletion.
  2022 IMC      Exploring the security and privacy risks of chatbots in messaging services.
  2022 PETS     Checking Websites’ GDPR Consent Compliance for Marketing Emails
  2022 PETS     Developers Say the Darnedest Things: Privacy Compliance Processes Followed by Developers of Child-Directed Apps
  2022 PETS     How Can and Would People Protect From Online Tracking?
  2022 PETS     “We may share the number of diaper changes”: A Privacy and Security Analysis of Mobile Child Care Applications
  2022 USENIX   SkillDetective: Automated Policy-Violation Detection of Voice Assistant Applications in the Wild
  2022 WWW      Et tu, Brute? Privacy Analysis of Government Websites and Mobile Apps.
  2022 WWW      Measuring Alexa Skill Privacy Practices across Three Years.
  2023 CCS      SkillScanner: Detecting Policy-Violating Voice Applications Through Static Analysis at the Development Phase.
  2023 IMC      Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem.
  2023 NDSS     CHKPLUG: Checking GDPR Compliance of WordPress Plugins via Cross-language Code Property Graph
  2023 PETS     Comparing Large-Scale Privacy and Security Notifications
  2023 USENIX   Are You Spying on Me? Large-Scale Analysis on IoT Data Exposure through Companion Apps
  2023 USENIX   The Digital-Safety Risks of Financial Technologies for Survivors of Intimate Partner Violence
  2024 CCS      A First Look at Security and Privacy Risks in the RapidAPI Ecosystem.
  2024 IEEE-SP  Wear's my Data? Understanding the Cross-Device Runtime Permission Model in Wearables.
  2024 IEEE-SP  Understanding the Privacy Practices of Political Campaigns: A Perspective from the 2020 US Election Websites.
  2024 NDSS     MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
  2024 PETS     The Medium is the Message: How Secure Messaging Apps Leak Sensitive Data to Push Notification Services
  2024 PETS     Two Steps Forward and One Step Back: The Right to Opt-out of Sale under CPRA
  2024 PETS     Connecting the Dots: Tracing Data Endpoints in IoT Devices
  2024 USENIX   Arcanum: Detecting and Evaluating the Privacy Risks of Browser Extensions on Web Pages and Web Content
  2025 CCS      The Odyssey of robots.txt Governance: Measuring Convention Implications of Web Bots in Large Language Model Services.
  2025 IEEE-SP  SoK: A Privacy Framework for Security Research Using Social Media Data.
  2025 IEEE-SP  On the (In)Security of LLM App Stores.
  2025 IMC      An In-Depth Investigation of Data Collection in LLM App Ecosystems.
  2025 PETS     Understanding Privacy Norms through Web Forms
  2025 PETS     The Effect of Platform Policies on App Privacy Compliance: A Study of Child-Directed Apps
  2025 PETS     Privacy Settings of Third-Party Libraries in Android Apps: A Study of Facebook SDKs
  2025 PETS     Who’s Watching You Zoom? Investigating Privacy of Third-Party Zoom Apps
  2025 PETS     Surveillance Disguised as Protection: A Comparative Analysis of Sideloaded and In-Store Parental Control Apps
  2025 USENIX   AUTOVR: Automated UI Exploration for Detecting Sensitive Data Flow Exposures in Virtual Reality Apps
  2025 USENIX   I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps
  2026 PETS     Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law’s Impact
  2026 PETS     Chatbot Confessions:~Large-Scale Analysis of Private Data Disclosure in Shared AI Chatbot Conversations
 
--- classification[].target, whole corpus, papers (enum — publishable)
target                 papers  share of 4,439 classified
---------------------  ------  -------------------------
other                  2594    58.4%
vulnerability          883     19.9%
website-category       424     9.6%
user-generated-text    419     9.4%
network-traffic        383     8.6%
domain                 351     7.9%
ip-address             295     6.6%
mobile-app             282     6.4%
web-request            262     5.9%
malware                160     3.6%
privacy-policy         102     2.3%
sdk-or-library         77      1.7%
email-message          54      1.2%
cookie                 53      1.2%
javascript             44      1.0%
consent-notice         39      0.9%
fingerprinting-script  32      0.7%
website-popularity     16      0.4%
dark-pattern           13      0.3%
 
--- POLICY and UNION per year (2026 PROVISIONAL: CCS/IMC 2026 not held, IEEE S&P/WWW 2026 under-selected)
Year   corpus  POLICY  PROBE  UNION  POLICY share of corpus  UNION share of corpus
-----  ------  ------  -----  -----  ----------------------  ---------------------
2014   166     1       1      1      0.6%                    0.6%
2016   182     1       2      2      0.5%                    1.1%
2017   231     3       1      3      1.3%                    1.3%
2018   254     2       2      2      0.8%                    0.8%
2019   402     6       3      6      1.5%                    1.5%
2020   404     8       5      9      2.0%                    2.2%
2021   379     7       8      9      1.8%                    2.4%
2022   546     17      10     19     3.1%                    3.5%
2023   719     14      10     16     1.9%                    2.2%
2024   690     19      16     24     2.8%                    3.5%
2025*  770     16      9      20     2.1%                    2.6%
2026*  415     8       10     12     1.9%                    2.9%
 
--- UNION per venue (denominator: that venue's whole corpus slice)
Venue    UNION  POLICY  venue papers  POLICY share of venue  UNION share of venue
-------  -----  ------  ------------  ---------------------  --------------------
PETS     54     43      510           8.4%                   10.6%
USENIX   27     25      1410          1.8%                   1.9%
CCS      12     8       990           0.8%                   1.2%
NDSS     9      8       701           1.1%                   1.3%
WWW      9      7       843           0.8%                   1.1%
IEEE-SP  7      6       767           0.8%                   0.9%
IMC      5      5       638           0.8%                   0.8%
top venue over the next, POLICY share: PETS / USENIX = 4.8x
top venue over the next, UNION share : PETS / USENIX = 5.5x
 
--- UNION by platform measured (multi-valued: an app+web paper is in two rows)
platform              papers  share of UNION
--------------------  ------  --------------
mobile                60      48.8%
web                   56      45.5%
other-online-service  34      27.6%
offline               10      8.1%
iot                   7       5.7%
not-applicable        1       0.8%
 
web only   39
mobile only 43
both        17
neither     24
 
==============================================================================
2. HOW THE FIELD LABELS POLICY TEXT  (population: POLICY, n=102)
==============================================================================
classification[].method is an ENUM. Multi-valued: a paper with two
privacy-policy tuples using two methods appears in two rows.
 
--- method, all years
method               papers  share of POLICY
-------------------  ------  ---------------
manual-labelling     43      42.2%
heuristic-rules      33      32.4%
supervised-ml        26      25.5%
third-party-service  13      12.7%
llm                  12      11.8%
regex-or-signature   9       8.8%
other                8       7.8%
static-analysis      4       3.9%
unsupervised-ml      3       2.9%
graph-analysis       2       2.0%
curated-database     1       1.0%
 
--- method by era — this is the CURRENCY table the page leans on
method               all  2014-2018 n=7  2019-2021 n=21  2022-2024 n=50  2025-2026* n=24
-------------------  ---  -------------  --------------  --------------  ---------------
manual-labelling     43   4 (57.1%)      9 (42.9%)       19 (38.0%)      11 (45.8%)
heuristic-rules      33   3 (42.9%)      6 (28.6%)       20 (40.0%)      4 (16.7%)
supervised-ml        26   3 (42.9%)      8 (38.1%)       12 (24.0%)      3 (12.5%)
third-party-service  13   0 (0.0%)       3 (14.3%)       7 (14.0%)       3 (12.5%)
llm                  12   0 (0.0%)       0 (0.0%)        1 (2.0%)        11 (45.8%)
regex-or-signature   9    1 (14.3%)      2 (9.5%)        3 (6.0%)        3 (12.5%)
other                8    0 (0.0%)       2 (9.5%)        4 (8.0%)        2 (8.3%)
static-analysis      4    0 (0.0%)       1 (4.8%)        2 (4.0%)        1 (4.2%)
unsupervised-ml      3    0 (0.0%)       0 (0.0%)        3 (6.0%)        0 (0.0%)
graph-analysis       2    0 (0.0%)       0 (0.0%)        2 (4.0%)        0 (0.0%)
curated-database     1    0 (0.0%)       1 (4.8%)        0 (0.0%)        0 (0.0%)
 
--- classification[].validation on privacy-policy tuples (enum; none-reported is a REAL value, not a gap in the data)
validation                  papers  share of POLICY
--------------------------  ------  ---------------
manual-validation           60      58.8%
none-reported               40      39.2%
held-out-test-set           18      17.6%
comparison-to-other-method  11      10.8%
not-applicable              7       6.9%
cross-validation            3       2.9%
 
BASE RATE, all 4,439 papers that classified anything:
validation                  papers  share of classified
--------------------------  ------  -------------------
manual-validation           2265    51.0%
none-reported               1881    42.4%
not-applicable              1160    26.1%
comparison-to-other-method  791     17.8%
held-out-test-set           547     12.3%
cross-validation            314     7.1%
 
--- LLM as the labelling method — POLICY papers whose privacy-policy tuple has method=="llm", listed in full
n=12 of 102 POLICY papers
  2024 IMC      Analyzing Corporate Privacy Policies using AI Chatbots.
  2025 IMC      An In-Depth Investigation of Data Collection in LLM App Ecosystems.
  2025 PETS     Privacy Settings of Third-Party Libraries in Android Apps: A Study of Facebook SDKs
  2025 USENIX   Evaluating Privacy Policies under Modern Privacy Laws At Scale: An LLM-Based Automated Approach
  2025 PETS     Automating Governing Knowledge Commons and Contextual Integrity (GKC-CI) Privacy Policy Annotations with Large Language Models
  2025 PETS     BehaVR: User Identification Based on VR Sensor Data
  2026 PETS     AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents
  2026 PETS     Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law’s Impact
  2026 PETS     Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Language Models
  2026 PETS     Designing Reflective Thinking-Based Contextual Privacy Policy for Mobile Applications
  2026 PETS     Personal Data Flows and Privacy Policy Traceability in Third-party LLM Apps in the GPT Ecosystem
  2026 PETS     Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale
 
==============================================================================
3. WHERE THE POLICIES COME FROM  (population: UNION, n=123)
==============================================================================
 
--- population[].unit (enum)
unit                papers
------------------  ------
mobile-apps         56
documents           45
other               30
websites            28
human-participants  27
domains             10
web-pages           7
code-repositories   7
iot-devices         4
emails              3
browser-extensions  2
network-flows       1
 
--- population[].sourceList, FOLDED through policy_fold.mjs (free text — a ranking, not percentages)
folded source                    papers
-------------------------------  ------
Google Play                      40
custom / hand-built list         32
Alexa list                       16
participant panel                14
OPP-115                          14
other app ecosystem              13
Apple App Store                  10
Tranco                           9
Wayback Machine                  4
Majestic / Umbrella / Quantcast  3
Alexa Skills Store               2
APP-350                          1
Princeton Policies-over-Time     1
 
RESIDUE — sourceList strings matching no fold rule (138 distinct, printed in full):
   4  PoliCheck Dataset
   2  Amazon market
   2  Google market
   2  top.gg
   2  Dcp
   2  TheInternetBackup public domain list
   2  10 mainstream VR platforms
   2  IoT Inspector
   2  Alexa skill marketplaces
   2  third-party GPT stores and OpenAI's official GPT store
   2  Zoom Marketplace
   2  Federal Reserve list of the largest commercial banks
   2  privacy policy corpus
   2  local Craigslist and sub-Reddit forums
   2  3u
   2  social media, LinkedIn groups, Discord servers, university bulletin boards, and snowball sampling
   1  Quantified Self community's Guide to Self-Tracking Tools
   1  Twitter
   1  three websites specialized in aggregating, recommending, and classifying pornographic content
   1  sanitized dataset combining three pornographic-web sources
   1  government documents and observed practices
   1  company websites
   1  APKPure
   1  prior-research list compiled by crawling the web
   1  AppCensus
   1  final filtered dataset
   1  Collaborative List of Open-Source iOS Apps and other public repositories
   1  Upwork, developer websites, Reddit, and iOSoho
   1  SDK vendor websites
   1  Who-TracksMe
   1  Disconnect Tracking Protection
   1  Evidon Global Opt-out
   1  DuckDuckGo Tracker Radar
   1  merged tracker databases
   1  SimilarWeb top websites in the US
   1  Evidon Global Opt-out list
   1  Discord privacy-policy pages
   1  SimilarWeb
   1  vendor websites and other policy-search resources
   1  device privacy policies identified in the availability analysis
   1  smart-home privacy policies
   1  vendor websites
   1  Singanamalla et al. government website dataset
   1  SkillExplorer dataset from US marketplace
   1  US skills store
   1  PPCrawl
   1  dataset released in [30]
   1  Alexa skill marketplace
   1  IoT companion apps collected in the wild (Dcp)
   1  research papers and news reports
   1  ACM Digital Library, IEEE Xplore, and USENIX Paper Proceedings
   1  AlternativeTo, Top Best Alternatives, and Games Like
   1  website [1] and research paper [31]
   1  CCPA, GDPR, PIPEDA, and VCDPA
   1  government websites
   1  CONLL2012
   1  Apidog
   1  apideck
   1  dataset used in a recent study [81]
   1  obtained privacy-policy links
   1  randomly selected privacy-policy links
   1  randomly selected privacy policies
   1  randomly selected sentences
   1  Vanguard Russell 3000 ETF
   1  Wagner's corpus
   1  Free Company Dataset
   1  UNSW IoT Analytics
   1  YourThings IoTFinder
   1  joinmastodon.org and instances.social
   1  centralized registries
   1  referral network
   1  referral network and centralized registries
   1  Apple Developer Documentation
   1  randomly selected SDKs
   1  Google search engine and crowd-knowledge platforms
   1  PoliCheck test set
   1  three existing GDPR privacy-policy datasets
   1  manual skill-output collection
   1  four academic databases
   1  GPT Store
   1  FlowGPT
   1  Poe
   1  Coze
   1  Cici
   1  Character.AI
   1  USENIX Security, IEEE S&P, ACM CCS, and NDSS
   1  AliPay daily active users
   1  AliPay volunteer participants
   1  AliPay consented volunteer dataset
   1  Fortune 2024 top 1K US companies
   1  Fortune 2024 top 500 EU companies
   1  PPGDPR
   1  C3PA
   1  Anthropic, OpenAI, Gemini, and DeepSeek privacy policies
   1  Promptfoo
   1  Presidio-research
   1  Apple's public sitemap
   1  apps identified through the longitudinal analysis
   1  email survey respondents willing to participate
   1  FDIC BankFind Suite
   1  university students
   1  Google web search results
   1  personal contacts, snowball sampling, and targeted email invitations
   1  WeChat groups
   1  CPP4APP
   1  CA4P-483
   1  MAPP Corpus
   1  GPTStore.ai
   1  OpenAI official store
   1  BeeTrove
   1  combined GPTStore.ai/OpenAI and BeeTrove dataset
   1  web archive snapshots and search engine indexes
   1  Domains with robots.txt files
   1  Domains with English policy documents
   1  Slack channels, mailing groups, and flyers
   1  WearBench dataset
   1  FEC
   1  Ballotpedia
   1  on-campus bookstore
   1  random passers-by and visitors to the bookstore
   1  MobiPurpose
   1  Google Scholar, Semantic Scholar, and WorldWideScience
   1  authors of 125 papers
   1  iOS SDK Ranking
   1  Apptopia
   1  APKCombo
   1  Princeton Privacy Crawl (PPCrawl)
   1  ground truth data set employed by Cui et al.
   1  local Discord server
   1  professional network and OpenDP mailing list
   1  enrolled participants' peer recommendations
   1  FCWs dataset
   1  Fraudulent and Legitimate Online Shops Dataset
   1  clusters of unfavorable and benign terms
   1  traffic analysis dataset [116]
   1  Similarweb
   1  research teams
   1  research team workshop
 
==============================================================================
4. THE TOOL LINEAGE  (population: whole corpus, so a tool used outside UNION is visible)
==============================================================================
Matched against tools[].name, otherToolsMentioned[].name,
classification[].resourceName/targetDetail, population[].sourceList,
detection[].phenomenon/technique. Counts are PAPERS.
artefact (first paper)                  papers, corpus  of those in UNION  year range of use  used 2024+
--------------------------------------  --------------  -----------------  -----------------  ----------
Privee (USENIX 2014)                    1               1                  2014-2014          0
OPP-115 corpus (ACL 2016)               14              14                 2017-2026          6
Polisis / PriBot (USENIX 2018)          13              13                 2018-2025          5
PolicyLint (USENIX 2019)                18              17                 2019-2025          3
MAPS (PETS 2019)                        2               2                  2019-2024          1
APP-350 corpus (2019)                   1               1                  2023-2023          0
PoliCheck (USENIX 2020)                 9               9                  2020-2025          4
PurPliance (2021)                       5               5                  2021-2023          0
PrivBERT (2021)                         3               3                  2024-2024          3
Calpric (USENIX 2023)                   1               1                  2023-2023          0
PoliGraph / PoliGraph-er (USENIX 2023)  6               6                  2023-2026          5
PolicyChecker (CCS 2023)                1               1                  2023-2023          0
Lalaine (USENIX 2023)                   2               2                  2023-2024          1
PolicyComp (USENIX 2023)                1               1                  2023-2023          0
PrivaSeer                               0               0                  —                  0
 
--- tools[] + otherToolsMentioned in UNION, FOLDED (free text — ranking only)
folded tool family           papers in UNION
---------------------------  ---------------
Selenium                     29
spaCy                        23
LLM (commercial API)         21
BERT family (non-privacy)    18
PolicyLint                   17
Polisis / PriBot             12
NLTK                         10
boilerplate stripper         10
language detection           10
LLM (open weights)           10
Stanford CoreNLP / AllenNLP  9
BeautifulSoup                7
PoliCheck                    7
OpenWPM                      6
Playwright                   6
PoliGraph                    5
PurPliance                   4
Puppeteer                    3
PrivBERT                     3
readability metric           2
PolicyChecker                1
 
RESIDUE — 572 distinct tool names matching no fold rule. Top 40 by frequency:
  11  Frida
   8  VirusTotal
   7  Qualtrics
   6  Google Chrome
   6  Firefox
   6  Apktool
   6  Amazon Mechanical Turk
   6  mitmproxy
   6  Prolific
   5  Python
   5  EasyList
   4  scikit-learn
   4  Fiddler
   4  adb
   4  WHOIS
   4  fastText
   4  Wayback Machine
   4  Docker
   4  FlowDroid
   4  Soot
   4  GloVe
   4  tldextract
   4  google-play-scraper
   4  Google Search
   4  Zoom
   3  Prolific Academic
   3  Monkey
   3  fuzzywuzzy
   3  Chrome
   3  TF-IDF
   3  Google Translate
   3  logistic regression
   3  LibRadar
   3  Random Forest
   3  CocoaPods
   3  Chromium
   3  MobSF
   3  Crunchbase
   3  custom web crawler
   2  Porter stemmer
  … and 532 more, each in 1-2 papers.
 
==============================================================================
5. MEASURED RESULTS: POLICY AVAILABILITY BY ECOSYSTEM
==============================================================================
Hand-keyed from detection[].prevalence tuples in UNION papers. The map
below is INSIDE this script on purpose: the page must not carry a
per-paper figure the script cannot print. Every row was read back
against the paper's own quote; see report_policies_quotecheck.mjs.
citekey              venue        ecosystem                                denominator (paper's own)                                                                                                                    policy availability
-------------------  -----------  ---------------------------------------  -------------------------------------------------------------------------------------------------------------------------------------------  ------------------------------------------------------------------------------------------------------------------------------
degeling2019_value   NDSS 2019    web, EU                                  6,357 (the total of the paper's own availability table); the paper also states 6,759 domains in its January lists and 6,579 in its abstract  84.5% had a policy after 25 May 2018, up from 79.6% in January
vallina2019_porn     IMC 2019     web, adult sites                         6,843 pornographic websites                                                                                                                  only 16% had an accessible privacy policy
cui2025_privacy      PETS 2025    web, sites with a PI-collecting form     10,143 websites                                                                                                                              94.2% (9,559) had a privacy-policy link
zimmeck2019_maps     PETS 2019    Android                                  1,035,853 analysed apps, of 1,049,790 retrieved                                                                                              only 50.5% had a policy link on the Play Store page
pan2024_trap         USENIX 2024  Android                                  99,194 usable apps — the paper divides link failures by its app count                                                                        37.5% (37,150/99,194) of policy links led to an unavailable page; separately 15.7% (15,572/99,194) of apps have no link at all
manandhar2022_smart  USENIX 2022  smart-home vendors                       596 vendors on 7 platforms                                                                                                                   48.99% had a device-applicable policy; 10.57% had none at all
lentzsch2021_alexa   NDSS 2021    Alexa skills                             150,708 skills across 7 country stores                                                                                                       36,475 (24.2%) provided a policy link
yan2024_quality      PETS 2024    Alexa skills                             65,195 skills                                                                                                                                21,063 of 65,195 provided a policy link
zhan2024_vpvet       CCS 2024     VR apps                                  11,923 apps on 10 VR platforms                                                                                                               29.5% had a findable privacy policy
edu2022_exploring    IMC 2022     Discord chatbots requesting permissions  15,525 unique active chatbots (the paper's own Table 2 total); its 14,852-without figure does not reconcile with it                          676 (4.35%) had a policy; 14,852 (95.67%) did not
wu2025_depth         IMC 2025     GPT Actions                              Actions declaring a legal_info_url                                                                                                           93.96% of those policies were reachable
 
rows: 11; every one resolves to a corpus paper.
 
==============================================================================
6. MEASURED RESULTS: POLICY-VERSUS-BEHAVIOUR CONSISTENCY
==============================================================================
citekey                  venue        denominator (paper's own)                       finding
-----------------------  -----------  ----------------------------------------------  ----------------------------------------------------------------------------------------------------------------------------
zimmeck2017_automated    NDSS 2017    9,050 apps with policies                        mean 1.83 potential inconsistencies per app
libert2018_automated     WWW 2018     1,807,491 identified third-party transmissions  only 14.80% were disclosed in the policy
andow2019_policylint     USENIX 2019  11,430 policies                                 14.2% (1,618) contained logical contradictions; 17.7% (2,028) contradictions or narrowing definitions
andow2020_actions        USENIX 2020  13,796 applications / 45,603 data flows         42.4% of apps had an omitted or incorrect disclosure; 31.1% of flows were omitted; only 0.5% of flows were clearly disclosed
trimananda2022_ovrseen   USENIX 2022  1,135 data flows in Oculus VR apps              68% (776) inconsistent disclosures
cui2023_poligraph        USENIX 2023  1,566 mapped statement pairs                    13.5% (211) conflicting; 25.5% (1,339/5,255) of policies define a term differently from the CCPA-based ontology
xiang2023_policychecker  CCS 2023     163,068 analysable policies                     99.3% incomplete under GDPR; 98.1% had at least one mandatory-requirement violation
xiao2023_lalaine         USENIX 2023  5,102 fully tested iOS apps                     3,423 (67.1%) non-compliant privacy labels
samarin2023_lessons      PETS 2023    69 apps with CCPA disclosures                   80% (55) collected an identifier they did not disclose
ali2024_honesty          PETS 2024    iOS apps labelled "Data Not Collected"          97% had policy statements indicating data collection
cui2025_privacy          PETS 2025    websites collecting PI via web forms            phi coefficient between observed collection and PoliGraph-er disclosure was < 0.20 for every PI type
 
rows: 11; every one resolves to a corpus paper.
 
==============================================================================
7. WHAT ELSE THE UNION MEASURES
==============================================================================
 
--- UNION papers that are also in the `legal` population (legal[] non-empty)
70 of 123 UNION papers assess a law; base rate 402 of 5859 (6.9%) corpus-wide.
law (free text, unfolded — ranking only)       papers
---------------------------------------------  ------
GDPR                                           52
CCPA                                           24
COPPA                                          17
CalOPPA                                        4
HIPAA                                          3
General Data Protection Regulation (GDPR)      2
ePrivacy Directive                             2
FTC Act                                        2
California Consumer Privacy Act (CCPA)         2
CPRA                                           2
Directive 95/46/EC                             1
DOPPA                                          1
Data Protection Directive 95/46/EC (DPD'25.1)  1
ePrivacy directive                             1
Digital Economy Act 2017                       1
 
--- human annotation in UNION — policy labelling is a hand-coding literature
UNION papers with humanAnnotation[]            120 of 123  97.6%
  …stating an agreement metric                 45  37.5% of annotated
  …stating an annotator count                  85  70.8% of annotated
 
BASE RATE, all 3318 papers that coded data by hand:
  …stating an agreement metric                 512  15.4%
 
--- temporal[].mode in UNION (enum) vs corpus base rate — is this a longitudinal literature?
UNION papers with a temporal[] tuple 118 of 123; corpus 5342 of 5859
temporal.mode       UNION  share of UNION w/ temporal  corpus  share of corpus w/ temporal
------------------  -----  --------------------------  ------  ---------------------------
live-crawl          84     71.2%                       1261    23.6%
existing-dataset    34     28.8%                       2534    47.4%
active-probing      17     14.4%                       1780    33.3%
web-archive         12     10.2%                       68      1.3%
passive-collection  10     8.5%                        1012    18.9%
 
--- artifact availability in UNION vs corpus base rate (enum)
availability                UNION  share  corpus  share
--------------------------  -----  -----  ------  -----
public                      77     63.1%  2845    51.4%
none-mentioned              36     29.5%  2183    39.4%
promised-not-yet-available  6      4.9%   289     5.2%
on-request                  1      0.8%   91      1.6%
restricted                  0      0.0%   79      1.4%
explicitly-withheld         2      1.6%   52      0.9%
 
==============================================================================
7b. SIGNIFICANCE INPUTS — parsed by policies_significance.py
==============================================================================
SIG|temporal.mode == web-archive|POLICY|11|101|57|5241|68|5342
SIG|temporal.mode == web-archive|UNION|12|118|56|5224|68|5342
SIG|humanAnnotation states an agreement metric|POLICY|38|101|474|3217|512|3318
SIG|humanAnnotation states an agreement metric|UNION|45|120|467|3198|512|3318
SIG|artifacts.availability == public|POLICY|63|101|2782|5438|2845|5539
SIG|artifacts.availability == public|UNION|77|122|2768|5417|2845|5539
SIG|assesses a law (legal[] non-empty)|POLICY|64|102|338|5757|402|5859
SIG|assesses a law (legal[] non-empty)|UNION|70|123|332|5736|402|5859
SIG|classification.validation == none-reported|POLICY|40|102|1823|4337|1881|4439
SIG|classification.validation == none-reported|UNION|40|102|1818|4321|1881|4439
 
Columns: SIG|comparison|population|sub hits|sub n|rest-of-corpus hits|rest n|whole-corpus hits|whole n
"rest" excludes the subgroup itself; "whole" is the figure a naive base rate would use.
 
==============================================================================
8. THE UNION, IN FULL — the audit surface for every count above
==============================================================================
BOTH  2014 USENIX   Privee: An Architecture for Automatically Analyzing Web Privacy Policies
BOTH  2016 PETS     Privacy Challenges in the Quantified Self Movement – An EU Perspective
PROBE 2016 PETS     Crowdsourcing for Context: Regarding Privacy in Beacon Encounters via Contextual Integrity
ENUM  2017 NDSS     Automated Analysis of Privacy Requirements for Mobile Apps
ENUM  2017 PETS     Analyzing Remote Server Locations for Personal Data Transfers in Mobile Apps
BOTH  2017 PETS     Cross-Device Tracking: Measurement and Disclosures
BOTH  2018 USENIX   Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep Learning
BOTH  2018 WWW      An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies.
ENUM  2019 IMC      Tales from the Porn: A Comprehensive Privacy Analysis of the Web Porn Ecosystem.
BOTH  2019 NDSS     we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy
ENUM  2019 PETS     MAPS: Scaling Privacy Compliance Analysis to a Million Apps
BOTH  2019 USENIX   HideMyApp: Hiding the Presence of Sensitive Apps on Android
BOTH  2019 USENIX   PolicyLint: Investigating Internal Privacy Policy Contradictions on Google Play
ENUM  2019 WWW      Understanding the Evolution of Mobile App Ecosystems: A Longitudinal Measurement Study of Google Play.
BOTH  2020 PETS     An Analysis of the Current State of the Consumer Credit Reporting System in China
ENUM  2020 PETS     Angel or Devil? A Privacy Study of Mobile Parental Control Apps
ENUM  2020 PETS     CanaryTrap: Detecting Data Misuse by Third-Party Apps on Online Social Networks
BOTH  2020 PETS     How private is your period?: A systematic analysis of menstrual app privacy policies
ENUM  2020 PETS     The Price is (Not) Right: Comparing Privacy in Free and Paid Apps
BOTH  2020 PETS     The Privacy Policy Landscape After the GDPR
BOTH  2020 USENIX   Actions Speak Louder than Words: Entity-Sensitive Privacy Policy and Data Flow Analysis with PoliCheck
ENUM  2020 USENIX   SkillExplorer: Understanding the Behavior of Skills in Large Scale
PROBE 2020 WWW      Finding a Choice in a Haystack: Automatic Extraction of Opt-Out Statements from Privacy Policy Text.
BOTH  2021 CCS      Automated Privacy Policy Annotation with Information Highlighting Made Practical Using Deep Representations.
PROBE 2021 CCS      Consistency Analysis of Data-Usage Purposes in Mobile Apps.
ENUM  2021 NDSS     Hey Alexa, is this Skill Safe?: Taking a Closer Look at the Alexa Skill Ecosystem
BOTH  2021 NDSS     PrivacyFlash Pro: Automating Privacy Policy Generation for Mobile Apps
BOTH  2021 PETS     Automated Extraction and Presentation of Data Practices in Privacy Policies
PROBE 2021 PETS     Defining Privacy: How Users Interpret Technical Terms in Privacy Policies
BOTH  2021 USENIX   Understanding Malicious Cross-library Data Harvesting on Android
BOTH  2021 WWW      Have You been Properly Notified? Automatic Compliance Analysis of Privacy Policy Text with GDPR Article 13.
BOTH  2021 WWW      Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset.
BOTH  2022 CCS      Do Opt-Outs Really Opt Me Out?
ENUM  2022 IEEE-SP  Scraping Sticky Leftovers: App User Information Left on Servers After Account Deletion.
ENUM  2022 IMC      Exploring the security and privacy risks of chatbots in messaging services.
ENUM  2022 PETS     Checking Websites’ GDPR Consent Compliance for Marketing Emails
ENUM  2022 PETS     Developers Say the Darnedest Things: Privacy Compliance Processes Followed by Developers of Child-Directed Apps
BOTH  2022 PETS     Exploring the Privacy Concerns of Bystanders in Smart Homes from the Perspectives of Both Owners and Bystanders
ENUM  2022 PETS     How Can and Would People Protect From Online Tracking?
BOTH  2022 PETS     Leave No Data Behind – Empirical Insights into Data Erasure from Online Services
ENUM  2022 PETS     “We may share the number of diaper changes”: A Privacy and Security Analysis of Mobile Child Care Applications
BOTH  2022 PETS     Who Knows I Like Jelly Beans? An Investigation Into Search Privacy
PROBE 2022 PETS     How Usable Are iOS App Privacy Labels?
PROBE 2022 PETS     Keeping Privacy Labels Honest
BOTH  2022 USENIX   A Large-scale Investigation into Geodifferences in Mobile Apps
BOTH  2022 USENIX   Electronic Monitoring Smartphone Apps: An Analysis of Risks from Technical, Human-Centered, and Legal Perspectives
BOTH  2022 USENIX   OVRseen: Auditing Network Traffic and Privacy Policies in Oculus VR
ENUM  2022 USENIX   SkillDetective: Automated Policy-Violation Detection of Voice Assistant Applications in the Wild
BOTH  2022 USENIX   Smart Home Privacy Policies Demystified: A Study of Availability, Content, and Coverage
ENUM  2022 WWW      Et tu, Brute? Privacy Analysis of Government Websites and Mobile Apps.
ENUM  2022 WWW      Measuring Alexa Skill Privacy Practices across Three Years.
BOTH  2023 CCS      PolicyChecker: Analyzing the GDPR Completeness of Mobile Apps' Privacy Policies.
ENUM  2023 CCS      SkillScanner: Detecting Policy-Violating Voice Applications Through Static Analysis at the Development Phase.
PROBE 2023 CCS      Poster: Longitudinal Measurement of the Adoption Dynamics in Apple's Privacy Label Ecosystem.
BOTH  2023 IEEE-SP  Detection of Inconsistencies in Privacy Practices of Browser Extensions.
ENUM  2023 IMC      Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem.
ENUM  2023 NDSS     CHKPLUG: Checking GDPR Compliance of WordPress Plugins via Cross-language Code Property Graph
BOTH  2023 PETS     Evolution of Composition, Readability, and Structure of Privacy Policies over Two Decades
BOTH  2023 PETS     Lessons in VCR Repair: Compliance of Android App Developers with the California Consumer Privacy Act (CCPA)
ENUM  2023 PETS     Comparing Large-Scale Privacy and Security Notifications
PROBE 2023 PETS     Researchers’ Experiences in Analyzing Privacy Policies: Challenges and Opportunities
ENUM  2023 USENIX   Are You Spying on Me? Large-Scale Analysis on IoT Data Exposure through Companion Apps
ENUM  2023 USENIX   The Digital-Safety Risks of Financial Technologies for Survivors of Intimate Partner Violence
BOTH  2023 USENIX   Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active Learning
BOTH  2023 USENIX   Lalaine: Measuring and Characterizing Non-Compliance of Apple Privacy Labels
BOTH  2023 USENIX   PoliGraph: Automated Privacy Policy Analysis using Knowledge Graphs
BOTH  2023 USENIX   POLICYCOMP: Counterpart Comparison of Privacy Policies Uncovers Overbroad Personal Data Collection Practices
ENUM  2024 CCS      A First Look at Security and Privacy Risks in the RapidAPI Ecosystem.
BOTH  2024 CCS      VPVet: Vetting Privacy Policies of Virtual Reality Apps.
PROBE 2024 CCS      Measuring Compliance Implications of Third-party Libraries' Privacy Label Disclosure Guidelines.
PROBE 2024 CCS      Are We Getting Well-informed? An In-depth Study of Runtime Privacy Notice Practice in Mobile Apps.
ENUM  2024 IEEE-SP  Wear's my Data? Understanding the Cross-Device Runtime Permission Model in Wearables.
ENUM  2024 IEEE-SP  Understanding the Privacy Practices of Political Campaigns: A Perspective from the 2020 US Election Websites.
BOTH  2024 IMC      Analyzing Corporate Privacy Policies using AI Chatbots.
ENUM  2024 NDSS     MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots
BOTH  2024 NDSS     Towards Automated Regulation Analysis for Effective Privacy Compliance
BOTH  2024 PETS     Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps' Privacy Policies
BOTH  2024 PETS     On the Quality of Privacy Policy Documents of Virtual Personal Assistant Applications
ENUM  2024 PETS     The Medium is the Message: How Secure Messaging Apps Leak Sensitive Data to Push Notification Services
ENUM  2024 PETS     Two Steps Forward and One Step Back: The Right to Opt-out of Sale under CPRA
BOTH  2024 PETS     A Bilingual Longitudinal Analysis of Privacy Policies Measuring the Impacts of the GDPR and the CCPA/CPRA
ENUM  2024 PETS     Connecting the Dots: Tracing Data Endpoints in IoT Devices
BOTH  2024 PETS     Privacy Policies on the Fediverse: A Case Study of Mastodon Instances
PROBE 2024 PETS     Data Safety vs. App Privacy: Comparing the Usability of Android and iOS Privacy Labels
ENUM  2024 USENIX   Arcanum: Detecting and Evaluating the Privacy Risks of Browser Extensions on Web Pages and Web Content
BOTH  2024 USENIX   iHunter: Hunting Privacy Violations at Scale in the Software Supply Chain on iOS
BOTH  2024 USENIX   Swipe Left for Identity Theft: An Analysis of User Data Privacy Risks on Location-based Dating Apps
BOTH  2024 USENIX   Is It a Trap? A Large-scale Empirical Study And Comprehensive Assessment of Online Automated Privacy Policy Generators for Mobile Apps
PROBE 2024 USENIX   Abandon All Hope Ye Who Enter Here: A Dynamic, Longitudinal Investigation of Android's Data Safety Section
PROBE 2024 USENIX   Unpacking Privacy Labels: A Measurement and Developer Perspective on Google's Data Safety Section
BOTH  2024 WWW      Understanding GDPR Non-Compliance in Privacy Policies of Alexa Skills in European Marketplaces.
BOTH  2025 CCS      Layered, Overlapping, and Inconsistent: A Large-Scale Analysis of the Multiple Privacy Policies and Controls of U.S. Banks.
ENUM  2025 CCS      The Odyssey of robots.txt Governance: Measuring Convention Implications of Web Bots in Large Language Model Services.
ENUM  2025 IEEE-SP  SoK: A Privacy Framework for Security Research Using Social Media Data.
ENUM  2025 IEEE-SP  On the (In)Security of LLM App Stores.
PROBE 2025 IEEE-SP  Let's Get Visual - Testing Visual Analogies and Metaphors for Conveying Privacy Policies and Data Handling Information.
ENUM  2025 IMC      An In-Depth Investigation of Data Collection in LLM App Ecosystems.
BOTH  2025 NDSS     SKILLPoV: Towards Accessible and Effective Privacy Notice for Amazon Alexa Skills
PROBE 2025 NDSS     PolicyPulse: Precision Semantic Role Extraction for Enhanced Privacy Policy Comprehension
ENUM  2025 PETS     Understanding Privacy Norms through Web Forms
ENUM  2025 PETS     The Effect of Platform Policies on App Privacy Compliance: A Study of Child-Directed Apps
ENUM  2025 PETS     Privacy Settings of Third-Party Libraries in Android Apps: A Study of Facebook SDKs
ENUM  2025 PETS     Who’s Watching You Zoom? Investigating Privacy of Third-Party Zoom Apps
BOTH  2025 PETS     Automating Governing Knowledge Commons and Contextual Integrity (GKC-CI) Privacy Policy Annotations with Large Language Models
BOTH  2025 PETS     BehaVR: User Identification Based on VR Sensor Data
ENUM  2025 PETS     Surveillance Disguised as Protection: A Comparative Analysis of Sideloaded and In-Store Parental Control Apps
PROBE 2025 PETS     "Free WiFi is not ultimately free": Privacy Perceptions of Users in the US regarding City-wide WiFi Services
ENUM  2025 USENIX   AUTOVR: Automated UI Exploration for Detecting Sensitive Data Flow Exposures in Virtual Reality Apps
ENUM  2025 USENIX   I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps
BOTH  2025 USENIX   Evaluating Privacy Policies under Modern Privacy Laws At Scale: An LLM-Based Automated Approach
PROBE 2025 WWW      Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale.
BOTH  2026 PETS     AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents
BOTH  2026 PETS     ``Because I didn't touch these and even don't know why I should to change these'': Why App Developers Do (Not) Update Apple’s Privacy Labels
ENUM  2026 PETS     Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law’s Impact
BOTH  2026 PETS     Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Language Models
BOTH  2026 PETS     Designing Reflective Thinking-Based Contextual Privacy Policy for Mobile Applications
BOTH  2026 PETS     Personal Data Flows and Privacy Policy Traceability in Third-party LLM Apps in the GPT Ecosystem
BOTH  2026 PETS     Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale
ENUM  2026 PETS     Chatbot Confessions:~Large-Scale Analysis of Private Data Disclosure in Shared AI Chatbot Conversations
PROBE 2026 PETS     ``We Need a Standard'': Toward an Expert–Informed Privacy Label for Differential Privacy
PROBE 2026 PETS     Privacy by Voice: Designing Usable Privacy Notices for the Voice Interface
PROBE 2026 PETS     From Lines of Code to Lines of Policy? Exploring Software Developers’ Perceptions of Their Privacy Policy–Related Activities
PROBE 2026 PETS     Are Bite-Size Data Safety Details a Healthy Diet for Android Telehealth App Users? Impacts of Privacy Nutrition Labels on Users’ Privacy Perceptions

Output of ''policies_fulltext_probe.mjs''

policies_fulltext_probe-output.txt
UNION = 123 papers; full text present for 123; missing 0.
Every count below is over the 123 papers with full text.
 
probe (narrow form is what the page quotes)                                                                         narrow  share  wide  share
------------------------------------------------------------------------------------------------------------------  ------  -----  ----  -----
finds the policy by LINK TEXT / anchor keyword                                                                      12      9.8%   112   91.1%
names a link-detection SEED PHRASE list                                                                             1       0.8%   2     1.6%
follows the policy link and reports FAILURES (404 / dead / unreachable)                                             7       5.7%   64    52.0%
handles a policy served as a PDF                                                                                    9       7.3%   22    17.9%
states the LANGUAGE of the policies it analysed  <-- not a superset: different question                             27      22.0%  16    13.0%
analyses policies in more than one language                                                                         30      24.4%  67    54.5%
strips BOILERPLATE / extracts the main content of the policy page                                                   13      10.6%  20    16.3%
measures READABILITY of the policy                                                                                  17      13.8%  68    55.3%
compares the policy against OBSERVED BEHAVIOUR (traffic, code, or storage)  <-- not a superset: different question  80      65.0%  29    23.6%
handles the policy being VERSIONED / changing under it                                                              68      55.3%  87    70.7%
reports the sentence/segment SEGMENTATION step                                                                      24      19.5%  56    45.5%
says which policy applies (app vs developer vs platform vs layered)  <-- not a superset: different question         27      22.0%  13    10.6%
deduplicates identical / templated policies                                                                         34      27.6%  75    61.0%
 
--- The two probes the page leans on hardest, with the matching papers named
 
follows the policy link and reports FAILURES (404 / dead / unreachable)  —  7 papers
  2018 WWW      An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Priva
  2019 NDSS     we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy
  2019 USENIX   PolicyLint: Investigating Internal Privacy Policy Contradictions on Google Play
  2022 WWW      Measuring Alexa Skill Privacy Practices across Three Years.
  2024 PETS     Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps' Privac
  2024 PETS     Privacy Policies on the Fediverse: A Case Study of Mastodon Instances
  2026 PETS     Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Langua
 
analyses policies in more than one language  —  30 papers
  2019 NDSS     we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy
  2019 PETS     MAPS: Scaling Privacy Compliance Analysis to a Million Apps
  2020 PETS     The Privacy Policy Landscape After the GDPR
  2021 USENIX   Understanding Malicious Cross-library Data Harvesting on Android
  2021 CCS      Consistency Analysis of Data-Usage Purposes in Mobile Apps.
  2021 PETS     Defining Privacy: How Users Interpret Technical Terms in Privacy Policies
  2022 CCS      Do Opt-Outs Really Opt Me Out?
  2022 PETS     Checking Websites’ GDPR Consent Compliance for Marketing Emails
  2022 PETS     “We may share the number of diaper changes”: A Privacy and Security Analysis of Mobile Child
  2023 USENIX   Are You Spying on Me? Large-Scale Analysis on IoT Data Exposure through Companion Apps
  2023 USENIX   POLICYCOMP: Counterpart Comparison of Privacy Policies Uncovers Overbroad Personal Data Coll
  2023 PETS     Researchers’ Experiences in Analyzing Privacy Policies: Challenges and Opportunities
  2024 NDSS     Towards Automated Regulation Analysis for Effective Privacy Compliance
  2024 PETS     On the Quality of Privacy Policy Documents of Virtual Personal Assistant Applications
  2024 PETS     A Bilingual Longitudinal Analysis of Privacy Policies Measuring the Impacts of the GDPR and 
  2024 USENIX   Is It a Trap? A Large-scale Empirical Study And Comprehensive Assessment of Online Automated
  2024 WWW      Understanding GDPR Non-Compliance in Privacy Policies of Alexa Skills in European Marketplac
  2024 USENIX   Unpacking Privacy Labels: A Measurement and Developer Perspective on Google's Data Safety Se
  2025 PETS     Understanding Privacy Norms through Web Forms
  2025 USENIX   Evaluating Privacy Policies under Modern Privacy Laws At Scale: An LLM-Based Automated Appro
  2025 PETS     Surveillance Disguised as Protection: A Comparative Analysis of Sideloaded and In-Store Pare
  2025 CCS      The Odyssey of robots.txt Governance: Measuring Convention Implications of Web Bots in Large
  2025 WWW      Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and
  2025 IEEE-SP  Let's Get Visual - Testing Visual Analogies and Metaphors for Conveying Privacy Policies and
  2026 PETS     AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents
  2026 PETS     Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law’s Impact
  2026 PETS     Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Langua
  2026 PETS     Designing Reflective Thinking-Based Contextual Privacy Policy for Mobile Applications
  2026 PETS     Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale
  2026 PETS     From Lines of Code to Lines of Policy? Exploring Software Developers’ Perceptions of Their P
 
 
--- LINEAGE ARTEFACTS: extraction fold vs FULL-TEXT mention, whole corpus
Full-text denominator: 5869 papers with a readable paper.cols.txt.
"extraction" is the column report_policies.mjs section 4 prints.
artefact                                full-text pattern   papers whose FULL TEXT names it
--------------------------------------  ------------------  -------------------------------
Privee (USENIX 2014)                    /\bPrivee\b/        25
OPP-115 corpus (ACL 2016)               /OPP-?115/i         30
Polisis / PriBot (USENIX 2018)          /polisis|pribot/i   68
PolicyLint (USENIX 2019)                /policylint/i       63
MAPS (PETS 2019)                        /MAPS:\s*Scaling/i  50
APP-350 corpus (2019)                   /APP-?350/i         9
PoliCheck (USENIX 2020)                 /policheck/i        63
PurPliance (2021)                       /purpliance/i       14
PrivBERT (2021)                         /privbert/i         3
Calpric (USENIX 2023)                   /calpric/i          3
PoliGraph / PoliGraph-er (USENIX 2023)  /poligraph/i        14
PolicyChecker (CCS 2023)                /policychecker/i    9
Lalaine (USENIX 2023)                   /lalaine/i          28
PolicyComp (USENIX 2023)                /policycomp/i       10
PrivaSeer                               /privaseer/i        12
 
Full-text patterns that differ from the extraction pattern, and why:
  MAPS (PETS 2019): /MAPS:\s*Scaling/i — /\bMAPS\b/ also matches "Google MAPS" and the Play category "MAPS & NAVIGATION"; the paper is always cited by its title
 
The two artefacts where the gap changes what may be said:
 
PrivaSeer — 12 papers:
  2021 PETS     automated-extraction-and-presentation-of-data-practices-in-privacy-policies
  2021 WWW      privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset
  2022 PETS     setting-the-bar-low-are-websites-complying-with-the-minimum-requirements-of-the
  2023 PETS     researchers-experiences-in-analyzing-privacy-policies-challenges-and-opportuniti
  2024 CCS      vpvet-vetting-privacy-policies-of-virtual-reality-apps
  2024 PETS     connecting-the-dots-tracing-data-endpoints-in-iot-devices
  2024 PETS     honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a
  2024 PETS     on-the-quality-of-privacy-policy-documents-of-virtual-personal-assistant-applica
  2025 CCS      layered-overlapping-and-inconsistent-a-large-scale-analysis-of-the-multiple-priv
  2025 CCS      the-odyssey-of-robots-txt-governance-measuring-convention-implications-of-web-bo
  2025 USENIX   evaluating-privacy-policies-under-modern-privacy-laws-at-scale-an-llm-based-auto
  2026 PETS     word-level-annotation-of-gdpr-transparency-compliance-in-privacy-policies-using
 
Calpric (USENIX 2023) — 3 papers:
  2023 USENIX   calpric-inclusive-and-fine-grain-labeling-of-privacy-policies-with-crowdsourcing
  2025 PETS     understanding-privacy-norms-through-web-forms
  2025 USENIX   evaluating-privacy-policies-under-modern-privacy-laws-at-scale-an-llm-based-auto

Output of ''policies_significance.py''

policies_significance-output.txt
parsed 10 SIG rows from scripts/report_policies-output.txt
 
comparison                                     pop          subgroup   rest of corpus  p (Fisher, 2-sided)      naive base
--------------------------------------------------------------------------------------------------------------------------
temporal.mode == web-archive                   POLICY 11/101 (10.9%) 57/5241 ( 1.1%)      3.99e-08  significant   [68/5342 =  1.3%]
temporal.mode == web-archive                   UNION  12/118 (10.2%) 56/5224 ( 1.1%)      1.97e-08  significant   [68/5342 =  1.3%]
humanAnnotation states an agreement metric     POLICY 38/101 (37.6%) 474/3217 (14.7%)      2.65e-08  significant   [512/3318 = 15.4%]
humanAnnotation states an agreement metric     UNION  45/120 (37.5%) 467/3198 (14.6%)      1.44e-09  significant   [512/3318 = 15.4%]
artifacts.availability == public               POLICY 63/101 (62.4%) 2782/5438 (51.2%)         0.027  significant   [2845/5539 = 51.4%]
artifacts.availability == public               UNION  77/122 (63.1%) 2768/5417 (51.1%)        0.0101  significant   [2845/5539 = 51.4%]
assesses a law (legal[] non-empty)             POLICY 64/102 (62.7%) 338/5757 ( 5.9%)      3.62e-50  significant   [402/5859 =  6.9%]
assesses a law (legal[] non-empty)             UNION  70/123 (56.9%) 332/5736 ( 5.8%)      9.63e-51  significant   [402/5859 =  6.9%]
classification.validation == none-reported     POLICY 40/102 (39.2%) 1823/4337 (42.0%)         0.613  NOT SIGNIFICANT — do not publish as a movement   [1881/4439 = 42.4%]
classification.validation == none-reported     UNION  40/102 (39.2%) 1818/4321 (42.1%)         0.612  NOT SIGNIFICANT — do not publish as a movement   [1881/4439 = 42.4%]
 
"rest of corpus" removes the subgroup's own papers from the base; the
bracketed column is the naive base rate that leaves them in.

Output of ''policies_table_check.mjs''

policies_table_check-output.txt
OK        method-by-era table: 11 rows
OK        per-year table: 12 rows
OK        per-venue table: 7 rows
OK        significance table: 5 rows
 
All four tables match the report cell by cell.

Output of ''policies_quotecheck.mjs''

policies_quotecheck-output.txt
OK  cols                     NDSS 2019  web policy availability 84.5% / 79.6%
OK  cols                     NDSS 2019  Degeling sample 6,579 sites, 500 per member state
OK  cols                     IMC 2019  adult sites: only 16% have an accessible policy, of 6,843
OK  cols                     PETS 2025  web forms 94.2% policy link
OK  cols                     PETS 2019  Play policy links 50.5%
OK  cols                     USENIX 2024  APPG: 37.5% of links unavailable
OK  cols                     USENIX 2024  APPG: 20.5% non-English
OK  cols                     USENIX 2024  APPG: 22.3% low quality under 2KB/200 words
OK  cols                     USENIX 2022  smart-home vendors 48.99% (292/596)
OK  cols                     NDSS 2021  Alexa skills 24.2% policy link
OK  cols                     CCS 2024  VR apps 29.5% have a policy
OK  cols                     IMC 2022  chatbots 95.67% lack a policy
OK  cols                     IMC 2019  adult sample 6,843 sites
OK  cols                     PETS 2025  web-forms denominator 10,143
OK  cols                     PETS 2019  MAPS denominator 1,049,790 retrieved / 1,035,853 analysed
OK  cols                     USENIX 2024  APPG link denominator 37,150/99,194
OK  cols                     USENIX 2022  smart home 10.57% no policy at all
OK  cols                     NDSS 2021  Alexa denominator 150,708 skills, 36,475 with a link
OK  cols                     PETS 2024  Alexa 21,063 of 65,195
OK  cols                     CCS 2024  VR denominator 11,923 apps
OK  cols                     IMC 2022  Discord chatbots 676 (4.35%) have a policy
OK  cols                     WWW 2018  only 14.80% of transmissions disclosed
OK  cols                     USENIX 2019  PolicyLint 14.2% (1,618/11,430) contradictions
OK  cols                     USENIX 2020  PoliCheck 42.4% of apps
OK  cols                     USENIX 2020  PoliCheck only 0.5% of flows clearly disclosed
OK  cols                     USENIX 2022  OVRseen 68% inconsistent
OK  cols                     USENIX 2023  PoliGraph 70.6% recall / 96.9% precision
OK  cols                     USENIX 2023  PoliGraph 25.5% of policies redefine a term
OK  cols                     CCS 2023  PolicyChecker 99.3% incomplete
OK  cols                     CCS 2023  PolicyChecker 163,068 analysable of 205,973
OK  cols                     USENIX 2023  Lalaine 3,423 of 5,102 apps
OK  cols                     PETS 2023  CCPA VCR 80% undisclosed identifier
OK  cols                     PETS 2024  Apple labels: 97% of "Data Not Collected" contradicted
OK  cols                     PETS 2025  phi < 0.20 policy-vs-form association
OK  cols                     WWW 2021  Princeton corpus 1,071,488 policies / 130,000 sites
OK  cols                     WWW 2021  median length 876 -> 1,522 words
OK  cols                     WWW 2021  FKGL 11.9 -> 13.2
OK  cols                     WWW 2021  beacons: 25.8% of policies vs 94.6% of top-10K sites
OK  cols                     WWW 2018  84.7 minutes to read applicable policies
OK  cols                     PETS 2020  EU policies gained 35% words / 33% sentences, Global 25% / 22%
OK  cols                     CCS 2024  VR policy reuse 54.5% (1,919/3,521)
OK  cols                     PETS 2024  template reuse 65%
OK  cols                     USENIX 2018  Polisis trained on 65 of the OPP-115 policies, 50 held out
OK  cols                     USENIX 2018  Polisis average F1 0.84
OK  cols                     USENIX 2023  Calpric 16,856 labelled segments
OK  cols                     PETS 2023  no best practices have emerged (26 interviews)
OK  cols                     PETS 2023  26 researchers interviewed
OK  cols                     PETS 2023  User Choice/Control semantic change 26%
OK  cols                     IMC 2024  GPT-4 chatbot annotation of corporate policies
OK  cols                     IMC 2025  LLM disclosure classifier 87.44% accuracy
OK  cols                     IMC 2025  only 5.8% of Actions clearly disclose
OK  cols                     USENIX 2024  APPG low-quality count 10,375/46,472
OK  cols                     USENIX 2024  APPG non-English count 9,523/46,472
OK  cols                     USENIX 2024  APPG 15.7% of apps provide no policy
OK  cols                     USENIX 2020  PoliCheck entity-insensitive false-consistency 37.1%
OK  cols                     USENIX 2020  PoliCheck 31.1% (14,409/45,603) omitted flows
OK  cols                     USENIX 2020  PoliCheck 14,409 omitted disclosures (the 31.1% numerator)
OK  cols                     PETS 2024  Apple labels: 228,539 apps policy-but-not-label
OK  cols                     PETS 2024  template count n=306,404 behind the 65%
OK  cols                     CCS 2023  PolicyChecker 98.1% mandatory-requirement violation
OK  cols                     NDSS 2017  Zimmeck mean 1.83 inconsistencies per app
OK  cols                     WWW 2021  Princeton-Leuven corpus reaches back to 1997
OK  pypdf (NOT in .cols)     WWW 2021  the length/readability series is 2009-2019
OK  cols                     NDSS 2019  Degeling: 6357 is the total of the availability table
OK  cols                     NDSS 2019  Degeling: the same paper also says 6,759 domains
OK  cols                     PETS 2019  MAPS 50.5% is over the analysed set, not the retrieved set
OK  cols                     PETS 2019  MAPS analysed 1,035,853 of 1,049,790 retrieved
OK  cols                     PETS 2026  cory2026 is word-level GDPR transparency annotation by LLM
OK  cols                     IMC 2022  Discord: the paper states 15,525 unique active chatbots
OK  cols                     IMC 2022  Discord: and separately 14,852 (95.67%) without a policy
 
SPECIFICITY of the 70 needles
  needles with no letters (pure number/punctuation): 30
  of those, also present in another check paper     : 4
  needles under 12 characters                       : 22
 
  WEAK — numeric-only and not unique to the paper they are attributed to:
     1 other check papers also contain "1,071,488"  (Princeton corpus 1,071,488 policies / 130,000 sites)
     1 other check papers also contain "84.7"  (84.7 minutes to read applicable policies)
     1 other check papers also contain "54.5%"  (VR policy reuse 54.5% (1,919/3,521))
     1 other check papers also contain "1,035,853"  (MAPS analysed 1,035,853 of 1,049,790 retrieved)
 
  These are not wrong — each was read in context — but they are the needles
  a future edit could break without this check noticing.
 
70 needles, 70 located, 0 MISSING.
69 located in paper.cols.txt; 1 located only outside paper.cols.txt.

Output of ''policies_external_checks.sh''

policies_external_checks-output.txt
=== GitHub repositories named on the page (commits API on the default branch) ===
citp/privacy-policy-historical                 default|license|archived|pushed_at: master | <no license file> | archived=False | pushed_at=2023-10-12T15:59:19Z
                                               last commit on default branch: 2023-10-12T15:59:18Z
citp/PrivacyPoliciesOverTime                   default|license|archived|pushed_at: master | <no license file> | archived=False | pushed_at=2022-06-06T16:25:20Z
                                               last commit on default branch: 2022-06-06T16:25:20Z
benandow/PrivacyPolicyAnalysis                 default|license|archived|pushed_at: master | NOASSERTION | archived=False | pushed_at=2022-10-05T15:45:48Z
                                               last commit on default branch: 2022-10-05T15:45:48Z
UCI-Networking-Group/PoliGraph                 default|license|archived|pushed_at: master | MIT | archived=False | pushed_at=2023-06-21T23:54:12Z
                                               last commit on default branch: 2023-06-21T23:53:52Z
AndyXiang945/PolicyChecker                     default|license|archived|pushed_at: main | <no license file> | archived=False | pushed_at=2023-11-20T22:17:26Z
                                               last commit on default branch: 2023-11-20T22:17:24Z
xiaoyue10131748/Lalaine                        default|license|archived|pushed_at: main | MIT | archived=False | pushed_at=2023-09-25T20:15:37Z
                                               last commit on default branch: 2023-09-25T20:15:37Z
dlgroupuoft/Calpric                            default|license|archived|pushed_at: main | <no license file> | archived=False | pushed_at=2023-06-21T10:37:11Z
                                               last commit on default branch: 2023-06-21T10:37:11Z
ducalpha/PurPlianceOpenSource                  default|license|archived|pushed_at: main | NOASSERTION | archived=False | pushed_at=2024-03-11T00:39:10Z
                                               last commit on default branch: 2024-03-11T00:39:10Z
SmartDataAnalytics/Polisis_Benchmark           default|license|archived|pushed_at: master | <no license file> | archived=False | pushed_at=2023-02-02T05:13:35Z
                                               last commit on default branch: 2020-02-13T10:56:26Z
quanmou/polisis                                default|license|archived|pushed_at: master | <no license file> | archived=False | pushed_at=2020-07-27T18:58:49Z
                                               last commit on default branch: 2020-07-27T18:46:21Z
 
=== pushed_at is not the default branch: per-branch last commit where they differ ===
  Polisis_Benchmark dependabot/pip/bleach-3.3.0              2021-02-02T22:28:16Z
  Polisis_Benchmark dependabot/pip/ipython-7.16.3            2022-01-21T19:56:39Z
  Polisis_Benchmark dependabot/pip/jinja2-2.11.3             2021-03-20T02:55:27Z
  Polisis_Benchmark dependabot/pip/mistune-2.0.3             2022-07-29T23:02:03Z
  Polisis_Benchmark dependabot/pip/nbconvert-6.5.1           2022-08-23T18:02:27Z
  Polisis_Benchmark dependabot/pip/nltk-3.4.5                2020-02-13T10:44:32Z
  Polisis_Benchmark dependabot/pip/notebook-6.4.12           2022-06-16T23:39:21Z
  Polisis_Benchmark dependabot/pip/numpy-1.22.0              2022-06-22T01:09:38Z
  Polisis_Benchmark dependabot/pip/pillow-9.3.0              2022-11-22T03:17:29Z
  Polisis_Benchmark dependabot/pip/protobuf-3.18.3           2022-09-23T22:36:10Z
  Polisis_Benchmark dependabot/pip/pygments-2.7.4            2021-03-29T21:56:42Z
  Polisis_Benchmark dependabot/pip/tensorflow-2.9.3          2022-11-21T21:21:28Z
  Polisis_Benchmark dependabot/pip/werkzeug-0.15.5           2023-02-02T05:13:29Z
  Polisis_Benchmark master                                   2020-02-13T10:56:26Z
 
=== LICENSE.txt of PurPliance (GitHub reports NOASSERTION) ===
  Copyright (c) 2022, the University of Michigan
  All rights reserved.
 
=== PurPliance: last three commits on the default branch ===
  2024-03-11T00:39:10Z  Update README.rst
  2022-12-25T07:13:07Z  release privacy-statement extractor
  2021-09-13T05:17:05Z  Update README.md
 
=== Is each lineage repo findable by GitHub search on the tool name? ===
PoliGraph        total=67  top=UCI-Networking-Group/PoliGraph
PolicyChecker    total=12  top=AndyXiang945/PolicyChecker
Lalaine          total=77  top=xiaoyue10131748/Lalaine
Calpric          total=9  top=dlgroupuoft/Calpric
Polisis          total=15  top=quanmou/polisis
PolicyLint       total=2  top=ShadowGuardAI/spea-policylint
 
=== LICENSE files GitHub cannot auto-detect (first line of each) ===
benandow/PrivacyPolicyAnalysis           LICENSE -> HTTP 404
benandow/PrivacyPolicyAnalysis           LICENSE.txt -> HTTP 200, first 2 non-empty lines:
    Copyright (c) 2019, North Carolina State University
    All rights reserved.
benandow/PrivacyPolicyAnalysis           LICENSE.md -> HTTP 404
 
=== Last-Modified on the W3C P3P pages (the page claims a date here) ===
https://www.w3.org/P3P/            <no last-modified header>
https://www.w3.org/TR/P3P11/       <no last-modified header>
 
=== PyPI: is there an installable package for the lineage? ===
poligraph                pypi.org/pypi/poligraph/json -> HTTP 404
poligraph-er             pypi.org/pypi/poligraph-er/json -> HTTP 404
policylint               pypi.org/pypi/policylint/json -> HTTP 404
policheck                pypi.org/pypi/policheck/json -> HTTP 404
privbert                 pypi.org/pypi/privbert/json -> HTTP 404
polisis                  pypi.org/pypi/polisis/json -> HTTP 404
 
=== Top GitHub name-search hits — is the FIRST hit the real artefact? ===
=== Polisis   total=15
    quanmou/polisis                                stars=  16 pushed=2020-07-27T18:58:49Z  Automated Analysis of Privacy Policies
    SmartDataAnalytics/Polisis_Benchmark           stars=  24 pushed=2023-02-02T05:13:35Z  Reproducing state-of-the-art results
    Maxikilliane/polisis-classifiers               stars=   7 pushed=2020-10-22T13:46:46Z  
=== PolicyLint   total=2
    ShadowGuardAI/spea-policylint                  stars=   0 pushed=2025-04-19T16:27:04Z  A command-line tool to lint security policies written in
    fatihkaplanfk/policyLint-dlp                   stars=   0 pushed=2026-08-12T13:05:48Z  
=== PoliCheck   total=13
    UvinduBro/PoliCheck                            stars=   0 pushed=2026-09-02T13:13:40Z  
    fuyukihatune-rgb/policheck                     stars=   0 pushed=2026-06-19T15:08:15Z  
    surelywang/PoliCheck                           stars=   0 pushed=2018-11-03T17:56:10Z  Political Bias Detection Tool
=== PoliGraph   total=67
    UCI-Networking-Group/PoliGraph                 stars=  34 pushed=2023-06-21T23:54:12Z  PoliGraph: Automated Privacy Policy Analysis using Knowl
    ironlam/poligraph                              stars=  39 pushed=2026-09-10T15:43:22Z  Observatoire citoyen de la transparence politique frança
    ben-cunningham/poligraph                       stars=   3 pushed=2017-08-03T05:51:04Z  Source code for http://poligraph.io
=== PolicyChecker   total=12
    AndyXiang945/PolicyChecker                     stars=   5 pushed=2023-11-20T22:17:26Z  
    ryanwakely/policychecker                       stars=   0 pushed=2014-07-15T05:57:42Z  
    KamilW-git/PolicyChecker                       stars=   0 pushed=2026-06-15T21:47:15Z  Aplikacja webowa do sprawdzania, czy wnioski zakupowe i 
=== Lalaine   total=77
    xiaoyue10131748/Lalaine                        stars=   9 pushed=2023-09-25T20:15:37Z  
    dauntl1ss/Lalaine                              stars=   0 pushed=2022-02-08T16:45:59Z  
    yujishen6-design/lalaine                       stars=   0 pushed=2025-11-21T06:55:43Z  
=== Calpric   total=9
    dlgroupuoft/Calpric                            stars=   2 pushed=2023-06-21T10:37:11Z  
    hernandess31/calpriceBETA                      stars=   0 pushed=2025-08-13T14:40:35Z  Calprice é um aplicativo em Python criado para auxiliar 
    sysjoma/calprice                               stars=   0 pushed=2020-08-05T04:15:34Z  Recalcular precios a la tasa del dólar
=== PurPliance   total=1
    ducalpha/PurPlianceOpenSource                  stars=  17 pushed=2024-03-11T00:39:10Z  Source code of PurPliance analysis tool.
=== PolicyComp   total=44
    policycompass/policycompass-frontend           stars=   3 pushed=2016-11-11T03:24:25Z  Angular frontend for the policycompass web application
    policycompass/policycompass-services           stars=   2 pushed=2016-11-09T13:09:05Z  web services for the policy compass frontend
    policycompass/policycompass                    stars=   5 pushed=2016-10-07T13:20:55Z  Policy Compass central repo
 
=== Polisis: the paper has no code repository; the demo site is claimed alive ===
pribot.org/polisis (the Polisis demo)          HTTP 200  redirects=1  bytes=8471
  final: http://pribot.org/polisis/
  <title>: <title>Polisis</title>
 
=== Dataset and model hosts ===
OPP-115 / APP-350 (usableprivacy.org/data)     HTTP 200  redirects=0  bytes=26541
  final: https://usableprivacy.org/data
PrivaSeer search engine                        HTTP 200  redirects=0  bytes=11585
  final: https://privaseer.ist.psu.edu/
PrivaSeer data + licence page                  HTTP 200  redirects=0  bytes=11838
  final: https://privaseer.ist.psu.edu/data
  The PrivaSeer corpus is a collection of 3,967,487 privacy policies
  the corpus is available under a CC BY-NC-SA license
PrivBERT model card (Hugging Face)             HTTP 200  redirects=0  bytes=147253
  final: https://huggingface.co/mukund/privbert
 
=== Standards and platform policy sources ===
W3C P3P home                                   HTTP 403  redirects=0  bytes=5583
  final: https://www.w3.org/P3P/
W3C P3P 1.1 (obsoleted note)                   HTTP 403  redirects=0  bytes=5598
  final: https://www.w3.org/TR/P3P11/
Apple: third-party SDK requirements            HTTP 200  redirects=0  bytes=113972
  final: https://developer.apple.com/news/?id=pvszzano
Google Play User Data policy                   HTTP 200  redirects=0  bytes=1646973
  final: https://support.google.com/googleplay/android-developer/answer/10144311
USENIX Sec 2023 artifact index                 HTTP 200  redirects=0  bytes=324140
  final: https://secartifacts.github.io/usenixsec2023/results
 
All repository lookups returned a date, not an API error.

Output of ''policies_w3c_p3p_check.mjs''

policies_w3c_p3p_check-output.txt
=== https://www.w3.org/P3P/
  final URL      : https://www.w3.org/P3P/
  HTTP status    : 200
  last-modified  : Fri, 02 Feb 2018 14:13:43 GMT
  body chars     : 8675
  contains "Retired 30 August 2018": false
  contains "should not be referenced in this form or implemented as-is": false
  years appearing in the visible text: 2002 2007 2018
 
=== https://www.w3.org/TR/P3P11/
  final URL      : https://www.w3.org/TR/P3P11/
  HTTP status    : 200
  last-modified  : Tue, 09 Oct 2018 13:16:04 GMT
  body chars     : 369540
  contains "Retired 30 August 2018": true
  contains "should not be referenced in this form or implemented as-is": true
  years appearing in the visible text: 1981 1995 1996 1997 1998 1999 2000 2001 2002 2003 2005 2006 2018
  copyright line : Copyright © 2006 W3C® (MIT, ERCIM, Keio), All Rights Reserved. W3C liability, trademark and document
 
=== w3.org/P3P/ — every sentence containing a year
  Platform for Privacy Preferences (P3P) Project Enabling smarter Privacy Tools for the Web PLING - W3C Policy Languages Interest Group 3 October 2007: The Policy Languages Interest Group (PLING) was created.
  Background P3P 1.1 is a direct consequence of the first Privacy Workshop that took place 2002 in Dulles/Virginia and targets short term improvements like the User Agent Guidelines.
  P3PToolbox.org, with lots of complementary information P3P Validator to test the results The www-p3p-policy mailing-list to discuss issues P3P Software and Tools that may help Other P3P Documents and Notes Working Draft:A P3P Preference Exchange Language 1.0 (APPEL1.0) A P3P Assurance Signature Profile An RDF Schema for P3P 1.0 Mailing lists www-p3p-dev is a mailing list for P3P software developers www-p3p-policy is a mailing list for people who are responsible for creating P3P policies for web sites Background Resources for Developers Feedback and Discussions Papers & Presentations about P3P Critiques of P3P Selected P3P Media Coverage Historical documents and things Working Group Pages P3P Group page[Member] P3P Specification WG Homepage Charter Contact: Lorrie Cranor (Chair) & Rigo Wenning (W3C) Last updated $Date: 2018/02/02 14:13:43 $ by $Author: rigo $

References

[1]
Ali, Mir Masood; Balash, David G.; Kodwani, Monica; Kanich, Chris; Aviv, Adam J. (2024): "Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps' Privacy Policies", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[2]
Amos, Ryan; Acar, Gunes; Lucherini, Eli; Kshirsagar, Mihir; Narayanan, Arvind; Mayer, Jonathan (2021): "Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset", in: Proceedings of the Web Conference 2021, pp. 2165–2176. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
[3]
Adhikari, Andrick; Das, Sanchari; Dewri, Rinku (2023): "Evolution of Composition, Readability, and Structure of Privacy Policies over Two Decades", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[4]
Cui, Hao; Trimananda, Rahmadi; Markopoulou, Athina (2025): "Understanding Privacy Norms through Web Forms", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[5]
Pan, Shidong; Zhang, Dawen; Staples, Mark; Xing, Zhenchang; Chen, Jieshan; Xu, Xiwei; Hoang, Thong (2024): "Is It a Trap? A Large-scale Empirical Study And Comprehensive Assessment of Online Automated Privacy Policy Generators for Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)
[6]
Andow, Benjamin; Mahmud, Samin Yaseer; Whitaker, Justin; Enck, William; Reaves, Bradley; Singh, Kapil; Egelman, Serge (2020): "Actions Speak Louder than Words: Entity-Sensitive Privacy Policy and Data Flow Analysis with PoliCheck", in: Proceedings of the USENIX Security Symposium. (Link)
[7]
Xiang, Anhao; Pei, Weiping; Yue, Chuan (2023): "PolicyChecker: Analyzing the GDPR Completeness of Mobile Apps' Privacy Policies", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[8]
Zimmeck, Sebastian; Wang, Ziqi; Zou, Lieyong; Iyengar, Roger; Liu, Bin; Schaub, Florian; Wilson, Shomir; Sadeh, Norman; Bellovin, Steven M.; Reidenberg, Joel (2017): "Automated Analysis of Privacy Requirements for Mobile Apps", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[9]
Khatun, Mst Eshita; Noureddine, Lamine; Bello, Sideeq; Ali-Gombe, Aisha (2026): "Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale", Proceedings on Privacy Enhancing Technologies 2026(4):213-231. (DOI)
[10]
Nenadić, Luka; Rodriguez, David; Calandrino, Joseph A. (2026): "Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact", Proceedings on Privacy Enhancing Technologies 2026(4):703-723. (DOI)
[11]
Cory, Thomas; Rieder, Wolf; Krämer, Julia; Raschke, Philip; Herbke, Patrick; Küpper, Axel (2026): "Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Language Models", Proceedings on Privacy Enhancing Technologies 2026(1):509-528. (DOI)
[12]
Chanenson, Jake; Pickering, Madison; Apthorpe, Noah (2025): "Automating Governing Knowledge Commons and Contextual Integrity (GKC-CI) Privacy Policy Annotations with Large Language Models", in: Proceedings on Privacy Enhancing Technologies. (DOI)
provenance/privacy/policies.txt · Last modified: by karel.kubicek.claude