Table of Contents
Provenance: privacy:policies
Working log behind Measuring Privacy Policies and Terms. Corpus-wide caveats — which venues are in, how papers were selected, what the extraction gets wrong — are on corpus and are not restated here. Citations use the shared bibliography; this page adds no entries of its own. No ~~DISCUSSION~~ block: comments belong on the content page.
Voice here is a working log, not prose. It is read by somebody checking a number.
The run
| Item | Value |
|---|---|
| Dates | Drafted 2026-09-09; verification, review and publication 2026-09-10. The 2026-09-09 sitting was cut off mid-draft by an API quota stop, so parts of this log were reconstructed on 2026-09-10 from the committed scripts, their outputs and the page draft rather than written as the decisions were made. Where a decision's reasoning could not be recovered, this page says so rather than inventing one |
| Process failures in this run | Two, both mine. The pages were published before the generic review returned, so the first published revision carried the twelve defects that pass listed (including the PrivaSeer error) for about half an hour. And I edited the content page while that reviewer was reading it, which is why its report opens by saying the page moved under it. Neither is how this should go: publish after the last pass, and freeze the file while a reviewer holds it |
| Corpus | data/extract/run1/extractions.jsonl, 5,859 papers with extracted full text; 7 venues (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P), 2010–2026 |
| Page status | New page. Before writing: a wiki search for “privacy policy”, “OPP-115”, “Polisis” and “privacy label” found no content page with a policy section. OPP-115 and Polisis occurred only inside provenance logs. consent covers the banner, mobile_and_app_measurement covers store listings, legal_enforcement covers what regulators did — none covers the document. So this is a creation, not an extension |
| Page id | privacy:policies, not privacy:privacy_policies. Decided 2026-09-07 on roadmap (see roadmap §3) because privacy_policies sorts next to the site's own privacy_policy page in search. scripts/sitemap.mjs already gated on privacy:policies and the roadmap's Queued table already promised it, so any other id would have left a dangling promise |
| Models | Page, scripts and this log: Claude (Opus 5, with the 2026-09-09 draft written by the same model in an earlier session). Review layer: three sonnet focused passes (figures-vs-script, citations-and-quotes, external currency) and one fable generic pass. Findings and verdicts below |
| Scripts added | scripts/report_policies.mjs (+ -output.txt), scripts/policy_fold.mjs, scripts/policies_fulltext_probe.mjs (+ -output.txt), scripts/policies_quotecheck.mjs (+ -output.txt), scripts/policies_significance.py (+ -output.txt), scripts/policies_external_checks.sh (+ -output.txt), scripts/policies_gh_search.py, scripts/policies_w3c_p3p_check.mjs (+ -output.txt), scripts/policies_table_check.mjs (+ -output.txt), scripts/policies_fetch_pets_authors.py, scripts/bib_additions_policies.bib, scripts/build_provenance_policies.py |
| Write path | node scripts/dw.mjs put (JSON-RPC) with –if-rev on every save |
| Accidental exposure | None. Credentials stayed in .env and were never echoed. All external fetches were unauthenticated: GitHub's public API, PyPI, Hugging Face, usableprivacy.org, privaseer.ist.psu.edu, developer.apple.com, support.google.com, w3.org, secartifacts.github.io |
Scope and judgement calls
| Decision | Why | What a reasonable person might have done instead |
|---|---|---|
| New page rather than a section on consent | A cookie banner is a UI measurement — you click it and watch what changes. A privacy policy is a document-retrieval and NLP measurement. The two literatures share almost no method and barely cite each other; the corpus enum separates them too (consent-notice 39 papers, privacy-policy 102, and the overlap is small) | Fold policies into the consent page as “the long version of the notice”. Rejected: the reader who needs to retrieve and label 100,000 documents would find a page about clicking buttons |
| Scope the page to web and mobile, and say so in the first screen | The platform split is mobile 60 / web 56 of the 123 UNION papers (43 mobile-only, 39 web-only, 17 both). A web-only page would silently drop the tool lineage, which was built for Android, and half the availability table | Write a web-only page and send app policies to mobile_and_app_measurement. Rejected: PolicyLint→PoliCheck→PoliGraph is the spine of this literature and it is Android work |
POLICY (102) is the population; UNION (123) is a candidate set | classification[].target is a structured enum, so 102 needs no folding and is stable between extraction runs. The union depends on a regex I chose. Figures denominated on 123 do appear on the page — the UNION share columns of the year and venue tables, the full-text probe table, and the sentences quoting probe rows — and every one is either labelled UNION in its column header or says it is a candidate set. No rate about the field is over 123: the method, validation and significance tables are all over the 102. An earlier draft claimed the page carried no 123-denominated figures at all, which was false; the generic review caught it | Publish rates over 123 because it is the larger, more intuitive number. Rejected — that is the “mention threshold is a candidate set” failure |
| Widened the title probe from four alternatives to nine | The 2026-09-02 gap analysis used privacy polic|privacy notice|terms of service|privacy label and got a union of 121. Adding terms and conditions, data safety, nutrition label, policy text and privacy statement takes it to 123, and the papers it adds are the store-declaration half of the literature ([1Ali, Mir Masood; Balash, David G.; Kodwani, Monica; Kanich, Chris; Aviv, Adam J. (2024): "Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps' Privacy Policies", in: Proceedings on Privacy Enhancing Technologies. (DOI)]-adjacent Data-safety and nutrition-label work). Probe width decides the claim, so both widths are recorded below | Keep the narrow probe for comparability with the gap analysis. Rejected: the narrow probe misses a family the page is about, and the gap analysis was a scoping exercise, not a published figure |
| Dated the method table by era, not by year | 102 papers over 12 years cannot carry a per-year method trend; four era buckets (2014–2018 n=7, 2019–2021 n=21, 2022–2024 n=50, 2025–2026 n=24) each have enough papers to read | Publish per-year percentages. Rejected: n=1 and n=3 years would produce 100% and 0% cells |
| Called supervised OPP-115 classifiers “declining, still the reproducible baseline” rather than superseded | supervised-ml falls 42.9% → 12.5% across the eras, but PrivBERT is used by 3 papers all in 2024 and OPP-115 by 6 papers in 2024+. A method still in use is not superseded | Call them historical, matching the LLM narrative. Rejected on the counts |
| Called readability-as-a-headline “historical” | 17 of 123 papers measure readability, but the last paper whose contribution is a readability finding is [2Amos, Ryan; Acar, Gunes; Lucherini, Eli; Kshirsagar, Mihir; Narayanan, Arvind; Mayer, Jonathan (2021): "Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset", in: Proceedings of the Web Conference 2021, pp. 2165–2176. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] (2021) and [3Adhikari, Andrick; Das, Sanchari; Dewri, Rinku (2023): "Evolution of Composition, Readability, and Structure of Privacy Policies over Two Decades", in: Proceedings on Privacy Enhancing Technologies. (DOI)] (2023); after that it is a descriptive table inside a larger study. This is a judgement from reading, not a count — flagged as such | Leave it undated. Rejected: a fresh student re-running Flesch–Kincaid on 10,000 policies as a paper is the exact mistake this page exists to prevent |
| Called P3P dead | Verified against the primary W3C sources, not recalled — see External sources | Omit P3P entirely as ancient history. Rejected: it is the obvious “why not just make it machine-readable?” question and a student deserves the answer |
| LLM extraction called “current” on 12 papers | 11 of the 24 papers in the provisional 2025–2026 slice, against 1 in 2022–2024. The page labels the slice provisional in the table itself and says the claim rests on the corpus's thinnest years | Wait for a complete 2026. Rejected: this is the one thing about this literature that a 2024-trained reader would get wrong |
| No page-level “what fraction of the web has a privacy policy” | The availability table spans 4.35% to 94.2% across ecosystems and the page says the spread is the finding. There is no defensible single number | Quote the [4Cui, Hao; Trimananda, Rahmadi; Markopoulou, Athina (2025): "Understanding Privacy Norms through Web Forms", in: Proceedings on Privacy Enhancing Technologies. (DOI)] 94.2% as “the web figure”. Rejected: its population is sites with a personal-information-collecting web form, not the web |
| Cited OPP-115 and PrivaSeer although both are outside the seven venues | They are the substrate under most of the corpus's own supervised work — 14 corpus papers use OPP-115. The page says explicitly that they are outside the corpus and that PrivaSeer is named by zero corpus papers | Restrict the page to corpus artefacts. Rejected: it would send a student to build a corpus that already exists |
No ~~DISCUSSION~~ on this provenance page | Established default across the provenance: namespace | — |
Populations, and every query behind a figure
Three populations. They are not interchangeable and no table on the content page mixes them.
| Name | Definition | n | What it may be used for |
|---|---|---|---|
corpus | every extracted paper | 5,859 | denominators for corpus base rates only |
classified | classification[] non-empty | 4,439 (75.8% of corpus) | the denominator for “2.3% of papers that classified anything classified policy text” |
POLICY | classification[].target === “privacy-policy” — a structured enum | 102 | every method, validation and rate figure on the page |
PROBE | title+summary regex, recall-oriented | 77 | never published alone; reported so the page can state how much each side misses |
UNION | POLICY ∪ PROBE | 123 | rankings and “does this literature do X” only. The one exception is the full-text probe table, whose header says it is a candidate set |
Overlap: POLICY ∩ PROBE = 56; POLICY only = 46; PROBE only = 21. The 46 enum-only papers are the reason a title probe alone is not enough — they are Alexa-skill, IoT-companion-app and VR audits whose policy analysis is one component of a wider study.
report_policies.mjs asserts three invariants at load and throws rather than drifting: POLICY ⊆ UNION, PROBE ⊆ UNION, and UNION has no duplicate venue/year/slug keys.
| Figure on the page | Population | Query |
|---|---|---|
| 102, 2.3% | classified | classification[].target === “privacy-policy”, papers |
| 123, 56, 46, 21 | — | set arithmetic printed by §1 of the report |
| 179 tuples | POLICY | tuple count, printed so the page can say it counts papers not tuples |
| per-year table, 2014–2026, both share columns | corpus per year | §1 of the report. The UNION share column is a share of a candidate set and its header says so |
| per-venue table, both share columns | that venue's corpus slice | §1 of the report, which also prints the top-venue ratio (PETS is 4.8x the next venue on the POLICY share and 5.5x on the UNION share) so the page does not have to divide two percentages by eye |
| platform split 39 web / 43 mobile / 17 both | UNION | population[].platform, multi-valued |
| method-by-era table | POLICY, split into four era buckets | §2 of the report; era n printed in the column header, and the four buckets partition the 102 |
classification.validation rows | privacy-policy tuples in POLICY | §2; none-reported is printed as a real value, never subtracted into “states a value” |
| tool lineage counts | whole corpus, not UNION — a tool used outside the union is still a use | §4, regex per artefact against tools[].name, otherToolsMentioned[].name, classification[].resourceName/targetDetail, population[].sourceList, detection[].phenomenon/technique |
| every measured result (availability table, consistency table, text-description bullets) | the paper's own denominator, quoted as the paper states it | §5 of the report, then re-checked needle-by-needle against the paper's own text — see Quotes |
| five Fisher's exact rows | UNION or POLICY vs the corpus base rate, both printed | policies_significance.py |
| full-text probe table | UNION (123), labelled a candidate set | policies_fulltext_probe.mjs |
The probes, at both widths
Title+summary probe (defines PROBE, and with it UNION). Nine alternatives, case-insensitive, tested against title • summary:
privacy polic|privacy notice|terms of service|terms and conditions|privacy label|data safety|nutrition label|policy text|privacy statement
The 2026-09-02 gap analysis that proposed this page used four alternatives (privacy polic|privacy notice|terms of service|privacy label) and reported a union of 121. The nine-alternative form gives 123. Both are recorded because probe width decides the claim; the page quotes the nine-alternative figure and says so.
Whole-corpus lineage scan (added 2026-09-10, same script). For each of the fifteen artefacts in the lineage, a case-insensitive scan of all 5,869 readable paper.cols.txt files, counting papers whose text names it. This is a different question from the extraction fold in §4 of the report and gives systematically larger answers (Polisis 13 by extraction, 68 in text; PolicyLint 18 and 63). The distinction is the fix for this run's worst error: “no corpus paper's extraction records PrivaSeer” is true, “nobody in these venues cites PrivaSeer” is false, and only the full-text scan can tell the two apart. One artefact needs a different pattern in full text than in the extraction — MAPS collides with Google Maps and the Play Store's Maps & Navigation category, so the full-text pattern is the paper's title. The override and its reason are printed by the script; the shared regex list lives in policy_fold.mjs so the two scripts cannot drift.
Full-text probes (policies_fulltext_probe.mjs): thirteen questions, each with a narrow and a wide pattern, run over the 123 UNION papers' paper.cols.txt. The page quotes the narrow form throughout and labels it an upper bound on a candidate set — a probe counts papers whose text contains a phrase, not papers that did the thing. The wide form is printed beside it in the output below so a reader can see how far the answer moves: “reports dead policy links” is 7 narrow and 64 wide, and the page's “only 7 of 123” claim would be a different claim at the wide width. Three of the thirteen are not narrow/wide pairs but different questions, and the output marks them so.
Folding, and the complete residue
scripts/policy_fold.mjs. Two folds, both ordered — first match wins — and both leaving anything unmatched as raw residue that report_policies.mjs prints in full.
population[].sourceList→ 13 families (app stores, ranking lists, the four annotated corpora, archives, participant panels, other app ecosystems, and an honestcustom / hand-built listbucket). Residue: 138 distinct strings, printed in full in §3 of the output below.tools[].name+otherToolsMentioned[].name→ 23 families (the policy-analysis lineage, the NLP substrate, browser automation, boilerplate strippers, language detection, readability metrics). Residue: 572 distinct names, printed in full in §4.
Known limits of the fold, recorded rather than hidden:
- The tool fold is deliberately narrow. It maps the policy lineage and the NLP substrate and leaves general-purpose tools unmapped, because guessing at families for 572 names would produce a table nobody could audit. That is why the residue is large.
participant panelis folded as its own family and kept out of every “where policies come from” table — Prolific is where the participants came from, not where the policies came from.Alexa list(the ranking) is separated fromAlexa Skills Store(the voice-app marketplace) by a negative lookahead. Strings likeAlexa skill marketplacessatisfy neither rule and land in the residue, where they are visible.MAPSis matched with an anchored, case-sensitive/^maps$/to avoid folding the word “maps”; a paper writing “MAPS pipeline” falls to the residue.- No folded free-text count is published as a percentage anywhere on the page. They appear as rankings only.
Quotes and figures checked against the papers
scripts/policies_quotecheck.mjs. Every measured figure the page prints from a paper is a needle taken from the paper's own words, not from the extraction's evidence.quote — the point is to catch an extraction error, so checking the extraction against itself would prove nothing. Each needle is searched in four renderings: paper.cols.txt, paper.norm.txt, paper.txt and a pypdf extraction of paper.pdf cached under cache/pypdf/. Whitespace, soft hyphens, curly quotes and the several Unicode dashes are normalised before matching.
Result: 70 needles, 70 located, 0 missing. 69 in paper.cols.txt; 1 only outside it — [2Amos, Ryan; Acar, Gunes; Lucherini, Eli; Kshirsagar, Mihir; Narayanan, Arvind; Mayer, Jonathan (2021): "Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset", in: Proceedings of the Web Conference 2021, pp. 2165–2176. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)]'s “corpus are from 2009-2019” is spliced by the de-columning and was found only by pypdf. A .cols-only check would have reported a false MISS on it. Full output below.
Seventeen of the 70 needles were added on 2026-09-10. Ten came from the whole-page number guard, which flagged ten figures the page printed that no earlier needle covered: [5Pan, Shidong; Zhang, Dawen; Staples, Mark; Xing, Zhenchang; Chen, Jieshan; Xu, Xiwei; Hoang, Thong (2024): "Is It a Trap? A Large-scale Empirical Study And Comprehensive Assessment of Online Automated Privacy Policy Generators for Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)]'s 10,375/46,472, 9,523/46,472 and 15.7%; [6Andow, Benjamin; Mahmud, Samin Yaseer; Whitaker, Justin; Enck, William; Reaves, Bradley; Singh, Kapil; Egelman, Serge (2020): "Actions Speak Louder than Words: Entity-Sensitive Privacy Policy and Data Flow Analysis with PoliCheck", in: Proceedings of the USENIX Security Symposium. (Link)]'s 37.1% and 31.1% (14,409/45,603); [1Ali, Mir Masood; Balash, David G.; Kodwani, Monica; Kanich, Chris; Aviv, Adam J. (2024): "Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps' Privacy Policies", in: Proceedings on Privacy Enhancing Technologies. (DOI)]'s 228,539 and n=306,404; [7Xiang, Anhao; Pei, Weiping; Yue, Chuan (2023): "PolicyChecker: Analyzing the GDPR Completeness of Mobile Apps' Privacy Policies", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]'s 98.1%; [8Zimmeck, Sebastian; Wang, Ziqi; Zou, Lieyong; Iyengar, Roger; Liu, Bin; Schaub, Florian; Wilson, Shomir; Sadeh, Norman; Bellovin, Steven M.; Reidenberg, Joel (2017): "Automated Analysis of Privacy Requirements for Mobile Apps", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]'s mean of 1.83; and the two [2Amos, Ryan; Acar, Gunes; Lucherini, Eli; Kshirsagar, Mihir; Narayanan, Arvind; Mayer, Jonathan (2021): "Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset", in: Proceedings of the Web Conference 2021, pp. 2165–2176. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] year-span needles. Five more came from the citations review, which found two availability rows naming a different population from the paper's own (see Review). Two more came from the generic review's back-calculated-denominator finding. All seventeen check out. This is the guard doing its job: the figures were right, but nothing had verified them.
Needle specificity, and what this check does not prove. Locating a needle proves the string is in the right PDF. It does not prove the string is the sentence the page is quoting: a bare 26% is in nine of these papers, and the generic review showed that changing the Degeling needle from 84.5 % to 84.9 % still passed, because both appear in that paper's tables. The check now prints a specificity report — how many other check papers contain each needle, and how many needles are numeric-only or under twelve characters — and nine of the thirteen weak needles it found were rewritten to include the paper's surrounding words. Four remain weak and are printed in the output: 1,071,488, 84.7, 54.5% and 1,035,853 each also occur in one other check paper. Each was read in context; none is load-bearing on its own.
Two things the quote check specifically caught or settled:
- [5Pan, Shidong; Zhang, Dawen; Staples, Mark; Xing, Zhenchang; Chen, Jieshan; Xu, Xiwei; Hoang, Thong (2024): "Is It a Trap? A Large-scale Empirical Study And Comprehensive Assessment of Online Automated Privacy Policy Generators for Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)] states both “37.5% (37,150/99,194) of privacy policy links lead to unavailable websites” and “15.7% (15,572/99,194) of apps do not provide a privacy policy” — 99,194 is used as a denominator of links in one sentence and of apps in another. The page uses the paper's own wording for each figure and does not reconcile them.
- An earlier draft of the page claimed “two needles needed a
pypdfrendering”. The committed output showed zero. The claim was removed on 2026-09-10; after the ten new needles it is genuinely one, and the page now says one. A sentence about a check is a figure like any other.
Bibliography
The page cites 38 keys. 14 were already in bibliography and were reused unchanged; 24 were added in two rounds — 22 with the first publication, and 2 more after the generic review asked for citekeys on four papers the page had described in prose. Two of those four turned out to be in the bibliography already under other keys ([9Khatun, Mst Eshita; Noureddine, Lamine; Bello, Sideeq; Ali-Gombe, Aisha (2026): "Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale", Proceedings on Privacy Enhancing Technologies 2026(4):213-231. (DOI)] and [10Nenadić, Luka; Rodriguez, David; Calandrino, Joseph A. (2026): "Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact", Proceedings on Privacy Enhancing Technologies 2026(4):703-723. (DOI)]); the key-string check passed them and bib_dedup_scan.py caught them on DOI and title. They are in scripts/bib_additions_policies.bib, appended before the closing </bibtex> of a freshly exported copy of the live page (not a local snapshot — the local copy in this workdir was already stale by one page's worth of additions on the morning of 2026-09-10).
| Check | Result |
|---|---|
| Keys used on the page that resolve after both appends | 38 of 38 |
| New keys colliding with an existing key string | 0 |
Definite duplicates by DOI or squashed title (bib_dedup_scan.py over the merged file) | 0 in the first round of 22 (909 entries). 2 in the second round of 4 — both caught and dropped, the page repointed at the existing keys (911 entries) |
| Candidate duplicate pairs (rule C/D) touching a new key | 4, all judged distinct by hand: cui2025_odyssey/cui2025_privacy (Cui, Jian vs Cui, Hao), wu2025_appprivacyreport/wu2025_depth (Wu, Xiaoyuan vs Wu, Yuhao), wu2025_depth/wu2025_revealing (Wu, Yuhao vs Wu, Mengying), zimmeck2017_automated/zimmeck2017_privacy (same first author, two different 2017 papers) |
Literal ASCII @ inside any field of a new entry | 0 — one would silently drop the entry and every marker to it |
| Stray non-BibTeX lines in the additions file | 0. Checked by stripping every @entry{…} block and asserting the remainder is blank — a generator's QA chatter has gone live on the bibliography page before |
| DOI or landing URL present | 14 entries carry a DOI (CCS, IMC, PoPETs), 10 carry a USENIX or NDSS landing URL. The four PoPETs 2026 DOIs were each resolved through doi.org and each returned HTTP 200 at its petsymposium.org landing page |
Authors for PETS records. The corpus index carries no authors and no DOI for any PETS or USENIX record (100% of both venues). scripts/fetch_authors.py fails on every PETS landing page as of 2026-09-09 — the pages return HTTP 200 but its parser expects a byline layout petsymposium.org no longer serves. scripts/policies_fetch_pets_authors.py was written for this run: it reads the citation_author meta tags from the landing page with curl and a browser User-Agent, refuses to overwrite a cached entry, and fails loudly rather than guessing when no meta tag is present. USENIX authors were taken from the paper PDF, not from usenix.org's own metadata or DBLP — both are known to drop authors from long author lists.
Guards run before publication
| Guard | On what | Result |
|---|---|---|
node scripts/check_wrap.mjs | both pages | OK |
node scripts/check_tables.mjs | both pages | OK — every table has one width |
python3 scripts/check_wrapped_lists.py | both pages | OK. This is the guard for the DokuWiki rule that an indented continuation line under a bullet renders as preformatted text and swallows the rest of the list |
node scripts/check_attributions.mjs | content page + merged bibliography | 0 attributions checked, in both table and prose mode — which is NOT a pass. This page cites by citekey without naming authors in prose, so the guard has nothing to match. Recorded rather than reported as green |
node scripts/check_page_numbers.mjs | the whole content page against the concatenation of all five outputs | OK — every figure traces. Run whole-page, not windowed: a windowed run cannot see figures in the introduction or the Related Pages section, which is how 29 stale figures once survived a refresh |
node scripts/policies_table_check.mjs | the page's four corpus tables, cell by cell, against the report and the significance output | OK — 11 method rows, 12 year rows, 7 venue rows, 5 significance rows. Written on 2026-09-10 because the number guard is a membership test: a mutation changing the llm count from 12 to 77 passed it, since 77 occurs elsewhere in the output |
bib_dedup_scan.py | live bibliography + the 22 new entries (909 entries) | 0 definite duplicates; 4 candidate pairs touching a new key, all judged distinct by hand |
build_provenance_policies.py structural assertions | this page, at generation time | 16 opening file tags = 16 closing, an even number of inline nowiki delimiters in the prose, no unescaped discussion macro, ≥10 level-2 headings |
What the number guard does not prove: it checks that each numeral on the page appears somewhere in the script output, not that it appears in the right sentence, and it prints only the first occurrence's context. A stale figure that happens to collide with a live one passes it — demonstrated, not assumed: changing the llm row from 12 to 77 passes the number guard. policies_table_check.mjs was written to close that hole for the four tables it can re-derive, and it catches that mutation and six others. Everything outside those four tables still rests on reading; that is how the availability-denominator findings were caught, not by a guard.
The number guard's ALLOW map is shared across every page on this wiki, so whitelisting a value here would silently bless it elsewhere. Nothing was added to it for this page. The four figures that would have needed an entry were instead given real evidence: two became quote-check needles against the paper text, and two are now printed by policies_external_checks.sh.
External sources
scripts/policies_external_checks.sh re-fetches every external fact the content page states; its unedited output is at the foot of this page. Two rules it encodes, both from earlier mistakes on this wiki: print %{http_code} and the effective URL before any byte count, because a 302 recorded as an empty body has been published here as “empty 200”; and never date a repository from /releases/latest, /tags or pushed_at — ask the commits API on the repository's own default branch.
| Source | How it was verified | Verdict |
|---|---|---|
OPP-115 and APP-350 (usableprivacy.org/data) | Fetched 2026-09-10, HTTP 200. Both entries present; OPP-115 = 115 policies with a documented commercial-licence path via CMU Flintbox; APP-350 = 350 app policies with the same research wording but no commercial path | Accepted, with the licence asymmetry stated on the page rather than glossed as “same licence” |
| PrivaSeer corpus size and licence | Fetched privaseer.ist.psu.edu/data, HTTP 200; the script prints the site's own sentence, “The PrivaSeer corpus is a collection of 3,967,487 privacy policies”, and “the corpus is available under a CC BY-NC-SA license” | Accepted. Note the site advertises a different, smaller figure for its live search index; that is a different object and the page quotes the corpus |
mukund/privbert on Hugging Face | Fetched, HTTP 200 | Accepted |
| Princeton–Leuven repositories | GitHub commits API on each default branch: citp/privacy-policy-historical master, last commit 2023-10-12; citp/PrivacyPoliciesOverTime master, last commit 2022-06-06; no licence file on either | Accepted. The earlier draft named only one of the two repositories; both are now named |
| PolicyLint / PoliCheck repository | benandow/PrivacyPolicyAnalysis, master, last commit 2022-10-05. GitHub reports the licence as “Other”; LICENSE.txt fetched directly is a three-clause BSD naming NC State | Accepted, with “GitHub says Other, the file says BSD-3” stated rather than picking one |
| PoliGraph | UCI-Networking-Group/PoliGraph, master, MIT, last commit 2023-06-21. PyPI queried for both poligraph and poligraph-er: HTTP 404 for each | Accepted |
| PoliGraph “needs a GPU” | README fetched: “A GPU is required to enable hardware acceleration… Note that PoliGraph-er can run without a GPU, but the performance would be significantly lower” | Corrected. The draft said “needs a GPU”; the page now states what the README states |
| Lalaine, Calpric, PolicyChecker repositories | xiaoyue10131748/Lalaine (MIT, 2023-09-25), dlgroupuoft/Calpric (no licence, 2023-06-21), AndyXiang945/PolicyChecker (no licence, 2023-11-20) | Accepted |
| PurPliance repository | ducalpha/PurPlianceOpenSource, main, last commit 2024-03-11 — but that commit is “Update README.rst”; the last code commit is 2022-12-25. LICENSE.txt is BSD-style, U. Michigan 2022 | Accepted, and it corrected the page. The draft said every tool “stopped being maintained within a year of its paper” and called PoliGraph the newest; PurPliance is newer, and the page now says so and distinguishes a README edit from a code commit |
| “Polisis has no official repo; the maintained lineage is a community reproduction, last commit 2023-02-02” | GitHub name search returns quanmou/polisis (master last commit 2020-07-27) and SmartDataAnalytics/Polisis_Benchmark (pushed_at 2023-02-02, but master last commit 2020-02-13). Every branch of the latter was enumerated: the 2023-02-02 activity is on dependabot/pip/werkzeug-0.15.5 | Rejected as written. pushed_at counts activity on any branch, and a Dependabot security bump is not maintenance. The page now gives the default-branch dates |
| “Calpric, Lalaine and PolicyChecker are effectively undiscoverable; only reachable through the USENIX artifact index” | Tested. secartifacts.github.io/usenixsec2023/results (HTTP 200) lists Calpric and Lalaine but not PolicyChecker, which is a CCS paper. GitHub name search returns the right repository first for PolicyChecker, Calpric, Lalaine and PoliGraph — but for PolicyLint and PoliCheck it returns only unrelated projects, because the repository is named after the paper series | Rejected and replaced. The claim was true of the wrong tools. PolicyChecker was removed from it; PolicyLint/PoliCheck were added, and the page now says which search finds what |
“w3.org/P3P has not changed since 2007” | w3.org answers curl with a Cloudflare HTTP 403, so this was checked with Playwright's chromium (scripts/policies_w3c_p3p_check.mjs). w3.org/P3P/ returns 200 with Last-Modified: Fri, 02 Feb 2018; its newest news item is “3 October 2007: The Policy Languages Interest Group (PLING) was created” and its footer reads “Last updated $Date: 2018/02/02” | Corrected. “Unchanged since 2007” was wrong; the page now states both dates and what each one is |
| W3C P3P 1.1 retirement | Same Playwright fetch of w3.org/TR/P3P11/ (HTTP 200): the visible text contains “Retired 30 August 2018” and “should not be referenced in this form or implemented as-is”, both asserted by the script | Accepted |
| Apple privacy manifests | developer.apple.com/news/?id=pvszzano fetched, HTTP 200, dated 26 April 2024, effective 1 May 2024, containing “will expand to include the entire app binary”. The reviewer additionally fetched Apple's live upcoming requirements page and found no newer deadline superseding it | Accepted, still current |
| Google Play Data safety and the User Data policy | support.google.com/…/answer/10144311 fetched, HTTP 200, still containing “Apps that do not access any personal and sensitive user data must still submit a privacy policy.” verbatim | Accepted for the quote. The “mandatory since 20 July 2022” date is not on the current Google page — it rests on contemporaneous reporting, and is flagged below as something the primary source no longer states |
| ADPC status | Search only, no primary fetch: still a proposal, not in production | Accepted as a weak claim, and it carries no figure |
| Any SEO listicle, vendor blog or “top 10 privacy policy tools” page | — | Rejected on sight. None was consulted or cited |
What could not be established
- Google's own current pages do not state when Data safety became mandatory. The 20 July 2022 date rests on contemporaneous third-party reporting. The quote the page uses is on the live primary page; the date is not. Closing this needs an archived snapshot of the Play Console announcement.
- Whether the visible text of
w3.org/P3P/has changed since 2007. Only the server'sLast-Modified(2018-02-02) and the page's own newest news item (2007-10-03) could be established; the Wayback Machine rate-limited both attempts to diff them. - How much of the 16%–94.2% availability spread is ecosystem and how much is method. Nobody has run five link-detection heuristics over one site sample. This is on the page as an open question, not resolved.
- Whether LLM extraction actually beats PoliGraph or a PrivBERT baseline. No head-to-head evaluation on the same documents with the same ground truth was found in these seven venues. The page says the switch is currently a preference, not a finding.
- Anything about venues outside the seven. OPP-115 (ACL 2016), APP-350 and PrivaSeer (ACL 2021) are cited as artefacts because corpus papers use them, but ACL, EMNLP, CHI, SOUPS and the law reviews are not in the corpus and no count here covers them.
- The ~20% free-text run-to-run stability and 0.9% unlocatable-quote rates quoted in this project's documentation were measured on the previous 4,322-paper extraction run and have not been re-measured on this one. They are used here only as an order of magnitude, and are the reason nothing on the content page publishes a percentage over a folded free-text field.
- Figures deliberately not published: no single “fraction of the web with a privacy policy”; no per-year method trend (the era buckets are used instead); no percentage over any folded free-text field; no rate over the 123-paper UNION except the full-text probe table and the two prose sentences that quote its largest row, all three labelled.
Review
Four reviewers, all told explicitly that the author's context might not be exhaustive, and all handed the page text, every script and its unedited output. The three focused passes ran in parallel on the first revision; the generic pass ran afterwards on the corrected page and on this log.
Pass 1 — figures against the script (''sonnet'')
Re-ran all four scripts and diffed them against the committed outputs: byte-identical, no drift. Then checked every numeral on the page.
| # | Finding | Verdict |
|---|---|---|
| 1 | The Fisher's-exact table states rates over the 123-paper UNION, which the page's own methodology section says never happens outside the probe table | Accepted. The comparisons were moved to POLICY (102). While fixing it, two further improvements: the counts are no longer hand-keyed — report_policies.mjs now emits SIG| lines that policies_significance.py parses — and the base rate now excludes the subgroup, since comparing a set against a population containing it shrinks the difference. The validation row's p moved 0.545 → 0.613 and is still not significant |
| 2 | “65% of the papers in this literature run a policy-versus-behaviour comparison” is a full-text probe result stated twice as fact, outside the table that carries the caveat | Accepted. Both sentences now say “80 of the 123 candidate papers” and name it as an upper bound, and the methodology bullet now lists all three places a 123-denominated figure appears |
| 3 | The EU availability row prints 6,579 while the paper's own availability table totals 6,357 | Accepted, and the fix went further than the finding: the paper states three numbers (6,759 January domains, 6,579 in its abstract, 6,357 in Table II). The page's cell now says which is which, and two quote-check needles were added |
| 4 | “99,194 policy links” is really 99,194 apps | Accepted. Both the page and the script's hand-keyed availability map now say “usable apps”, and the page points out that the paper divides link failures by its app count |
| 5 | “Every measured figure above was checked… 51 needles” overstates the coverage: several published figures had no needle | Accepted. Ten needles were added (pan2024_trap ×3, andow2020_actions ×3, ali2024_honesty ×2, xiang2023_policychecker, zimmeck2017_automated), then five more from pass 2. The check now runs 68 needles, 68 located |
| 6 | The method table silently drops the other (8) and curated-database (1) rows | Accepted. Both restored. A dropped enum row is exactly where something hides |
| 7 | The per-year table starts at 2019, silently dropping 8 UNION papers | Accepted. 2014, 2016, 2017 and 2018 restored |
The pass also independently re-derived a long list of figures it found correct, including the whole tool-lineage table, the per-venue table, the platform split and the probe table.
Pass 2 — citations and quotes (''sonnet'')
Independently re-extracted the page's citekeys and quoted spans rather than working from a supplied list, and verified all 22 new BibTeX entries against the corpus index, the paper PDFs and the PETS landing pages.
| # | Finding | Verdict |
|---|---|---|
| 1 | The page says two PETS 2026 papers compare policy text against store-declared labels; [11Cory, Thomas; Rieder, Wolf; Krämer, Julia; Raschke, Philip; Herbke, Patrick; Küpper, Axel (2026): "Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Language Models", Proceedings on Privacy Enhancing Technologies 2026(1):509-528. (DOI)] does not — it is word-level GDPR-transparency annotation, and privacy label, data safety and nutrition label appear zero times in it | Accepted. Re-checked directly (0 hits for each phrase against 108 for GDPR as a positive control). The sentence now names only the papers that do this, and cory2026_wordlevel was moved to the LLM row, where it belongs |
| 2 | MAPS's 50.5% is over the 1,035,853 analysed apps, not the 1,049,790 retrieved | Accepted. Verified against the paper (“our analysis reveals that only 50.5% of apps have links”, in a section operating on the analysed set). Page and script both corrected, plus two needles |
| 3 | “across CS, NLP and law” is not the interview paper's own discipline breakdown | Accepted. The page now gives the paper's own breakdown |
| 4 | The Degeling denominator is internally inconsistent in the source paper; flagged for awareness, not as a page defect | Accepted as information, and merged with pass 1's finding 3 |
It also confirmed: 34 of 34 citekeys resolve, 0 key collisions, 0 DOI or title duplicates against the live bibliography, and all 22 new entries correct in author order, venue, year and identifier — including three where the PDF's front matter splits the byline across columns and a careless reader would drop authors.
Pass 3 — external currency (''sonnet'')
Fetched rather than recalled. Its findings and their verdicts are in the External sources table above rather than duplicated here; in summary it corrected the P3P currency claim, the PoliGraph GPU claim, the APP-350 licence symmetry and the artifact-discoverability footnote, confirmed every dataset host, every repository date and licence, and both platform-policy quotes, and could not verify two things now listed under What could not be established.
One of its corrections was itself incomplete and was corrected in turn: it proposed replacing the discoverability footnote with a claim about Calpric and Lalaine, but a direct GitHub name search showed the genuinely unfindable pair is PolicyLint and PoliCheck, whose repository is named after neither tool. A reviewer finding is a lead, not a verdict.
Pass 4 — generic (''fable'')
No checklist. It ran on the corrected page and on this log, and it was the most productive of the four. Its own disclosures, recorded because they matter: one of its mutation-test commands was denied, so a later sed ran against the committed scripts/policies_quotecheck.mjs instead of a copy; it reverted the change and said so. Verified independently afterwards — the script reproduces its committed output byte for byte. It also noted that the content page was edited while it read, which is true and is a process failure of mine, recorded below.
| # | Finding | Verdict |
|---|---|---|
| 1 | “PrivaSeer is named by zero corpus papers” and “Nobody in these seven venues cites it” are false: 12 corpus papers name it in their full text, including four the page itself cites | Accepted, and it is the worst error the four reviews found. Re-derived independently: a scan of all 5,869 paper.cols.txt finds exactly 12. The zero was a fold over tools/sourceList — a fact about the extraction, published as a fact about the literature. Fixed by adding a second count column to the lineage table (extraction vs full text) and a whole-corpus full-text scan to the probe script, and by rewording all three places. The scan also found a homonym: /\bMAPS\b/ matches “Google MAPS” and the Play category “Maps & Navigation”, so the full-text pattern for MAPS is its title, and the override is printed |
| 2 | The page publishes UNION-denominated shares (the “share of corpus” and “share of venue” columns, and 70/123 in an open question) while claiming it never does | Accepted. The report now prints a POLICY share beside every UNION share, both are on the page with UNION named in the column header, the open question uses 64/102, and the methodology bullet now enumerates the 123-denominated figures instead of denying they exist. The earlier claim was written when the Fisher table was the only offender and was not revisited after pass 1 fixed that one |
| 3 | Probe counts stated as facts, several with “at all” | Accepted. Six sentences rewritten to name the probe and its width. “Only 7 papers say anything at all” became “a narrow probe matches 7 of 123, the wide form 64, and the gap is how much probe width decides the answer” |
| 4 | The Discord denominator 15,528 is arithmetic on the paper's own non-reconciling numbers; the paper states 15,525 | Accepted. The page now gives 15,525 as the paper states it and says the paper's own 14,852 does not reconcile. This is a back-calculated denominator, which is a named failure mode here |
| 5 | “heuristic-rules peaked at 40% in 2022–2024” contradicts the table two screens above (42.9% in 2014–2018) | Accepted, reworded |
| 6 | “three of the four corpora are still downloadable” while every row says live | Accepted, “all four were still reachable on 2026-09-10” |
| 7 | This log was stale after pass 2 (63 needles, “per-year 2019–2026”) | Accepted, corrected |
| 8 | [[roadmap]] on this page resolves to provenance:privacy:roadmap and renders red | Accepted, and confirmed in the rendered DOM of the first published revision: exactly one wikilink2 span. Now [[:roadmap]] |
| 9 | The Apple December 2020 date is not primary-sourced while the parallel Google date is flagged | Accepted, flagged in the same way |
| 10 | Generator guards: deleting two whole sections passed (floor of ≥10 headings), emptying an output passed, and the file-block count assertion was tautological | Accepted, all three. The heading check now names the thirteen expected sections; each output must be over 200 bytes, contain a known terminal string, and be newer than the script that produced it; and OUTPUTS is compared against glob('scripts/policies_*-output.txt') rather than against itself. All five mutations are now caught |
| 11 | Quote-check needles are weak: 26 of 68 are bare numbers, and a wrong digit passes | Accepted. See Quotes above: a specificity report was added and nine needles rewritten. Four remain weak and are printed |
| 12 | The 46 enum-only papers are characterised from memory as “mostly Alexa-skill, IoT and VR” | Accepted. The report now lists all 46 with their platform enum (other-online-service 21, mobile 20, web 18, iot 4, offline 1) and the page says what is actually there |
| 13 | “Half of this literature is about Android apps” — the store-label half is iOS-heavy | Accepted, “mobile apps” |
| 14 | “Both are current and both are enforced” — no source for enforcement | Accepted. The page now says both requirements are current, that neither store publishes enforcement figures, and to treat “required” as a rule rather than a fact about the population |
| 15 | “every tool in the lineage is abandoned code” overstates last-commit dates | Accepted, softened to unmaintained code, with the distinction stated |
| 16 | Four corpus papers referred to by description with no citekey, so a reader cannot find them | Accepted. Keys added for all four. Two of them turned out to be already in the bibliography under other keys ([9Khatun, Mst Eshita; Noureddine, Lamine; Bello, Sideeq; Ali-Gombe, Aisha (2026): "Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale", Proceedings on Privacy Enhancing Technologies 2026(4):213-231. (DOI)], [10Nenadić, Luka; Rodriguez, David; Calandrino, Joseph A. (2026): "Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact", Proceedings on Privacy Enhancing Technologies 2026(4):703-723. (DOI)]) — caught by bib_dedup_scan.py, not by the key-string check, which is exactly what that scan exists for |
| 17 | “four in 2024–2026” store-label comparison papers is a reading, not a count | Accepted, the number removed and the reason given |
| 18 | “the two most-starred” reproductions — the search is by relevance, not stars | Accepted, “the top two hits” |
| 19 | Licence wordings are tighter than the sources: OPP-115 says “in the spirit of” CC BY-NC and its commercial licence covers the annotation files | Accepted for OPP-115, APP-350 and PrivaSeer. The PolicyLint and PurPliance licence descriptions were left as they are: the external-check output prints the first lines of both LICENSE.txt files and the page says what they are |
| 20 | [12Chanenson, Jake; Pickering, Madison; Apthorpe, Noah (2025): "Automating Governing Knowledge Commons and Contextual Integrity (GKC-CI) Privacy Policy Annotations with Large Language Models", in: Proceedings on Privacy Enhancing Technologies. (DOI)]'s comparison arm is a custom RNN, not an OPP-115-trained classifier, so “the two to copy” misdescribes it | Accepted, split into one model and one companion read |
| 21 | A footnote published this page's own draft history | Accepted, moved here |
| 22 | P3P did define a well-known location, which the page's “no fixed address” bullet invites as an objection | Accepted, one clause added |
| — | Two gaps it named but did not call defects: no reusable retrieval code is pointed at, and nothing on the cost of LLM extraction at scale | Not fixed. Both are real. The retrieval gap is already the page's first open question; the cost question has no corpus source and would be a vendor-price claim with a shelf life of months. Recorded here rather than guessed at |
Mutation tests of the guards this page publishes
Reading a guard does not tell you whether it asserts anything. Both the generic reviewer and I broke things on purpose in /tmp copies and checked that the guard failed.
| Mutation | Guard | Result |
|---|---|---|
| A needle that is nowhere in the paper | policies_quotecheck.mjs | caught |
| A needle attributed to the wrong paper | policies_quotecheck.mjs | caught |
A needle with the wrong digit (84.5 → 84.9) | policies_quotecheck.mjs | not caught — both strings are in that paper. Hence the specificity report |
An empty needle, or % | policies_quotecheck.mjs | not caught. Left as it is: the specificity report now makes short needles visible |
llm count 12 → 77 (collides with a live value) | check_page_numbers.mjs | not caught → policies_table_check.mjs written, which catches it |
| An era percentage changed to another value from the same table | policies_table_check.mjs | caught |
| A method row silently deleted | policies_table_check.mjs | caught |
| A venue share changed to a value used elsewhere | policies_table_check.mjs | caught |
| A significance base swapped for the naive base rate | policies_table_check.mjs | caught |
| A p-value exponent moved from 10⁻⁸ to 10⁻⁵ | policies_table_check.mjs | not caught at first — the check compared three digits and ignored the exponent. Now compares the numbers, and catches it |
| Odd inline nowiki count in the prose | build_provenance_policies.py | caught |
| An unescaped discussion macro | build_provenance_policies.py | caught |
| A published script containing a closing file tag | build_provenance_policies.py | caught |
| Two whole sections deleted from the prose | build_provenance_policies.py | not caught (floor of ≥10 headings) → now names all thirteen, and catches it |
| An output file emptied, or truncated | build_provenance_policies.py | not caught → now caught |
| An output dropped from the published list | build_provenance_policies.py | not caught (tautological) → now compared against the filesystem, and caught |
| A script edited without re-running it | build_provenance_policies.py | not caught → now caught by an mtime comparison |
Seven of seventeen mutations survived the guards as first written. That ratio is the argument for mutation-testing every published check rather than reading it.
The scripts, as committed
Every block below is the file itself, inserted by scripts/build_provenance_policies.py at build time — not a sample, not an abridgement. Re-running the generator re-inserts whatever is on disk.
The report script — ''report_policies.mjs''
- report_policies.mjs
// Every figure on privacy:policies, with its denominator. // // node scripts/report_policies.mjs > scripts/report_policies-output.txt // // THREE populations are used and they are NOT interchangeable. Every table // says which one it is on, and no table mixes them: // // POLICY classification[].target === 'privacy-policy' — a STRUCTURED ENUM. // "the extractor recorded this paper as classifying privacy-policy // text". No folding needed, stable between runs. This is the page's // primary population. // PROBE /privacy polic|.../ over title + summary only — a CANDIDATE SET, // recall-oriented, NOT a population. Reported so the page can say how // much the enum misses and vice versa. Never published as a rate. // UNION POLICY ∪ PROBE — used only for "which tools/sources appear in this // literature at all" rankings, never for a percentage. // // Counts are of PAPERS, never tuples. Sentinels are printed but never counted // as an answer. Free-text sourceList and tool names are folded through // policy_fold.mjs and the unmapped residue is printed in full at the end. import { loadExtractions, pct, table, isSentinel } from './lib.mjs'; import { foldSource, foldTool, LINEAGE } from './policy_fold.mjs'; const P = loadExtractions(); const key = (p) => `${p.venue}/${p.year}/${p.slug}`; // The probe. Widened from the 2026-09-02 brainstorm's four alternatives to nine: // 'privacy label', 'data safety', 'nutrition label' and 'privacy statement' are // the store-declaration half of this literature and the narrower probe missed // them. Recorded here because probe width decides the claim. const PROBE_RE = /privacy polic|privacy notice|terms of service|terms and conditions|privacy label|data safety|nutrition label|policy text|privacy statement/i; const titleSummary = (p) => `${p.title ?? ''} • ${p.summary ?? ''}`; const POLICY = P.filter((p) => p.classification.some((c) => c.target === 'privacy-policy')); const PROBE = P.filter((p) => PROBE_RE.test(titleSummary(p))); const UNION = [...new Set([...POLICY, ...PROBE])]; const CLASSIFIED = P.filter((p) => p.classification.length > 0); const h = (s) => console.log(`\n${'='.repeat(78)}\n${s}\n${'='.repeat(78)}`); const sub = (s) => console.log(`\n--- ${s}`); // Sanity invariant: every published population must be reproducible from the // definitions above. If the corpus moves, this throws rather than drifting. if (POLICY.some((p) => !UNION.includes(p))) throw new Error('POLICY not a subset of UNION'); if (PROBE.some((p) => !UNION.includes(p))) throw new Error('PROBE not a subset of UNION'); if (UNION.length !== new Set(UNION.map(key)).size) throw new Error('UNION has duplicate keys'); // ------------------------------------------------------------ 1. populations h('1. POPULATIONS'); console.log(`corpus ${P.length}`); console.log(`classified classification[] non-empty ${CLASSIFIED.length} ${pct(CLASSIFIED.length, P.length)} of corpus`); console.log(`POLICY classification[].target=='privacy-policy' ${POLICY.length} ${pct(POLICY.length, CLASSIFIED.length)} of classified`); console.log(`PROBE title+summary regex (candidate set) ${PROBE.length}`); console.log(`UNION POLICY u PROBE ${UNION.length}`); console.log(` POLICY n PROBE ${POLICY.filter((p) => PROBE.includes(p)).length}`); console.log(` POLICY only (enum fires, title silent) ${POLICY.filter((p) => !PROBE.includes(p)).length}`); console.log(` PROBE only (title fires, enum silent) ${PROBE.filter((p) => !POLICY.includes(p)).length}`); console.log(`\nprivacy-policy tuples in POLICY ${POLICY.reduce((n, p) => n + p.classification.filter((c) => c.target === 'privacy-policy').length, 0)}`); // The enum-only papers are the argument for not using a title probe alone, so // print them rather than characterising them from memory. A 2026-09-10 review // found the page describing this set as "mostly Alexa-skill, IoT and VR studies" // on no evidence; the list below is the evidence. sub('POLICY-only: the enum fires and the title probe is silent (why a title probe is not enough)'); { const only = POLICY.filter((p) => !PROBE.includes(p)) .sort((a, b) => a.year - b.year || a.venue.localeCompare(b.venue)); console.log(`${only.length} papers. Platform measured (enum, multi-valued):`); const plat = {}; for (const p of only) for (const x of new Set(p.platforms)) plat[x] = (plat[x] ?? 0) + 1; console.log(table(['platform', 'papers of the enum-only set'], Object.entries(plat).sort((a, b) => b[1] - a[1]))); if (!Object.keys(plat).length) throw new Error('enum-only platform table came out empty'); for (const p of only) console.log(` ${p.year} ${p.venue.padEnd(8)} ${p.title}`); } sub('classification[].target, whole corpus, papers (enum — publishable)'); { const m = {}; for (const p of CLASSIFIED) for (const t of new Set(p.classification.map((c) => c.target))) m[t] = (m[t] ?? 0) + 1; console.log(table(['target', 'papers', 'share of 4,439 classified'], Object.entries(m).sort((a, b) => b[1] - a[1]).map(([k, v]) => [k, v, pct(v, CLASSIFIED.length)]))); } sub('POLICY and UNION per year (2026 PROVISIONAL: CCS/IMC 2026 not held, IEEE S&P/WWW 2026 under-selected)'); { const y = {}; for (const p of UNION) { y[p.year] ??= { u: 0, e: 0, t: 0 }; y[p.year].u += 1; if (POLICY.includes(p)) y[p.year].e += 1; if (PROBE.includes(p)) y[p.year].t += 1; } const allY = {}; for (const p of P) allY[p.year] = (allY[p.year] ?? 0) + 1; // Both shares are printed. UNION/corpus is a share of a regex-widened // candidate set and must be labelled as such wherever it is published; // POLICY/corpus is the share of the stable enum. console.log(table(['Year', 'corpus', 'POLICY', 'PROBE', 'UNION', 'POLICY share of corpus', 'UNION share of corpus'], Object.keys(y).sort().map((k) => [k + (k >= '2025' ? '*' : ''), allY[k], y[k].e, y[k].t, y[k].u, pct(y[k].e, allY[k]), pct(y[k].u, allY[k])]))); } sub('UNION per venue (denominator: that venue\'s whole corpus slice)'); { const v = {}, vAll = {}; for (const p of P) vAll[p.venue] = (vAll[p.venue] ?? 0) + 1; for (const p of UNION) v[p.venue] = (v[p.venue] ?? 0) + 1; console.log(table(['Venue', 'UNION', 'POLICY', 'venue papers', 'POLICY share of venue', 'UNION share of venue'], Object.entries(v).sort((a, b) => b[1] - a[1]).map(([k, n]) => { const pol = POLICY.filter((p) => p.venue === k).length; return [k, n, pol, vAll[k], pct(pol, vAll[k]), pct(n, vAll[k])]; }))); // The page says "PETS by a factor of five over the next venue"; print the ratio // rather than leaving the reader to divide two percentages by eye. { const byPolicy = Object.keys(vAll) .map((k) => [k, POLICY.filter((p) => p.venue === k).length / vAll[k]]) .sort((a, b) => b[1] - a[1]); const byUnion = Object.entries(v).map(([k, n]) => [k, n / vAll[k]]).sort((a, b) => b[1] - a[1]); console.log(`top venue over the next, POLICY share: ${byPolicy[0][0]} / ${byPolicy[1][0]} = ${(byPolicy[0][1] / byPolicy[1][1]).toFixed(1)}x`); console.log(`top venue over the next, UNION share : ${byUnion[0][0]} / ${byUnion[1][0]} = ${(byUnion[0][1] / byUnion[1][1]).toFixed(1)}x`); } } sub('UNION by platform measured (multi-valued: an app+web paper is in two rows)'); { const m = {}; for (const p of UNION) for (const x of new Set(p.platforms)) m[x] = (m[x] ?? 0) + 1; console.log(table(['platform', 'papers', 'share of UNION'], Object.entries(m).sort((a, b) => b[1] - a[1]).map(([k, v]) => [k, v, pct(v, UNION.length)]))); const web = UNION.filter((p) => p.platforms.includes('web')); const mob = UNION.filter((p) => p.platforms.includes('mobile')); console.log(`\nweb only ${web.filter((p) => !p.platforms.includes('mobile')).length}`); console.log(`mobile only ${mob.filter((p) => !p.platforms.includes('web')).length}`); console.log(`both ${web.filter((p) => p.platforms.includes('mobile')).length}`); console.log(`neither ${UNION.filter((p) => !p.platforms.includes('web') && !p.platforms.includes('mobile')).length}`); } // ------------------------------------------------- 2. how policies are labelled h('2. HOW THE FIELD LABELS POLICY TEXT (population: POLICY, n=' + POLICY.length + ')'); console.log('classification[].method is an ENUM. Multi-valued: a paper with two'); console.log('privacy-policy tuples using two methods appears in two rows.'); const polTuples = (p) => p.classification.filter((c) => c.target === 'privacy-policy'); const methodSet = (p) => new Set(polTuples(p).map((c) => c.method)); sub('method, all years'); { const m = {}; for (const p of POLICY) for (const x of methodSet(p)) m[x] = (m[x] ?? 0) + 1; console.log(table(['method', 'papers', 'share of POLICY'], Object.entries(m).sort((a, b) => b[1] - a[1]).map(([k, v]) => [k, v, pct(v, POLICY.length)]))); } // ERAS. 2025-2026 is starred everywhere it appears. const ERAS = [ ['2014-2018', (y) => y <= 2018], ['2019-2021', (y) => y >= 2019 && y <= 2021], ['2022-2024', (y) => y >= 2022 && y <= 2024], ['2025-2026*', (y) => y >= 2025], ]; sub('method by era — this is the CURRENCY table the page leans on'); { const methods = [...new Set(POLICY.flatMap((p) => [...methodSet(p)]))]; const rows = methods.map((mm) => { const cells = ERAS.map(([, f]) => { const sub2 = POLICY.filter((p) => f(p.year)); const n = sub2.filter((p) => methodSet(p).has(mm)).length; return `${n} (${pct(n, sub2.length)})`; }); const total = POLICY.filter((p) => methodSet(p).has(mm)).length; return [mm, total, ...cells]; }).sort((a, b) => b[1] - a[1]); console.log(table(['method', 'all', ...ERAS.map(([l]) => `${l} n=${POLICY.filter((p) => ERAS.find(([ll]) => ll === l)[1](p.year)).length}`)], rows)); } sub('classification[].validation on privacy-policy tuples (enum; none-reported is a REAL value, not a gap in the data)'); { const m = {}; for (const p of POLICY) for (const x of new Set(polTuples(p).map((c) => c.validation))) m[x] = (m[x] ?? 0) + 1; console.log(table(['validation', 'papers', 'share of POLICY'], Object.entries(m).sort((a, b) => b[1] - a[1]).map(([k, v]) => [k, v, pct(v, POLICY.length)]))); // Same row for the whole classified corpus, so the page can print the base rate // beside the subgroup share instead of implying the subgroup is unusual. const base = {}; for (const p of CLASSIFIED) for (const x of new Set(p.classification.map((c) => c.validation))) base[x] = (base[x] ?? 0) + 1; console.log('\nBASE RATE, all 4,439 papers that classified anything:'); console.log(table(['validation', 'papers', 'share of classified'], Object.entries(base).sort((a, b) => b[1] - a[1]).map(([k, v]) => [k, v, pct(v, CLASSIFIED.length)]))); } sub('LLM as the labelling method — POLICY papers whose privacy-policy tuple has method=="llm", listed in full'); { const llm = POLICY.filter((p) => methodSet(p).has('llm')).sort((a, b) => a.year - b.year); console.log(`n=${llm.length} of ${POLICY.length} POLICY papers`); for (const p of llm) console.log(` ${p.year} ${p.venue.padEnd(8)} ${p.title}`); } // --------------------------------------------------- 3. where policies come from h('3. WHERE THE POLICIES COME FROM (population: UNION, n=' + UNION.length + ')'); sub('population[].unit (enum)'); { const m = {}; for (const p of UNION) for (const x of new Set(p.population.map((x2) => x2.unit))) if (!isSentinel(x)) m[x] = (m[x] ?? 0) + 1; console.log(table(['unit', 'papers'], Object.entries(m).sort((a, b) => b[1] - a[1]))); } sub('population[].sourceList, FOLDED through policy_fold.mjs (free text — a ranking, not percentages)'); { const m = {}, residue = {}; for (const p of UNION) { const seen = new Set(); for (const s of p.population) { if (isSentinel(s.sourceList) || !s.sourceList) continue; const f = foldSource(s.sourceList); if (f === null) { residue[s.sourceList] = (residue[s.sourceList] ?? 0) + 1; continue; } seen.add(f); } for (const f of seen) m[f] = (m[f] ?? 0) + 1; } console.log(table(['folded source', 'papers'], Object.entries(m).sort((a, b) => b[1] - a[1]))); console.log(`\nRESIDUE — sourceList strings matching no fold rule (${Object.keys(residue).length} distinct, printed in full):`); for (const [k, v] of Object.entries(residue).sort((a, b) => b[1] - a[1])) console.log(` ${String(v).padStart(2)} ${k}`); } // ------------------------------------------------------------- 4. the lineage h('4. THE TOOL LINEAGE (population: whole corpus, so a tool used outside UNION is visible)'); console.log('Matched against tools[].name, otherToolsMentioned[].name,'); console.log('classification[].resourceName/targetDetail, population[].sourceList,'); console.log('detection[].phenomenon/technique. Counts are PAPERS.'); { const NAMED = LINEAGE; // defined in policy_fold.mjs, shared with the full-text probe const hay = (p) => { const a = []; for (const t of [...p.tools, ...p.otherToolsMentioned]) a.push(t.name ?? '', t.purpose ?? ''); for (const c of p.classification) a.push(c.resourceName ?? '', c.targetDetail ?? ''); for (const s of p.population) a.push(s.sourceList ?? ''); for (const d of p.detection) a.push(d.phenomenon ?? '', d.technique ?? ''); return a.join(' | '); }; const rows = NAMED.map(([label, re]) => { const all = P.filter((p) => re.test(hay(p))); const yrs = all.map((p) => p.year).sort(); return [label, all.length, all.filter((p) => UNION.includes(p)).length, yrs.length ? `${yrs[0]}-${yrs[yrs.length - 1]}` : '—', all.filter((p) => p.year >= 2024).length]; }); console.log(table(['artefact (first paper)', 'papers, corpus', 'of those in UNION', 'year range of use', 'used 2024+'], rows)); } sub('tools[] + otherToolsMentioned in UNION, FOLDED (free text — ranking only)'); { const m = {}, residue = {}; for (const p of UNION) { const seen = new Set(); for (const t of [...p.tools, ...p.otherToolsMentioned]) { if (!t.name) continue; const f = foldTool(t.name); if (f === null) { residue[t.name] = (residue[t.name] ?? 0) + 1; continue; } seen.add(f); } for (const f of seen) m[f] = (m[f] ?? 0) + 1; } console.log(table(['folded tool family', 'papers in UNION'], Object.entries(m).sort((a, b) => b[1] - a[1]))); const res = Object.entries(residue).sort((a, b) => b[1] - a[1]); console.log(`\nRESIDUE — ${res.length} distinct tool names matching no fold rule. Top 40 by frequency:`); for (const [k, v] of res.slice(0, 40)) console.log(` ${String(v).padStart(2)} ${k}`); console.log(` … and ${Math.max(0, res.length - 40)} more, each in 1-2 papers.`); } // --------------------------------------------------- 5. measured results h('5. MEASURED RESULTS: POLICY AVAILABILITY BY ECOSYSTEM'); console.log('Hand-keyed from detection[].prevalence tuples in UNION papers. The map'); console.log('below is INSIDE this script on purpose: the page must not carry a'); console.log('per-paper figure the script cannot print. Every row was read back'); console.log('against the paper\'s own quote; see report_policies_quotecheck.mjs.'); { // [citekey, venue/year/slug, ecosystem, denominator as the paper states it, figure] const AVAIL = [ ['degeling2019_value', 'NDSS/2019/we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy', 'web, EU', '6,357 (the total of the paper\'s own availability table); the paper also states 6,759 domains in its January lists and 6,579 in its abstract', '84.5% had a policy after 25 May 2018, up from 79.6% in January'], ['vallina2019_porn', 'IMC/2019/tales-from-the-porn-a-comprehensive-privacy-analysis-of-the-web-porn-ecosystem', 'web, adult sites', '6,843 pornographic websites', 'only 16% had an accessible privacy policy'], ['cui2025_privacy', 'PETS/2025/understanding-privacy-norms-through-web-forms', 'web, sites with a PI-collecting form', '10,143 websites', '94.2% (9,559) had a privacy-policy link'], ['zimmeck2019_maps', 'PETS/2019/maps-scaling-privacy-compliance-analysis-to-a-million-apps', 'Android', '1,035,853 analysed apps, of 1,049,790 retrieved', 'only 50.5% had a policy link on the Play Store page'], ['pan2024_trap', 'USENIX/2024/is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin', 'Android', '99,194 usable apps — the paper divides link failures by its app count', '37.5% (37,150/99,194) of policy links led to an unavailable page; separately 15.7% (15,572/99,194) of apps have no link at all'], ['manandhar2022_smart', 'USENIX/2022/smart-home-privacy-policies-demystified-a-study-of-availability-content-and-cove', 'smart-home vendors', '596 vendors on 7 platforms', '48.99% had a device-applicable policy; 10.57% had none at all'], ['lentzsch2021_alexa', 'NDSS/2021/hey-alexa-is-this-skill-safe-taking-a-closer-look-at-the-alexa-skill-ecosystem', 'Alexa skills', '150,708 skills across 7 country stores', '36,475 (24.2%) provided a policy link'], ['yan2024_quality', 'PETS/2024/on-the-quality-of-privacy-policy-documents-of-virtual-personal-assistant-applica', 'Alexa skills', '65,195 skills', '21,063 of 65,195 provided a policy link'], ['zhan2024_vpvet', 'CCS/2024/vpvet-vetting-privacy-policies-of-virtual-reality-apps', 'VR apps', '11,923 apps on 10 VR platforms', '29.5% had a findable privacy policy'], ['edu2022_exploring', 'IMC/2022/exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services', 'Discord chatbots requesting permissions', "15,525 unique active chatbots (the paper's own Table 2 total); its 14,852-without figure does not reconcile with it", '676 (4.35%) had a policy; 14,852 (95.67%) did not'], ['wu2025_depth', 'IMC/2025/an-in-depth-investigation-of-data-collection-in-llm-app-ecosystems', 'GPT Actions', 'Actions declaring a legal_info_url', '93.96% of those policies were reachable'], ]; const byKey = new Map(P.map((p) => [key(p), p])); const rows = AVAIL.map(([ck, k, eco, den, fig]) => { if (!byKey.has(k)) throw new Error(`AVAIL row names a paper not in the corpus: ${k}`); const p = byKey.get(k); return [ck, `${p.venue} ${p.year}`, eco, den, fig]; }); console.log(table(['citekey', 'venue', 'ecosystem', 'denominator (paper\'s own)', 'policy availability'], rows)); console.log(`\nrows: ${rows.length}; every one resolves to a corpus paper.`); } h('6. MEASURED RESULTS: POLICY-VERSUS-BEHAVIOUR CONSISTENCY'); { const CONS = [ ['zimmeck2017_automated', 'NDSS/2017/automated-analysis-of-privacy-requirements-for-mobile-apps', '9,050 apps with policies', 'mean 1.83 potential inconsistencies per app'], ['libert2018_automated', 'WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w', '1,807,491 identified third-party transmissions', 'only 14.80% were disclosed in the policy'], ['andow2019_policylint', 'USENIX/2019/policylint-investigating-internal-privacy-policy-contradictions-on-google-play', '11,430 policies', '14.2% (1,618) contained logical contradictions; 17.7% (2,028) contradictions or narrowing definitions'], ['andow2020_actions', 'USENIX/2020/actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an', '13,796 applications / 45,603 data flows', '42.4% of apps had an omitted or incorrect disclosure; 31.1% of flows were omitted; only 0.5% of flows were clearly disclosed'], ['trimananda2022_ovrseen', 'USENIX/2022/ovrseen-auditing-network-traffic-and-privacy-policies-in-oculus-vr', '1,135 data flows in Oculus VR apps', '68% (776) inconsistent disclosures'], ['cui2023_poligraph', 'USENIX/2023/poligraph-automated-privacy-policy-analysis-using-knowledge-graphs', '1,566 mapped statement pairs', '13.5% (211) conflicting; 25.5% (1,339/5,255) of policies define a term differently from the CCPA-based ontology'], ['xiang2023_policychecker', 'CCS/2023/policychecker-analyzing-the-gdpr-completeness-of-mobile-apps-privacy-policies', '163,068 analysable policies', '99.3% incomplete under GDPR; 98.1% had at least one mandatory-requirement violation'], ['xiao2023_lalaine', 'USENIX/2023/lalaine-measuring-and-characterizing-non-compliance-of-apple-privacy-labels', '5,102 fully tested iOS apps', '3,423 (67.1%) non-compliant privacy labels'], ['samarin2023_lessons', 'PETS/2023/lessons-in-vcr-repair-compliance-of-android-app-developers-with-the-california-c', '69 apps with CCPA disclosures', '80% (55) collected an identifier they did not disclose'], ['ali2024_honesty', 'PETS/2024/honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a', 'iOS apps labelled "Data Not Collected"', '97% had policy statements indicating data collection'], ['cui2025_privacy', 'PETS/2025/understanding-privacy-norms-through-web-forms', 'websites collecting PI via web forms', 'phi coefficient between observed collection and PoliGraph-er disclosure was < 0.20 for every PI type'], ]; const byKey = new Map(P.map((p) => [key(p), p])); const rows = CONS.map(([ck, k, den, fig]) => { if (!byKey.has(k)) throw new Error(`CONS row names a paper not in the corpus: ${k}`); const p = byKey.get(k); return [ck, `${p.venue} ${p.year}`, den, fig]; }); console.log(table(['citekey', 'venue', 'denominator (paper\'s own)', 'finding'], rows)); console.log(`\nrows: ${rows.length}; every one resolves to a corpus paper.`); } // ------------------------------------------------------- 7. reading / language h('7. WHAT ELSE THE UNION MEASURES'); sub('UNION papers that are also in the `legal` population (legal[] non-empty)'); { const L = UNION.filter((p) => p.legal.length > 0); const LC = P.filter((p) => p.legal.length > 0); console.log(`${L.length} of ${UNION.length} UNION papers assess a law; base rate ${LC.length} of ${P.length} (${pct(LC.length, P.length)}) corpus-wide.`); const laws = {}; for (const p of L) for (const x of new Set(p.legal.map((l) => l.law).filter((x2) => x2 && !isSentinel(x2)))) laws[x] = (laws[x] ?? 0) + 1; console.log(table(['law (free text, unfolded — ranking only)', 'papers'], Object.entries(laws).sort((a, b) => b[1] - a[1]).slice(0, 15))); } sub('human annotation in UNION — policy labelling is a hand-coding literature'); { const ann = UNION.filter((p) => p.humanAnnotation.length > 0); const withMetric = ann.filter((p) => p.humanAnnotation.some((a) => !isSentinel(a.agreementMetric))); const withCount = ann.filter((p) => p.humanAnnotation.some((a) => !isSentinel(a.annotatorCount))); console.log(`UNION papers with humanAnnotation[] ${ann.length} of ${UNION.length} ${pct(ann.length, UNION.length)}`); console.log(` …stating an agreement metric ${withMetric.length} ${pct(withMetric.length, ann.length)} of annotated`); console.log(` …stating an annotator count ${withCount.length} ${pct(withCount.length, ann.length)} of annotated`); const ANN = P.filter((p) => p.humanAnnotation.length > 0); const bMetric = ANN.filter((p) => p.humanAnnotation.some((a) => !isSentinel(a.agreementMetric))); console.log(`\nBASE RATE, all ${ANN.length} papers that coded data by hand:`); console.log(` …stating an agreement metric ${bMetric.length} ${pct(bMetric.length, ANN.length)}`); } sub('temporal[].mode in UNION (enum) vs corpus base rate — is this a longitudinal literature?'); { const count = (set) => { const m = {}; for (const p of set) for (const x of new Set(p.temporal.map((t) => t.mode))) if (!isSentinel(x)) m[x] = (m[x] ?? 0) + 1; return m; }; const u = count(UNION), b = count(P); const uT = UNION.filter((p) => p.temporal.length > 0).length; const bT = P.filter((p) => p.temporal.length > 0).length; console.log(`UNION papers with a temporal[] tuple ${uT} of ${UNION.length}; corpus ${bT} of ${P.length}`); console.log(table(['temporal.mode', 'UNION', 'share of UNION w/ temporal', 'corpus', 'share of corpus w/ temporal'], [...new Set([...Object.keys(u), ...Object.keys(b)])].sort((x, y) => (u[y] ?? 0) - (u[x] ?? 0)) .map((k) => [k, u[k] ?? 0, pct(u[k] ?? 0, uT), b[k] ?? 0, pct(b[k] ?? 0, bT)]))); } sub('artifact availability in UNION vs corpus base rate (enum)'); { const a = {}, b = {}; for (const p of UNION) if (p.artifacts) a[p.artifacts.availability] = (a[p.artifacts.availability] ?? 0) + 1; for (const p of P) if (p.artifacts) b[p.artifacts.availability] = (b[p.artifacts.availability] ?? 0) + 1; const withArt = P.filter((p) => p.artifacts); const uWithArt = UNION.filter((p) => p.artifacts); console.log(table(['availability', 'UNION', 'share', 'corpus', 'share'], [...new Set([...Object.keys(a), ...Object.keys(b)])].sort((x, y) => (b[y] ?? 0) - (b[x] ?? 0)) .map((k) => [k, a[k] ?? 0, pct(a[k] ?? 0, uWithArt.length), b[k] ?? 0, pct(b[k] ?? 0, withArt.length)]))); } // --------------------------------------- 7b. inputs for the significance test // // policies_significance.py PARSES these lines. They are emitted here, from the // same population objects every other table uses, so the p-values can never be // computed over a hand-keyed count that has drifted from the report. // // Each comparison is emitted TWICE: once over POLICY (the stable enum, 102) and // once over UNION (the regex-widened candidate set, 123). The page publishes the // POLICY rows; the UNION rows are published on the provenance page so a reader // can see whether the choice of population changes any conclusion. h('7b. SIGNIFICANCE INPUTS — parsed by policies_significance.py'); { const hasTemporalArchive = (p) => p.temporal.some((x) => x.mode === 'web-archive'); const hasTemporal = (p) => p.temporal.length > 0; const annotated = (p) => p.humanAnnotation.length > 0; const statesMetric = (p) => p.humanAnnotation.some((a) => !isSentinel(a.agreementMetric)); const hasArtifacts = (p) => Boolean(p.artifacts); const publicArtifacts = (p) => Boolean(p.artifacts) && p.artifacts.availability === 'public'; const assessesLaw = (p) => p.legal.length > 0; const policyTuples = (p) => p.classification.filter((c) => c.target === 'privacy-policy'); const noValidation = (p) => policyTuples(p).some((c) => c.validation === 'none-reported'); const hasPolicyTuple = (p) => policyTuples(p).length > 0; // The validation comparison is the one asymmetric case: on the subgroup side // the question is about the paper's PRIVACY-POLICY tuples, on the corpus side // it is about ANY classification tuple. Giving each comparison its own pair of // predicates keeps that visible instead of hiding it in one shared predicate. const anyNoValidation = (p) => p.classification.some((c) => c.validation === 'none-reported'); const classifiedAnything = (p) => p.classification.length > 0; // [label, subgroup restrict, subgroup property, base restrict, base property] const COMPARISONS = [ ['temporal.mode == web-archive', hasTemporal, hasTemporalArchive, hasTemporal, hasTemporalArchive], ['humanAnnotation states an agreement metric', annotated, statesMetric, annotated, statesMetric], ['artifacts.availability == public', hasArtifacts, publicArtifacts, hasArtifacts, publicArtifacts], ['assesses a law (legal[] non-empty)', () => true, assessesLaw, () => true, assessesLaw], ['classification.validation == none-reported', hasPolicyTuple, noValidation, classifiedAnything, anyNoValidation], ]; for (const [label, sRestrict, sProp, bRestrict, bProp] of COMPARISONS) { const base = P.filter(bRestrict); for (const [popName, pop] of [['POLICY', POLICY], ['UNION', UNION]]) { const sub2 = pop.filter(sRestrict); // The base rate must EXCLUDE the subgroup's own papers, otherwise the // subgroup is compared against a population that contains it. const rest = base.filter((p) => !pop.includes(p)); console.log(`SIG|${label}|${popName}|${sub2.filter(sProp).length}|${sub2.length}|` + `${rest.filter(bProp).length}|${rest.length}|` + `${base.filter(bProp).length}|${base.length}`); } } console.log('\nColumns: SIG|comparison|population|sub hits|sub n|rest-of-corpus hits|rest n|whole-corpus hits|whole n'); console.log('"rest" excludes the subgroup itself; "whole" is the figure a naive base rate would use.'); } // --------------------------------------------------------------- 8. full list h('8. THE UNION, IN FULL — the audit surface for every count above'); for (const p of [...UNION].sort((a, b) => a.year - b.year || a.venue.localeCompare(b.venue))) { const tag = POLICY.includes(p) ? (PROBE.includes(p) ? 'BOTH ' : 'ENUM ') : 'PROBE'; console.log(`${tag} ${p.year} ${p.venue.padEnd(8)} ${p.title}`); }
The fold — ''policy_fold.mjs''
- policy_fold.mjs
// Fold the free-text names that privacy:policies aggregates, and nothing else. // // WHY: two fields on that page are free text and ~20% stable run-to-run by exact // string (data/extract/README.md), so aggregating them raw undercounts: // // population[].sourceList — "Google Play Store" (15 papers) and "Google Play" // (15) are one source under two spellings; "OPP-115" (11) and "OPP-115 // corpus" (2) are one dataset; five different Alexa spellings are one list. // tools[].name / otherToolsMentioned[].name — "Beautiful Soup" / "BeautifulSoup" // / "bs4", "PoliGraph" / "PoliGraph-er", "Selenium" / "Selenium WebDriver". // // Rules are ordered; the first match wins. Anything matching no rule keeps its // raw name and is printed as residue by report_policies.mjs, so the part this // file cannot classify stays visible instead of vanishing. // [canonical family, matcher tested against the raw name] export const SOURCE_FAMILIES = [ // --- app stores and app-set sources ['Google Play', /google ?play|play ?store|androzoo|playdrone|google-play-scraper/i], ['Apple App Store', /app ?store(?! optimization)|apple ?app|ios app store|app annie|appfigures/i], ['Alexa list', /\balexa\b(?!.*skill)/i], ['Alexa Skills Store', /alexa (skill|store).*skill|skill store|alexa skills/i], ['Tranco', /\btranco\b/i], ['Majestic / Umbrella / Quantcast', /majestic|umbrella|quantcast|chrome ux|\bcrux\b/i], // --- annotated policy corpora ['OPP-115', /opp-?115|usable ?privacy|acl\/coling 2014|ramanath/i], ['APP-350', /app-?350/i], ['PrivaSeer', /privaseer/i], ['Princeton Policies-over-Time', /princeton.*polic|policies over time|million-document/i], // --- archives ['Wayback Machine', /wayback|internet archive|common ?crawl/i], // --- participant panels (these are NOT policy sources; kept separate so they // cannot be silently mixed into a "where policies come from" table) ['participant panel', /prolific|mechanical turk|\bmturk\b|qualtrics panel|clickworker|respondi/i], // --- other app / extension / package ecosystems ['other app ecosystem', /sidequest|oculus|meta quest|steam|chrome web store|cocoapods|rapidapi|wordpress|github|npm|pypi/i], // --- the honest bucket: the paper made its own list ['custom / hand-built list', /custom|hand-?(built|picked|curated)|manual(ly)? (selected|compiled)|own (list|corpus|selection)|seed list/i], ]; export const TOOL_FAMILIES = [ ['Polisis / PriBot', /polisis|pribot/i], ['PolicyLint', /policylint/i], ['PoliCheck', /policheck/i], ['PoliGraph', /poligraph/i], ['PolicyChecker', /policychecker/i], ['PurPliance', /purpliance/i], ['PrivBERT', /privbert/i], ['MAPS', /^maps$/], ['spaCy', /^spacy/i], ['NLTK', /^nltk/i], ['Stanford CoreNLP / AllenNLP', /corenlp|allennlp|stanza/i], ['BERT family (non-privacy)', /^bert|roberta|distilbert|legal-?bert|^albert/i], ['LLM (commercial API)', /gpt-?[345]|chatgpt|openai|claude|gemini|\bbard\b/i], ['LLM (open weights)', /llama|mistral|qwen|vicuna|falcon|deepseek/i], ['boilerplate stripper', /boilerpipe|readability|trafilatura|html2text|justext|goose/i], ['BeautifulSoup', /beautiful ?soup|^bs4$/i], ['Selenium', /^selenium/i], ['Playwright', /^playwright/i], ['Puppeteer', /^puppeteer/i], ['OpenWPM', /openwpm/i], ['language detection', /langdetect|langid|\bcld[23]?\b|fasttext.*lang/i], ['readability metric', /flesch|kincaid|gunning|smog|coleman|dale-?chall|textstat/i], ]; function fold(raw, families) { const s = String(raw ?? '').trim(); if (!s) return null; for (const [canon, re] of families) if (re.test(s)) return canon; return null; // residue: caller keeps the raw string and prints it } export const foldSource = (raw) => fold(raw, SOURCE_FAMILIES); export const foldTool = (raw) => fold(raw, TOOL_FAMILIES); // The policy-analysis lineage, one regex per artefact. Exported so that // report_policies.mjs (which matches it against EXTRACTION fields) and // policies_fulltext_probe.mjs (which matches it against the paper's FULL TEXT) // cannot drift apart. The two ask different questions and give different // answers: PrivaSeer is named as a tool or data source by no corpus paper and // cited in the text of twelve. export const LINEAGE = [ ['Privee (USENIX 2014)', /\bPrivee\b/], ['OPP-115 corpus (ACL 2016)', /OPP-?115/i], ['Polisis / PriBot (USENIX 2018)', /polisis|pribot/i], ['PolicyLint (USENIX 2019)', /policylint/i], ['MAPS (PETS 2019)', /\bMAPS\b/], ['APP-350 corpus (2019)', /APP-?350/i], ['PoliCheck (USENIX 2020)', /policheck/i], ['PurPliance (2021)', /purpliance/i], ['PrivBERT (2021)', /privbert/i], ['Calpric (USENIX 2023)', /calpric/i], ['PoliGraph / PoliGraph-er (USENIX 2023)', /poligraph/i], ['PolicyChecker (CCS 2023)', /policychecker/i], ['Lalaine (USENIX 2023)', /lalaine/i], ['PolicyComp (USENIX 2023)', /policycomp/i], ['PrivaSeer', /privaseer/i], ]; // Matching a tool name against a paper's FULL TEXT is not the same problem as // matching it against the extraction's tool fields. In full text an acronym // collides with unrelated uses: /\bMAPS\b/ matches "Google MAPS abuse" and the // Play Store's "MAPS & NAVIGATION" category. Where that happens, the full-text // scan uses the artefact's own title instead of its acronym, and the override is // recorded here rather than buried in the scanning script. export const LINEAGE_FULLTEXT_OVERRIDE = new Map([ ['MAPS (PETS 2019)', { re: /MAPS:\s*Scaling/i, why: '/\\bMAPS\\b/ also matches "Google MAPS" and the Play category "MAPS & NAVIGATION"; the paper is always cited by its title', }], ]);
The full-text probe — ''policies_fulltext_probe.mjs''
- policies_fulltext_probe.mjs
// Full-text probes for the methodological questions the extraction schema has no // field for: how a paper FINDS a privacy policy, what it does about language, and // what it does about the fact that the policy is a moving target. // // node scripts/policies_fulltext_probe.mjs > scripts/policies_fulltext_probe-output.txt // // Population: the 123 UNION papers from report_policies.mjs, re-derived here from // the same two definitions so the two scripts cannot drift apart. // // A probe count is a CANDIDATE SET, not a measurement: it says the phrase is in // the paper, not that the paper did the thing. Two forms are printed for every // probe, because probe width decides the claim. Whitespace is collapsed and // hyphenation rejoined first — a PDF line break inside a phrase otherwise // silently undercounts. // // CAVEAT ON THE SECOND COLUMN: 'wide' is a looser phrasing of the same question // for most rows, and there it is a superset. For three rows it is a DIFFERENT // question (a tool-name probe rather than a phrase probe), so it can be smaller // than 'narrow'. Those rows are flagged in the output; do not read them as a // narrow-vs-wide comparison. import fs from 'node:fs'; import path from 'node:path'; import { loadExtractions, dataRoot, pct, table } from './lib.mjs'; import { LINEAGE, LINEAGE_FULLTEXT_OVERRIDE } from './policy_fold.mjs'; const ROOT = path.join(dataRoot(), 'fulltext'); const P = loadExtractions(); const PROBE_RE = /privacy polic|privacy notice|terms of service|terms and conditions|privacy label|data safety|nutrition label|policy text|privacy statement/i; const ts = (p) => `${p.title ?? ''} • ${p.summary ?? ''}`; const UNION = [...new Set([ ...P.filter((p) => p.classification.some((c) => c.target === 'privacy-policy')), ...P.filter((p) => PROBE_RE.test(ts(p))), ])]; const norm = (s) => s.replace(//g, '').replace(/-\n/g, '').replace(/\s+/g, ' '); function text(p) { const f = path.join(ROOT, String(p.year), p.venue, p.slug, 'paper.cols.txt'); if (!fs.existsSync(f)) return null; return norm(fs.readFileSync(f, 'utf8')); } const TEXTS = new Map(); let noText = 0; for (const p of UNION) { const t = text(p); if (t === null) { noText += 1; continue; } TEXTS.set(p, t); } console.log(`UNION = ${UNION.length} papers; full text present for ${TEXTS.size}; missing ${noText}.`); console.log('Every count below is over the ' + TEXTS.size + ' papers with full text.\n'); // [label, narrow regex, wide regex]. Narrow is what the page quotes. const PROBES = [ ['finds the policy by LINK TEXT / anchor keyword', /(link|anchor)s? (text|label)|matching the (link|anchor)|keyword(s)? (in|on) the (link|footer)|footer link/i, /link text|anchor text|keyword/i], ['names a link-detection SEED PHRASE list', /seed phrase|seed keyword|list of (candidate )?(keywords|phrases) (used )?to (find|locate|identify)/i, /seed (phrase|term|keyword)/i], ['follows the policy link and reports FAILURES (404 / dead / unreachable)', /polic(y|ies)[^.]{0,80}(404|dead link|broken link|unreachable|did not resolve|failed to (load|download|retrieve))/i, /(404|broken link|dead link|unreachable)/i], ['handles a policy served as a PDF', /polic(y|ies)[^.]{0,60}\bPDF\b|\bPDF\b[^.]{0,60}polic(y|ies)/i, /\bPDF\b/], ['states the LANGUAGE of the policies it analysed', /english[- ]language (privacy )?polic|polic(y|ies)[^.]{0,60}in english|non-?english polic|language of the polic/i, /langdetect|langid|\bcld[23]\b|language detection/i], ['analyses policies in more than one language', /multilingual|bilingual|(translat(e|ed|ing|ion))[^.]{0,50}polic|polic(y|ies)[^.]{0,50}translat/i, /multilingual|bilingual|translat/i], ['strips BOILERPLATE / extracts the main content of the policy page', /boilerpipe|trafilatura|readability(-lxml| library|\.js)|justext|main content extraction|html to text|html2text/i, /boilerplate|extract(ed|ing)? the (main )?(text|content)/i], ['measures READABILITY of the policy', /flesch|kincaid|gunning fog|\bSMOG\b|coleman-liau|dale-chall|reading (ease|grade|level|time)/i, /readab(le|ility)/i], ['compares the policy against OBSERVED BEHAVIOUR (traffic, code, or storage)', /(consisten|inconsisten|discrepan|mismatch|divergen|misalign)[a-z]*[^.]{0,90}(polic|label|disclos)/i, /consistency analysis|policy[- ]to[- ]?(code|flow|traffic)/i], ['handles the policy being VERSIONED / changing under it', /polic(y|ies)[^.]{0,80}(version|snapshot|revision|updated?|chang(e|ed|es))|wayback|internet archive/i, /longitudinal|over time|snapshot/i], ['reports the sentence/segment SEGMENTATION step', /segment(ation|ed|ing)? (the )?polic|polic[^.]{0,40}(into|by) (sentences|segments|paragraphs)|sentence[- ]level/i, /segment/i], ['says which policy applies (app vs developer vs platform vs layered)', /(which|applicable|relevant|correct) privacy polic|polic(y|ies) (that )?appl(y|ies)|layered polic|multiple privacy polic|generic (company|corporate) polic/i, /applicable polic|multiple polic/i], ['deduplicates identical / templated policies', /(duplicate|template|boilerplate|reuse[d]?|identical)[^.]{0,60}polic|polic[^.]{0,60}(duplicat|templat|reus)/i, /duplicat|templat/i], ]; const rows = []; for (const [label, narrow, wide] of PROBES) { let n = 0, w = 0; for (const [, t] of TEXTS) { if (narrow.test(t)) n += 1; if (wide.test(t)) w += 1; } const flag = w < n ? ' <-- not a superset: different question' : ''; rows.push([label + flag, n, pct(n, TEXTS.size), w, pct(w, TEXTS.size)]); } console.log(table(['probe (narrow form is what the page quotes)', 'narrow', 'share', 'wide', 'share'], rows)); console.log('\n--- The two probes the page leans on hardest, with the matching papers named'); for (const [label, narrow] of [ ['follows the policy link and reports FAILURES (404 / dead / unreachable)', PROBES[2][1]], ['analyses policies in more than one language', PROBES[5][1]], ]) { const hits = [...TEXTS.entries()].filter(([, t]) => narrow.test(t)).map(([p]) => p) .sort((a, b) => a.year - b.year); console.log(`\n${label} — ${hits.length} papers`); for (const p of hits) console.log(` ${p.year} ${p.venue.padEnd(8)} ${p.title.slice(0, 92)}`); } // --------------------------------------------------------------------------- // The lineage, counted a second way: WHOLE CORPUS full text rather than the // extraction's tool/source fields. // // These are different questions and they give different answers. "No corpus // paper's extraction names PrivaSeer" is a fact about the extraction. "Nobody // in these seven venues cites PrivaSeer" would be a fact about the literature, // and it is false. A fold over a structured field cannot support an "at all" // claim; only a full-text scan can even try. // // Scans every paper.cols.txt in the corpus, not just the UNION. { const ROOTDIR = ROOT; const all = []; for (const y of fs.readdirSync(ROOTDIR)) { const yd = path.join(ROOTDIR, y); if (!fs.statSync(yd).isDirectory()) continue; for (const v of fs.readdirSync(yd)) { const vd = path.join(yd, v); if (!fs.statSync(vd).isDirectory()) continue; for (const s of fs.readdirSync(vd)) { const f = path.join(vd, s, 'paper.cols.txt'); if (fs.existsSync(f)) all.push([`${y} ${v.padEnd(8)} ${s}`, f]); } } } console.log(`\n\n--- LINEAGE ARTEFACTS: extraction fold vs FULL-TEXT mention, whole corpus`); console.log(`Full-text denominator: ${all.length} papers with a readable paper.cols.txt.`); console.log(`"extraction" is the column report_policies.mjs section 4 prints.`); const PATTERNS = LINEAGE.map(([label, re]) => { const o = LINEAGE_FULLTEXT_OVERRIDE.get(label); return [label, o ? o.re : re, o ? o.why : '']; }); const seen = new Map(PATTERNS.map(([label]) => [label, []])); for (const [id, f] of all) { const t2 = norm(fs.readFileSync(f, 'utf8').replace(/\0/g, '')); for (const [label, re] of PATTERNS) if (re.test(t2)) seen.get(label).push(id); } console.log(table(['artefact', 'full-text pattern', 'papers whose FULL TEXT names it'], PATTERNS.map(([label, re]) => [label, String(re), seen.get(label).length]))); const overridden = PATTERNS.filter(([, , why]) => why); console.log('\nFull-text patterns that differ from the extraction pattern, and why:'); for (const [label, re, why] of overridden) console.log(` ${label}: ${re} — ${why}`); console.log('\nThe two artefacts where the gap changes what may be said:'); for (const label of ['PrivaSeer', 'Calpric (USENIX 2023)']) { console.log(`\n${label} — ${seen.get(label).length} papers:`); for (const id of seen.get(label)) console.log(' ' + id); } }
The quote check — ''policies_quotecheck.mjs''
- policies_quotecheck.mjs
// Verify every measured figure privacy:policies publishes against the paper's // own text — not against the extraction's evidence.quote, which can be a // different sentence from the one carrying the number. // // node scripts/policies_quotecheck.mjs > scripts/policies_quotecheck-output.txt // // Checked against FOUR renderings of each PDF, because they fail on different // sentences: paper.cols.txt (column order repaired — the one to quote from), // paper.norm.txt, paper.txt, and a pypdf extraction cached under cache/pypdf/. // A needle matching any of them is LOCATED, and the rendering that matched is // named, so a .cols-only miss is visible. The pypdf pass is not decoration: // PETS 2024 'Honesty is the Best Policy' states its 65% template figure in a // sentence that .cols splices with an unrelated one, so .cols alone reports a // false MISS on a figure the paper does state. // // Whitespace is collapsed and hyphenation rejoined before matching: a PDF line // break inside a phrase otherwise produces a false MISS. // // Exit status is non-zero if any needle is unlocated, so this cannot pass by // being ignored. import fs from 'node:fs'; import { spawnSync } from 'node:child_process'; import path from 'node:path'; import { dataRoot } from './lib.mjs'; const ROOT = path.join(dataRoot(), 'fulltext'); const norm = (s) => s .replace(//g, '') .replace(/-\n/g, '') .replace(/[‘’ʼ]/g, "'") .replace(/[“”]/g, '"') .replace(/[‐-―−]/g, '-') .replace(/\s+/g, ' ') .trim(); const RENDERINGS = ['paper.cols.txt', 'paper.norm.txt', 'paper.txt']; const PYCACHE = 'cache/pypdf'; // the dataset mount is read-only, so cache locally function pypdfText(dir, venue, year, slug) { const out = path.join(PYCACHE, `${year}_${venue}_${slug}.txt`); if (!fs.existsSync(out)) { const pdf = path.join(dir, 'paper.pdf'); if (!fs.existsSync(pdf)) return null; fs.mkdirSync(PYCACHE, { recursive: true }); const r = spawnSync('python3', ['-c', 'import sys,pypdf;print("".join(p.extract_text() or "" for p in pypdf.PdfReader(sys.argv[1]).pages))', pdf], { encoding: 'utf8', maxBuffer: 64 * 1024 * 1024 }); if (r.status !== 0) return null; fs.writeFileSync(out, r.stdout); } return fs.readFileSync(out, 'utf8'); } const cache = new Map(); function renderings(venue, year, slug) { const k = `${venue}/${year}/${slug}`; if (!cache.has(k)) { const dir = path.join(ROOT, String(year), venue, slug); const out = {}; for (const r of RENDERINGS) { const f = path.join(dir, r); if (fs.existsSync(f)) out[r] = norm(fs.readFileSync(f, 'utf8')); } const py = pypdfText(dir, venue, year, slug); if (py !== null) out['pypdf'] = norm(py); if (Object.keys(out).length === 0) throw new Error(`no full text at all: ${dir}`); cache.set(k, out); } return cache.get(k); } // [venue, year, slug, label as it appears on the page, needle from the PAPER] // The needle is the paper's own words carrying the number the page prints. const CHECKS = [ // --- section: how many sites/apps have a policy at all ['NDSS', 2019, 'we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy', 'web policy availability 84.5% / 79.6%', '84.5 %'], ['NDSS', 2019, 'we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy', 'Degeling sample 6,579 sites, 500 per member state', '500 most popular websites'], ['IMC', 2019, 'tales-from-the-porn-a-comprehensive-privacy-analysis-of-the-web-porn-ecosystem', 'adult sites: only 16% have an accessible policy, of 6,843', 'only 16% of the analyzed websites have an accessible privacy policy'], ['PETS', 2025, 'understanding-privacy-norms-through-web-forms', 'web forms 94.2% policy link', '94.2% (9,559)'], ['PETS', 2019, 'maps-scaling-privacy-compliance-analysis-to-a-million-apps', 'Play policy links 50.5%', '50.5%'], ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin', 'APPG: 37.5% of links unavailable', 'we found that 37.5%'], ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin', 'APPG: 20.5% non-English', '20.5%'], ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin', 'APPG: 22.3% low quality under 2KB/200 words', 'we identify 22.3% (10,375/46,472)'], ['USENIX', 2022, 'smart-home-privacy-policies-demystified-a-study-of-availability-content-and-cove', 'smart-home vendors 48.99% (292/596)', '48.99%'], ['NDSS', 2021, 'hey-alexa-is-this-skill-safe-taking-a-closer-look-at-the-alexa-skill-ecosystem', 'Alexa skills 24.2% policy link', '36,475 (24.2 %)'], ['CCS', 2024, 'vpvet-vetting-privacy-policies-of-virtual-reality-apps', 'VR apps 29.5% have a policy', '(i.e., 29.5%) of privacy policies were successfully found'], ['IMC', 2022, 'exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services', 'chatbots 95.67% lack a policy', '95.67%'], ['IMC', 2019, 'tales-from-the-porn-a-comprehensive-privacy-analysis-of-the-web-porn-ecosystem', 'adult sample 6,843 sites', '6,843 pornographic websites'], ['PETS', 2025, 'understanding-privacy-norms-through-web-forms', 'web-forms denominator 10,143', 'leaving 10,143 websites for this analysis'], ['PETS', 2019, 'maps-scaling-privacy-compliance-analysis-to-a-million-apps', 'MAPS denominator 1,049,790 retrieved / 1,035,853 analysed', '1,049,790 retrieved apps, 1,035,853'], ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin', 'APPG link denominator 37,150/99,194', '(37,150/99,194)'], ['USENIX', 2022, 'smart-home-privacy-policies-demystified-a-study-of-availability-content-and-cove', 'smart home 10.57% no policy at all', '10.57% not providing privacy policies at all'], ['NDSS', 2021, 'hey-alexa-is-this-skill-safe-taking-a-closer-look-at-the-alexa-skill-ecosystem', 'Alexa denominator 150,708 skills, 36,475 with a link', 'Combined 150,708 36,475 (24.2 %)'], ['PETS', 2024, 'on-the-quality-of-privacy-policy-documents-of-virtual-personal-assistant-applica', 'Alexa 21,063 of 65,195', '65,195 skills, of which 21,063'], ['CCS', 2024, 'vpvet-vetting-privacy-policies-of-virtual-reality-apps', 'VR denominator 11,923 apps', '11,923 VR apps'], ['IMC', 2022, 'exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services', 'Discord chatbots 676 (4.35%) have a policy', '676 (4.35%)'], // --- section: policy versus behaviour ['WWW', 2018, 'an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w', 'only 14.80% of transmissions disclosed', '14.80%'], ['USENIX', 2019, 'policylint-investigating-internal-privacy-policy-contradictions-on-google-play', 'PolicyLint 14.2% (1,618/11,430) contradictions', '1,618/11,430'], ['USENIX', 2020, 'actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an', 'PoliCheck 42.4% of apps', '42.4%'], ['USENIX', 2020, 'actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an', 'PoliCheck only 0.5% of flows clearly disclosed', 'Only 0.5% of data flows were explicitly discussed'], ['USENIX', 2022, 'ovrseen-auditing-network-traffic-and-privacy-policies-in-oculus-vr', 'OVRseen 68% inconsistent', '68% (776/1,135)'], ['USENIX', 2023, 'poligraph-automated-privacy-policy-analysis-using-knowledge-graphs', 'PoliGraph 70.6% recall / 96.9% precision', '70.6% recall'], ['USENIX', 2023, 'poligraph-automated-privacy-policy-analysis-using-knowledge-graphs', 'PoliGraph 25.5% of policies redefine a term', '25.5%'], ['CCS', 2023, 'policychecker-analyzing-the-gdpr-completeness-of-mobile-apps-privacy-policies', 'PolicyChecker 99.3% incomplete', '99.3%'], ['CCS', 2023, 'policychecker-analyzing-the-gdpr-completeness-of-mobile-apps-privacy-policies', 'PolicyChecker 163,068 analysable of 205,973', '163,068'], ['USENIX', 2023, 'lalaine-measuring-and-characterizing-non-compliance-of-apple-privacy-labels', 'Lalaine 3,423 of 5,102 apps', '3,423'], ['PETS', 2023, 'lessons-in-vcr-repair-compliance-of-android-app-developers-with-the-california-c', 'CCPA VCR 80% undisclosed identifier', '(55 apps, 80%)'], ['PETS', 2024, 'honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a', 'Apple labels: 97% of "Data Not Collected" contradicted', 'almost all (97%) apps that indicate in their privacy'], ['PETS', 2025, 'understanding-privacy-norms-through-web-forms', 'phi < 0.20 policy-vs-form association', 'the association appears to be weak (< 0.20) for all PI types'], // --- section: what the policy text itself looks like ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset', 'Princeton corpus 1,071,488 policies / 130,000 sites', '1,071,488'], ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset', 'median length 876 -> 1,522 words', '1,522'], ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset', 'FKGL 11.9 -> 13.2', 'to 2019B (13.2)'], ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset', 'beacons: 25.8% of policies vs 94.6% of top-10K sites', '25.8%'], ['WWW', 2018, 'an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w', '84.7 minutes to read applicable policies', '84.7'], ['PETS', 2020, 'the-privacy-policy-landscape-after-the-gdpr', 'EU policies gained 35% words / 33% sentences, Global 25% / 22%', '(+35%, +25%)'], ['CCS', 2024, 'vpvet-vetting-privacy-policies-of-virtual-reality-apps', 'VR policy reuse 54.5% (1,919/3,521)', '54.5%'], ['PETS', 2024, 'honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a', 'template reuse 65%', 'privacy policies of 65%'], // --- section: the annotated corpora ['USENIX', 2018, 'polisis-automated-analysis-and-presentation-of-privacy-policies-using-deep-learn', 'Polisis trained on 65 of the OPP-115 policies, 50 held out', 'we used the data from 65 policies in the OPP-115 dataset, and we kept 50 policies as a testing set'], ['USENIX', 2018, 'polisis-automated-analysis-and-presentation-of-privacy-policies-using-deep-learn', 'Polisis average F1 0.84', 'Average 0.87 0.83 0.84 0.84'], ['USENIX', 2023, 'calpric-inclusive-and-fine-grain-labeling-of-privacy-policies-with-crowdsourcing', 'Calpric 16,856 labelled segments', '16856 labeled text segments'], ['PETS', 2023, 'researchers-experiences-in-analyzing-privacy-policies-challenges-and-opportuniti', 'no best practices have emerged (26 interviews)', 'no clear best practices or methodologies for privacy policy analysis have emerged'], ['PETS', 2023, 'researchers-experiences-in-analyzing-privacy-policies-challenges-and-opportuniti', '26 researchers interviewed', 'semi-structured interviews with 26 researchers'], ['PETS', 2023, 'evolution-of-composition-readability-and-structure-of-privacy-policies-over-two', 'User Choice/Control semantic change 26%', 'semantic modifications are made to 26% of the policies'], // --- section: LLM extraction (2024-2026) ['IMC', 2024, 'analyzing-corporate-privacy-policies-using-ai-chatbots', 'GPT-4 chatbot annotation of corporate policies', 'GPT-4'], ['IMC', 2025, 'an-in-depth-investigation-of-data-collection-in-llm-app-ecosystems', 'LLM disclosure classifier 87.44% accuracy', '87.44%'], ['IMC', 2025, 'an-in-depth-investigation-of-data-collection-in-llm-app-ecosystems', 'only 5.8% of Actions clearly disclose', 'only 5.8% of Actions clearly disclosing'], // --- added 2026-09-10 after the whole-page number guard flagged these as // unaccounted: figures the page prints that no earlier needle covered. ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin', 'APPG low-quality count 10,375/46,472', '22.3% (10,375/46,472)'], ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin', 'APPG non-English count 9,523/46,472', '(9,523/46,472)'], ['USENIX', 2024, 'is-it-a-trap-a-large-scale-empirical-study-and-comprehensive-assessment-of-onlin', 'APPG 15.7% of apps provide no policy', '15.7% (15,572/99,194)'], ['USENIX', 2020, 'actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an', 'PoliCheck entity-insensitive false-consistency 37.1%', '37.1% of inconsistent third-party flows as consistent'], ['USENIX', 2020, 'actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an', 'PoliCheck 31.1% (14,409/45,603) omitted flows', '31.1% (14,409/45,603)'], ['USENIX', 2020, 'actions-speak-louder-than-words-entity-sensitive-privacy-policy-and-data-flow-an', 'PoliCheck 14,409 omitted disclosures (the 31.1% numerator)', 'Of the 14,409 omitted'], ['PETS', 2024, 'honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a', 'Apple labels: 228,539 apps policy-but-not-label', 'an additional 228,539 apps'], ['PETS', 2024, 'honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a', 'template count n=306,404 behind the 65%', '306, 404'], ['CCS', 2023, 'policychecker-analyzing-the-gdpr-completeness-of-mobile-apps-privacy-policies', 'PolicyChecker 98.1% mandatory-requirement violation', '98.1% of them had at least one'], ['NDSS', 2017, 'automated-analysis-of-privacy-requirements-for-mobile-apps', 'Zimmeck mean 1.83 inconsistencies per app', 'mean of 1.83'], ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset', 'Princeton-Leuven corpus reaches back to 1997', 'policies from as early as 1997'], ['WWW', 2021, 'privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset', 'the length/readability series is 2009-2019', 'corpus are from 2009-2019'], // --- added 2026-09-10 after the citations reviewer found two availability // denominators that named a different population from the paper's own. ['NDSS', 2019, 'we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy', 'Degeling: 6357 is the total of the availability table', 'Total 6357 79.6 % 84.5 %'], ['NDSS', 2019, 'we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy', 'Degeling: the same paper also says 6,759 domains', 'contained 6,759 different domains'], ['PETS', 2019, 'maps-scaling-privacy-compliance-analysis-to-a-million-apps', 'MAPS 50.5% is over the analysed set, not the retrieved set', 'our analysis reveals that only 50.5% of apps have links to privacy policies'], ['PETS', 2019, 'maps-scaling-privacy-compliance-analysis-to-a-million-apps', 'MAPS analysed 1,035,853 of 1,049,790 retrieved', '1,035,853'], ['PETS', 2026, 'word-level-annotation-of-gdpr-transparency-compliance-in-privacy-policies-using', 'cory2026 is word-level GDPR transparency annotation by LLM', 'word-level'], // --- added 2026-09-10 after the generic review found a back-calculated // denominator: 15,528 was 14,852 + 676, not a number the paper states. ['IMC', 2022, 'exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services', 'Discord: the paper states 15,525 unique active chatbots', 'Unique active chatbots 15,525 100%'], ['IMC', 2022, 'exploring-the-security-and-privacy-risks-of-chatbots-in-messaging-services', 'Discord: and separately 14,852 (95.67%) without a policy', 'the remaining 14,852 (95.67%)'], ]; let miss = 0; let colsOnlyMiss = 0; for (const [venue, year, slug, label, needle] of CHECKS) { const rs = renderings(venue, year, slug); const n = norm(needle).toLowerCase(); const hits = Object.entries(rs).filter(([, t]) => t.toLowerCase().includes(n)).map(([r]) => r); if (hits.length === 0) { miss += 1; console.log(`MISS ${venue} ${year} ${label}\n needle: ${needle}`); } else { if (!hits.includes('paper.cols.txt')) colsOnlyMiss += 1; const tag = hits.includes('paper.cols.txt') ? 'OK cols' : `OK ${hits[0]} (NOT in .cols)`; console.log(`${tag.padEnd(28)} ${venue} ${year} ${label}`); } } // SPECIFICITY. Locating a needle proves the string is in the right PDF; it does // not prove the string is the sentence the page is quoting. A bare "26%" is in // nine of these papers. A 2026-09-10 review showed that changing the Degeling // needle from "84.5 %" to "84.9 %" still passed, because both appear in that // paper's tables. So: report, for every needle, how many OTHER check papers also // contain it, and how long it is. A needle that is short, numeric and shared is // weak evidence and the page should not lean on it alone. { const papers = [...new Set(CHECKS.map(([v, y, s]) => `${v}/${y}/${s}`))]; const texts = new Map(papers.map((k) => { const [v, y, s] = k.split('/'); return [k, Object.values(renderings(v, Number(y), s)).join(' \n ').toLowerCase()]; })); const rows = []; for (const [venue, year, slug, label, needle] of CHECKS) { const n = norm(needle).toLowerCase(); const own = `${venue}/${year}/${slug}`; const elsewhere = papers.filter((k) => k !== own && texts.get(k).includes(n)); const numericOnly = !/[a-z]/i.test(needle); rows.push([needle.length, numericOnly, elsewhere.length, label, needle]); } const weak = rows.filter(([len, num, el]) => num && el > 0); console.log(`\nSPECIFICITY of the ${CHECKS.length} needles`); console.log(` needles with no letters (pure number/punctuation): ${rows.filter((r) => r[1]).length}`); console.log(` of those, also present in another check paper : ${weak.length}`); console.log(` needles under 12 characters : ${rows.filter((r) => r[0] < 12).length}`); if (weak.length) { console.log('\n WEAK — numeric-only and not unique to the paper they are attributed to:'); for (const [len, , el, label, needle] of weak.sort((a, b) => b[2] - a[2])) { console.log(` ${String(el).padStart(2)} other check papers also contain ${JSON.stringify(needle)} (${label})`); } console.log('\n These are not wrong — each was read in context — but they are the needles'); console.log(' a future edit could break without this check noticing.'); } } console.log(`\n${CHECKS.length} needles, ${CHECKS.length - miss} located, ${miss} MISSING.`); console.log(`${CHECKS.length - miss - colsOnlyMiss} located in paper.cols.txt; ${colsOnlyMiss} located only outside paper.cols.txt.`); if (miss > 0) process.exitCode = 1;
The significance test — ''policies_significance.py''
- policies_significance.py
"""Fisher's exact test for every subgroup-vs-base-rate claim privacy:policies makes. node scripts/report_policies.mjs > scripts/report_policies-output.txt python3 scripts/policies_significance.py > scripts/policies_significance-output.txt WHY: five of the page's comparisons are of the form "policy papers do X more often than the rest of the corpus does". Four survive; one does not, and the one that does not (validation reporting) is the one that reads most like a finding. Counts are NOT hand-keyed. report_policies.mjs emits them as `SIG|...` lines from the same population objects every other table in that report uses, and this script parses those lines — so a p-value here cannot drift away from the report the way a re-typed count can. Each comparison is tested twice, once over POLICY (the stable enum, 102 papers) and once over UNION (the regex-widened candidate set, 123). The page publishes the POLICY rows because a rate over UNION is a rate over a set a regex chose; both are printed here so the reader can see whether that choice changes anything. The base rate is the REST of the corpus — the subgroup's own papers removed — because comparing a subgroup against a population that contains it shrinks the difference. The "naive" column shows what the un-removed base rate would have been, so the size of that effect stays visible. No scipy in this container, so the two-sided Fisher p-value is computed directly from the hypergeometric distribution (sum of all tables at most as probable as the observed one). """ import sys from math import comb REPORT = 'scripts/report_policies-output.txt' def fisher_two_sided(a, b, c, d): n = a + b + c + d def p(x): return comb(a + b, x) * comb(c + d, a + c - x) / comb(n, a + c) obs = p(a) lo, hi = max(0, a + c - (c + d)), min(a + b, a + c) return sum(p(x) for x in range(lo, hi + 1) if p(x) <= obs * (1 + 1e-9)) rows = [] for line in open(REPORT, encoding='utf-8'): if not line.startswith('SIG|'): continue _, label, pop, a, an, c, cn, wc, wn = line.rstrip('\n').split('|') rows.append((label, pop, int(a), int(an), int(c), int(cn), int(wc), int(wn))) if not rows: sys.exit(f'no SIG| lines in {REPORT} — re-run report_policies.mjs first') print(f'parsed {len(rows)} SIG rows from {REPORT}\n') hdr = (f"{'comparison':46s} {'pop':6s} {'subgroup':>14s} {'rest of corpus':>16s} " f"{'p (Fisher, 2-sided)':>20s} {'naive base':>14s}") print(hdr) print('-' * len(hdr)) for label, pop, a, an, c, cn, wc, wn in rows: p = fisher_two_sided(a, an - a, c, cn - c) verdict = 'significant' if p < 0.05 else 'NOT SIGNIFICANT — do not publish as a movement' print(f'{label:46s} {pop:6s} {a}/{an} ({100*a/an:4.1f}%) {c}/{cn} ({100*c/cn:4.1f}%) ' f'{p:>13.3g} {verdict} [{wc}/{wn} = {100*wc/wn:4.1f}%]') print('\n"rest of corpus" removes the subgroup\'s own papers from the base; the') print('bracketed column is the naive base rate that leaves them in.')
The cell-by-cell table check — ''policies_table_check.mjs''
- policies_table_check.mjs
// Assert that privacy:policies' four corpus tables match report_policies.mjs // CELL BY CELL, not merely that each numeral occurs somewhere in the output. // // node scripts/policies_table_check.mjs // // WHY this exists in addition to check_page_numbers.mjs: that guard tests // whether every numeral on the page appears anywhere in the concatenated script // output. A mutation test on 2026-09-10 changed the `llm` row's count from 12 to // 77 and the guard still passed, because 77 occurs elsewhere in the output (the // PROBE size, and the public-artifacts count). A wrong figure that collides with // a live one is invisible to a membership test. This one re-derives each row // from the report and compares positionally. // // Exits non-zero on the first mismatch, printing the row it disagrees with. import fs from 'node:fs'; const PAGE = process.argv[2] ?? 'pages/privacy_policies.txt'; const REPORT = process.argv[3] ?? 'scripts/report_policies-output.txt'; const SIG = process.argv[4] ?? 'scripts/policies_significance-output.txt'; const page = fs.readFileSync(PAGE, 'utf8'); const report = fs.readFileSync(REPORT, 'utf8'); const sig = fs.readFileSync(SIG, 'utf8'); let failures = 0; const fail = (what, expected, got) => { failures += 1; console.log(`MISMATCH ${what}\n report: ${expected}\n page : ${got}`); }; const ok = (what) => console.log(`OK ${what}`); // A page table row -> array of trimmed cells, markup stripped. const cells = (line) => line.replace(/^\||\|$/g, '').split('|') .map((c) => c.replace(/\*\*|''|\/\//g, '').trim()); const pageRows = (headerMatch) => { const lines = page.split('\n'); const i = lines.findIndex((l) => l.startsWith('^') && headerMatch.test(l)); if (i < 0) throw new Error(`no page table whose header matches ${headerMatch}`); const out = []; for (let j = i + 1; j < lines.length && lines[j].startsWith('|'); j += 1) out.push(cells(lines[j])); if (!out.length) throw new Error(`page table ${headerMatch} has no rows`); return out; }; const reportSection = (marker) => { const i = report.indexOf(marker); if (i < 0) throw new Error(`report has no section ${JSON.stringify(marker)}`); const rest = report.slice(i + marker.length); // A section ends at the next "--- " subsection header or "===" banner. The // column-rule line under each table is also dashes, so match the space. const end = rest.search(/\n(--- |=====)/); return (end < 0 ? rest : rest.slice(0, end)).split('\n').filter((l) => l.trim()); }; // ---------------------------------------------------------------- 1. by era { const want = new Map(); for (const l of reportSection('--- method by era')) { const m = l.match(/^(\S+)\s+(\d+)\s+\d+ \(\s*([\d.]+)%\)\s+\d+ \(\s*([\d.]+)%\)\s+\d+ \(\s*([\d.]+)%\)\s+\d+ \(\s*([\d.]+)%\)/); if (m) want.set(m[1], m.slice(2)); } if (want.size < 5) throw new Error('parsed too few method-by-era rows from the report'); const rows = pageRows(/classification\.method/); if (rows.length !== want.size) fail('method-by-era row count', want.size, rows.length); for (const r of rows) { const [name, all, ...eras] = r; if (!want.has(name)) { fail(`method row "${name}"`, '(no such method in the report)', r.join(' | ')); continue; } const w = want.get(name); const got = [all, ...eras.map((e) => e.replace('%', ''))]; const exp = [w[0], ...w.slice(1).map((x) => String(parseFloat(x)))]; const norm = got.map((x) => String(parseFloat(x))); if (norm.join(',') !== exp.join(',')) fail(`method row "${name}"`, exp.join(' | '), norm.join(' | ')); } ok(`method-by-era table: ${rows.length} rows`); } // ---------------------------------------------------------------- 2. by year { const want = new Map(); for (const l of reportSection('--- POLICY and UNION per year')) { const m = l.match(/^(\d{4})\*?\s+(\d+)\s+(\d+)\s+(\d+)\s+(\d+)\s+([\d.]+)%/); if (m) want.set(m[1], [m[2], m[3], m[5], m[6]]); // corpus, POLICY, UNION, share } const rows = pageRows(/\^ Year \^/); for (const r of rows) { const year = r[0].replace('*', ''); if (!want.has(year)) { fail(`year row ${year}`, '(not in the report)', r.join(' | ')); continue; } const w = want.get(year); const got = [r[1].replace(/,/g, ''), r[2], r[3], r[4].replace('%', '')]; const exp = [w[0], w[1], w[2], String(parseFloat(w[3]))]; if (got.map((x) => String(parseFloat(x))).join(',') !== exp.join(',')) { fail(`year row ${year}`, exp.join(' | '), got.join(' | ')); } } if (rows.length !== want.size) fail('per-year row count', want.size, rows.length); ok(`per-year table: ${rows.length} rows`); } // --------------------------------------------------------------- 3. by venue { const want = new Map(); for (const l of reportSection('--- UNION per venue')) { const m = l.match(/^(\S+)\s+(\d+)\s+(\d+)\s+(\d+)\s+([\d.]+)%/); if (m) want.set(m[1], [m[2], m[3], m[4], m[5]]); } const alias = { 'USENIX Sec': 'USENIX', TheWebConf: 'WWW', 'IEEE S&P': 'IEEE-SP' }; const rows = pageRows(/\^ Venue \^/); for (const r of rows) { const v = alias[r[0]] ?? r[0]; if (!want.has(v)) { fail(`venue row ${r[0]}`, '(not in the report)', r.join(' | ')); continue; } const w = want.get(v); const got = [r[1], r[2], r[3].replace(/,/g, ''), r[4].replace('%', '')]; const exp = [w[0], w[1], w[2], String(parseFloat(w[3]))]; if (got.map((x) => String(parseFloat(x))).join(',') !== exp.join(',')) { fail(`venue row ${r[0]}`, exp.join(' | '), got.join(' | ')); } } if (rows.length !== want.size) fail('per-venue row count', want.size, rows.length); ok(`per-venue table: ${rows.length} rows`); } // ---------------------------------------------------------- 4. Fisher's exact { // Only the POLICY rows are published; the UNION rows stay on the provenance page. const want = []; for (const l of sig.split('\n')) { const m = l.match(/^(.+?)\s+POLICY\s+(\d+)\/(\d+) \(\s*[\d.]+%\)\s+(\d+)\/(\d+) \(\s*[\d.]+%\)\s+(\S+)/); if (m) want.push({ sub: `${m[2]}/${m[3]}`, base: `${m[4]}/${m[5]}`, p: m[6] }); } if (want.length !== 5) throw new Error(`expected 5 POLICY significance rows, parsed ${want.length}`); const rows = pageRows(/\^ Property \^/); if (rows.length !== want.length) fail('significance row count', want.length, rows.length); // The page orders rows for reading; match on the subgroup fraction, not position. const bySub = new Map(want.map((w) => [w.sub, w])); for (const r of rows) { const sub = r[1].replace(/,/g, '').split(' ')[0]; if (!bySub.has(sub)) { fail(`significance row "${r[0]}"`, '(no POLICY row with that numerator/denominator)', r[1]); continue; } const w = bySub.get(sub); const base = r[2].replace(/,/g, '').split(' ')[0]; if (base !== w.base) fail(`significance base for "${r[0]}"`, w.base, base); // p is printed on the page in scientific form with superscript digits // ("2.65 x 10^-8"); the script prints "2.65e-08". Normalise BOTH to a number // and compare the value, exponent included. An earlier form of this check // compared only the first three digits and a mutation test on 2026-09-10 // showed it could not see an exponent moved from -8 to -5. const SUP = { '⁰': '0', '¹': '1', '²': '2', '³': '3', '⁴': '4', '⁵': '5', '⁶': '6', '⁷': '7', '⁸': '8', '⁹': '9', '⁻': '-' }; const pageP = (() => { const s = r[3].split('—')[0].trim(); const sci = s.match(/^([\d.]+)\s*×\s*10([⁻⁰¹²³⁴⁵⁶⁷⁸⁹]+)$/); if (sci) { const exp = [...sci[2]].map((c) => SUP[c] ?? c).join(''); return Number(sci[1]) * 10 ** Number(exp); } return Number(s); })(); const repP = Number(w.p); if (!Number.isFinite(pageP)) { fail(`significance p for "${r[0]}" is unparseable`, w.p, r[3]); } else if (Math.abs(pageP - repP) > Math.abs(repP) * 0.02) { fail(`significance p for "${r[0]}"`, `${w.p} (${repP})`, `${r[3]} (${pageP})`); } } ok(`significance table: ${rows.length} rows`); } console.log(failures ? `\n${failures} MISMATCHES` : '\nAll four tables match the report cell by cell.'); process.exitCode = failures ? 1 : 0;
The external checks — ''policies_external_checks.sh''
- policies_external_checks.sh
#!/usr/bin/env bash # Every external, non-corpus fact privacy:policies states, re-fetched. # # bash scripts/policies_external_checks.sh > scripts/policies_external_checks-output.txt 2>&1 # # Unauthenticated GitHub API: 60 requests/hour. A re-run that prints <none> # everywhere has been rate-limited, not found an empty repository — check # https://api.github.com/rate_limit before believing a negative result. # # Rules this script exists to enforce: # - print %{http_code} and %{redirect_url} before any byte count, because a # 302 recorded as an empty body has been published as "empty 200" here before; # - never ask GitHub for /releases/latest or /tags to date a repository — the # first 404s on tag-only repos and the second is unsorted. Ask the commits # API on the repository's own default branch; # - print the raw JSON field, not a summary of it. set -u API_ERRORS=0 UA='Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0 Safari/537.36' GH='https://api.github.com/repos' hdr () { printf '\n=== %s ===\n' "$1"; } probe () { # probe <label> <url> local out; out=$(curl -sL -A "$UA" -o /tmp/_ext.$$ -w '%{http_code} %{num_redirects} %{url_effective}' "$2") printf '%-46s HTTP %s redirects=%s bytes=%s\n final: %s\n' \ "$1" "${out%% *}" "$(echo "$out" | cut -d' ' -f2)" "$(wc -c < /tmp/_ext.$$)" "$(echo "$out" | cut -d' ' -f3-)" rm -f /tmp/_ext.$$ } # GitHub's unauthenticated API allows 60 requests/hour and this script makes # more than that. Use `gh api` when it is available (it supplies its own # credentials and never puts a token on a command line); fall back to curl. # `gh auth status` can exit non-zero while gh is perfectly usable (a stale # secondary account is enough), so probe the thing we actually need instead. if command -v gh >/dev/null 2>&1 && gh api rate_limit >/dev/null 2>&1; then ghapi () { gh api "$1"; } else ghapi () { curl -s -A "$UA" "https://api.github.com/$1"; } fi repo () { # repo <owner/name> local meta last meta=$(ghapi "repos/$1" | python3 -c 'import json,sys d = json.load(sys.stdin) if "full_name" not in d: print("<API error: %s>" % d.get("message", d)) else: print("%s | %s | archived=%s | pushed_at=%s" % (d["default_branch"], (d.get("license") or {}).get("spdx_id", "<no license file>"), d["archived"], d["pushed_at"]))' 2>&1) last=$(ghapi "repos/$1/commits?per_page=1" | python3 -c 'import json,sys d = json.load(sys.stdin) print(d[0]["commit"]["committer"]["date"] if isinstance(d, list) and d else "<none>")' 2>&1) printf '%-46s default|license|archived|pushed_at: %s\n%-46s last commit on default branch: %s\n' \ "$1" "$meta" "" "$last" # A rate-limited run prints an error string where a date belongs. That has been # published as provenance on this wiki before; fail the script instead. case "$meta$last" in *"API error"*|*"rate limit"*) API_ERRORS=$((API_ERRORS + 1)) ;; esac } hdr 'GitHub repositories named on the page (commits API on the default branch)' for r in citp/privacy-policy-historical citp/PrivacyPoliciesOverTime \ benandow/PrivacyPolicyAnalysis UCI-Networking-Group/PoliGraph \ AndyXiang945/PolicyChecker xiaoyue10131748/Lalaine \ dlgroupuoft/Calpric ducalpha/PurPlianceOpenSource \ SmartDataAnalytics/Polisis_Benchmark quanmou/polisis; do repo "$r"; done hdr 'pushed_at is not the default branch: per-branch last commit where they differ' # SmartDataAnalytics/Polisis_Benchmark reports pushed_at 2023-02-02 while its # master last moved in 2020. Reading pushed_at as "last commit" would call a # Dependabot security bump on a side branch evidence of maintenance. for b in $(ghapi "repos/SmartDataAnalytics/Polisis_Benchmark/branches?per_page=100" | python3 -c 'import json,sys; print(" ".join(b["name"] for b in json.load(sys.stdin)))' 2>/dev/null); do d=$(ghapi "repos/SmartDataAnalytics/Polisis_Benchmark/commits?sha=$b&per_page=1" | python3 -c 'import json,sys; d=json.load(sys.stdin); print(d[0]["commit"]["committer"]["date"] if isinstance(d,list) and d else "<none>")' 2>/dev/null) printf ' Polisis_Benchmark %-40s %s\n' "$b" "$d" done hdr 'LICENSE.txt of PurPliance (GitHub reports NOASSERTION)' curl -s -A "$UA" 'https://raw.githubusercontent.com/ducalpha/PurPlianceOpenSource/main/LICENSE.txt' | head -2 | sed 's/^/ /' hdr 'PurPliance: last three commits on the default branch' ghapi "repos/ducalpha/PurPlianceOpenSource/commits?per_page=3" | python3 -c 'import json,sys for c in json.load(sys.stdin): print(" %s %s" % (c["commit"]["committer"]["date"], c["commit"]["message"].split(chr(10))[0][:70]))' 2>/dev/null hdr 'Is each lineage repo findable by GitHub search on the tool name?' for q in PoliGraph PolicyChecker Lalaine Calpric Polisis PolicyLint; do printf '%-16s ' "$q" curl -s -A "$UA" "https://api.github.com/search/repositories?q=$q+in:name" | python3 -c 'import json,sys d = json.load(sys.stdin) items = d.get("items", []) top = items[0]["full_name"] if items else "<none>" print("total={} top={}".format(d.get("total_count"), top))' done hdr 'LICENSE files GitHub cannot auto-detect (first line of each)' for r in benandow/PrivacyPolicyAnalysis; do for f in LICENSE LICENSE.txt LICENSE.md; do u="https://raw.githubusercontent.com/$r/master/$f" code=$(curl -s -o /tmp/_lic.$$ -w '%{http_code}' -A "$UA" "$u") if [ "$code" = 200 ]; then printf '%-40s %s -> HTTP 200, first 2 non-empty lines:\n' "$r" "$f" grep -v '^[[:space:]]*$' /tmp/_lic.$$ | head -2 | sed 's/^/ /' else printf '%-40s %s -> HTTP %s\n' "$r" "$f" "$code" fi rm -f /tmp/_lic.$$ done done hdr 'Last-Modified on the W3C P3P pages (the page claims a date here)' for u in https://www.w3.org/P3P/ https://www.w3.org/TR/P3P11/; do printf '%-34s ' "$u" curl -sI -A "$UA" "$u" | tr -d '\r' | grep -i '^last-modified' || echo '<no last-modified header>' done hdr 'PyPI: is there an installable package for the lineage?' for p in poligraph poligraph-er policylint policheck privbert polisis; do code=$(curl -s -o /dev/null -w '%{http_code}' "https://pypi.org/pypi/$p/json") printf '%-24s pypi.org/pypi/%s/json -> HTTP %s\n' "$p" "$p" "$code" done hdr 'Top GitHub name-search hits — is the FIRST hit the real artefact?' python3 scripts/policies_gh_search.py hdr 'Polisis: the paper has no code repository; the demo site is claimed alive' probe 'pribot.org/polisis (the Polisis demo)' 'https://pribot.org/polisis' printf ' <title>: ' curl -sL -A "$UA" 'https://pribot.org/polisis' | grep -o -i '<title>[^<]*</title>' | head -1 hdr 'Dataset and model hosts' probe 'OPP-115 / APP-350 (usableprivacy.org/data)' 'https://usableprivacy.org/data' probe 'PrivaSeer search engine' 'https://privaseer.ist.psu.edu/' probe 'PrivaSeer data + licence page' 'https://privaseer.ist.psu.edu/data' # The page quotes the corpus size and licence from this sentence; print it, do # not summarise it. The site's own homepage advertises a DIFFERENT, smaller # number for its live search index — the two are not the same thing. curl -sL -A "$UA" 'https://privaseer.ist.psu.edu/data' | python3 -c 'import html, re, sys txt = re.sub(r"\s+", " ", html.unescape(re.sub(r"<[^>]+>", " ", sys.stdin.read()))) for pat in (r"The PrivaSeer corpus is a collection of [\d,]+ privacy policies", r"the corpus is available under a [A-Z][^.]{0,40}license"): m = re.search(pat, txt) print(" " + (m.group(0) if m else "<NOT FOUND: " + pat + ">"))' probe 'PrivBERT model card (Hugging Face)' 'https://huggingface.co/mukund/privbert' hdr 'Standards and platform policy sources' probe 'W3C P3P home' 'https://www.w3.org/P3P/' probe 'W3C P3P 1.1 (obsoleted note)' 'https://www.w3.org/TR/P3P11/' probe 'Apple: third-party SDK requirements' 'https://developer.apple.com/news/?id=pvszzano' probe 'Google Play User Data policy' 'https://support.google.com/googleplay/android-developer/answer/10144311' probe 'USENIX Sec 2023 artifact index' 'https://secartifacts.github.io/usenixsec2023/results' if [ "$API_ERRORS" -gt 0 ]; then printf '\nFAILED: %s repository lookups returned an API error rather than a date.\n' "$API_ERRORS" printf 'This output is NOT evidence of anything. Re-run with gh authenticated,\n' printf 'or wait for the unauthenticated limit to reset (api.github.com/rate_limit).\n' exit 1 fi printf '\nAll repository lookups returned a date, not an API error.\n'
The GitHub name-search check — ''policies_gh_search.py''
- policies_gh_search.py
#!/usr/bin/env python3 """Is the first GitHub name-search hit for each tool the real artefact? The page claims some of the policy-analysis lineage is hard to find. That claim is testable: search GitHub for the tool name and look at what comes back. python3 scripts/policies_gh_search.py """ import json import shutil import subprocess import urllib.request def api(path): """GitHub API. Prefer `gh` (it supplies its own credentials and raises the rate limit from 60/hour to 5,000); fall back to an unauthenticated fetch.""" if shutil.which('gh'): r = subprocess.run(['gh', 'api', path], capture_output=True, text=True) if r.returncode == 0: return json.loads(r.stdout) req = urllib.request.Request('https://api.github.com/' + path, headers={'User-Agent': 'curl'}) return json.load(urllib.request.urlopen(req)) TOOLS = ['Polisis', 'PolicyLint', 'PoliCheck', 'PoliGraph', 'PolicyChecker', 'Lalaine', 'Calpric', 'PurPliance', 'PolicyComp'] for q in TOOLS: d = api(f'search/repositories?q={q}+in:name&per_page=3') print(f'=== {q} total={d.get("total_count")}') if not d.get('items'): print(' <no repository has this name>') for i in d['items']: print(' %-46s stars=%4d pushed=%s %s' % (i['full_name'], i['stargazers_count'], i['pushed_at'], (i.get('description') or '')[:56]))
The W3C P3P check (Playwright: w3.org 403s curl) — ''policies_w3c_p3p_check.mjs''
- policies_w3c_p3p_check.mjs
// w3.org sits behind a Cloudflare challenge that answers curl with HTTP 403, // so the P3P claims on privacy:policies are checked with Playwright's own // chromium (PLAYWRIGHT_BROWSERS_PATH=/workspace/.playwright). // // node scripts/policies_w3c_p3p_check.mjs import { chromium } from 'playwright'; const URLS = ['https://www.w3.org/P3P/', 'https://www.w3.org/TR/P3P11/']; const NEEDLES = [ 'Retired 30 August 2018', 'should not be referenced in this form or implemented as-is', ]; const browser = await chromium.launch(); const page = await browser.newPage(); for (const url of URLS) { const resp = await page.goto(url, { waitUntil: 'domcontentloaded', timeout: 60000 }); const headers = resp.headers(); const body = await page.evaluate(() => document.body.innerText); console.log(`\n=== ${url}`); console.log(` final URL : ${page.url()}`); console.log(` HTTP status : ${resp.status()}`); console.log(` last-modified : ${headers['last-modified'] ?? '<none>'}`); console.log(` body chars : ${body.length}`); for (const n of NEEDLES) { console.log(` contains ${JSON.stringify(n)}: ${body.includes(n)}`); } // Any four-digit year the page states about itself, in document order. const years = [...new Set((body.match(/\b(19|20)\d\d\b/g) ?? []))].sort(); console.log(` years appearing in the visible text: ${years.join(' ')}`); const copy = body.match(/.{0,90}(Copyright|©).{0,90}/); if (copy) console.log(` copyright line : ${copy[0].replace(/\s+/g, ' ').trim()}`); } await browser.close(); // Second pass: print the sentences of w3.org/P3P/ that carry a year, so the // page's claim about the home page's own currency rests on its words. const b2 = await chromium.launch(); const pg2 = await b2.newPage(); await pg2.goto('https://www.w3.org/P3P/', { waitUntil: 'domcontentloaded', timeout: 60000 }); const txt = (await pg2.evaluate(() => document.body.innerText)).replace(/\s+/g, ' '); console.log('\n=== w3.org/P3P/ — every sentence containing a year'); for (const s of txt.split(/(?<=[.!?]) /)) if (/\b(19|20)\d\d\b/.test(s)) console.log(' ' + s.trim()); await b2.close();
The PETS author fetch — ''policies_fetch_pets_authors.py''
- policies_fetch_pets_authors.py
#!/usr/bin/env python3 """Fill out/authors.json for the PETS papers privacy:policies cites. scripts/fetch_authors.py fails on every PETS landing page as of 2026-09-09 (its parser expects a byline layout petsymposium.org no longer serves) while the pages themselves return HTTP 200 and carry the authors in `citation_author` meta tags. This reads those tags instead. It is deliberately narrow: PETS only, and it refuses to overwrite an entry that is already cached. python3 scripts/policies_fetch_pets_authors.py PETS/2020/the-privacy-... [...] """ import json import os import re import subprocess import sys UA = ('Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 ' '(KHTML, like Gecko) Chrome/131.0 Safari/537.36') IDX = '/workspace/publications_dataset/data/corpus2/.meta' OUT = 'out/authors.json' def landing(venue, year, slug): f = os.path.join(IDX, f'{venue}-{year}.json') for rec in json.load(open(f))['papers']: if rec['slug'] == slug: return rec['landingUrl'] raise KeyError(f'{venue}/{year}/{slug} not in {f}') def main(): cache = json.load(open(OUT)) if os.path.exists(OUT) else {} added = 0 for k in sys.argv[1:]: venue, year, slug = k.split('/', 2) if venue != 'PETS': raise SystemExit(f'{k}: this script is PETS-only') if k in cache: print(f'CACHED {k} {len(cache[k])} authors') continue url = landing(venue, int(year), slug) html = subprocess.run(['curl', '-sL', '-A', UA, url], capture_output=True, text=True, check=True).stdout authors = re.findall(r'name="citation_author"\s+content="([^"]+)"', html) if not authors: raise SystemExit(f'FAILED {k} {url} (no citation_author meta tags in ' f'{len(html)} bytes — do not guess, read the page)') cache[k] = authors added += 1 print(f'FETCHED {k} {len(authors)} authors: {"; ".join(authors)}') json.dump(cache, open(OUT, 'w'), indent=1, ensure_ascii=False) print(f'\n{added} added, {len(cache)} cached total') main()
This page's own generator — ''build_provenance_policies.py''
- build_provenance_policies.py
#!/usr/bin/env python3 """Assemble pages/provenance_privacy_policies.txt from the prose log + the real script sources and their real, unedited outputs. python3 scripts/build_provenance_policies.py Why a generator rather than a hand-written page: a published script must BE the committed script, and a published output must be the output that script actually produced. Pasting either by hand is how an abridged sample ends up reproducing four of nine table rows. The structural assertions at the end exist because an in-place generator can silently drop a section. """ import glob import os import re import sys PROSE = 'scripts/provenance_policies_prose.txt' OUT = 'pages/provenance_privacy_policies.txt' # (heading, path, dokuwiki <file> language) SCRIPTS = [ ('The report script', 'scripts/report_policies.mjs', 'javascript'), ('The fold', 'scripts/policy_fold.mjs', 'javascript'), ('The full-text probe', 'scripts/policies_fulltext_probe.mjs', 'javascript'), ('The quote check', 'scripts/policies_quotecheck.mjs', 'javascript'), ('The significance test', 'scripts/policies_significance.py', 'python'), ('The cell-by-cell table check', 'scripts/policies_table_check.mjs', 'javascript'), ('The external checks', 'scripts/policies_external_checks.sh', 'bash'), ('The GitHub name-search check', 'scripts/policies_gh_search.py', 'python'), ('The W3C P3P check (Playwright: w3.org 403s curl)', 'scripts/policies_w3c_p3p_check.mjs', 'javascript'), ('The PETS author fetch', 'scripts/policies_fetch_pets_authors.py', 'python'), ('This page\'s own generator', 'scripts/build_provenance_policies.py', 'python'), ] # (producing script, output path, a substring the output MUST end with). # The third field exists because an empty or truncated output file is the exact # way this generator can publish nothing and still pass: a 2026-09-10 mutation # test emptied one of these files and the page shipped an empty <file> block. OUTPUTS = [ ('report_policies.mjs', 'scripts/report_policies-output.txt', 'PROBE 2026 PETS'), ('policies_fulltext_probe.mjs', 'scripts/policies_fulltext_probe-output.txt', 'Calpric (USENIX 2023) —'), ('policies_significance.py', 'scripts/policies_significance-output.txt', 'bracketed column is the naive base rate that leaves them in.'), ('policies_table_check.mjs', 'scripts/policies_table_check-output.txt', 'All four tables match the report cell by cell.'), ('policies_quotecheck.mjs', 'scripts/policies_quotecheck-output.txt', 'located only outside paper.cols.txt.'), ('policies_external_checks.sh', 'scripts/policies_external_checks-output.txt', 'All repository lookups returned a date, not an API error.'), ('policies_w3c_p3p_check.mjs', 'scripts/policies_w3c_p3p_check-output.txt', 'Last updated'), ] def read(path): with open(path, encoding='utf-8') as f: return f.read() # A literal closing <file> tag inside a <file> block terminates it and renders # the rest of the page as markup. The sentinel is assembled from pieces so that # THIS file — which is itself published in a <file> block below — does not # contain the literal it is guarding against. CLOSE_TAG = '</' + 'file' + '>' def block(lang, name, body): if CLOSE_TAG in body: sys.exit(f'{name}: contains a literal closing file tag; it cannot go in a <file> block') return f'<file {lang} {name}>\n{body.rstrip()}\n{CLOSE_TAG}\n' parts = [read(PROSE).rstrip(), ''] parts.append('===== The scripts, as committed =====\n') parts.append('Every block below is the file itself, inserted by ' "''scripts/build_provenance_policies.py'' at build time — not a sample, not an " 'abridgement. Re-running the generator re-inserts whatever is on disk.\n') for heading, path, lang in SCRIPTS: parts.append(f"==== {heading} — ''{os.path.basename(path)}'' ====\n") parts.append(block(lang, os.path.basename(path), read(path))) parts.append('===== The outputs, unedited =====\n') for name, path, tail in OUTPUTS: parts.append(f"==== Output of ''{name}'' ====\n") parts.append(block('text', os.path.basename(path), read(path))) parts.append('====== References ======\n') parts.append('<bibtex bibliography></bibtex>\n') # --- Is every output actually an output? An empty or truncated file publishes an # empty <file> block and every other assertion still passes. for _name, _path, _tail in OUTPUTS: _body = read(_path) # 100 bytes is a floor against an emptied file, not a quality bar — one of # these outputs is legitimately six lines long. The terminal-string check # below is what catches truncation. if len(_body.strip()) < 100: sys.exit(f'{_path}: {len(_body)} bytes — that is not an output, it is a stub') if _tail not in _body: sys.exit(f'{_path}: does not contain {_tail!r} — truncated, or the script it ' f'came from changed and this expectation was not updated') if os.path.getmtime(_path) < os.path.getmtime(f'scripts/{_name}'): sys.exit(f'{_path} is older than scripts/{_name} — re-run the script before publishing') # --- Does OUTPUTS cover everything on disk? Comparing len(OUTPUTS) against a # count derived from OUTPUTS is tautological; compare it against the filesystem. _on_disk = set(glob.glob('scripts/policies_*-output.txt')) | {'scripts/report_policies-output.txt'} _listed = {o[1] for o in OUTPUTS} if _on_disk != _listed: sys.exit('OUTPUTS does not match the outputs on disk.\n' f' not published: {sorted(_on_disk - _listed)}\n' f' listed but absent: {sorted(_listed - _on_disk)}') page = '\n'.join(parts).rstrip() + '\n' # --- structural assertions: an in-place generator deletes silently. h1 = page.count('\n====== ') h2 = page.count('\n===== ') # Count only tags at the start of a line: the generator's own source is # published below and mentions the tag in prose several times. files = len([l for l in page.split('\n') if l.startswith('<file ')]) closes = len([l for l in page.split('\n') if l == CLOSE_TAG]) tables = page.count('\n^ ') assert files == len(SCRIPTS) + len(OUTPUTS), f'{files} <file> blocks, expected {len(SCRIPTS) + len(OUTPUTS)}' assert files == closes, f'{files} opening vs {closes} closing file tags' # %% inside a <file> block is literal and harmless; only the prose matters. An # odd count in the prose silently kills every heading, table and citation below it. prose_only = re.sub(r'(?ms)^<file [^>]*>.*?^' + re.escape(CLOSE_TAG) + r'$', '', page) assert prose_only.count('%%') % 2 == 0, 'odd number of %% delimiters in the prose' # The macro fires even inside ''monospace'' — only %%...%% escapes it — so the # guard looks for an UNESCAPED occurrence, not for the string. DISCUSSION = '~~' + 'DISCUSSION' + '~~' # split so this file can be published assert not re.search(r'(?<!%%)' + re.escape(DISCUSSION), prose_only), \ 'an unescaped discussion macro would put a comment box on a provenance page' # `h2 >= 10` was the first form of this check and a mutation test on 2026-09-10 # showed it asserts almost nothing: deleting the whole second half of the prose # still leaves ten headings. Name them instead. EXPECTED_H2 = [ 'The run', 'Scope and judgement calls', 'Populations, and every query behind a figure', 'The probes, at both widths', 'Folding, and the complete residue', 'Quotes and figures checked against the papers', 'Bibliography', 'Guards run before publication', 'External sources', 'What could not be established', 'Review', 'The scripts, as committed', 'The outputs, unedited', ] found = re.findall(r'(?m)^=====\s*(.+?)\s*=====$', page) missing_h2 = [h for h in EXPECTED_H2 if h not in found] assert not missing_h2, f'sections missing from the page: {missing_h2}' assert h2 == len(EXPECTED_H2), f'{h2} level-2 headings, expected {len(EXPECTED_H2)}: {found}' with open(OUT, 'w', encoding='utf-8') as f: f.write(page) print(f'wrote {OUT}') print(f' bytes {len(page):,}') print(f' level-1 headings {h1}') print(f' level-2 headings {h2}') print(f' <file> blocks {files}') print(f' table rows {tables}')
The outputs, unedited
Output of ''report_policies.mjs''
- report_policies-output.txt
============================================================================== 1. POPULATIONS ============================================================================== corpus 5859 classified classification[] non-empty 4439 75.8% of corpus POLICY classification[].target=='privacy-policy' 102 2.3% of classified PROBE title+summary regex (candidate set) 77 UNION POLICY u PROBE 123 POLICY n PROBE 56 POLICY only (enum fires, title silent) 46 PROBE only (title fires, enum silent) 21 privacy-policy tuples in POLICY 179 --- POLICY-only: the enum fires and the title probe is silent (why a title probe is not enough) 46 papers. Platform measured (enum, multi-valued): platform papers of the enum-only set -------------------- --------------------------- other-online-service 21 mobile 20 web 18 iot 4 offline 1 2017 NDSS Automated Analysis of Privacy Requirements for Mobile Apps 2017 PETS Analyzing Remote Server Locations for Personal Data Transfers in Mobile Apps 2019 IMC Tales from the Porn: A Comprehensive Privacy Analysis of the Web Porn Ecosystem. 2019 PETS MAPS: Scaling Privacy Compliance Analysis to a Million Apps 2019 WWW Understanding the Evolution of Mobile App Ecosystems: A Longitudinal Measurement Study of Google Play. 2020 PETS Angel or Devil? A Privacy Study of Mobile Parental Control Apps 2020 PETS CanaryTrap: Detecting Data Misuse by Third-Party Apps on Online Social Networks 2020 PETS The Price is (Not) Right: Comparing Privacy in Free and Paid Apps 2020 USENIX SkillExplorer: Understanding the Behavior of Skills in Large Scale 2021 NDSS Hey Alexa, is this Skill Safe?: Taking a Closer Look at the Alexa Skill Ecosystem 2022 IEEE-SP Scraping Sticky Leftovers: App User Information Left on Servers After Account Deletion. 2022 IMC Exploring the security and privacy risks of chatbots in messaging services. 2022 PETS Checking Websites’ GDPR Consent Compliance for Marketing Emails 2022 PETS Developers Say the Darnedest Things: Privacy Compliance Processes Followed by Developers of Child-Directed Apps 2022 PETS How Can and Would People Protect From Online Tracking? 2022 PETS “We may share the number of diaper changes”: A Privacy and Security Analysis of Mobile Child Care Applications 2022 USENIX SkillDetective: Automated Policy-Violation Detection of Voice Assistant Applications in the Wild 2022 WWW Et tu, Brute? Privacy Analysis of Government Websites and Mobile Apps. 2022 WWW Measuring Alexa Skill Privacy Practices across Three Years. 2023 CCS SkillScanner: Detecting Policy-Violating Voice Applications Through Static Analysis at the Development Phase. 2023 IMC Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem. 2023 NDSS CHKPLUG: Checking GDPR Compliance of WordPress Plugins via Cross-language Code Property Graph 2023 PETS Comparing Large-Scale Privacy and Security Notifications 2023 USENIX Are You Spying on Me? Large-Scale Analysis on IoT Data Exposure through Companion Apps 2023 USENIX The Digital-Safety Risks of Financial Technologies for Survivors of Intimate Partner Violence 2024 CCS A First Look at Security and Privacy Risks in the RapidAPI Ecosystem. 2024 IEEE-SP Wear's my Data? Understanding the Cross-Device Runtime Permission Model in Wearables. 2024 IEEE-SP Understanding the Privacy Practices of Political Campaigns: A Perspective from the 2020 US Election Websites. 2024 NDSS MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots 2024 PETS The Medium is the Message: How Secure Messaging Apps Leak Sensitive Data to Push Notification Services 2024 PETS Two Steps Forward and One Step Back: The Right to Opt-out of Sale under CPRA 2024 PETS Connecting the Dots: Tracing Data Endpoints in IoT Devices 2024 USENIX Arcanum: Detecting and Evaluating the Privacy Risks of Browser Extensions on Web Pages and Web Content 2025 CCS The Odyssey of robots.txt Governance: Measuring Convention Implications of Web Bots in Large Language Model Services. 2025 IEEE-SP SoK: A Privacy Framework for Security Research Using Social Media Data. 2025 IEEE-SP On the (In)Security of LLM App Stores. 2025 IMC An In-Depth Investigation of Data Collection in LLM App Ecosystems. 2025 PETS Understanding Privacy Norms through Web Forms 2025 PETS The Effect of Platform Policies on App Privacy Compliance: A Study of Child-Directed Apps 2025 PETS Privacy Settings of Third-Party Libraries in Android Apps: A Study of Facebook SDKs 2025 PETS Who’s Watching You Zoom? Investigating Privacy of Third-Party Zoom Apps 2025 PETS Surveillance Disguised as Protection: A Comparative Analysis of Sideloaded and In-Store Parental Control Apps 2025 USENIX AUTOVR: Automated UI Exploration for Detecting Sensitive Data Flow Exposures in Virtual Reality Apps 2025 USENIX I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps 2026 PETS Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law’s Impact 2026 PETS Chatbot Confessions:~Large-Scale Analysis of Private Data Disclosure in Shared AI Chatbot Conversations --- classification[].target, whole corpus, papers (enum — publishable) target papers share of 4,439 classified --------------------- ------ ------------------------- other 2594 58.4% vulnerability 883 19.9% website-category 424 9.6% user-generated-text 419 9.4% network-traffic 383 8.6% domain 351 7.9% ip-address 295 6.6% mobile-app 282 6.4% web-request 262 5.9% malware 160 3.6% privacy-policy 102 2.3% sdk-or-library 77 1.7% email-message 54 1.2% cookie 53 1.2% javascript 44 1.0% consent-notice 39 0.9% fingerprinting-script 32 0.7% website-popularity 16 0.4% dark-pattern 13 0.3% --- POLICY and UNION per year (2026 PROVISIONAL: CCS/IMC 2026 not held, IEEE S&P/WWW 2026 under-selected) Year corpus POLICY PROBE UNION POLICY share of corpus UNION share of corpus ----- ------ ------ ----- ----- ---------------------- --------------------- 2014 166 1 1 1 0.6% 0.6% 2016 182 1 2 2 0.5% 1.1% 2017 231 3 1 3 1.3% 1.3% 2018 254 2 2 2 0.8% 0.8% 2019 402 6 3 6 1.5% 1.5% 2020 404 8 5 9 2.0% 2.2% 2021 379 7 8 9 1.8% 2.4% 2022 546 17 10 19 3.1% 3.5% 2023 719 14 10 16 1.9% 2.2% 2024 690 19 16 24 2.8% 3.5% 2025* 770 16 9 20 2.1% 2.6% 2026* 415 8 10 12 1.9% 2.9% --- UNION per venue (denominator: that venue's whole corpus slice) Venue UNION POLICY venue papers POLICY share of venue UNION share of venue ------- ----- ------ ------------ --------------------- -------------------- PETS 54 43 510 8.4% 10.6% USENIX 27 25 1410 1.8% 1.9% CCS 12 8 990 0.8% 1.2% NDSS 9 8 701 1.1% 1.3% WWW 9 7 843 0.8% 1.1% IEEE-SP 7 6 767 0.8% 0.9% IMC 5 5 638 0.8% 0.8% top venue over the next, POLICY share: PETS / USENIX = 4.8x top venue over the next, UNION share : PETS / USENIX = 5.5x --- UNION by platform measured (multi-valued: an app+web paper is in two rows) platform papers share of UNION -------------------- ------ -------------- mobile 60 48.8% web 56 45.5% other-online-service 34 27.6% offline 10 8.1% iot 7 5.7% not-applicable 1 0.8% web only 39 mobile only 43 both 17 neither 24 ============================================================================== 2. HOW THE FIELD LABELS POLICY TEXT (population: POLICY, n=102) ============================================================================== classification[].method is an ENUM. Multi-valued: a paper with two privacy-policy tuples using two methods appears in two rows. --- method, all years method papers share of POLICY ------------------- ------ --------------- manual-labelling 43 42.2% heuristic-rules 33 32.4% supervised-ml 26 25.5% third-party-service 13 12.7% llm 12 11.8% regex-or-signature 9 8.8% other 8 7.8% static-analysis 4 3.9% unsupervised-ml 3 2.9% graph-analysis 2 2.0% curated-database 1 1.0% --- method by era — this is the CURRENCY table the page leans on method all 2014-2018 n=7 2019-2021 n=21 2022-2024 n=50 2025-2026* n=24 ------------------- --- ------------- -------------- -------------- --------------- manual-labelling 43 4 (57.1%) 9 (42.9%) 19 (38.0%) 11 (45.8%) heuristic-rules 33 3 (42.9%) 6 (28.6%) 20 (40.0%) 4 (16.7%) supervised-ml 26 3 (42.9%) 8 (38.1%) 12 (24.0%) 3 (12.5%) third-party-service 13 0 (0.0%) 3 (14.3%) 7 (14.0%) 3 (12.5%) llm 12 0 (0.0%) 0 (0.0%) 1 (2.0%) 11 (45.8%) regex-or-signature 9 1 (14.3%) 2 (9.5%) 3 (6.0%) 3 (12.5%) other 8 0 (0.0%) 2 (9.5%) 4 (8.0%) 2 (8.3%) static-analysis 4 0 (0.0%) 1 (4.8%) 2 (4.0%) 1 (4.2%) unsupervised-ml 3 0 (0.0%) 0 (0.0%) 3 (6.0%) 0 (0.0%) graph-analysis 2 0 (0.0%) 0 (0.0%) 2 (4.0%) 0 (0.0%) curated-database 1 0 (0.0%) 1 (4.8%) 0 (0.0%) 0 (0.0%) --- classification[].validation on privacy-policy tuples (enum; none-reported is a REAL value, not a gap in the data) validation papers share of POLICY -------------------------- ------ --------------- manual-validation 60 58.8% none-reported 40 39.2% held-out-test-set 18 17.6% comparison-to-other-method 11 10.8% not-applicable 7 6.9% cross-validation 3 2.9% BASE RATE, all 4,439 papers that classified anything: validation papers share of classified -------------------------- ------ ------------------- manual-validation 2265 51.0% none-reported 1881 42.4% not-applicable 1160 26.1% comparison-to-other-method 791 17.8% held-out-test-set 547 12.3% cross-validation 314 7.1% --- LLM as the labelling method — POLICY papers whose privacy-policy tuple has method=="llm", listed in full n=12 of 102 POLICY papers 2024 IMC Analyzing Corporate Privacy Policies using AI Chatbots. 2025 IMC An In-Depth Investigation of Data Collection in LLM App Ecosystems. 2025 PETS Privacy Settings of Third-Party Libraries in Android Apps: A Study of Facebook SDKs 2025 USENIX Evaluating Privacy Policies under Modern Privacy Laws At Scale: An LLM-Based Automated Approach 2025 PETS Automating Governing Knowledge Commons and Contextual Integrity (GKC-CI) Privacy Policy Annotations with Large Language Models 2025 PETS BehaVR: User Identification Based on VR Sensor Data 2026 PETS AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents 2026 PETS Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law’s Impact 2026 PETS Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Language Models 2026 PETS Designing Reflective Thinking-Based Contextual Privacy Policy for Mobile Applications 2026 PETS Personal Data Flows and Privacy Policy Traceability in Third-party LLM Apps in the GPT Ecosystem 2026 PETS Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale ============================================================================== 3. WHERE THE POLICIES COME FROM (population: UNION, n=123) ============================================================================== --- population[].unit (enum) unit papers ------------------ ------ mobile-apps 56 documents 45 other 30 websites 28 human-participants 27 domains 10 web-pages 7 code-repositories 7 iot-devices 4 emails 3 browser-extensions 2 network-flows 1 --- population[].sourceList, FOLDED through policy_fold.mjs (free text — a ranking, not percentages) folded source papers ------------------------------- ------ Google Play 40 custom / hand-built list 32 Alexa list 16 participant panel 14 OPP-115 14 other app ecosystem 13 Apple App Store 10 Tranco 9 Wayback Machine 4 Majestic / Umbrella / Quantcast 3 Alexa Skills Store 2 APP-350 1 Princeton Policies-over-Time 1 RESIDUE — sourceList strings matching no fold rule (138 distinct, printed in full): 4 PoliCheck Dataset 2 Amazon market 2 Google market 2 top.gg 2 Dcp 2 TheInternetBackup public domain list 2 10 mainstream VR platforms 2 IoT Inspector 2 Alexa skill marketplaces 2 third-party GPT stores and OpenAI's official GPT store 2 Zoom Marketplace 2 Federal Reserve list of the largest commercial banks 2 privacy policy corpus 2 local Craigslist and sub-Reddit forums 2 3u 2 social media, LinkedIn groups, Discord servers, university bulletin boards, and snowball sampling 1 Quantified Self community's Guide to Self-Tracking Tools 1 Twitter 1 three websites specialized in aggregating, recommending, and classifying pornographic content 1 sanitized dataset combining three pornographic-web sources 1 government documents and observed practices 1 company websites 1 APKPure 1 prior-research list compiled by crawling the web 1 AppCensus 1 final filtered dataset 1 Collaborative List of Open-Source iOS Apps and other public repositories 1 Upwork, developer websites, Reddit, and iOSoho 1 SDK vendor websites 1 Who-TracksMe 1 Disconnect Tracking Protection 1 Evidon Global Opt-out 1 DuckDuckGo Tracker Radar 1 merged tracker databases 1 SimilarWeb top websites in the US 1 Evidon Global Opt-out list 1 Discord privacy-policy pages 1 SimilarWeb 1 vendor websites and other policy-search resources 1 device privacy policies identified in the availability analysis 1 smart-home privacy policies 1 vendor websites 1 Singanamalla et al. government website dataset 1 SkillExplorer dataset from US marketplace 1 US skills store 1 PPCrawl 1 dataset released in [30] 1 Alexa skill marketplace 1 IoT companion apps collected in the wild (Dcp) 1 research papers and news reports 1 ACM Digital Library, IEEE Xplore, and USENIX Paper Proceedings 1 AlternativeTo, Top Best Alternatives, and Games Like 1 website [1] and research paper [31] 1 CCPA, GDPR, PIPEDA, and VCDPA 1 government websites 1 CONLL2012 1 Apidog 1 apideck 1 dataset used in a recent study [81] 1 obtained privacy-policy links 1 randomly selected privacy-policy links 1 randomly selected privacy policies 1 randomly selected sentences 1 Vanguard Russell 3000 ETF 1 Wagner's corpus 1 Free Company Dataset 1 UNSW IoT Analytics 1 YourThings IoTFinder 1 joinmastodon.org and instances.social 1 centralized registries 1 referral network 1 referral network and centralized registries 1 Apple Developer Documentation 1 randomly selected SDKs 1 Google search engine and crowd-knowledge platforms 1 PoliCheck test set 1 three existing GDPR privacy-policy datasets 1 manual skill-output collection 1 four academic databases 1 GPT Store 1 FlowGPT 1 Poe 1 Coze 1 Cici 1 Character.AI 1 USENIX Security, IEEE S&P, ACM CCS, and NDSS 1 AliPay daily active users 1 AliPay volunteer participants 1 AliPay consented volunteer dataset 1 Fortune 2024 top 1K US companies 1 Fortune 2024 top 500 EU companies 1 PPGDPR 1 C3PA 1 Anthropic, OpenAI, Gemini, and DeepSeek privacy policies 1 Promptfoo 1 Presidio-research 1 Apple's public sitemap 1 apps identified through the longitudinal analysis 1 email survey respondents willing to participate 1 FDIC BankFind Suite 1 university students 1 Google web search results 1 personal contacts, snowball sampling, and targeted email invitations 1 WeChat groups 1 CPP4APP 1 CA4P-483 1 MAPP Corpus 1 GPTStore.ai 1 OpenAI official store 1 BeeTrove 1 combined GPTStore.ai/OpenAI and BeeTrove dataset 1 web archive snapshots and search engine indexes 1 Domains with robots.txt files 1 Domains with English policy documents 1 Slack channels, mailing groups, and flyers 1 WearBench dataset 1 FEC 1 Ballotpedia 1 on-campus bookstore 1 random passers-by and visitors to the bookstore 1 MobiPurpose 1 Google Scholar, Semantic Scholar, and WorldWideScience 1 authors of 125 papers 1 iOS SDK Ranking 1 Apptopia 1 APKCombo 1 Princeton Privacy Crawl (PPCrawl) 1 ground truth data set employed by Cui et al. 1 local Discord server 1 professional network and OpenDP mailing list 1 enrolled participants' peer recommendations 1 FCWs dataset 1 Fraudulent and Legitimate Online Shops Dataset 1 clusters of unfavorable and benign terms 1 traffic analysis dataset [116] 1 Similarweb 1 research teams 1 research team workshop ============================================================================== 4. THE TOOL LINEAGE (population: whole corpus, so a tool used outside UNION is visible) ============================================================================== Matched against tools[].name, otherToolsMentioned[].name, classification[].resourceName/targetDetail, population[].sourceList, detection[].phenomenon/technique. Counts are PAPERS. artefact (first paper) papers, corpus of those in UNION year range of use used 2024+ -------------------------------------- -------------- ----------------- ----------------- ---------- Privee (USENIX 2014) 1 1 2014-2014 0 OPP-115 corpus (ACL 2016) 14 14 2017-2026 6 Polisis / PriBot (USENIX 2018) 13 13 2018-2025 5 PolicyLint (USENIX 2019) 18 17 2019-2025 3 MAPS (PETS 2019) 2 2 2019-2024 1 APP-350 corpus (2019) 1 1 2023-2023 0 PoliCheck (USENIX 2020) 9 9 2020-2025 4 PurPliance (2021) 5 5 2021-2023 0 PrivBERT (2021) 3 3 2024-2024 3 Calpric (USENIX 2023) 1 1 2023-2023 0 PoliGraph / PoliGraph-er (USENIX 2023) 6 6 2023-2026 5 PolicyChecker (CCS 2023) 1 1 2023-2023 0 Lalaine (USENIX 2023) 2 2 2023-2024 1 PolicyComp (USENIX 2023) 1 1 2023-2023 0 PrivaSeer 0 0 — 0 --- tools[] + otherToolsMentioned in UNION, FOLDED (free text — ranking only) folded tool family papers in UNION --------------------------- --------------- Selenium 29 spaCy 23 LLM (commercial API) 21 BERT family (non-privacy) 18 PolicyLint 17 Polisis / PriBot 12 NLTK 10 boilerplate stripper 10 language detection 10 LLM (open weights) 10 Stanford CoreNLP / AllenNLP 9 BeautifulSoup 7 PoliCheck 7 OpenWPM 6 Playwright 6 PoliGraph 5 PurPliance 4 Puppeteer 3 PrivBERT 3 readability metric 2 PolicyChecker 1 RESIDUE — 572 distinct tool names matching no fold rule. Top 40 by frequency: 11 Frida 8 VirusTotal 7 Qualtrics 6 Google Chrome 6 Firefox 6 Apktool 6 Amazon Mechanical Turk 6 mitmproxy 6 Prolific 5 Python 5 EasyList 4 scikit-learn 4 Fiddler 4 adb 4 WHOIS 4 fastText 4 Wayback Machine 4 Docker 4 FlowDroid 4 Soot 4 GloVe 4 tldextract 4 google-play-scraper 4 Google Search 4 Zoom 3 Prolific Academic 3 Monkey 3 fuzzywuzzy 3 Chrome 3 TF-IDF 3 Google Translate 3 logistic regression 3 LibRadar 3 Random Forest 3 CocoaPods 3 Chromium 3 MobSF 3 Crunchbase 3 custom web crawler 2 Porter stemmer … and 532 more, each in 1-2 papers. ============================================================================== 5. MEASURED RESULTS: POLICY AVAILABILITY BY ECOSYSTEM ============================================================================== Hand-keyed from detection[].prevalence tuples in UNION papers. The map below is INSIDE this script on purpose: the page must not carry a per-paper figure the script cannot print. Every row was read back against the paper's own quote; see report_policies_quotecheck.mjs. citekey venue ecosystem denominator (paper's own) policy availability ------------------- ----------- --------------------------------------- ------------------------------------------------------------------------------------------------------------------------------------------- ------------------------------------------------------------------------------------------------------------------------------ degeling2019_value NDSS 2019 web, EU 6,357 (the total of the paper's own availability table); the paper also states 6,759 domains in its January lists and 6,579 in its abstract 84.5% had a policy after 25 May 2018, up from 79.6% in January vallina2019_porn IMC 2019 web, adult sites 6,843 pornographic websites only 16% had an accessible privacy policy cui2025_privacy PETS 2025 web, sites with a PI-collecting form 10,143 websites 94.2% (9,559) had a privacy-policy link zimmeck2019_maps PETS 2019 Android 1,035,853 analysed apps, of 1,049,790 retrieved only 50.5% had a policy link on the Play Store page pan2024_trap USENIX 2024 Android 99,194 usable apps — the paper divides link failures by its app count 37.5% (37,150/99,194) of policy links led to an unavailable page; separately 15.7% (15,572/99,194) of apps have no link at all manandhar2022_smart USENIX 2022 smart-home vendors 596 vendors on 7 platforms 48.99% had a device-applicable policy; 10.57% had none at all lentzsch2021_alexa NDSS 2021 Alexa skills 150,708 skills across 7 country stores 36,475 (24.2%) provided a policy link yan2024_quality PETS 2024 Alexa skills 65,195 skills 21,063 of 65,195 provided a policy link zhan2024_vpvet CCS 2024 VR apps 11,923 apps on 10 VR platforms 29.5% had a findable privacy policy edu2022_exploring IMC 2022 Discord chatbots requesting permissions 15,525 unique active chatbots (the paper's own Table 2 total); its 14,852-without figure does not reconcile with it 676 (4.35%) had a policy; 14,852 (95.67%) did not wu2025_depth IMC 2025 GPT Actions Actions declaring a legal_info_url 93.96% of those policies were reachable rows: 11; every one resolves to a corpus paper. ============================================================================== 6. MEASURED RESULTS: POLICY-VERSUS-BEHAVIOUR CONSISTENCY ============================================================================== citekey venue denominator (paper's own) finding ----------------------- ----------- ---------------------------------------------- ---------------------------------------------------------------------------------------------------------------------------- zimmeck2017_automated NDSS 2017 9,050 apps with policies mean 1.83 potential inconsistencies per app libert2018_automated WWW 2018 1,807,491 identified third-party transmissions only 14.80% were disclosed in the policy andow2019_policylint USENIX 2019 11,430 policies 14.2% (1,618) contained logical contradictions; 17.7% (2,028) contradictions or narrowing definitions andow2020_actions USENIX 2020 13,796 applications / 45,603 data flows 42.4% of apps had an omitted or incorrect disclosure; 31.1% of flows were omitted; only 0.5% of flows were clearly disclosed trimananda2022_ovrseen USENIX 2022 1,135 data flows in Oculus VR apps 68% (776) inconsistent disclosures cui2023_poligraph USENIX 2023 1,566 mapped statement pairs 13.5% (211) conflicting; 25.5% (1,339/5,255) of policies define a term differently from the CCPA-based ontology xiang2023_policychecker CCS 2023 163,068 analysable policies 99.3% incomplete under GDPR; 98.1% had at least one mandatory-requirement violation xiao2023_lalaine USENIX 2023 5,102 fully tested iOS apps 3,423 (67.1%) non-compliant privacy labels samarin2023_lessons PETS 2023 69 apps with CCPA disclosures 80% (55) collected an identifier they did not disclose ali2024_honesty PETS 2024 iOS apps labelled "Data Not Collected" 97% had policy statements indicating data collection cui2025_privacy PETS 2025 websites collecting PI via web forms phi coefficient between observed collection and PoliGraph-er disclosure was < 0.20 for every PI type rows: 11; every one resolves to a corpus paper. ============================================================================== 7. WHAT ELSE THE UNION MEASURES ============================================================================== --- UNION papers that are also in the `legal` population (legal[] non-empty) 70 of 123 UNION papers assess a law; base rate 402 of 5859 (6.9%) corpus-wide. law (free text, unfolded — ranking only) papers --------------------------------------------- ------ GDPR 52 CCPA 24 COPPA 17 CalOPPA 4 HIPAA 3 General Data Protection Regulation (GDPR) 2 ePrivacy Directive 2 FTC Act 2 California Consumer Privacy Act (CCPA) 2 CPRA 2 Directive 95/46/EC 1 DOPPA 1 Data Protection Directive 95/46/EC (DPD'25.1) 1 ePrivacy directive 1 Digital Economy Act 2017 1 --- human annotation in UNION — policy labelling is a hand-coding literature UNION papers with humanAnnotation[] 120 of 123 97.6% …stating an agreement metric 45 37.5% of annotated …stating an annotator count 85 70.8% of annotated BASE RATE, all 3318 papers that coded data by hand: …stating an agreement metric 512 15.4% --- temporal[].mode in UNION (enum) vs corpus base rate — is this a longitudinal literature? UNION papers with a temporal[] tuple 118 of 123; corpus 5342 of 5859 temporal.mode UNION share of UNION w/ temporal corpus share of corpus w/ temporal ------------------ ----- -------------------------- ------ --------------------------- live-crawl 84 71.2% 1261 23.6% existing-dataset 34 28.8% 2534 47.4% active-probing 17 14.4% 1780 33.3% web-archive 12 10.2% 68 1.3% passive-collection 10 8.5% 1012 18.9% --- artifact availability in UNION vs corpus base rate (enum) availability UNION share corpus share -------------------------- ----- ----- ------ ----- public 77 63.1% 2845 51.4% none-mentioned 36 29.5% 2183 39.4% promised-not-yet-available 6 4.9% 289 5.2% on-request 1 0.8% 91 1.6% restricted 0 0.0% 79 1.4% explicitly-withheld 2 1.6% 52 0.9% ============================================================================== 7b. SIGNIFICANCE INPUTS — parsed by policies_significance.py ============================================================================== SIG|temporal.mode == web-archive|POLICY|11|101|57|5241|68|5342 SIG|temporal.mode == web-archive|UNION|12|118|56|5224|68|5342 SIG|humanAnnotation states an agreement metric|POLICY|38|101|474|3217|512|3318 SIG|humanAnnotation states an agreement metric|UNION|45|120|467|3198|512|3318 SIG|artifacts.availability == public|POLICY|63|101|2782|5438|2845|5539 SIG|artifacts.availability == public|UNION|77|122|2768|5417|2845|5539 SIG|assesses a law (legal[] non-empty)|POLICY|64|102|338|5757|402|5859 SIG|assesses a law (legal[] non-empty)|UNION|70|123|332|5736|402|5859 SIG|classification.validation == none-reported|POLICY|40|102|1823|4337|1881|4439 SIG|classification.validation == none-reported|UNION|40|102|1818|4321|1881|4439 Columns: SIG|comparison|population|sub hits|sub n|rest-of-corpus hits|rest n|whole-corpus hits|whole n "rest" excludes the subgroup itself; "whole" is the figure a naive base rate would use. ============================================================================== 8. THE UNION, IN FULL — the audit surface for every count above ============================================================================== BOTH 2014 USENIX Privee: An Architecture for Automatically Analyzing Web Privacy Policies BOTH 2016 PETS Privacy Challenges in the Quantified Self Movement – An EU Perspective PROBE 2016 PETS Crowdsourcing for Context: Regarding Privacy in Beacon Encounters via Contextual Integrity ENUM 2017 NDSS Automated Analysis of Privacy Requirements for Mobile Apps ENUM 2017 PETS Analyzing Remote Server Locations for Personal Data Transfers in Mobile Apps BOTH 2017 PETS Cross-Device Tracking: Measurement and Disclosures BOTH 2018 USENIX Polisis: Automated Analysis and Presentation of Privacy Policies Using Deep Learning BOTH 2018 WWW An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies. ENUM 2019 IMC Tales from the Porn: A Comprehensive Privacy Analysis of the Web Porn Ecosystem. BOTH 2019 NDSS we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy ENUM 2019 PETS MAPS: Scaling Privacy Compliance Analysis to a Million Apps BOTH 2019 USENIX HideMyApp: Hiding the Presence of Sensitive Apps on Android BOTH 2019 USENIX PolicyLint: Investigating Internal Privacy Policy Contradictions on Google Play ENUM 2019 WWW Understanding the Evolution of Mobile App Ecosystems: A Longitudinal Measurement Study of Google Play. BOTH 2020 PETS An Analysis of the Current State of the Consumer Credit Reporting System in China ENUM 2020 PETS Angel or Devil? A Privacy Study of Mobile Parental Control Apps ENUM 2020 PETS CanaryTrap: Detecting Data Misuse by Third-Party Apps on Online Social Networks BOTH 2020 PETS How private is your period?: A systematic analysis of menstrual app privacy policies ENUM 2020 PETS The Price is (Not) Right: Comparing Privacy in Free and Paid Apps BOTH 2020 PETS The Privacy Policy Landscape After the GDPR BOTH 2020 USENIX Actions Speak Louder than Words: Entity-Sensitive Privacy Policy and Data Flow Analysis with PoliCheck ENUM 2020 USENIX SkillExplorer: Understanding the Behavior of Skills in Large Scale PROBE 2020 WWW Finding a Choice in a Haystack: Automatic Extraction of Opt-Out Statements from Privacy Policy Text. BOTH 2021 CCS Automated Privacy Policy Annotation with Information Highlighting Made Practical Using Deep Representations. PROBE 2021 CCS Consistency Analysis of Data-Usage Purposes in Mobile Apps. ENUM 2021 NDSS Hey Alexa, is this Skill Safe?: Taking a Closer Look at the Alexa Skill Ecosystem BOTH 2021 NDSS PrivacyFlash Pro: Automating Privacy Policy Generation for Mobile Apps BOTH 2021 PETS Automated Extraction and Presentation of Data Practices in Privacy Policies PROBE 2021 PETS Defining Privacy: How Users Interpret Technical Terms in Privacy Policies BOTH 2021 USENIX Understanding Malicious Cross-library Data Harvesting on Android BOTH 2021 WWW Have You been Properly Notified? Automatic Compliance Analysis of Privacy Policy Text with GDPR Article 13. BOTH 2021 WWW Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset. BOTH 2022 CCS Do Opt-Outs Really Opt Me Out? ENUM 2022 IEEE-SP Scraping Sticky Leftovers: App User Information Left on Servers After Account Deletion. ENUM 2022 IMC Exploring the security and privacy risks of chatbots in messaging services. ENUM 2022 PETS Checking Websites’ GDPR Consent Compliance for Marketing Emails ENUM 2022 PETS Developers Say the Darnedest Things: Privacy Compliance Processes Followed by Developers of Child-Directed Apps BOTH 2022 PETS Exploring the Privacy Concerns of Bystanders in Smart Homes from the Perspectives of Both Owners and Bystanders ENUM 2022 PETS How Can and Would People Protect From Online Tracking? BOTH 2022 PETS Leave No Data Behind – Empirical Insights into Data Erasure from Online Services ENUM 2022 PETS “We may share the number of diaper changes”: A Privacy and Security Analysis of Mobile Child Care Applications BOTH 2022 PETS Who Knows I Like Jelly Beans? An Investigation Into Search Privacy PROBE 2022 PETS How Usable Are iOS App Privacy Labels? PROBE 2022 PETS Keeping Privacy Labels Honest BOTH 2022 USENIX A Large-scale Investigation into Geodifferences in Mobile Apps BOTH 2022 USENIX Electronic Monitoring Smartphone Apps: An Analysis of Risks from Technical, Human-Centered, and Legal Perspectives BOTH 2022 USENIX OVRseen: Auditing Network Traffic and Privacy Policies in Oculus VR ENUM 2022 USENIX SkillDetective: Automated Policy-Violation Detection of Voice Assistant Applications in the Wild BOTH 2022 USENIX Smart Home Privacy Policies Demystified: A Study of Availability, Content, and Coverage ENUM 2022 WWW Et tu, Brute? Privacy Analysis of Government Websites and Mobile Apps. ENUM 2022 WWW Measuring Alexa Skill Privacy Practices across Three Years. BOTH 2023 CCS PolicyChecker: Analyzing the GDPR Completeness of Mobile Apps' Privacy Policies. ENUM 2023 CCS SkillScanner: Detecting Policy-Violating Voice Applications Through Static Analysis at the Development Phase. PROBE 2023 CCS Poster: Longitudinal Measurement of the Adoption Dynamics in Apple's Privacy Label Ecosystem. BOTH 2023 IEEE-SP Detection of Inconsistencies in Privacy Practices of Browser Extensions. ENUM 2023 IMC Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem. ENUM 2023 NDSS CHKPLUG: Checking GDPR Compliance of WordPress Plugins via Cross-language Code Property Graph BOTH 2023 PETS Evolution of Composition, Readability, and Structure of Privacy Policies over Two Decades BOTH 2023 PETS Lessons in VCR Repair: Compliance of Android App Developers with the California Consumer Privacy Act (CCPA) ENUM 2023 PETS Comparing Large-Scale Privacy and Security Notifications PROBE 2023 PETS Researchers’ Experiences in Analyzing Privacy Policies: Challenges and Opportunities ENUM 2023 USENIX Are You Spying on Me? Large-Scale Analysis on IoT Data Exposure through Companion Apps ENUM 2023 USENIX The Digital-Safety Risks of Financial Technologies for Survivors of Intimate Partner Violence BOTH 2023 USENIX Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active Learning BOTH 2023 USENIX Lalaine: Measuring and Characterizing Non-Compliance of Apple Privacy Labels BOTH 2023 USENIX PoliGraph: Automated Privacy Policy Analysis using Knowledge Graphs BOTH 2023 USENIX POLICYCOMP: Counterpart Comparison of Privacy Policies Uncovers Overbroad Personal Data Collection Practices ENUM 2024 CCS A First Look at Security and Privacy Risks in the RapidAPI Ecosystem. BOTH 2024 CCS VPVet: Vetting Privacy Policies of Virtual Reality Apps. PROBE 2024 CCS Measuring Compliance Implications of Third-party Libraries' Privacy Label Disclosure Guidelines. PROBE 2024 CCS Are We Getting Well-informed? An In-depth Study of Runtime Privacy Notice Practice in Mobile Apps. ENUM 2024 IEEE-SP Wear's my Data? Understanding the Cross-Device Runtime Permission Model in Wearables. ENUM 2024 IEEE-SP Understanding the Privacy Practices of Political Campaigns: A Perspective from the 2020 US Election Websites. BOTH 2024 IMC Analyzing Corporate Privacy Policies using AI Chatbots. ENUM 2024 NDSS MASTERKEY: Automated Jailbreaking of Large Language Model Chatbots BOTH 2024 NDSS Towards Automated Regulation Analysis for Effective Privacy Compliance BOTH 2024 PETS Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps' Privacy Policies BOTH 2024 PETS On the Quality of Privacy Policy Documents of Virtual Personal Assistant Applications ENUM 2024 PETS The Medium is the Message: How Secure Messaging Apps Leak Sensitive Data to Push Notification Services ENUM 2024 PETS Two Steps Forward and One Step Back: The Right to Opt-out of Sale under CPRA BOTH 2024 PETS A Bilingual Longitudinal Analysis of Privacy Policies Measuring the Impacts of the GDPR and the CCPA/CPRA ENUM 2024 PETS Connecting the Dots: Tracing Data Endpoints in IoT Devices BOTH 2024 PETS Privacy Policies on the Fediverse: A Case Study of Mastodon Instances PROBE 2024 PETS Data Safety vs. App Privacy: Comparing the Usability of Android and iOS Privacy Labels ENUM 2024 USENIX Arcanum: Detecting and Evaluating the Privacy Risks of Browser Extensions on Web Pages and Web Content BOTH 2024 USENIX iHunter: Hunting Privacy Violations at Scale in the Software Supply Chain on iOS BOTH 2024 USENIX Swipe Left for Identity Theft: An Analysis of User Data Privacy Risks on Location-based Dating Apps BOTH 2024 USENIX Is It a Trap? A Large-scale Empirical Study And Comprehensive Assessment of Online Automated Privacy Policy Generators for Mobile Apps PROBE 2024 USENIX Abandon All Hope Ye Who Enter Here: A Dynamic, Longitudinal Investigation of Android's Data Safety Section PROBE 2024 USENIX Unpacking Privacy Labels: A Measurement and Developer Perspective on Google's Data Safety Section BOTH 2024 WWW Understanding GDPR Non-Compliance in Privacy Policies of Alexa Skills in European Marketplaces. BOTH 2025 CCS Layered, Overlapping, and Inconsistent: A Large-Scale Analysis of the Multiple Privacy Policies and Controls of U.S. Banks. ENUM 2025 CCS The Odyssey of robots.txt Governance: Measuring Convention Implications of Web Bots in Large Language Model Services. ENUM 2025 IEEE-SP SoK: A Privacy Framework for Security Research Using Social Media Data. ENUM 2025 IEEE-SP On the (In)Security of LLM App Stores. PROBE 2025 IEEE-SP Let's Get Visual - Testing Visual Analogies and Metaphors for Conveying Privacy Policies and Data Handling Information. ENUM 2025 IMC An In-Depth Investigation of Data Collection in LLM App Ecosystems. BOTH 2025 NDSS SKILLPoV: Towards Accessible and Effective Privacy Notice for Amazon Alexa Skills PROBE 2025 NDSS PolicyPulse: Precision Semantic Role Extraction for Enhanced Privacy Policy Comprehension ENUM 2025 PETS Understanding Privacy Norms through Web Forms ENUM 2025 PETS The Effect of Platform Policies on App Privacy Compliance: A Study of Child-Directed Apps ENUM 2025 PETS Privacy Settings of Third-Party Libraries in Android Apps: A Study of Facebook SDKs ENUM 2025 PETS Who’s Watching You Zoom? Investigating Privacy of Third-Party Zoom Apps BOTH 2025 PETS Automating Governing Knowledge Commons and Contextual Integrity (GKC-CI) Privacy Policy Annotations with Large Language Models BOTH 2025 PETS BehaVR: User Identification Based on VR Sensor Data ENUM 2025 PETS Surveillance Disguised as Protection: A Comparative Analysis of Sideloaded and In-Store Parental Control Apps PROBE 2025 PETS "Free WiFi is not ultimately free": Privacy Perceptions of Users in the US regarding City-wide WiFi Services ENUM 2025 USENIX AUTOVR: Automated UI Exploration for Detecting Sensitive Data Flow Exposures in Virtual Reality Apps ENUM 2025 USENIX I Can Tell Your Secrets: Inferring Privacy Attributes from Mini-app Interaction History in Super-apps BOTH 2025 USENIX Evaluating Privacy Policies under Modern Privacy Laws At Scale: An LLM-Based Automated Approach PROBE 2025 WWW Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and Conditions in Shopping Websites at Scale. BOTH 2026 PETS AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents BOTH 2026 PETS ``Because I didn't touch these and even don't know why I should to change these'': Why App Developers Do (Not) Update Apple’s Privacy Labels ENUM 2026 PETS Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law’s Impact BOTH 2026 PETS Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Language Models BOTH 2026 PETS Designing Reflective Thinking-Based Contextual Privacy Policy for Mobile Applications BOTH 2026 PETS Personal Data Flows and Privacy Policy Traceability in Third-party LLM Apps in the GPT Ecosystem BOTH 2026 PETS Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale ENUM 2026 PETS Chatbot Confessions:~Large-Scale Analysis of Private Data Disclosure in Shared AI Chatbot Conversations PROBE 2026 PETS ``We Need a Standard'': Toward an Expert–Informed Privacy Label for Differential Privacy PROBE 2026 PETS Privacy by Voice: Designing Usable Privacy Notices for the Voice Interface PROBE 2026 PETS From Lines of Code to Lines of Policy? Exploring Software Developers’ Perceptions of Their Privacy Policy–Related Activities PROBE 2026 PETS Are Bite-Size Data Safety Details a Healthy Diet for Android Telehealth App Users? Impacts of Privacy Nutrition Labels on Users’ Privacy Perceptions
Output of ''policies_fulltext_probe.mjs''
- policies_fulltext_probe-output.txt
UNION = 123 papers; full text present for 123; missing 0. Every count below is over the 123 papers with full text. probe (narrow form is what the page quotes) narrow share wide share ------------------------------------------------------------------------------------------------------------------ ------ ----- ---- ----- finds the policy by LINK TEXT / anchor keyword 12 9.8% 112 91.1% names a link-detection SEED PHRASE list 1 0.8% 2 1.6% follows the policy link and reports FAILURES (404 / dead / unreachable) 7 5.7% 64 52.0% handles a policy served as a PDF 9 7.3% 22 17.9% states the LANGUAGE of the policies it analysed <-- not a superset: different question 27 22.0% 16 13.0% analyses policies in more than one language 30 24.4% 67 54.5% strips BOILERPLATE / extracts the main content of the policy page 13 10.6% 20 16.3% measures READABILITY of the policy 17 13.8% 68 55.3% compares the policy against OBSERVED BEHAVIOUR (traffic, code, or storage) <-- not a superset: different question 80 65.0% 29 23.6% handles the policy being VERSIONED / changing under it 68 55.3% 87 70.7% reports the sentence/segment SEGMENTATION step 24 19.5% 56 45.5% says which policy applies (app vs developer vs platform vs layered) <-- not a superset: different question 27 22.0% 13 10.6% deduplicates identical / templated policies 34 27.6% 75 61.0% --- The two probes the page leans on hardest, with the matching papers named follows the policy link and reports FAILURES (404 / dead / unreachable) — 7 papers 2018 WWW An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Priva 2019 NDSS we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy 2019 USENIX PolicyLint: Investigating Internal Privacy Policy Contradictions on Google Play 2022 WWW Measuring Alexa Skill Privacy Practices across Three Years. 2024 PETS Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps' Privac 2024 PETS Privacy Policies on the Fediverse: A Case Study of Mastodon Instances 2026 PETS Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Langua analyses policies in more than one language — 30 papers 2019 NDSS we-value-your-privacy-now-take-some-cookies-measuring-the-gdprs-impact-on-web-privacy 2019 PETS MAPS: Scaling Privacy Compliance Analysis to a Million Apps 2020 PETS The Privacy Policy Landscape After the GDPR 2021 USENIX Understanding Malicious Cross-library Data Harvesting on Android 2021 CCS Consistency Analysis of Data-Usage Purposes in Mobile Apps. 2021 PETS Defining Privacy: How Users Interpret Technical Terms in Privacy Policies 2022 CCS Do Opt-Outs Really Opt Me Out? 2022 PETS Checking Websites’ GDPR Consent Compliance for Marketing Emails 2022 PETS “We may share the number of diaper changes”: A Privacy and Security Analysis of Mobile Child 2023 USENIX Are You Spying on Me? Large-Scale Analysis on IoT Data Exposure through Companion Apps 2023 USENIX POLICYCOMP: Counterpart Comparison of Privacy Policies Uncovers Overbroad Personal Data Coll 2023 PETS Researchers’ Experiences in Analyzing Privacy Policies: Challenges and Opportunities 2024 NDSS Towards Automated Regulation Analysis for Effective Privacy Compliance 2024 PETS On the Quality of Privacy Policy Documents of Virtual Personal Assistant Applications 2024 PETS A Bilingual Longitudinal Analysis of Privacy Policies Measuring the Impacts of the GDPR and 2024 USENIX Is It a Trap? A Large-scale Empirical Study And Comprehensive Assessment of Online Automated 2024 WWW Understanding GDPR Non-Compliance in Privacy Policies of Alexa Skills in European Marketplac 2024 USENIX Unpacking Privacy Labels: A Measurement and Developer Perspective on Google's Data Safety Se 2025 PETS Understanding Privacy Norms through Web Forms 2025 USENIX Evaluating Privacy Policies under Modern Privacy Laws At Scale: An LLM-Based Automated Appro 2025 PETS Surveillance Disguised as Protection: A Comparative Analysis of Sideloaded and In-Store Pare 2025 CCS The Odyssey of robots.txt Governance: Measuring Convention Implications of Web Bots in Large 2025 WWW Harmful Terms and Where to Find Them: Measuring and Modeling Unfavorable Financial Terms and 2025 IEEE-SP Let's Get Visual - Testing Visual Analogies and Metaphors for Conveying Privacy Policies and 2026 PETS AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents 2026 PETS Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law’s Impact 2026 PETS Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Langua 2026 PETS Designing Reflective Thinking-Based Contextual Privacy Policy for Mobile Applications 2026 PETS Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale 2026 PETS From Lines of Code to Lines of Policy? Exploring Software Developers’ Perceptions of Their P --- LINEAGE ARTEFACTS: extraction fold vs FULL-TEXT mention, whole corpus Full-text denominator: 5869 papers with a readable paper.cols.txt. "extraction" is the column report_policies.mjs section 4 prints. artefact full-text pattern papers whose FULL TEXT names it -------------------------------------- ------------------ ------------------------------- Privee (USENIX 2014) /\bPrivee\b/ 25 OPP-115 corpus (ACL 2016) /OPP-?115/i 30 Polisis / PriBot (USENIX 2018) /polisis|pribot/i 68 PolicyLint (USENIX 2019) /policylint/i 63 MAPS (PETS 2019) /MAPS:\s*Scaling/i 50 APP-350 corpus (2019) /APP-?350/i 9 PoliCheck (USENIX 2020) /policheck/i 63 PurPliance (2021) /purpliance/i 14 PrivBERT (2021) /privbert/i 3 Calpric (USENIX 2023) /calpric/i 3 PoliGraph / PoliGraph-er (USENIX 2023) /poligraph/i 14 PolicyChecker (CCS 2023) /policychecker/i 9 Lalaine (USENIX 2023) /lalaine/i 28 PolicyComp (USENIX 2023) /policycomp/i 10 PrivaSeer /privaseer/i 12 Full-text patterns that differ from the extraction pattern, and why: MAPS (PETS 2019): /MAPS:\s*Scaling/i — /\bMAPS\b/ also matches "Google MAPS" and the Play category "MAPS & NAVIGATION"; the paper is always cited by its title The two artefacts where the gap changes what may be said: PrivaSeer — 12 papers: 2021 PETS automated-extraction-and-presentation-of-data-practices-in-privacy-policies 2021 WWW privacy-policies-over-time-curation-and-analysis-of-a-million-document-dataset 2022 PETS setting-the-bar-low-are-websites-complying-with-the-minimum-requirements-of-the 2023 PETS researchers-experiences-in-analyzing-privacy-policies-challenges-and-opportuniti 2024 CCS vpvet-vetting-privacy-policies-of-virtual-reality-apps 2024 PETS connecting-the-dots-tracing-data-endpoints-in-iot-devices 2024 PETS honesty-is-the-best-policy-on-the-accuracy-of-apple-privacy-labels-compared-to-a 2024 PETS on-the-quality-of-privacy-policy-documents-of-virtual-personal-assistant-applica 2025 CCS layered-overlapping-and-inconsistent-a-large-scale-analysis-of-the-multiple-priv 2025 CCS the-odyssey-of-robots-txt-governance-measuring-convention-implications-of-web-bo 2025 USENIX evaluating-privacy-policies-under-modern-privacy-laws-at-scale-an-llm-based-auto 2026 PETS word-level-annotation-of-gdpr-transparency-compliance-in-privacy-policies-using Calpric (USENIX 2023) — 3 papers: 2023 USENIX calpric-inclusive-and-fine-grain-labeling-of-privacy-policies-with-crowdsourcing 2025 PETS understanding-privacy-norms-through-web-forms 2025 USENIX evaluating-privacy-policies-under-modern-privacy-laws-at-scale-an-llm-based-auto
Output of ''policies_significance.py''
- policies_significance-output.txt
parsed 10 SIG rows from scripts/report_policies-output.txt comparison pop subgroup rest of corpus p (Fisher, 2-sided) naive base -------------------------------------------------------------------------------------------------------------------------- temporal.mode == web-archive POLICY 11/101 (10.9%) 57/5241 ( 1.1%) 3.99e-08 significant [68/5342 = 1.3%] temporal.mode == web-archive UNION 12/118 (10.2%) 56/5224 ( 1.1%) 1.97e-08 significant [68/5342 = 1.3%] humanAnnotation states an agreement metric POLICY 38/101 (37.6%) 474/3217 (14.7%) 2.65e-08 significant [512/3318 = 15.4%] humanAnnotation states an agreement metric UNION 45/120 (37.5%) 467/3198 (14.6%) 1.44e-09 significant [512/3318 = 15.4%] artifacts.availability == public POLICY 63/101 (62.4%) 2782/5438 (51.2%) 0.027 significant [2845/5539 = 51.4%] artifacts.availability == public UNION 77/122 (63.1%) 2768/5417 (51.1%) 0.0101 significant [2845/5539 = 51.4%] assesses a law (legal[] non-empty) POLICY 64/102 (62.7%) 338/5757 ( 5.9%) 3.62e-50 significant [402/5859 = 6.9%] assesses a law (legal[] non-empty) UNION 70/123 (56.9%) 332/5736 ( 5.8%) 9.63e-51 significant [402/5859 = 6.9%] classification.validation == none-reported POLICY 40/102 (39.2%) 1823/4337 (42.0%) 0.613 NOT SIGNIFICANT — do not publish as a movement [1881/4439 = 42.4%] classification.validation == none-reported UNION 40/102 (39.2%) 1818/4321 (42.1%) 0.612 NOT SIGNIFICANT — do not publish as a movement [1881/4439 = 42.4%] "rest of corpus" removes the subgroup's own papers from the base; the bracketed column is the naive base rate that leaves them in.
Output of ''policies_table_check.mjs''
- policies_table_check-output.txt
OK method-by-era table: 11 rows OK per-year table: 12 rows OK per-venue table: 7 rows OK significance table: 5 rows All four tables match the report cell by cell.
Output of ''policies_quotecheck.mjs''
- policies_quotecheck-output.txt
OK cols NDSS 2019 web policy availability 84.5% / 79.6% OK cols NDSS 2019 Degeling sample 6,579 sites, 500 per member state OK cols IMC 2019 adult sites: only 16% have an accessible policy, of 6,843 OK cols PETS 2025 web forms 94.2% policy link OK cols PETS 2019 Play policy links 50.5% OK cols USENIX 2024 APPG: 37.5% of links unavailable OK cols USENIX 2024 APPG: 20.5% non-English OK cols USENIX 2024 APPG: 22.3% low quality under 2KB/200 words OK cols USENIX 2022 smart-home vendors 48.99% (292/596) OK cols NDSS 2021 Alexa skills 24.2% policy link OK cols CCS 2024 VR apps 29.5% have a policy OK cols IMC 2022 chatbots 95.67% lack a policy OK cols IMC 2019 adult sample 6,843 sites OK cols PETS 2025 web-forms denominator 10,143 OK cols PETS 2019 MAPS denominator 1,049,790 retrieved / 1,035,853 analysed OK cols USENIX 2024 APPG link denominator 37,150/99,194 OK cols USENIX 2022 smart home 10.57% no policy at all OK cols NDSS 2021 Alexa denominator 150,708 skills, 36,475 with a link OK cols PETS 2024 Alexa 21,063 of 65,195 OK cols CCS 2024 VR denominator 11,923 apps OK cols IMC 2022 Discord chatbots 676 (4.35%) have a policy OK cols WWW 2018 only 14.80% of transmissions disclosed OK cols USENIX 2019 PolicyLint 14.2% (1,618/11,430) contradictions OK cols USENIX 2020 PoliCheck 42.4% of apps OK cols USENIX 2020 PoliCheck only 0.5% of flows clearly disclosed OK cols USENIX 2022 OVRseen 68% inconsistent OK cols USENIX 2023 PoliGraph 70.6% recall / 96.9% precision OK cols USENIX 2023 PoliGraph 25.5% of policies redefine a term OK cols CCS 2023 PolicyChecker 99.3% incomplete OK cols CCS 2023 PolicyChecker 163,068 analysable of 205,973 OK cols USENIX 2023 Lalaine 3,423 of 5,102 apps OK cols PETS 2023 CCPA VCR 80% undisclosed identifier OK cols PETS 2024 Apple labels: 97% of "Data Not Collected" contradicted OK cols PETS 2025 phi < 0.20 policy-vs-form association OK cols WWW 2021 Princeton corpus 1,071,488 policies / 130,000 sites OK cols WWW 2021 median length 876 -> 1,522 words OK cols WWW 2021 FKGL 11.9 -> 13.2 OK cols WWW 2021 beacons: 25.8% of policies vs 94.6% of top-10K sites OK cols WWW 2018 84.7 minutes to read applicable policies OK cols PETS 2020 EU policies gained 35% words / 33% sentences, Global 25% / 22% OK cols CCS 2024 VR policy reuse 54.5% (1,919/3,521) OK cols PETS 2024 template reuse 65% OK cols USENIX 2018 Polisis trained on 65 of the OPP-115 policies, 50 held out OK cols USENIX 2018 Polisis average F1 0.84 OK cols USENIX 2023 Calpric 16,856 labelled segments OK cols PETS 2023 no best practices have emerged (26 interviews) OK cols PETS 2023 26 researchers interviewed OK cols PETS 2023 User Choice/Control semantic change 26% OK cols IMC 2024 GPT-4 chatbot annotation of corporate policies OK cols IMC 2025 LLM disclosure classifier 87.44% accuracy OK cols IMC 2025 only 5.8% of Actions clearly disclose OK cols USENIX 2024 APPG low-quality count 10,375/46,472 OK cols USENIX 2024 APPG non-English count 9,523/46,472 OK cols USENIX 2024 APPG 15.7% of apps provide no policy OK cols USENIX 2020 PoliCheck entity-insensitive false-consistency 37.1% OK cols USENIX 2020 PoliCheck 31.1% (14,409/45,603) omitted flows OK cols USENIX 2020 PoliCheck 14,409 omitted disclosures (the 31.1% numerator) OK cols PETS 2024 Apple labels: 228,539 apps policy-but-not-label OK cols PETS 2024 template count n=306,404 behind the 65% OK cols CCS 2023 PolicyChecker 98.1% mandatory-requirement violation OK cols NDSS 2017 Zimmeck mean 1.83 inconsistencies per app OK cols WWW 2021 Princeton-Leuven corpus reaches back to 1997 OK pypdf (NOT in .cols) WWW 2021 the length/readability series is 2009-2019 OK cols NDSS 2019 Degeling: 6357 is the total of the availability table OK cols NDSS 2019 Degeling: the same paper also says 6,759 domains OK cols PETS 2019 MAPS 50.5% is over the analysed set, not the retrieved set OK cols PETS 2019 MAPS analysed 1,035,853 of 1,049,790 retrieved OK cols PETS 2026 cory2026 is word-level GDPR transparency annotation by LLM OK cols IMC 2022 Discord: the paper states 15,525 unique active chatbots OK cols IMC 2022 Discord: and separately 14,852 (95.67%) without a policy SPECIFICITY of the 70 needles needles with no letters (pure number/punctuation): 30 of those, also present in another check paper : 4 needles under 12 characters : 22 WEAK — numeric-only and not unique to the paper they are attributed to: 1 other check papers also contain "1,071,488" (Princeton corpus 1,071,488 policies / 130,000 sites) 1 other check papers also contain "84.7" (84.7 minutes to read applicable policies) 1 other check papers also contain "54.5%" (VR policy reuse 54.5% (1,919/3,521)) 1 other check papers also contain "1,035,853" (MAPS analysed 1,035,853 of 1,049,790 retrieved) These are not wrong — each was read in context — but they are the needles a future edit could break without this check noticing. 70 needles, 70 located, 0 MISSING. 69 located in paper.cols.txt; 1 located only outside paper.cols.txt.
Output of ''policies_external_checks.sh''
- policies_external_checks-output.txt
=== GitHub repositories named on the page (commits API on the default branch) === citp/privacy-policy-historical default|license|archived|pushed_at: master | <no license file> | archived=False | pushed_at=2023-10-12T15:59:19Z last commit on default branch: 2023-10-12T15:59:18Z citp/PrivacyPoliciesOverTime default|license|archived|pushed_at: master | <no license file> | archived=False | pushed_at=2022-06-06T16:25:20Z last commit on default branch: 2022-06-06T16:25:20Z benandow/PrivacyPolicyAnalysis default|license|archived|pushed_at: master | NOASSERTION | archived=False | pushed_at=2022-10-05T15:45:48Z last commit on default branch: 2022-10-05T15:45:48Z UCI-Networking-Group/PoliGraph default|license|archived|pushed_at: master | MIT | archived=False | pushed_at=2023-06-21T23:54:12Z last commit on default branch: 2023-06-21T23:53:52Z AndyXiang945/PolicyChecker default|license|archived|pushed_at: main | <no license file> | archived=False | pushed_at=2023-11-20T22:17:26Z last commit on default branch: 2023-11-20T22:17:24Z xiaoyue10131748/Lalaine default|license|archived|pushed_at: main | MIT | archived=False | pushed_at=2023-09-25T20:15:37Z last commit on default branch: 2023-09-25T20:15:37Z dlgroupuoft/Calpric default|license|archived|pushed_at: main | <no license file> | archived=False | pushed_at=2023-06-21T10:37:11Z last commit on default branch: 2023-06-21T10:37:11Z ducalpha/PurPlianceOpenSource default|license|archived|pushed_at: main | NOASSERTION | archived=False | pushed_at=2024-03-11T00:39:10Z last commit on default branch: 2024-03-11T00:39:10Z SmartDataAnalytics/Polisis_Benchmark default|license|archived|pushed_at: master | <no license file> | archived=False | pushed_at=2023-02-02T05:13:35Z last commit on default branch: 2020-02-13T10:56:26Z quanmou/polisis default|license|archived|pushed_at: master | <no license file> | archived=False | pushed_at=2020-07-27T18:58:49Z last commit on default branch: 2020-07-27T18:46:21Z === pushed_at is not the default branch: per-branch last commit where they differ === Polisis_Benchmark dependabot/pip/bleach-3.3.0 2021-02-02T22:28:16Z Polisis_Benchmark dependabot/pip/ipython-7.16.3 2022-01-21T19:56:39Z Polisis_Benchmark dependabot/pip/jinja2-2.11.3 2021-03-20T02:55:27Z Polisis_Benchmark dependabot/pip/mistune-2.0.3 2022-07-29T23:02:03Z Polisis_Benchmark dependabot/pip/nbconvert-6.5.1 2022-08-23T18:02:27Z Polisis_Benchmark dependabot/pip/nltk-3.4.5 2020-02-13T10:44:32Z Polisis_Benchmark dependabot/pip/notebook-6.4.12 2022-06-16T23:39:21Z Polisis_Benchmark dependabot/pip/numpy-1.22.0 2022-06-22T01:09:38Z Polisis_Benchmark dependabot/pip/pillow-9.3.0 2022-11-22T03:17:29Z Polisis_Benchmark dependabot/pip/protobuf-3.18.3 2022-09-23T22:36:10Z Polisis_Benchmark dependabot/pip/pygments-2.7.4 2021-03-29T21:56:42Z Polisis_Benchmark dependabot/pip/tensorflow-2.9.3 2022-11-21T21:21:28Z Polisis_Benchmark dependabot/pip/werkzeug-0.15.5 2023-02-02T05:13:29Z Polisis_Benchmark master 2020-02-13T10:56:26Z === LICENSE.txt of PurPliance (GitHub reports NOASSERTION) === Copyright (c) 2022, the University of Michigan All rights reserved. === PurPliance: last three commits on the default branch === 2024-03-11T00:39:10Z Update README.rst 2022-12-25T07:13:07Z release privacy-statement extractor 2021-09-13T05:17:05Z Update README.md === Is each lineage repo findable by GitHub search on the tool name? === PoliGraph total=67 top=UCI-Networking-Group/PoliGraph PolicyChecker total=12 top=AndyXiang945/PolicyChecker Lalaine total=77 top=xiaoyue10131748/Lalaine Calpric total=9 top=dlgroupuoft/Calpric Polisis total=15 top=quanmou/polisis PolicyLint total=2 top=ShadowGuardAI/spea-policylint === LICENSE files GitHub cannot auto-detect (first line of each) === benandow/PrivacyPolicyAnalysis LICENSE -> HTTP 404 benandow/PrivacyPolicyAnalysis LICENSE.txt -> HTTP 200, first 2 non-empty lines: Copyright (c) 2019, North Carolina State University All rights reserved. benandow/PrivacyPolicyAnalysis LICENSE.md -> HTTP 404 === Last-Modified on the W3C P3P pages (the page claims a date here) === https://www.w3.org/P3P/ <no last-modified header> https://www.w3.org/TR/P3P11/ <no last-modified header> === PyPI: is there an installable package for the lineage? === poligraph pypi.org/pypi/poligraph/json -> HTTP 404 poligraph-er pypi.org/pypi/poligraph-er/json -> HTTP 404 policylint pypi.org/pypi/policylint/json -> HTTP 404 policheck pypi.org/pypi/policheck/json -> HTTP 404 privbert pypi.org/pypi/privbert/json -> HTTP 404 polisis pypi.org/pypi/polisis/json -> HTTP 404 === Top GitHub name-search hits — is the FIRST hit the real artefact? === === Polisis total=15 quanmou/polisis stars= 16 pushed=2020-07-27T18:58:49Z Automated Analysis of Privacy Policies SmartDataAnalytics/Polisis_Benchmark stars= 24 pushed=2023-02-02T05:13:35Z Reproducing state-of-the-art results Maxikilliane/polisis-classifiers stars= 7 pushed=2020-10-22T13:46:46Z === PolicyLint total=2 ShadowGuardAI/spea-policylint stars= 0 pushed=2025-04-19T16:27:04Z A command-line tool to lint security policies written in fatihkaplanfk/policyLint-dlp stars= 0 pushed=2026-08-12T13:05:48Z === PoliCheck total=13 UvinduBro/PoliCheck stars= 0 pushed=2026-09-02T13:13:40Z fuyukihatune-rgb/policheck stars= 0 pushed=2026-06-19T15:08:15Z surelywang/PoliCheck stars= 0 pushed=2018-11-03T17:56:10Z Political Bias Detection Tool === PoliGraph total=67 UCI-Networking-Group/PoliGraph stars= 34 pushed=2023-06-21T23:54:12Z PoliGraph: Automated Privacy Policy Analysis using Knowl ironlam/poligraph stars= 39 pushed=2026-09-10T15:43:22Z Observatoire citoyen de la transparence politique frança ben-cunningham/poligraph stars= 3 pushed=2017-08-03T05:51:04Z Source code for http://poligraph.io === PolicyChecker total=12 AndyXiang945/PolicyChecker stars= 5 pushed=2023-11-20T22:17:26Z ryanwakely/policychecker stars= 0 pushed=2014-07-15T05:57:42Z KamilW-git/PolicyChecker stars= 0 pushed=2026-06-15T21:47:15Z Aplikacja webowa do sprawdzania, czy wnioski zakupowe i === Lalaine total=77 xiaoyue10131748/Lalaine stars= 9 pushed=2023-09-25T20:15:37Z dauntl1ss/Lalaine stars= 0 pushed=2022-02-08T16:45:59Z yujishen6-design/lalaine stars= 0 pushed=2025-11-21T06:55:43Z === Calpric total=9 dlgroupuoft/Calpric stars= 2 pushed=2023-06-21T10:37:11Z hernandess31/calpriceBETA stars= 0 pushed=2025-08-13T14:40:35Z Calprice é um aplicativo em Python criado para auxiliar sysjoma/calprice stars= 0 pushed=2020-08-05T04:15:34Z Recalcular precios a la tasa del dólar === PurPliance total=1 ducalpha/PurPlianceOpenSource stars= 17 pushed=2024-03-11T00:39:10Z Source code of PurPliance analysis tool. === PolicyComp total=44 policycompass/policycompass-frontend stars= 3 pushed=2016-11-11T03:24:25Z Angular frontend for the policycompass web application policycompass/policycompass-services stars= 2 pushed=2016-11-09T13:09:05Z web services for the policy compass frontend policycompass/policycompass stars= 5 pushed=2016-10-07T13:20:55Z Policy Compass central repo === Polisis: the paper has no code repository; the demo site is claimed alive === pribot.org/polisis (the Polisis demo) HTTP 200 redirects=1 bytes=8471 final: http://pribot.org/polisis/ <title>: <title>Polisis</title> === Dataset and model hosts === OPP-115 / APP-350 (usableprivacy.org/data) HTTP 200 redirects=0 bytes=26541 final: https://usableprivacy.org/data PrivaSeer search engine HTTP 200 redirects=0 bytes=11585 final: https://privaseer.ist.psu.edu/ PrivaSeer data + licence page HTTP 200 redirects=0 bytes=11838 final: https://privaseer.ist.psu.edu/data The PrivaSeer corpus is a collection of 3,967,487 privacy policies the corpus is available under a CC BY-NC-SA license PrivBERT model card (Hugging Face) HTTP 200 redirects=0 bytes=147253 final: https://huggingface.co/mukund/privbert === Standards and platform policy sources === W3C P3P home HTTP 403 redirects=0 bytes=5583 final: https://www.w3.org/P3P/ W3C P3P 1.1 (obsoleted note) HTTP 403 redirects=0 bytes=5598 final: https://www.w3.org/TR/P3P11/ Apple: third-party SDK requirements HTTP 200 redirects=0 bytes=113972 final: https://developer.apple.com/news/?id=pvszzano Google Play User Data policy HTTP 200 redirects=0 bytes=1646973 final: https://support.google.com/googleplay/android-developer/answer/10144311 USENIX Sec 2023 artifact index HTTP 200 redirects=0 bytes=324140 final: https://secartifacts.github.io/usenixsec2023/results All repository lookups returned a date, not an API error.
Output of ''policies_w3c_p3p_check.mjs''
- policies_w3c_p3p_check-output.txt
=== https://www.w3.org/P3P/ final URL : https://www.w3.org/P3P/ HTTP status : 200 last-modified : Fri, 02 Feb 2018 14:13:43 GMT body chars : 8675 contains "Retired 30 August 2018": false contains "should not be referenced in this form or implemented as-is": false years appearing in the visible text: 2002 2007 2018 === https://www.w3.org/TR/P3P11/ final URL : https://www.w3.org/TR/P3P11/ HTTP status : 200 last-modified : Tue, 09 Oct 2018 13:16:04 GMT body chars : 369540 contains "Retired 30 August 2018": true contains "should not be referenced in this form or implemented as-is": true years appearing in the visible text: 1981 1995 1996 1997 1998 1999 2000 2001 2002 2003 2005 2006 2018 copyright line : Copyright © 2006 W3C® (MIT, ERCIM, Keio), All Rights Reserved. W3C liability, trademark and document === w3.org/P3P/ — every sentence containing a year Platform for Privacy Preferences (P3P) Project Enabling smarter Privacy Tools for the Web PLING - W3C Policy Languages Interest Group 3 October 2007: The Policy Languages Interest Group (PLING) was created. Background P3P 1.1 is a direct consequence of the first Privacy Workshop that took place 2002 in Dulles/Virginia and targets short term improvements like the User Agent Guidelines. P3PToolbox.org, with lots of complementary information P3P Validator to test the results The www-p3p-policy mailing-list to discuss issues P3P Software and Tools that may help Other P3P Documents and Notes Working Draft:A P3P Preference Exchange Language 1.0 (APPEL1.0) A P3P Assurance Signature Profile An RDF Schema for P3P 1.0 Mailing lists www-p3p-dev is a mailing list for P3P software developers www-p3p-policy is a mailing list for people who are responsible for creating P3P policies for web sites Background Resources for Developers Feedback and Discussions Papers & Presentations about P3P Critiques of P3P Selected P3P Media Coverage Historical documents and things Working Group Pages P3P Group page[Member] P3P Specification WG Homepage Charter Contact: Lorrie Cranor (Chair) & Rigo Wenning (W3C) Last updated $Date: 2018/02/02 14:13:43 $ by $Author: rigo $
References
- [1]
- Ali, Mir Masood; Balash, David G.; Kodwani, Monica; Kanich, Chris; Aviv, Adam J. (2024): "Honesty is the Best Policy: On the Accuracy of Apple Privacy Labels Compared to Apps' Privacy Policies", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [2]
- Amos, Ryan; Acar, Gunes; Lucherini, Eli; Kshirsagar, Mihir; Narayanan, Arvind; Mayer, Jonathan (2021): "Privacy Policies over Time: Curation and Analysis of a Million-Document Dataset", in: Proceedings of the Web Conference 2021, pp. 2165–2176. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
- [3]
- Adhikari, Andrick; Das, Sanchari; Dewri, Rinku (2023): "Evolution of Composition, Readability, and Structure of Privacy Policies over Two Decades", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [4]
- Cui, Hao; Trimananda, Rahmadi; Markopoulou, Athina (2025): "Understanding Privacy Norms through Web Forms", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [5]
- Pan, Shidong; Zhang, Dawen; Staples, Mark; Xing, Zhenchang; Chen, Jieshan; Xu, Xiwei; Hoang, Thong (2024): "Is It a Trap? A Large-scale Empirical Study And Comprehensive Assessment of Online Automated Privacy Policy Generators for Mobile Apps", in: Proceedings of the USENIX Security Symposium. (Link)
- [6]
- Andow, Benjamin; Mahmud, Samin Yaseer; Whitaker, Justin; Enck, William; Reaves, Bradley; Singh, Kapil; Egelman, Serge (2020): "Actions Speak Louder than Words: Entity-Sensitive Privacy Policy and Data Flow Analysis with PoliCheck", in: Proceedings of the USENIX Security Symposium. (Link)
- [7]
- Xiang, Anhao; Pei, Weiping; Yue, Chuan (2023): "PolicyChecker: Analyzing the GDPR Completeness of Mobile Apps' Privacy Policies", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [8]
- Zimmeck, Sebastian; Wang, Ziqi; Zou, Lieyong; Iyengar, Roger; Liu, Bin; Schaub, Florian; Wilson, Shomir; Sadeh, Norman; Bellovin, Steven M.; Reidenberg, Joel (2017): "Automated Analysis of Privacy Requirements for Mobile Apps", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [9]
- Khatun, Mst Eshita; Noureddine, Lamine; Bello, Sideeq; Ali-Gombe, Aisha (2026): "Disclosure Divergence: Measuring Privacy Policy and Data Safety Misalignment at Scale", Proceedings on Privacy Enhancing Technologies 2026(4):213-231. (DOI)
- [10]
- Nenadić, Luka; Rodriguez, David; Calandrino, Joseph A. (2026): "Overcoming Language Barriers: Multilingual Analysis of the 2023 Swiss Privacy Law's Impact", Proceedings on Privacy Enhancing Technologies 2026(4):703-723. (DOI)
- [11]
- Cory, Thomas; Rieder, Wolf; Krämer, Julia; Raschke, Philip; Herbke, Patrick; Küpper, Axel (2026): "Word-level Annotation of GDPR Transparency Compliance in Privacy Policies using Large Language Models", Proceedings on Privacy Enhancing Technologies 2026(1):509-528. (DOI)
- [12]
- Chanenson, Jake; Pickering, Madison; Apthorpe, Noah (2025): "Automating Governing Knowledge Commons and Contextual Integrity (GKC-CI) Privacy Policy Annotations with Large Language Models", in: Proceedings on Privacy Enhancing Technologies. (DOI)
