User Tools

Site Tools


provenance:privacy:cookies

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
provenance:privacy:cookies [2026/09/10 16:34] – Add review log (self-audit + figures-vs-script pass, with one rejection), correct custom-prefix count and citekey span arithmetic, record the bibliography cache purge. Authored by Claude karel.kubicek.claudeprovenance:privacy:cookies [2026/09/10 16:44] (current) – Replace the literal citekey span count with the invariant it was checking; a snapshot count is stale after the next edit. Authored by Claude karel.kubicek.claude
Line 169: Line 169:
 ^ Claim on the page ^ Primary source, and how it was checked ^ ^ Claim on the page ^ Primary source, and how it was checked ^
 | Cookiepedia holds **42,020,489** cookies; benchmark split 1% / 5% / 3% / 58% / **32% unknown** | ''cookiepedia.co.uk'' front page, read directly. Labelled on the page as a vendor self-report | | Cookiepedia holds **42,020,489** cookies; benchmark split 1% / 5% / 3% / 58% / **32% unknown** | ''cookiepedia.co.uk'' front page, read directly. Labelled on the page as a vendor self-report |
-| Cookiepedia rejects plain ''curl'' | ''cookiepedia.co.uk/cookies/user_id'' returns **403** to ''curl'' with a desktop User-Agent while the site root returns **200** |+| Cookiepedia rejects plain ''curl'' | ''cookiepedia.co.uk/cookies/user_id'' returns **403** to ''curl'' with a desktop User-Agent while the site root returns **200**. The currency pass reproduced this and additionally found that a headless Chromium from this datacentre address is blocked too, so the per-cookie pages could not be read at all — the ''user_id'' example on the page is now footnoted as unverified |
 | The CookieGraph Cookiepedia name table: **45,785** names, **39.3%** categorised, 35.1% ''Error'', 25.6% ''Unknown'' | Computed from the released artifact by ''scripts/cookiepedia_coverage.py'' (§7.1). Reproducible in one ''curl'' | | The CookieGraph Cookiepedia name table: **45,785** names, **39.3%** categorised, 35.1% ''Error'', 25.6% ''Unknown'' | Computed from the released artifact by ''scripts/cookiepedia_coverage.py'' (§7.1). Reproducible in one ''curl'' |
 | The CookieGraph ''cookiepedia.csv'' is **917,551 rows** of CMP-declared labels in the CookieBlock ''consent_data'' schema, not Cookiepedia classifications | Downloaded and counted; header is '',id,browser_id,visit_id,name,domain,cat_id,cat_name,purpose,expiry,type_name,type_id''. Confirmed against ''labelling_scripts/cookiepedia.py'', which merges it with ''tranco.csv'' into a ''declared_label'' | | The CookieGraph ''cookiepedia.csv'' is **917,551 rows** of CMP-declared labels in the CookieBlock ''consent_data'' schema, not Cookiepedia classifications | Downloaded and counted; header is '',id,browser_id,visit_id,name,domain,cat_id,cat_name,purpose,expiry,type_name,type_id''. Confirmed against ''labelling_scripts/cookiepedia.py'', which merges it with ''tranco.csv'' into a ''declared_label'' |
Line 366: Line 366:
 Checked before saving: the key does not exist in ''literature:bibliography'' (0 matches), the DOI does not exist (0 matches), and ''scripts/bib_dedup_scan.py'' reports no new A/B/C/D candidate pair involving it. The one pre-existing [C] pair the scan reports (''bratton2019_replication'' / ''sumner2014_exaggeration'') is unrelated to this run and untouched. Checked before saving: the key does not exist in ''literature:bibliography'' (0 matches), the DOI does not exist (0 matches), and ''scripts/bib_dedup_scan.py'' reports no new A/B/C/D candidate pair involving it. The one pre-existing [C] pair the scan reports (''bratton2019_replication'' / ''sumner2014_exaggeration'') is unrelated to this run and untouched.
  
-Every other citekey on the page already existed. **20 distinct keys, 62 markers carrying 65 citekey instances, 130 rendered ''bibtex_citekey'' spans** (the plugin emits two per instance) and **20 references**.+Every other citekey on the page already existed. **20 distinct keys and 20 references.** The check that matters is the invariantnot a snapshot: the plugin emits **two ''bibtex_citekey'' spans per citekey //instance//** — not per marker, because three markers here are multi-key — so rendered spans must equal twice the number of comma-separated keys in the source's ''{[…]}'' markers, and distinct rendered references must equal distinct source keys. Both held at every save. A literal span count is not published here because it changes with every edit and would be stale within the hour.
  
-**The new key did not render until the bibliography's own cache was purged.** Immediately after saving, ''privacy:cookies'' showed 128 spans and 19 references: ''cahn2016_cookies'' resolved to nothing, with no warning of any kind. ''?purge=true'' on ''privacy:cookies'' alone did not fix it; ''?purge=true'' on **''literature:bibliography''** did. Anyone adding a key should count rendered references against distinct source keys after purging both pages, not assume a save is enough.+**The new key did not render until the bibliography's own cache was purged.** Immediately after saving, ''privacy:cookies'' showed 19 references instead of 20: ''cahn2016_cookies'' resolved to nothing, with no warning of any kind. ''?purge=true'' on ''privacy:cookies'' alone did not fix it; ''?purge=true'' on **''literature:bibliography''** did. Anyone adding a key should count rendered references against distinct source keys after purging both pages, not assume a save is enough.
  
 ===== 11. The report script and its output ===== ===== 11. The report script and its output =====
Line 2184: Line 2184:
 ===== 12. Review log ===== ===== 12. Review log =====
  
-Four passes were planned: three focused (Sonnet) and one generic (Fable). The +Four passes were planned: three focused (Sonnet) and one generic (Fable). **All 
-figures-vs-script pass returned and was acted on. The citations-and-quotes+three focused passes returned and every finding is logged below with whether it 
-external-currency and generic passes were still running when this page was +was accepted or rejected. The generic pass was not run** — see the last 
-saved; **their findings are not yet reflected here.** That is recorded rather +subsectionwhich is the largest remaining gap in this page's review coverage. 
-than hiddentreat §5 and §7 as verified by the author onlyand re-read this + 
-section before trusting the review coverage.+Between the self-audit and the three passes**21 published figures, citations 
 +or claims were corrected**. Most are the same underlying mistake in two 
 +flavoursa number that the report script did not producehand-carried into 
 +the prose; and a number read out of a paper's table without reading the column 
 +header. Every figure of the first kind has since been moved into the script.
  
 ==== Self-audit, before any reviewer returned ==== ==== Self-audit, before any reviewer returned ====
Line 2237: Line 2241:
   * Everything else it checked — every cell of the methods, questions, categories, population, venue, year, coverage, quiet, crawl-config, consentAction and LLM tables, the 82% / 16pp / 38pp / "8 of the 10" arithmetic, sentinel handling, paper-vs-tuple counting, and the fold's throw-on-unmapped behaviour — reproduced exactly. It also independently re-verified six external sources.   * Everything else it checked — every cell of the methods, questions, categories, population, venue, year, coverage, quiet, crawl-config, consentAction and LLM tables, the 82% / 16pp / 38pp / "8 of the 10" arithmetic, sentinel handling, paper-vs-tuple counting, and the fold's throw-on-unmapped behaviour — reproduced exactly. It also independently re-verified six external sources.
  
-==== Passes 2, and 4 — not yet returned ====+==== Pass — citations and quotes (Sonnet) ==== 
 + 
 +Given both pagesthe bibliography, the report output and the papers; asked to 
 +check that every key resolves, every attributed claim is supported, and every 
 +quotation is verbatim. **Six findings accepted, four of them material.** Each 
 +was re-verified against the source before the page was changed. 
 + 
 +  * **Accepted, and the worst error on the page — {[lin2024_browsing]}'s "28.50%".** Its Table 1 columns are //Total / Mean / Median//, not Total / Percent. The row reads ''Unclassified 14619 28.50 18'', so 28.50 is a **mean per site**, not a share; the paper states no percentage. Summing the five category totals gives 18,074 cookies, of which 14,619 are unclassified — **80.9%**, not 28.5%. The page's own adjacent quote ("the vast majority of cookies are unclassified") contradicted its own figure and that did not get noticed before publication. Corrected, and the 80.9% is labelled as computed here rather than as stated by the paper. 
 +  * **Accepted — {[jiwani2024crumbling]}'s 67% named the wrong term.** 7% is the score for the original term "functional"; 67% is the score for the candidate "personalized experience". "Anonymous analytics" also scores 67%, but against "performance" (24%). Two identical numbers in adjacent paragraphs, and the page paired the wrong one. Corrected, and the paper's actual recommendations are now named. 
 +  * **Accepted after independent verification — the opening sentence's citations.** The page carried "between 80% in 2012 {[roesner2012_detecting]} and 90% in 2019 {[solomos2019_clash,sanchezrola2019can]}", a faithful reproduction of the introduction of {[bollinger2022automating]}, which says exactly that. The reviewer reported that neither primary contains its figure. **Checked before acting rather than taken on trust:** the Roesner NSDI 2012 PDF was fetched from ''usenix.org/system/files/conference/nsdi12/nsdi12-final17.pdf'' and scanned — it contains **no percentage between 70% and 99% anywhere in its text**. The sentence is rewritten around {[sanchezrola2019can]}, which does support "more than 90%", with a footnote recording where the trend line comes from and that it could not be confirmed. The two shaky citations are kept //in the footnote// rather than deleted, so the next reader can re-open the question instead of re-discovering it. 
 +  * **Accepted — the {[bollinger2022automating]} discrepancy footnote was itself wrong twice.** The 83.4% is in **§4.1 "Baseline"**, not §4.4, and the **abstract states neither figure** — it gives 84.4%, CookieBlock's own balanced accuracy, a different quantity. A footnote written to flag someone else's inconsistency contained two of its own. 
 +  * **Accepted — {[sanchezrola2021_journey]}'s population.** "137,997,677 cookies collected across 387K websites" pairs the total-cookie count with the **cookie-sharing** site count from a different section ("8.97M cookie sharing events over 387K websites"). The crawl is 6.2M pages from a 1M-domain seed. Corrected in both places. 
 +  * **Accepted, minor — the 7.2% label-noise figure** is a lower bound on **third-party** cookie labels; the page had dropped the scope. 
 +  * **Noted, no change — README quote casing.** The page quotes "its outputs should not be used with the CookieBlock extension directly"; the README has a capital I and a trailing exclamation mark. Mid-sentence quoting convention, not a misquote. 
 +  * **Noted, no change — the Chrome Web Store string.** This reviewer could not reach the store page (its fetch hit Google's consent interstitial) and confirmed the claim only via the CRX endpoint. §7.2 already treats that as the decisive evidence, and the currency pass did reach the page. 
 +  * Everything else reproduced: sixteen papers' worth of figures, the bibliography's freedom from duplicate keys, DOIs and titles, ''cahn2016_cookies'' against the corpus index, and the verbatim Chrome MV2, Privacy Sandbox and Cookiepedia quotations. 
 + 
 +**What this pass says about the run.** Three of its findings are the same 
 +mistake: **a number read out of a table without reading the column header or 
 +the caption.** The automated quote check in §5 passed all three, because the 
 +quote really is in the paper — it is the //interpretation// of the quote that 
 +was wrong, and no substring match can catch that. That is the argument for this 
 +reviewer slot existing, and it is recorded here rather than smoothed over. 
 + 
 +==== Pass 3 — external currency (Sonnet) ==== 
 + 
 +Asked to fetch, not recall, every external claim as of 2026-09-10, with the 
 +CookieBlock removal singled out for controls. 
 + 
 +  * **Accepted — "12 purposes by IAB" is stale.** Verified independently against the live policy: it now defines **Purposes 1–11, Special Purposes 1–3, Features 1–3 and Special Features 1–2**. The archive.org snapshot the page linked is from 2021 and has 10 + 2 = 12. TCF v2.2 added Purpose 11. The page now gives the current counts, links the live canonical, keeps the snapshot as the version most of the literature actually used, and says to cite the version rather than the number. 
 +  * **Accepted — the Cookiepedia ''user_id'' example cannot be re-verified.** The per-cookie pages sit behind a Cloudflare interstitial that returns 403 to ''curl'' **and** to a headless Chromium from this address. The claim is now footnoted as last-observed and unverified rather than left reading as a live fact. 
 +  * **Accepted — ''privacysandbox.com'' is retired.** It redirects to ''privacysandbox.google.com/blog/…'' with the quoted text unchanged. Footnote repointed and the redirect recorded. 
 +  * **Noted, no change here — the 2025-10-17 Privacy Sandbox retirements.** Real (Topics, Attribution Reporting, Protected Audience, Related Website Sets and six more retired; CHIPS, FedCM and Private State Tokens kept) and correctly out of scope for this page, which defers all of it to [[Privacy:Privacy sandbox]]. The reviewer's suggestion to confirm that neighbour actually carries the date is fair and is left as a TODO for whoever owns it — **not** something this run should reach across and edit. 
 +  * **Noted, no change — "There will be soon a new release based on December 2024 crawl."** The reviewer points out that "soon" is now 21 months old and no such Zenodo release is visible. That sentence is Karel's own note to readers and only he can retire it; flagged here rather than removed. 
 +  * Everything else it fetched reproduced exactly, including all three CookieBlock availability checks with their controls, the Chrome MV2 quote, the AMO and Edge listings, all four label-source sites, the Cookie-Script 404, every CookieGraph artifact count, the Classifier README quotes and the Zenodo file path. It also found the correct live URL for Mozilla's manifest-version distribution page, which §7 had recorded under a guessed slug.
  
-Citations and quotes (Sonnet), external currency (Sonnet) and the generic pass +==== Pass 4 — generic (Fable) — not run ====
-(Fable) were still running at save time. **This section is incomplete and the +
-page should be re-reviewed.** The specific things they were asked to test and +
-that therefore remain unconfirmed by a second reader:+
  
-  every ''{[key]}'' resolving, and every quoted figure appearing verbatim in the cited paper (the author checked 18 automatically and 4 by hand — §5); +The generic pass was **not run**: the three focused passes returned late in the 
-  * every external URL and vendor claim as of 2026-09-10, including the CookieBlock removal, the Chrome MV2 timeline, and the four label-source sites (the author checked these with controls — §7 — but they have not been independently re-fetched); +session and the run closed after acting on them. This page has therefore had no 
-  whatever a reader without a checklist would notice.+reader without a checklist — the pass that historically catches overstated 
 +claimsstructural problems and a page that does not answer its own question. 
 +**Treat that as the largest outstanding gap in this page's review coverage**, 
 +and run it before treating the page as settled.
  
 ===== 13. Related ===== ===== 13. Related =====
provenance/privacy/cookies.1789058053.txt.gz · Last modified: by karel.kubicek.claude