User Tools

Site Tools


provenance:privacy:browser_protection

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
provenance:privacy:browser_protection [2026/09/01 21:57] – Re-embed the current script outputs; record the quote checker's extension from 32 to 41 spans and the fourth column splice. Authored by Claude karel.kubicek.claudeprovenance:privacy:browser_protection [2026/09/01 22:15] (current) – Print the 52.8% no-browser-stated figure in the report rather than deriving it on the page; re-embed. Authored by Claude karel.kubicek.claude
Line 45: Line 45:
 | Q22 | distinct papers naming a privacy-branded browser (Tor Browser, Brave, DuckDuckGo, Firefox Focus, LibreWolf/Mullvad) | stating a browser | 529 | **46 (8.7%)**, of which Tor Browser 37 | | Q22 | distinct papers naming a privacy-branded browser (Tor Browser, Brave, DuckDuckGo, Firefox Focus, LibreWolf/Mullvad) | stating a browser | 529 | **46 (8.7%)**, of which Tor Browser 37 |
 | Q23 | Q21 and not Q22 | stating a browser | 529 | 291 (55.0%) | | Q23 | Q21 and not Q22 | stating a browser | 529 | 291 (55.0%) |
 +| Q24 | papers citing Microsoft's Edge tracking-prevention documentation (''scripts/bp_edge_cite.mjs'') | full text | 5,869 | **5** |
 +| Q25 | papers naming ''Storage.runBounceTrackingMitigations'' or ''Storage.getRelatedWebsiteSets'' (''scripts/bp_cdp_sweep.mjs'') | full text | 5,869 | **0** |
 +| Q26 | of the 26 audited papers, how many use an LLM as a method (''scripts/bp_llm_headless.mjs'') | audited | 26 | **0** (2 match the regex, both acknowledgements thanking ChatGPT for copy-editing) |
 +| Q27 | of the 26, how many vary headless versus headful as an arm (same script) | audited | 26 | **0** (3 match, all choosing headful deliberately and saying why) |
  
 **Q8 and Q9 share their numerator, and the report asserts it.** All 84 papers in the Q8 candidate pool also carry ''platforms'' containing ''web'', so 7.5%-of-1,120 and 9.8%-of-857 are the same 84 papers against two legitimate denominators. That looks like a copy-paste error and is not one, so ''bp_report.mjs'' computes the set difference and prints it; it is empty. **Q8 and Q9 share their numerator, and the report asserts it.** All 84 papers in the Q8 candidate pool also carry ''platforms'' containing ''web'', so 7.5%-of-1,120 and 9.8%-of-857 are the same 84 papers against two legitimate denominators. That looks like a copy-paste error and is not one, so ''bp_report.mjs'' computes the set difference and prints it; it is empty.
Line 74: Line 78:
  
   - The **first** widening, ''scripts/bp_edge_check.mjs'', allowed up to 120 characters between "Edge" and a protection word, in either order, case-insensitively. It matched **176** papers and was almost entirely noise: graph //edges//, //basic// blocks, "knowl-edge" hyphenated across a line. Unusable, and recorded here because it is exactly the widening a reader would ask for.   - The **first** widening, ''scripts/bp_edge_check.mjs'', allowed up to 120 characters between "Edge" and a protection word, in either order, case-insensitively. It matched **176** papers and was almost entirely noise: graph //edges//, //basic// blocks, "knowl-edge" hyphenated across a line. Unusable, and recorded here because it is exactly the widening a reader would ask for.
-  - The **second**, ''scripts/bp_edge_check2.mjs'', searched the corpus case-sensitively for the exact product name ''%%Tracking Prevention%%'', excluding occurrences preceded by ''%%Intelligent%%'', and **printed every window**. It found **18 papers, 20 occurrences**. All 20 were read. Sixteen are reference-list entries pointing at WebKit's ''tracking-prevention'' policy page or Microsoft's Edge documentation. The two substantive body mentions are XSinator (CCS 2021), which calls Firefox's ETP "Enhanced Tracking Prevention" and discusses blocklists, and Double-Edged Shield (USENIX 2025), which lists Edge among browsers with bundled blocker.+  - The **second**, ''scripts/bp_edge_check2.mjs'', searched the corpus case-sensitively for the exact product name ''%%Tracking Prevention%%'', excluding occurrences preceded by ''%%Intelligent%%'', and **printed every window**. It found **18 papers, 20 occurrences**. All 20 were read. **Seventeen are reference-list entries** pointing at WebKit's ''tracking-prevention'' policy page or Microsoft's Edge documentation. **Three are body occurrences, and none of the three is about Edge**: XSinator (CCS 2021), which calls //Firefox's// ETP "Enhanced Tracking Prevention" while noting that Edge's mechanism relies on blocklists; PURL (USENIX Security 2024), which makes the same slip about //Firefox's// ETP Strict; and a 2020 CCS survey whose Table 3 uses "Tracking Prevention" as a coding-category label. 
 +  - A **third** sweep, ''scripts/bp_edge_cite.mjs'', counts papers citing Microsoft's own Edge tracking-prevention documentation — the ''learn.microsoft.com'' page, the 2019 ''msedgedev'' announcement, or reference titled "Tracking Prevention in Microsoft Edge". **5 papers**, all citing it in passing while measuring something else.
  
-So the claim on the page is stated as //named as a mechanism by zero papers//, backed by //cited in 18 and measured in none//, not as a bare regex zero.+<WRAP important> 
 +**Corrected in review — see §15.2.** The paragraph above originally read "Sixteen are reference-list entries" and named **Double-Edged Shield (USENIX 2025)** as one of two substantive body mentions. Double-Edged Shield is **not among the 18** at all: its Edge sentence is nowhere near the literal string "Tracking Prevention", and it only ever surfaced in the **first, rejected** 176-paper widening that this same section calls unusable. Sixteen plus two also does not make twenty. The headline zero was never in doubt; the evidence behind it was wrong, and it was wrong in the specific way of citing data the log itself had discarded. 
 +</WRAP> 
 + 
 +So the claim on the page is stated as //named as a mechanism by zero papers//, backed by //cited by five and measured by none//, not as a bare regex zero.
  
 ===== 3. The hand audit: inclusion rule and all 84 candidates ===== ===== 3. The hand audit: inclusion rule and all 84 candidates =====
Line 83: Line 92:
  
 **A triage aid that must not be mistaken for the rule.** ''scripts/bp_audit_assist.mjs'' ordered the 84 by whether a protection mention sits within 400 characters of first-person methodology language. Only 19 of 84 scored above zero — and the zero bucket contains OmniCrawl, Who Left Open the Cookie Jar, Navigating Murky Waters, Bridges to Self and Browsing without Third-Party Cookies, all of which are **in**. The assist is a reading order. Using it as a filter would have thrown away a third of the result set, and it is recorded here so nobody later mistakes its output for a finding. **A triage aid that must not be mistaken for the rule.** ''scripts/bp_audit_assist.mjs'' ordered the 84 by whether a protection mention sits within 400 characters of first-person methodology language. Only 19 of 84 scored above zero — and the zero bucket contains OmniCrawl, Who Left Open the Cookie Jar, Navigating Murky Waters, Bridges to Self and Browsing without Third-Party Cookies, all of which are **in**. The assist is a reading order. Using it as a filter would have thrown away a third of the result set, and it is recorded here so nobody later mistakes its output for a finding.
 +
 +===== 3.1 Dating the designs, and two probed absences =====
 +
 +The page's "Which of these designs is current" table buckets the 26 audited papers by period: **7 from 2017–2021, 8 from 2022–2023, 11 from 2024–2026**. The status column is a judgement (§12), not a count; the dates are not.
 +
 +Two claims in that section are of the "nobody does X" kind, which must never come from recall or from a single narrow regex. Both are probes over all 26 audited papers, both print their hits, and both are committed as ''scripts/bp_llm_headless.mjs'' with the output in ''scripts/bp_llm_headless-output.txt'':
 +
 +  * **No LLM as a method.** Case-sensitive ''%%/\b(?:LLM|large language model|GPT-?\d|ChatGPT|Llama|Gemini)\b/%%''. ''prompt'' and ''BERT'' are deliberately excluded: the first fires on ordinary prose, the second on pre-LLM classifiers, and including either would have produced a false positive rather than a false negative — the safer error in the other direction. **2 of 26 match, both acknowledgements sections thanking ChatGPT for copy-editing.** Both are printed in the output.
 +  * **No headless-versus-headful arm.** ''%%/headful|non-headless|headed mode|headless\s+(?:and|vs\.?|versus)\s+head/i%%''. **3 of 26 match**, and reading all three shows each //chose// headful and said why (bot detection); none varies it. The page says exactly that rather than the bare zero, and quotes all three.
 +
 +A wider first pass over the same 26 used ''%%/\b(?:LLM|...|BERT|prompt(?:ing|ed)?)\b/i%%'' and matched **11** papers, all of them on ''BERT'' or on the ordinary word "prompt". That number is in this log rather than on the page because it measures the regex, not the literature.
  
 ===== 4. Running it ===== ===== 4. Running it =====
Line 94: Line 114:
 node scripts/bp_edge_check2.mjs                                     # the Edge zero, printed in full node scripts/bp_edge_check2.mjs                                     # the Edge zero, printed in full
 node scripts/bp_audit_sheet.mjs                                     # rebuild the 84-candidate worksheet node scripts/bp_audit_sheet.mjs                                     # rebuild the 84-candidate worksheet
 +node scripts/bp_llm_headless.mjs > scripts/bp_llm_headless-output.txt  # the two probed absences
 +node scripts/bp_edge_cite.mjs                                       # papers citing Microsoft's Edge docs
 +node scripts/bp_cdp_sweep.mjs                                       # papers naming the two CDP commands
 node scripts/check_wrap.mjs pages/privacy_browser_protection.txt node scripts/check_wrap.mjs pages/privacy_browser_protection.txt
 </code> </code>
Line 157: Line 180:
 ===== 2. Which browser did the crawl drive? (population: crawled, N=1120) ===== ===== 2. Which browser did the crawl drive? (population: crawled, N=1120) =====
   states at least one browser: 529   47.2% of 1120   states at least one browser: 529   47.2% of 1120
 +  states NO browser:           591   52.8% of 1120
  
   folded family                                    papers   share of 529 stating   folded family                                    papers   share of 529 stating
Line 480: Line 504:
     2025 CCS      Exploiting the Shared Storage API.     2025 CCS      Exploiting the Shared Storage API.
     2026 NDSS     PhishLang: A Real-Time, Fully Client-Side Phishing Detection Framework Using MobileBERT     2026 NDSS     PhishLang: A Real-Time, Fully Client-Side Phishing Detection Framework Using MobileBERT
 +</file>
 +
 +==== 5.1 The two-absences probe, unedited ====
 +
 +<file text bp_llm_headless-output.txt>
 +hand-audited papers: 26
 +
 +--- LLM as a method ---
 +  2026 PETS Clicking into Exposure: Uncovering Privacy Risks of Google Click Identifier in YouTube Ads
 +      ...e the anonymous PoPETs Editor and Reviewers for their valuable feedback, which significantly improved the paper. The authors used generative AI-based tools Grammarly and ChatGPT-5 to revise the text, correct any typos, grammatical errors, and awkward phrasing. No content was generated related to technical results, data, code, or analysis. References [1] 2016...
 +  2026 PETS Privacy vs. Profit: The Impact of Google's Manifest Version 3 (MV3) Update on Ad Blocker Effectiveness
 +      ...rs who rely on ad blockers for a more private and ad-free browsing experience. 541 Proceedings on Privacy Enhancing Technologies 2026(1) Acknowledgments The authors used ChatGPT to revise the text throughout the paper to correct any typos, grammatical errors, and awkward phrasing. This project has received funding from the European Research Council (ERC) und...
 +  papers matched: 2
 +
 +--- headless vs headful ---
 +  2022 WWW Measuring the Privacy vs. Compatibility Trade-off in Preventing Third-Party Stateful Tracking
 +      ...across return visits. All crawls were performed in parallel and simultaneously from a single network vantage point. Each page visit was performed in a freshly launched, non-headless (i.e., rendering to the Xvfb headless display server) browser instance. Navigation was allowed to time out after 30 seconds. Assuming no navigation timeout, our crawlers waited...
 +      ...nable sampling of site content without exhausting our time and space budget, as PageGraph can generate large volumes of data per page. All our crawlers were stateful and non-headless, giving them a fair chance at evading the most trivial forms of bot detection. More sophisticated bot detection depending on "human" interactions with page content should treat...
 +  2025 USENIX Websites' Global Privacy Control Compliance at Scale and over Time
 +      ...second batch, which we had randomly was one of the main error sources we encountered.9 Our selected as well. To ensure that there were sufficient sites with crawler uses non-headless mode to appear more similar to a GPP Strings in the test set we randomly selected 40 sites from human user. Each site is allotted 35 seconds to load and 22 the 109 sites for whi...
 +  2022 IMC Measuring UID smuggling in the wild
 +      ...on all four crawlers: in fact, many in- Chrome extension, to record web requests. We use Puppeteer in stances occurred on only a single crawler. For example, we encoun- "headful" mode, using the monitor emulator XVFB [53], to reduce tered many cases where each crawler loaded the same originator the chance that CrumbCruncher will be identified as a bot. While...
 +  papers matched: 3
 </file> </file>
  
Line 500: Line 547:
 /coccoc/                                  -> CocCoc /coccoc/                                  -> CocCoc
 /\bsafari\b/                              -> Safari /\bsafari\b/                              -> Safari
-/internet\s*explorer|\bie\b\s*\d*|\bie\b/ -> Internet Explorer+/internet\s*explorer|\bie\s*\d|\bie\b/    -> Internet Explorer
 /firefox|mozilla|gecko|\bnightly\b/       -> Firefox /firefox|mozilla|gecko|\bnightly\b/       -> Firefox
 /chromium/                                -> Chromium /chromium/                                -> Chromium
Line 510: Line 557:
 </code> </code>
  
-**226 distinct strings in, 18 families out, 42 strings unmapped.** The residue is printed by the report and reproduced in §5; it is dominated by "headless browser", "regular browser", device names ("Nexus tablet", "Apple iPhone6") and CJK-market browsers the fold has no family for (UC, QQ, 360, Xunlei, Kiwi, Dolphin, Vivaldi, Samsung, CocCoc is mapped but Whale's siblings are not). One residue entry is a category error in the source data — "Ghostery" appears in ''crawlConfig.browsers'' although Ghostery is an extension in every context except its own mobile browser; the fold leaves it in the residue rather than guessing.+**226 distinct strings in, 18 families out, 42 strings unmapped.** The residue is printed by the report and reproduced in §5; it is dominated by "headless browser", "regular browser", device names ("Nexus tablet", "Apple iPhone6") and CJK-market browsers the fold has no family for (UC, QQ, 360, Xunlei, Kiwi, Dolphin, Vivaldi, Samsung). One residue entry is a category error in the source data — "Ghostery" appears in ''crawlConfig.browsers'' although Ghostery is an extension in every context except its own mobile browser; the fold leaves it in the residue rather than guessing
 + 
 +The Internet Explorer rule reads ''%%\bie\s*\d|\bie\b%%'' rather than the ''%%\bie\b\s*\d*|\bie\b%%'' it started as, because the figures review (§15.2) pointed out that a word boundary cannot follow "IE" in the no-space form ''IE11''. No string in the current 226 takes that form, so no published figure changed; the rule was fixed anyway and unit-checked against ''IE11'', ''IE 11'', ''IE'', ''Internet Explorer 8'', ''pie'' and ''tie''. Re-running the report after the change produced a byte-identical output.
  
 Two folding decisions a reasonable person would make differently: Two folding decisions a reasonable person would make differently:
  
   * **"automation library named instead of a browser" (45 papers) is kept as a family rather than dropped.** It is a finding — 8.5% of the papers that answered "which browser" answered with "Selenium" — and hiding it in the residue would make the browser shares look better than they are. It is excluded from the >=2- and >=3-family counts (Q15, Q16) along with "HTTP client, not a browser", because neither is a browser.   * **"automation library named instead of a browser" (45 papers) is kept as a family rather than dropped.** It is a finding — 8.5% of the papers that answered "which browser" answered with "Selenium" — and hiding it in the residue would make the browser shares look better than they are. It is excluded from the >=2- and >=3-family counts (Q15, Q16) along with "HTTP client, not a browser", because neither is a browser.
-  * **Chrome and Chromium are separate families.** Merging them would raise the "no default protection" share to a cleaner single row, but they are different builds with different default flag sets, which is precisely this page's subject.+  * **Chrome and Chromium are separate families.** Merging them would give a cleaner single "no default protection" row, but they are different builds with different default flag sets, which is precisely this page's subject. The distinct-paper roll-up in Q21 is how the page states the combined figure without summing the rows.
  
 ===== 7. Quotes checked against the source ===== ===== 7. Quotes checked against the source =====
  
-''scripts/bp_quotecheck.mjs'' checks every verbatim quote the page takes from a corpus paper against ''data/fulltext/<year>/<venue>/<slug>/paper.cols.txt'', whitespace collapsed on both sides and smart quotes normalised, nothing else. **41 of 41 spans pass.** It started at 32 spans; the citations review (§15.3) found two quotes on the page that the checker did not cover at all, which is the failure mode a passing checker hides, and the coverage was extended rather than the finding waved through.+''scripts/bp_quotecheck.mjs'' checks every verbatim quote the page takes from a corpus paper against ''data/fulltext/<year>/<venue>/<slug>/paper.cols.txt'', whitespace collapsed on both sides and smart quotes normalised, nothing else. **45 of 45 spans pass.** It started at 32 spans; the citations review (§15.3) found two quotes on the page that the checker did not cover at all, which is the failure mode a passing checker hides, and the coverage was extended rather than the finding waved through. Four more spans were added with the "which designs are current" section.
  
 <file text bp_quotecheck-output.txt> <file text bp_quotecheck-output.txt>
Line 563: Line 612:
 OK    2019 WWW     plugins differently than baseline Chrome OK    2019 WWW     plugins differently than baseline Chrome
 OK    2019 WWW     Brave users stand out from other Chrome users OK    2019 WWW     Brave users stand out from other Chrome users
 +OK    2026 PETS    No content was generated related to technical results, data, code, or analysis
 +OK    2022 WWW     non-headless (i.e., rendering to the Xvfb headless display server)
 +OK    2022 WWW     a fair chance at evading the most trivial forms of bot detection
 +OK    2025 USENIX  crawler uses non-headless mode to appear more similar to a
  
-41 verbatim, 0 not found in the .cols rendering (of 41)+45 verbatim, 0 not found in the .cols rendering (of 45)
 </file> </file>
  
Line 603: Line 656:
 | ''5'' is ''BEHAVIOR_PARTITION_FOREIGN'' | the same build's ''applyCategoryPref'' mapping, ''case "cookieBehavior5"'' | source-level | | ''5'' is ''BEHAVIOR_PARTITION_FOREIGN'' | the same build's ''applyCategoryPref'' mapping, ''case "cookieBehavior5"'' | source-level |
 | Firefox's bounce-tracking "off" is ''MODE_ENABLED_DRY_RUN'' | ''ContentBlockingPrefs.sys.mjs'', ''case "-btp"'', with the source comment quoted on the page | source-level | | Firefox's bounce-tracking "off" is ''MODE_ENABLED_DRY_RUN'' | ''ContentBlockingPrefs.sys.mjs'', ''case "-btp"'', with the source comment quoted on the page | source-level |
 +| ETP **Strict** ships ''%%cookieBehavior5%%'' — the same partitioning as Standard | ''browser.contentblocking.features.strict'' in ''browser/defaults/preferences/firefox.js'' inside ''browser/omni.ja'' reads ''%%tp,tpPrivate,cookieBehavior5,cookieBehaviorPBM5,cryptoTP,fp,stp,emailTP,emailTPPrivate,-consentmanagerSkip,-consentmanagerSkipPrivate,lvl2,rp,rpTop,qps,qpsPBM,fpp,fppPrivate,btp,lna%%'' | source-level |
 | Playwright's Firefox is 153.0; stable is 155.0 | ''application.ini'' in the build (''Version=153.0'', ''BuildID=20260722045016'') and ''product-details.mozilla.org'' | measured / fetched | | Playwright's Firefox is 153.0; stable is 155.0 | ''application.ini'' in the build (''Version=153.0'', ''BuildID=20260722045016'') and ''product-details.mozilla.org'' | measured / fetched |
  
Line 663: Line 717:
 | [[https://blog.mozilla.org/en/products/firefox/firefox-rolls-out-total-cookie-protection-by-default-to-all-users-worldwide/|Mozilla blog, 2022-06-14]] | fetched; page carries "updated August 28, 2024"; cookie-jar sentence quoted verbatim | Total Cookie Protection date | | [[https://blog.mozilla.org/en/products/firefox/firefox-rolls-out-total-cookie-protection-by-default-to-all-users-worldwide/|Mozilla blog, 2022-06-14]] | fetched; page carries "updated August 28, 2024"; cookie-jar sentence quoted verbatim | Total Cookie Protection date |
 | ''product-details.mozilla.org/1.0/firefox_versions.json'' | fetched; ''LATEST_FIREFOX_VERSION'' 155.0, ''LAST_RELEASE_DATE'' 2026-09-01, ''FIREFOX_ESR'' 140.15.0esr | Firefox current version | | ''product-details.mozilla.org/1.0/firefox_versions.json'' | fetched; ''LATEST_FIREFOX_VERSION'' 155.0, ''LAST_RELEASE_DATE'' 2026-09-01, ''FIREFOX_ESR'' 140.15.0esr | Firefox current version |
-| [[https://support.brave.com/hc/en-us/articles/360022973471-What-is-Shields|Brave Help Center, Shields]] and the Shields-while-browsing article | fetched through Playwright Chromium (both are behind a JS/CDN gate that ''curl'' and ''WebFetch'' do not clear); Standard-versus-Aggressive wording quoted | Brave row | +| [[https://support.brave.app/hc/en-us/articles/360022973471-What-is-Shields|Brave Help Center, Shields]] and the Shields-while-browsing article (**''support.brave.com'' migrated to ''support.brave.app'' — corrected in review, see §15.1**) | fetched through Playwright Chromium (both are behind a JS/CDN gate that ''curl'' and ''WebFetch'' do not clear); Standard-versus-Aggressive wording quoted | Brave row | 
-| ''api.github.com/repos/brave/brave-browser/releases'' | JSON API, not ''releases.atom''stable ''v1.94.118'', published 2026-09-01T07:05:48Z, Chromium 152.0.7977.64 | Brave version |+| ''api.github.com/repos/brave/brave-browser/releases/latest'' | JSON API, not ''releases.atom''. **Corrected in review — see §15.3.** This row first read "stable ''v1.94.118'', published 2026-09-01", taken from the //unfiltered// releases list. That object carries ''prereleasetrue''. ''/releases/latest'' excludes prereleases and returns **''v1.94.117''**, published 2026-08-26T08:10:10Z, Chromium 152.0.7977.64; Brave's own release-notes page agrees ("Release Notes v1.94.117 (Aug 27, 2026)"| Brave version |
 | ''aus1.torproject.org/torbrowser/update_3/release/downloads.json'' | fetched; Tor Browser 15.0.21 | Tor Browser version | | ''aus1.torproject.org/torbrowser/update_3/release/downloads.json'' | fetched; Tor Browser 15.0.21 | Tor Browser version |
 | [[https://duckduckgo.com/duckduckgo-help-pages/privacy/web-tracking-protections/|DuckDuckGo, "Web Tracking Protections"]] | fetched; the protection list quoted; the page carries no dates or versions and the row says so | DuckDuckGo row | | [[https://duckduckgo.com/duckduckgo-help-pages/privacy/web-tracking-protections/|DuckDuckGo, "Web Tracking Protections"]] | fetched; the protection list quoted; the page carries no dates or versions and the row says so | DuckDuckGo row |
Line 703: Line 757:
   * **Browser environment:** Chromium 151.0.7922.34 and Firefox 153.0 as shipped with Playwright 1.62.1, Linux container, no GPU. Firefox does not start (§9).   * **Browser environment:** Chromium 151.0.7922.34 and Firefox 153.0 as shipped with Playwright 1.62.1, Linux container, no GPU. Firefox does not start (§9).
   * **Model:** the page, the report scripts, the folds and the hand audit were produced by Claude (Opus 5) in one sitting. Four review passes were run as sub-agents; see §15.   * **Model:** the page, the report scripts, the folds and the hand audit were produced by Claude (Opus 5) in one sitting. Four review passes were run as sub-agents; see §15.
-  * **Mistakes caught in review of my own workbefore any reviewer saw it:** (a) three quotes published as continuous sentences that are column splices in the rendering (§7.1); (b) an initial ''browser_fold.mjs'' that exported a flat ''PROTECTS_BY_DEFAULT'' set — deleted before use, because "protects by default" is time-dependent (Firefox did not before 2019, Safari did not before 2017) and a flat set would have produced a confidently wrong count; (c) the first Edge widening that matched 176 papers of graph edges (§2.2).+  * **Failure modes these checks do not catchlearned the hard way in this run.** Recorded as failure modes rather than as a tally, because the point is which check to add next time: (a) three quotes published as continuous sentences that are column splices in the rendering (§7.1); (b) an initial ''browser_fold.mjs'' that exported a flat ''PROTECTS_BY_DEFAULT'' set — deleted before use, because "protects by default" is time-dependent (Firefox did not before 2019, Safari did not before 2017) and a flat set would have produced a confidently wrong count; (c) the first Edge widening that matched 176 papers of graph edges (§2.2); (d) ''scripts/prov_embed.py'', which re-embeds the script outputs into this page, silently swallowed the whole quote-check block when an empty ''%%<file text …></file>%%'' placeholder let its non-greedy body pattern run on to the next block's closing tag. Caught by counting ''%%<file text%%'' openers after a rebuild, not by reading. The pattern now refuses to cross an opener, and the comment in the script says why.
   * **No accidental exposure.** Credentials live in ''.env'', gitignored, and are never logged by ''scripts/dw.mjs''. No content was sent to any external LLM provider.   * **No accidental exposure.** Credentials live in ''.env'', gitignored, and are never logged by ''scripts/dw.mjs''. No content was sent to any external LLM provider.
  
Line 747: Line 801:
 | **§15 of this log was an empty header** while §14 promised four review passes | accepted; you are reading the fix | | **§15 of this log was an empty header** while §14 promised four review passes | accepted; you are reading the fix |
  
-==== 15.4 What the focused reviewers did not catch ====+==== 15.4 What a narrow brief cannot catch ====
  
-Two errors were found by the authorbefore and between the review passes, and neither reviewer flagged either:+Three defects in this run fell between the focused reviewers' briefs. They are recorded as //classes// of blind spotbecause each one implies a check that is missing rather than a reviewer who was careless. 
 + 
 +  * **Correct arithmetic on the wrong quantity.** The 60.1% Chrome-or-Chromium share was a sum of multi-valued table rows; the figures reviewer re-derived it and confirmed the arithmetic, which was right. A number guard compares a page to a script; it cannot see that the script is computing something the sentence does not mean. Only the distinct-paper roll-up (Q21–Q23) closes it. 
 +  * **A claim with no fetchable source, no script output and no paper quote.** The ETP Strict row (§15.5) sat in exactly the hole between the three briefs: the currency reviewer had nothing to fetch because SUMO is CAPTCHA-walled, the figures reviewer had no script that produced it, and the citations reviewer had no paper to check it against — so it was the //only// mechanism cell in the browser table with no citation, and it was wrong. The check that closes this is mechanical: **grep the page for every factual cell and ask which of the three passes owns it; the ones nobody owns are the ones to re-derive.** 
 +  * **A passing checker that does not cover the thing you want checked.** ''bp_quotecheck.mjs'' reported 32 of 32 while two quotes on the page were absent from its list. Coverage was found by a reviewer reading the page against the checker, not by the checker. The lesson is the same one as for the probesa green check is a statement about its inputs.
  
-  * The **summed multi-valued column** (60.1%, and a "60 papers" partial sum) described in §2. The figures reviewer re-derived 60.1% from the table and confirmed the arithmetic — which was correct arithmetic on the wrong quantity. An arithmetic check cannot catch a category error, and this is the clearest illustration on the page of why the number guard is not enough. 
-  * The **three column-spliced quotes** in §7.1, caught by writing the quote checker rather than by reading. The fourth splice was caught only when the citations reviewer pointed out the two quotes the checker did not cover — so the tool and the reviewer each found what the other missed. 
  
 ==== 15.5 Generic pass (Fable) ==== ==== 15.5 Generic pass (Fable) ====
  
-//Pending: this section is filled after the generic reviewer runs against the published page.//+No checklist; briefed to find what the three focused passes were not looking for, and given the target reader. It found the run's worst error, in the one place the other three structurally could not. 
 + 
 +^ Finding ^ Action ^ 
 +| **The ETP Strict row was factually wrong, and was the only mechanism cell in the browser table with no citation.** It said Strict adds "blocking //all// cross-site cookies", and the published Firefox snippet's comment said ''cookieBehavior'' ''1'' was "closest to ETP Strict". The build the page reads for every other Firefox claim says otherwise''browser.contentblocking.features.strict'' is ''%%tp,tpPrivate,cookieBehavior5,cookieBehaviorPBM5,cryptoTP,fp,stp,emailTP,emailTPPrivate,-consentmanagerSkip,-consentmanagerSkipPrivate,lvl2,rp,rpTop,qps,qpsPBM,fpp,fppPrivate,btp,lna%%'' — **Strict's cookie behaviour is 5, identical to Standard** | **accepted, and the most serious finding of the run.** Verified independently from ''browser/omni.ja'' before changing anything. The row is rewritten from the strict string, the code comment is corrected (''cookieBehavior'' ''1'' is a custom state no ETP preset ships), the strict string is now evidence in §9 (Q-row-equivalent), and a footnote records that Mozilla's user-facing wording describes what partitioning //achieves// rather than a different ''cookieBehavior''. Blocked versus partitioned is precisely the distinction this page teaches, so getting it backwards was not a detail | 
 +| **This log's evidence sections still asserted three facts that §15 said had been corrected** — the Edge narrative in §2.2, the Brave version in §11, the Brave support domain in §11. Someone checking a number lands on the section, not the review log | **accepted.** All three corrected in place with a "corrected in review — see §15.x" marker, so the wrong version and the reason it was wrong both stay visible | 
 +| **An indented line rendered as literal preformatted text**, ''%%[...]%%'' markup and all, because DokuWiki does not parse inline markup inside an indented block. Verified in the live DOM. (The reviewer notes ''privacy:privacy_sandbox'' has the same bug, so "match the sibling" was the wrong instinct) | **accepted.** Replaced with a ''%%<code>%%'' block. The sibling's instance is left alone: not this page's to fix, and flagged here instead | 
 +| **OpenWPM — the field's dominant Firefox harness — was missing from the section whose thesis is "your library has already set the protection state"** | **accepted.** A row added saying plainly that it was not examined here and that a third pref-setting layer makes the read-back matter more, not less | 
 +| **No reusable-artifact section**, unlike the sibling ''privacy:privacy_sandbox'' | **accepted, and it paid.** Ten artefacts listed, every URL fetched: nine resolve, and **{[jueckstock2022_privacy]}'s Chromium storage-policy patches 404 at the ''wspr-ncsu'' URL printed in the paper** and live under the first author's account instead. The flagship paper's arm-building code was one dead link from being unfindable | 
 +| **The page preaches read-back but supplied a mechanism for only two of eight arms** | **accepted.** A new subsection gives the one recipe that generalises — a canary origin visited first in every arm — and says explicitly that nothing in the corpus does this, so it is a recommendation and not a measured practice | 
 +| **Three claims where confidence outran evidence**: the lead's "a good number are not in the state their authors think"; a trend row comparing 2023–2026 with its own subset 2024–2026; "sixteen years" for a seventeen-venue-year corpus | **all three accepted.** The lead now says what the evidence is (tool defaults plus one probe) and that nobody has checked published crawls; the subset row is deleted and the two remaining windows are disjoint; the count is corrected | 
 +| **"Four produced a control arm" flattens a real distinction** — three failed silently, one threw a protocol error, and the silence is the hazard | **accepted.** Both the box and the discussion now separate them | 
 +| **One zero with no committed probe**: the CDP-command sweep | **accepted.** ''scripts/bp_cdp_sweep.mjs'' was already written but had no Q-row and no cross-reference. Q25 added, and the page now points at itQ24, Q26 and Q27 added for the other three sweeps at the same time | 
 +| **Edge appears five times, twice near-verbatim**; and a fourth paragraph under "Three things in that table" is crawl tooling, not table fallout | **accepted.** The Open Questions bullet is cut to a pointer, and the CDP bounce-tracking paragraph is moved into the driving section under its own heading | 
 +| **The driving table does not say which rows are hearsay** — only Chromium was exercised, and the admission lived 200 lines away | **accepted.** One sentence above the table | 
 +| **Two sections of this log are scorekeeping**: §14's "mistakes caught before any reviewer saw it" and §15.4's "what the focused reviewers did not catch" | **accepted.** Both reframed as failure modes with the check that would close each, which is the reusable part. The reviewer is right that credit allocation is not what a working log is for | 
 +| "The page does answer its own question" | recorded, and the three gaps it named between that and a usable Monday morning (artifacts, read-back, OpenWPM) are the three accepted above | 
 + 
 +==== 15.6 What no reviewer caught ==== 
 + 
 +''scripts/prov_embed.py'', which re-embeds script output into this page, silently deleted §6 and §7 of this log during a rebuild: an empty ''%%<file text …></file>%%'' placeholder let its non-greedy body pattern run on to the next block's closing tag and swallow everything between. Caught by counting ''%%<file text%%'' openers and section headings after the rebuild, not by reading, and not by any review pass — all four had read the file before the damage. The pattern now refuses to cross an opener. **The general form: a generator that edits a published page in place needs a structural check after every run, because the failure is a deletion and a deletion leaves nothing to notice.**
  
 ====== References ====== ====== References ======
provenance/privacy/browser_protection.1788299827.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki