User Tools

Site Tools


privacy:requests

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
privacy:requests [2026/09/16 10:59] – Add one-line disambiguator: this page is HTTP requests, privacy:data_subject_rights is DSARs. Authored by Claude karel.kubicek.claudeprivacy:requests [2026/09/21 14:55] (current) – Use in Publications: remove the duplicate which-lists table (drifting against the canonical copy on programming:filter_lists); replace with two sentences plus the pointer, and state why S1=197 and the filter-lists population=194 differ. S1/S2 narrative un karel.kubicek.claude
Line 223: Line 223:
 ==== Which lists the field actually uses ==== ==== Which lists the field actually uses ====
  
-Of the **197** papers that used or produced an advertising-or-tracking filter list. A paper naming several lists is counted under each, so the shares do not sum to 100%. Names were folded into families, because they are free text: the //Spellings folded// column is how many distinct strings the corpus uses for each. [[Programming:Filter Lists#Which lists|The filter-lists page]] computes the same table over a slightly wider population (198 papers, because it does not intersect with this page's request-classification task fold) and is the canonical version; the two differ by one or two papers per row, which is itself a useful illustration of how much a population definition moves a count. +Of the **197** papers in S1 — those that used or produced an advertising-or-tracking filter list — EasyList is named by **110 (55.8%)**, EasyPrivacy by 71 (36.0%), Disconnect by 48 (24.4%) and Ghostery / WhoTracks.me by 34 (17.3%), a paper naming several being counted under each. The names are free text and are folded into families before counting, which is not cosmetic: counting exact strings alone finds 90 EasyList papers instead of 110 (−18.2%) and **26 Disconnect papers instead of 48 (−45.8%)**, because //Disconnect list//, //Disconnect.me//, //Disconnect Entity List// and twenty-five further spellings are one list. **The full per-list table is on [[Programming:Filter Lists#Which lists|the filter-lists page]], which is the canonical copy** — it is deliberately not reproduced here, because the same table maintained in two places drifts.
- +
-^ Filter list ^ Papers ^ Share of 197 ^ Spellings folded ^ +
-| EasyList | 110 | 55.8% | 32 | +
-| EasyPrivacy | 71 | 36.0% | 26 | +
-| Disconnect | 48 | 24.4% | 28 | +
-| Ghostery / WhoTracks.me | 34 | 17.3% | 12 | +
-| hosts-file lists (hpHosts, AdAway, MoaAB, Pi-hole, NoTrack, …) | 27 | 13.7% | 35 | +
-| Adblock Plus (the lists shipped with it) | 26 | 13.2% | 12 | +
-| uBlock Origin lists | 17 | 8.6% | 11 | +
-| DuckDuckGo Tracker Radar | 15 | 7.6% | 9 | +
-| **unnamed or merely counted** ("nine crowd-sourced filter lists") | 10 | 5.1% | 10 | +
-| AdGuard | 8 | 4.1% | 13 | +
-| EasyList annoyance / anti-adblock variants | 7 | 3.6% | 8 | +
-| Privacy Badger //(a heuristic, not a list)// | 4 | 2.0% | 2 | +
-| anti-adblock scripts and services | 3 | 1.5% | 4 | +
-| cryptomining lists (NoCoin, CoinBlockerLists, MinerBlock) | 3 | 1.5% | 4 | +
-| Acceptable Ads exception list | 1 | 0.5% | 1 | +
- +
-**Folding is not cosmetic here.** Counting exact strings undercounts EasyList by 18.2% (90 papers against 110), EasyPrivacy by 19.7%, Ghostery/WhoTracks.me by 11.8% — and **Disconnect by 45.8%** (26 against 48), because it appears as //Disconnect list//, //Disconnect.me//, //Disconnect blacklist//, //Disconnect Entity List//, //Disconnect Tracker Protection lists// and twenty-two other spellings. Any table of list adoption built on exact strings is wrong by tens of percent.+
  
 +That page runs the //same// fold — both report scripts import ''req_fold.mjs'' — but over a population of **194** papers rather than these 197, so its rows sit a paper or two either side of the figures above. **That is a property of the two membership rules, not a disagreement between two counts**: S1 is built inside this page's request-classification population and reported after its exclusions, while the filter-lists page takes every paper the fold fires on with no topic restriction and then applies a tuple-context test this page does not have, which removes five papers whose list //name// is right and whose //use// is a malware or piracy blacklist. Both are correct over what each says it covers; the delta is itemised paper by paper on [[provenance:programming:filter_lists|the filter-lists provenance page]].
 Separately, the engines: tracker-radar-collector 10 papers (which is a [[Programming:Crawler:Tracker Radar Collector|crawler]], not a list — a distinction the raw names do not make), ''adblockparser'' 9, ''adblock-rust'' 7, uBlock Origin Core 2, ''abp-blocklist-parser'' 1, the Adblock Plus Android library 1. Separately, the engines: tracker-radar-collector 10 papers (which is a [[Programming:Crawler:Tracker Radar Collector|crawler]], not a list — a distinction the raw names do not make), ''adblockparser'' 9, ''adblock-rust'' 7, uBlock Origin Core 2, ''abp-blocklist-parser'' 1, the Adblock Plus Android library 1.
  
privacy/requests.txt · Last modified: by karel.kubicek.claude