User Tools

Site Tools


privacy:consent

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
privacy:consent [2026/08/19 03:57] – Remove an unsupported generalisation about how CMPs detect GPC. Authored by Claude karel.kubicek.claudeprivacy:consent [2026/09/16 11:20] (current) – Add the boundary with privacy:data_subject_rights to the GPC section: this page sets the signal, that page owns it as an opt-out right. Authored by Claude karel.kubicek.claude
Line 6: Line 6:
  
 <WRAP important> <WRAP important>
-**The finding that should shape your methods section.** Of the **1,120** papers in [[literature:corpus|this corpus]] that ran an automated web crawl, **349 (31.2%)** say what they did about consent notices at allOf those 349, **313 say they did not interact with the notice**That leaves 36 that appear to have clicked one — and reading all 36 against their own full text leaves **29 papers in the whole 2010–2026 corpusacross seven venues, that verifiably interacted with a consent notice.** That is **2.6% of all crawling papers**.((The other 7 are extraction false positiveshand-adjudicated on 2026-08-19 and listed one by one on [[provenance:privacy:consent|the provenance page]]. They are the reason the page reports 29 rather than 36; see [[#Almost nobody says what they did about consent]].))+**The finding that should shape your methods section.** Of the **1,120** papers in [[literature:corpus|this corpus]] that ran an automated web crawl, **55 (4.9%)** say in their own words what they did about consent notices. **32** of those describe interacting with one; **22** state that they deliberately did not; 1 does both, in different arms.((Every one of the 349 papers the extraction credits with a stated consent action has now been read against its own full text — the 36 claiming an interaction on 2026-08-19the 313 labelled //no-interaction// on 2026-09-05. **279 of those 313 give no evidence of their own crawl's consent action**: the label is the extractor's default, not the paper's claim. (Most of the 279 do use the word //consent// somewhere — about an IRB form, an OAuth screen, a banner ad — which is exactly how the extractor came to fill the field.) Full verdicts, one line per paper, on [[provenance:privacy:consent|the provenance page]]; see [[#Almost nobody says what they did about consent]].)) A further 8 papers say they perform no page interaction whatsoever without ever mentioning consent, which takes the lenient count to **63 (5.6%)**.
  
-Two consequences. First, the reviewer question "what did you do about the banner?" has no established answer to point at, so **state yours explicitly** — you will be in a minority that does. Second, when you compare your prevalence number against a published one, check what that paper did with the notice before you conclude anything from the difference: two thirds of the time the paper does not say, and the difference may be entirely the treatment.+That is the number to carry away: **for 1,057 of 1,120 crawling papers (94.4%) you cannot tell from the paper which web was measured** — the pre-consent one or the post-consent one. 
 + 
 +Two consequences. First, the reviewer question "what did you do about the banner?" has no established answer to point at, so **state yours explicitly** — you will be in a small minority that does. Second, when you compare your prevalence number against a published one, check what that paper did with the notice before you conclude anything from the difference. Excluding even the 236 papers for which the question genuinely does not arise, **more than nine in ten still say nothing**, and the difference between your number and theirs may be entirely the treatment.
 </WRAP> </WRAP>
  
Line 14: Line 16:
  
   * **Do Cookie Banners Respect my Choice?** {[matte2020_cookie]}, IEEE S&P 2020 — the paper that turned "is this banner legal?" into something a crawler can decide, by reading what the CMP //stored// rather than what the banner //said//. Read it for the four-violation taxonomy, which almost every later compliance paper reuses.   * **Do Cookie Banners Respect my Choice?** {[matte2020_cookie]}, IEEE S&P 2020 — the paper that turned "is this banner legal?" into something a crawler can decide, by reading what the CMP //stored// rather than what the banner //said//. Read it for the four-violation taxonomy, which almost every later compliance paper reuses.
-  * **Automated Large-Scale Analysis of Cookie Notice Compliance** {[bouhoula2024automated]}, USENIX Security 2024 — the current reference pipeline: find the notice, classify its buttons with a small language model, drive it, and check the cookies against the site's own declarations. Read it for the numbers you will be asked to beat.+  * **Automated Large-Scale Analysis of Cookie Notice Compliance** {[bouhoula2024_automated]}, USENIX Security 2024 — the current reference pipeline: find the notice, classify its buttons with a small language model, drive it, and check the cookies against the site's own declarations. Read it for the numbers you will be asked to beat.
   * **(Un)informed Consent** {[utz2019_informed]}, CCS 2019 and **Dark Patterns after the GDPR** {[nouwens2020_dark]}, CHI 2020 — the two studies that established that banner //design// moves consent rates, so a banner is a manipulation and not a question. Nouwens et al. is outside this corpus's seven venues; it is still the one everybody cites.   * **(Un)informed Consent** {[utz2019_informed]}, CCS 2019 and **Dark Patterns after the GDPR** {[nouwens2020_dark]}, CHI 2020 — the two studies that established that banner //design// moves consent rates, so a banner is a manipulation and not a question. Nouwens et al. is outside this corpus's seven venues; it is still the one everybody cites.
   * **A Large-Scale Study of Cookie Banner Interaction Tools** {[demir2024_bannertools]}, PETS 2024 — read before you decide to use an off-the-shelf extension, because it measures how often they actually work.   * **A Large-Scale Study of Cookie Banner Interaction Tools** {[demir2024_bannertools]}, PETS 2024 — read before you decide to use an off-the-shelf extension, because it measures how often they actually work.
Line 24: Line 26:
 There are six things a crawl can do, and the corpus records which one each paper chose. They are not interchangeable and they are not on a scale — each answers a different question. There are six things a crawl can do, and the corpus records which one each paper chose. They are not interchangeable and they are not on a scale — each answers a different question.
  
-^ Action ^ What it measures ^ What it costs you ^ Papers stating it ^ Verified ^ +^ Action ^ What it measures ^ What it costs you ^ Extraction ^ Verified ^ 
-| **no interaction** — load and leave the banner alone | The **pre-consent** web: what a site does before it has any legal basis. This is the right treatment for a violation study, because tracking before consent is the violation | Not what a user experiences. Any "how much tracking is there" number from this treatment is a floor, not a level | **313** | not audited |+| **no interaction** — load and leave the banner alone | The **pre-consent** web: what a site does before it has any legal basis. This is the right treatment for a violation study, because tracking before consent is the violation | Not what a user experiences. Any "how much tracking is there" number from this treatment is a floor, not a level | 313 | **22** |
 | **accept all** | The **upper bound**: everything the site is prepared to do with permission. The right treatment for enumerating vendors, purposes and the full cookie set | Says nothing about compliance, and over-states what a typical user is exposed to | 15 | **11** | | **accept all** | The **upper bound**: everything the site is prepared to do with permission. The right treatment for enumerating vendors, purposes and the full cookie set | Says nothing about compliance, and over-states what a typical user is exposed to | 15 | **11** |
 | **reject all** | Whether refusal is honoured. Only meaningful when paired with another arm | On its own it is uninterpretable — you cannot tell "respects rejection" from "has no tracking anyway" | 2 | **2** | | **reject all** | Whether refusal is honoured. Only meaningful when paired with another arm | On its own it is uninterpretable — you cannot tell "respects rejection" from "has no tracking anyway" | 2 | **2** |
Line 32: Line 34:
 | **dismiss or remove** — close the banner, or delete it from the DOM | Nothing about consent. Useful only to unblock a crawl whose real subject is something else | Removing the banner from the DOM is **not** a consent choice: no consent string is written, and the site may behave as it does pre-consent. Say "we removed the overlay", never "we declined" | 3 | **0** | | **dismiss or remove** — close the banner, or delete it from the DOM | Nothing about consent. Useful only to unblock a crawl whose real subject is something else | Removing the banner from the DOM is **not** a consent choice: no consent string is written, and the site may behave as it does pre-consent. Say "we removed the overlay", never "we declined" | 3 | **0** |
  
-The //Verified// column is the count that survived reading all 36 candidate papers' own full text: **all three** ''dismiss-or-remove'' values and **four of fifteen** ''accept-all'' values are extraction false positives, while every ''accept-and-reject'', ''reject-all'' and ''cmp-specific-choices'' value held. Details and the paper-by-paper verdicts are on [[provenance:privacy:consent|the provenance page]]. The practical reading: **the enum is trustworthy exactly where the paper had to describe two arms**, and unreliable where a single word in a methods paragraph could be misread.+The //Verified// column comes from reading each of these 349 papers' own full text (see [[#Almost nobody says what they did about consent|below]] and [[provenance:privacy:consent|the provenance page]]). Two patterns in it are worth carrying away even if you never touch this corpus. **The enum is trustworthy exactly where the paper had to describe two arms**, and unreliable where a single ambiguous word in a methods paragraph could be misread. And **//no interaction// is not a finding about the field, it is a finding about the extractor**: 279 of those 313 papers say nothing about what their own crawl did with a notice, and a further **3** turned out to have driven one after all — the notice-driving crawl of CookieEnforcer {[khandelwal2023automated]}, the BannerClick accept/reject arms of Lin et al. {[lin2024_browsing]}, and the three TCF consent modes of Morel et al. {[morel2026_tcf]}. Read a paper before you count it in either direction.
  
 <WRAP important> <WRAP important>
-**Two arms or no claim.** If your result is "sites track users who rejected", you need the rejection arm //and// a baseline. If your result is "sites track before consent", you need the no-interaction arm. A single accept-all crawl supports neither. Only **14 papers in this corpus** run both an accept and a reject arm — and that is the one value the audit above confirmed at 14 out of 14. It is why so many consent findings are hard to compare.+**Two arms or no claim.** If your result is "sites track users who rejected", you need the rejection arm //and// a baseline. If your result is "sites track before consent", you need the no-interaction arm. A single accept-all crawl supports neither. Only **16 papers in this corpus** run both an accept and a reject arm — the 14 the extraction labels ''accept-and-reject'', which the audit confirmed at 14 out of 14, plus two the audit recovered from the //no interaction// pile. It is why so many consent findings are hard to compare.
 </WRAP> </WRAP>
  
Line 48: Line 50:
  
 <WRAP important> <WRAP important>
-Whether the reload can carry the decision at all is a [[Programming:Stateful Stateless|stateful/stateless]] question. A stateless crawl that clears the profile between visits cannot observe post-consent behaviour on a **later** visit, because the consent cookie went with the profile. In this corpus only **115 crawling papers state both a consent action and a statefulness**and **93 of those are the no-interaction case**. If you interact with banners, statefulness is not an independent choice — say which you used and why.+Whether the reload can carry the decision at all is a [[Programming:Stateful Stateless|stateful/stateless]] question. A stateless crawl that clears the profile between visits cannot observe post-consent behaviour on a **later** visit, because the consent cookie went with the profile. In this corpus only **32 crawling papers state, in their own words, both a consent action and a statefulness** — 20 that interacted and 12 that deliberately did not. (The extraction puts that figure at 115, but 93 of those are //no-interaction// labels the paper never makes; see below.) Of the 32, **18 are stateless**, which is the combination that most often makes the consent arm unobservable. If you interact with banners, statefulness is not an independent choice — say which you used and why. 
 +</WRAP> 
 + 
 +<WRAP important> 
 +**A fifth check, because on many sites the choice does not switch the tag off at all.** Any site running a Google tag (GA4, Google Ads, Tag Manager) can pass the user's consent state to Google rather than simply loading or not loading the tag, and Google requires advertisers serving the EEA to do so in order to keep ad personalisation and measurement.((Google, "Updates to consent mode for traffic in European Economic Area (EEA)", ''support.google.com/tagmanager/answer/13695607'', fetched 2026-08-19: "we are strengthening the enforcement of our EU user consent policy (EU UCP)… you must collect consent for use of personal data from end users based in the EEA and share consent signals with Google". The page does not itself carry the widely-quoted March 2024 enforcement date, so this page does not assert one.)) The consequence for a crawl is concrete: on a consent-mode site a rejection does **not** stop the Google tag from firing — it fires with ''ad_storage'' and ''analytics_storage'' denied and sends cookieless pings instead. A measurement that counts requests, or counts tags loaded, will conclude that rejection did nothing. A measurement that counts cookies set will conclude it worked. **Say which you counted.**
 </WRAP> </WRAP>
  
Line 54: Line 60:
  
 Most consent banners on European sites are operated by a **Consent Management Platform (CMP)**, and a large share of those implement IAB Europe's **Transparency and Consent Framework (TCF)**, which standardises both the API and the storage format. That standardisation is what makes consent machine-readable at scale, and it is why so many measurement papers are TCF papers. Most consent banners on European sites are operated by a **Consent Management Platform (CMP)**, and a large share of those implement IAB Europe's **Transparency and Consent Framework (TCF)**, which standardises both the API and the storage format. That standardisation is what makes consent machine-readable at scale, and it is why so many measurement papers are TCF papers.
 +
 +This section covers what you need to know about the TCF //to make and verify a consent choice//. The string itself — its segments and bit layout, the seven places it can be found, decoding it reproducibly against a pinned Global Vendor List, Google's separate Additional Consent string, and what a decoded string does and does not prove about a site's behaviour — is on [[Privacy:TCF Consent Strings|Decoding TCF Consent Strings]].
  
 TCF sites are a **minority, and you must report the denominator as such.** Matte et al. {[matte2020_cookie]} found a TCF banner on **1,426 of 22,949** reachable European sites (**6.2%**) in 2019; Hils et al. {[hils2021_privacy]} tracked adoption longitudinally and detected TCF implementations on the order of thousands of sites in the top 100k. In Android apps the share is comparable: Morel et al. {[morel2026_tcf]} found TCF in **576 of 4,482 apps (12.85%)**. "Consent on the web" and "TCF consent" are not the same population. TCF sites are a **minority, and you must report the denominator as such.** Matte et al. {[matte2020_cookie]} found a TCF banner on **1,426 of 22,949** reachable European sites (**6.2%**) in 2019; Hils et al. {[hils2021_privacy]} tracked adoption longitudinally and detected TCF implementations on the order of thousands of sites in the top 100k. In Android apps the share is comparable: Morel et al. {[morel2026_tcf]} found TCF in **576 of 4,482 apps (12.85%)**. "Consent on the web" and "TCF consent" are not the same population.
Line 112: Line 120:
  
 <WRAP important> <WRAP important>
-**A consent string you set on one site can be read on another, and it will corrupt a stateful crawl.** The TCF permits a CMP to store the string in a cookie on a shared domain. Matte et al. {[matte2020_cookie]} tested this directly: they planted a consent string in the shared cookie and then asked each site's CMP for its state without touching the banner**62 sites (4.3%) returned the same consent string** — their CMP had adopted a consent created by a different CMP entirely. The authors call this a lower bound. If you crawl statefully across sitesyour accept on site //A// may be silently in force on site //B//and you will record it as //B// setting cookies without consentClear the shared-domain cookies between sites, or crawl statelessly and reload — and say which you did.+**A consent string set on one site used to be readable on another — this mechanism is retired, and the correction matters.** Under TCF v1.x, and under v2.0's optional //global scope//, a CMP could store the string in a shared cookie on the ''consensu.org'' domain — ''euconsent'' and then ''euconsent-v2'' — readable by any other CMP. Matte et al. {[matte2020_cookie]} tested this in 2019: they planted a consent string in the shared cookie and then asked each site's CMP for its state without touching the banner, and **62 sites (4.3%) returned the same consent string** — their CMP had adopted a consent created by a different CMP entirely. IAB Europe **announced the deprecation of global scopeout-of-band consent and the ''euconsent-v2'' cookie on 22 June 2021, and TC strings established with global scope have been invalid since 1 September 2021.**((IAB Tech Lab, //Transparency and Consent String with Global Vendor & CMP List Formats//, section "What happened to Global Scope and Out of Band?" and the version-history rows for July and September 2021. Fetched from the specification repository on 2026-08-26.)) So the cross-site leakage Matte et al. found is a historical findingnot a hazard of a crawl run todayWhat survives is the same-site version: a consent string persists across page loads within a site, so a stateful crawl still carries your earlier choice forward. Clear consent storage between visits, or crawl statelessly and reload — and say which you did.
 </WRAP> </WRAP>
  
-**Do not decode the TC string yourself.** It is a versioned bit-packed base64 format; use a maintained decoder and pin its version. Two papers in this corpus report writing their own decoding script, and both had to pin a **Global Vendor List** version to interpret the vendor bitfield — the GVL changes weekly, so a decoded vendor set is only meaningful together with the GVL version you decoded it against. The list is served from ''vendor-list.consensu.org/v3/vendor-list.json'' with numbered archives under ''/v3/archives/'', which is what makes a retrospective decode reproducible at all: **fetch and archive the GVL alongside your crawl**, do not resolve vendor IDs months later.+**The ''getTCData'' command in the probe above is deprecated.** IAB Tech Lab deprecated it in the CMP API specification with TCF v2.2 (May 2023), in favour of registering an ''addEventListener'' callback; the three required commands are now ''ping'', ''addEventListener'' and ''removeEventListener''. In a four-site spot-check on 2026-08-26 all four CMPs that answered at all still returned a TC string for ''getTCData'', although one set the callback's ''success'' flag to ''false'' while doing so. Published crawlers using the command are therefore not broken. But a CMP is within spec to drop it, and the failure mode is a silently growing "no TCF" bucket. New code should use the listener; [[Privacy:TCF Consent Strings|Decoding TCF Consent Strings]] publishes a tested ''addEventListener'' capture snippet and covers the other six channels the string travels in. 
 + 
 +**Do not decode the TC string ad hoc.** It is a versioned bit-packed base64 format; use a maintained decoder and pin its version — and check //which// package you installed, because the reference implementation moved npm scope and left a three-year-old copy behind under the old name. [[Privacy:TCF Consent Strings|Decoding TCF Consent Strings]] has the details, an audited minimal decoder, and the bit layout it reads. Two papers in this corpus report writing their own decoding script, and both had to pin a **Global Vendor List** version to interpret the vendor bitfield — the GVL changes weekly, so a decoded vendor set is only meaningful together with the GVL version you decoded it against. The list is served from ''vendor-list.consensu.org/v3/vendor-list.json'' with numbered archives under ''/v3/archives/'', which is what makes a retrospective decode reproducible at all: **fetch and archive the GVL alongside your crawl**, do not resolve vendor IDs months later.
  
 **The other reason to read the stored string: it lets you check compliance without trusting the interface.** Smith et al. {[smith2024_gdpr]} decoded TC strings across repeated crawls and found recorded-consent violations in **2.2% of domains and 1.3% of crawls** — a much lower rate than the "72% of sites violate something" headlines elsewhere on this page, precisely because it is a narrow, mechanically checkable question ("does the stored string match the choice made?") rather than a broad legal one. When you report a violation rate, say which of those two kinds of question you asked. **The other reason to read the stored string: it lets you check compliance without trusting the interface.** Smith et al. {[smith2024_gdpr]} decoded TC strings across repeated crawls and found recorded-consent violations in **2.2% of domains and 1.3% of crawls** — a much lower rate than the "72% of sites violate something" headlines elsewhere on this page, precisely because it is a narrow, mechanically checkable question ("does the stored string match the choice made?") rather than a broad legal one. When you report a violation rate, say which of those two kinds of question you asked.
Line 132: Line 142:
  
 DNT sent a ''DNT: 1'' request header and exposed ''navigator.doNotTrack''. It failed because nothing obliged anyone to honour it: Libert {[libert2018_automated]} found that only **7%** of privacy policies even contained the string "do not track", and of the ones that did, **64.80% explicitly said they did not honour it** against **19.46%** that committed to honouring it. Among 25 third-party data collectors, nine mentioned DNT and **none offered unqualified support**. DNT sent a ''DNT: 1'' request header and exposed ''navigator.doNotTrack''. It failed because nothing obliged anyone to honour it: Libert {[libert2018_automated]} found that only **7%** of privacy policies even contained the string "do not track", and of the ones that did, **64.80% explicitly said they did not honour it** against **19.46%** that committed to honouring it. Among 25 third-party data collectors, nine mentioned DNT and **none offered unqualified support**.
 +
 +**Where this page stops on GPC.** This section is about **setting** the signal from a crawl and reading what the site recorded. [[Privacy:Data subject rights]] owns GPC as an //opt-out right//: which statutes make it binding, the US Privacy and GPP strings that carry the opt-out onward, what the corpus has measured about whether sites honour it, and the access and deletion requests that sit beside it. If your question is "does this site comply", start there; if it is "how do I make my browser say it", stay here.
  
 It is formally dead. The W3C Tracking Protection Working Group concluded its work and republished both specifications as **W3C Working Group Notes on 17 January 2019**, saying in the status section that "there has not been sufficient deployment of these extensions (as defined) to justify further advancement".((''w3.org/TR/tracking-dnt/'', //Tracking Preference Expression (DNT)//, W3C Working Group Note 17 January 2019. Fetched 2026-08-19.)) Browsers have since diverged rather than converged: **Safari** dropped DNT alongside ITP 2.1 in 2019, and **Firefox removed the checkbox in version 135 (4 February 2025)**, whose release notes point users at "Tell websites not to sell or share my data" — which is GPC.((Mozilla, //Firefox 135.0 release notes//: "The 'Do Not Track' checkbox has been removed from preferences. If you wish to ask websites to respect your privacy, you can use the 'Tell websites not to sell or share my data' setting instead. This option is built on top of the Global Privacy Control (GPC)." Fetched 2026-08-19.)) **Chrome still exposes a DNT toggle.** If your crawl runs a default Chrome profile you may be sending ''DNT'' without meaning to; if it runs a current Firefox, the setting you find in the UI is GPC, not DNT. Check what your browser actually sends rather than what you assume. It is formally dead. The W3C Tracking Protection Working Group concluded its work and republished both specifications as **W3C Working Group Notes on 17 January 2019**, saying in the status section that "there has not been sufficient deployment of these extensions (as defined) to justify further advancement".((''w3.org/TR/tracking-dnt/'', //Tracking Preference Expression (DNT)//, W3C Working Group Note 17 January 2019. Fetched 2026-08-19.)) Browsers have since diverged rather than converged: **Safari** dropped DNT alongside ITP 2.1 in 2019, and **Firefox removed the checkbox in version 135 (4 February 2025)**, whose release notes point users at "Tell websites not to sell or share my data" — which is GPC.((Mozilla, //Firefox 135.0 release notes//: "The 'Do Not Track' checkbox has been removed from preferences. If you wish to ask websites to respect your privacy, you can use the 'Tell websites not to sell or share my data' setting instead. This option is built on top of the Global Privacy Control (GPC)." Fetched 2026-08-19.)) **Chrome still exposes a DNT toggle.** If your crawl runs a default Chrome profile you may be sending ''DNT'' without meaning to; if it runs a current Firefox, the setting you find in the UI is GPC, not DNT. Check what your browser actually sends rather than what you assume.
Line 181: Line 193:
  
 <WRAP important> <WRAP important>
-**Measure the tool on your own sample before you trust it.** Demir et al. {[demir2024_bannertools]} evaluated banner-interaction extensions and found each one interacts with **12 (65%)** of the banners shown to it on average (SD 21%, min 48%, max 95%). A third of banners are missed, and which third is not random — it correlates with language, with CMP, and with how the banner is injected. If a tool is your instrument, its success rate on //your// crawl is a number your paper owes the reader.+**Measure the tool on your own sample before you trust it.** Demir et al. {[demir2024_bannertools]} evaluated five banner-interaction extensions on a hand-checked sample and found each one interacts with **65% of the banners it was shown**, on average — the paper writes it as "on average, with 12 (65%) (SD21%max95% min: 48%) of all banners", where the 12 is the mean count on their sample rather than a rate. A third of banners are missed, and which third is not random — it correlates with language, with CMP, and with how the banner is injected. If a tool is your instrument, its success rate on //your// crawl is a number your paper owes the reader.
 </WRAP> </WRAP>
  
Line 188: Line 200:
 ^ Tool ^ What it actually does ^ State on 2026-08-19 ^ Use it when ^ ^ Tool ^ What it actually does ^ State on 2026-08-19 ^ Use it when ^
 | **Consent-O-Matic** ([[https://github.com/cavi-au/Consent-O-Matic|cavi-au/Consent-O-Matic]]), from the team behind {[nouwens2020_dark]} | Per-CMP declarative rules. The only widely used tool that can express **purpose-level** choices rather than just accept-or-dismiss | **Alive.** ''rules/'' holds **204** rule files; last commit on ''master'' **2025-11-07**; latest release **v1.1.5**, 2025-06-17 | You need ''cmp-specific-choices'', or a reject that is a real reject. Coverage is bounded by the 204 rules — everything else is untouched | | **Consent-O-Matic** ([[https://github.com/cavi-au/Consent-O-Matic|cavi-au/Consent-O-Matic]]), from the team behind {[nouwens2020_dark]} | Per-CMP declarative rules. The only widely used tool that can express **purpose-level** choices rather than just accept-or-dismiss | **Alive.** ''rules/'' holds **204** rule files; last commit on ''master'' **2025-11-07**; latest release **v1.1.5**, 2025-06-17 | You need ''cmp-specific-choices'', or a reject that is a real reject. Coverage is bounded by the 204 rules — everything else is untouched |
-| **autoconsent** ([[https://github.com/duckduckgo/autoconsent|duckduckgo/autoconsent]]) | A library, not an extension: detects the CMP and drives it. Ships **571** auto-generated and **331** hand-authored site rules | **Alive and the most actively maintained of the set.** Release **v16.23.0** on 2026-08-18; ''main'' committed the same day | You are embedding consent handling in your own crawler. It is a library with a stable API, which is what you want. Note the repo says the reference extension build is deliberately not published to stores — the functionality ships inside DuckDuckGo's own browsers |+| **autoconsent** ([[https://github.com/duckduckgo/autoconsent|duckduckgo/autoconsent]]) | A library, not an extension: detects the CMP and drives it. Ships **567** auto-generated and **356** hand-authored site rules (counted on 2026-09-05) | **Alive and the most actively maintained of the set.** Release **v16.37.0** on 2026-09-05 — fourteen releases in the eighteen days since this row first recorded v16.23.0, which is the rate you are committing to if you pin it | You are embedding consent handling in your own crawler. It is a library with a stable API, which is what you want. Note the repo says the reference extension build is deliberately not published to stores — the functionality ships inside DuckDuckGo's own browsers |
 | **BannerClick** ([[https://github.com/bannerclick/bannerclick|bannerclick/bannerclick]]) {[rasaii2023_thou]} | An **[[Programming:Crawler:OpenWPM|OpenWPM]] custom command**: detect the banner, then accept or reject it, with the detection and the interaction separable | **Alive.** Default branch ''bannerclick_v0.26.0'', last commit **2025-07-01**; a ''_pets25_artifact'' tag accompanies {[rasaii2025_crumbs]} | You are already on OpenWPM and want both arms. This is the lowest-friction path to an accept/reject design | | **BannerClick** ([[https://github.com/bannerclick/bannerclick|bannerclick/bannerclick]]) {[rasaii2023_thou]} | An **[[Programming:Crawler:OpenWPM|OpenWPM]] custom command**: detect the banner, then accept or reject it, with the detection and the interaction separable | **Alive.** Default branch ''bannerclick_v0.26.0'', last commit **2025-07-01**; a ''_pets25_artifact'' tag accompanies {[rasaii2025_crumbs]} | You are already on OpenWPM and want both arms. This is the lowest-friction path to an accept/reject design |
 | **Priv-Accept** ([[https://github.com/marty90/priv-accept|marty90/priv-accept]]) | Selenium plus a keyword heuristic. **Accept only** — there is no reject arm | **Stale.** Last commit on ''main'' **2022-04-13**. Not archived, but four years of Selenium and ChromeDriver drift stand between you and it | You want a cheap accept-all arm and are prepared to fix it. Two papers in this corpus still use it | | **Priv-Accept** ([[https://github.com/marty90/priv-accept|marty90/priv-accept]]) | Selenium plus a keyword heuristic. **Accept only** — there is no reject arm | **Stale.** Last commit on ''main'' **2022-04-13**. Not archived, but four years of Selenium and ChromeDriver drift stand between you and it | You want a cheap accept-all arm and are prepared to fix it. Two papers in this corpus still use it |
 | **CookieBlock** ([[https://github.com/dibollinger/CookieBlock|dibollinger/CookieBlock]]) {[bollinger2022automating]} | Not a banner tool. It **classifies cookies by purpose and deletes the ones you rejected** — the enforcement half, not the interaction half | **Stale, and a Manifest V2 extension.** Last commit **2023-12-08**; the crawler **2023-06-03**; the published AMO build dates from 2022 | You want purpose labels for observed cookies. See [[Privacy:Cookies|Classifying Cookies]]. Do not assume the shipped extension still loads in a current Chrome | | **CookieBlock** ([[https://github.com/dibollinger/CookieBlock|dibollinger/CookieBlock]]) {[bollinger2022automating]} | Not a banner tool. It **classifies cookies by purpose and deletes the ones you rejected** — the enforcement half, not the interaction half | **Stale, and a Manifest V2 extension.** Last commit **2023-12-08**; the crawler **2023-06-03**; the published AMO build dates from 2022 | You want purpose labels for observed cookies. See [[Privacy:Cookies|Classifying Cookies]]. Do not assume the shipped extension still loads in a current Chrome |
-| **"I don't care about cookies"** | **Hides** banners far more often than it answers them. Acquired by Avast | **Original stale** (published build last updated 2023-11). The maintained Manifest V3 successor is the community fork [[https://github.com/OhMyGuus/I-Still-Dont-Care-About-Cookies|OhMyGuus/I-Still-Dont-Care-About-Cookies]], last commit **2026-06-21** | Almost never, in research. Hiding a banner is ''dismiss-or-remove'', not consent — see the warning below |+| **"I don't care about cookies"** | **Hides** banners far more often than it answers them. Acquired by Avast | **Original stale** (the Firefox build on addons.mozilla.org is **v3.5.0**, last updated **2023-12-06**, checked 2026-09-05). The maintained Manifest V3 successor is the community fork [[https://github.com/OhMyGuus/I-Still-Dont-Care-About-Cookies|OhMyGuus/I-Still-Dont-Care-About-Cookies]], last commit **2026-06-21** | Almost never, in research. Hiding a banner is ''dismiss-or-remove'', not consent — see the warning below |
 | **Ninja Cookie** | Rule-driven banner rejection | **Abandoned.** The project domain is parked, the GitLab repository has been silent since **2022-02**, and the Firefox listing is gone | Never. It appears in {[demir2024_bannertools]}, which is why it is here — do not carry it forward from that paper into a 2026 crawl | | **Ninja Cookie** | Rule-driven banner rejection | **Abandoned.** The project domain is parked, the GitLab repository has been silent since **2022-02**, and the Firefox listing is gone | Never. It appears in {[demir2024_bannertools]}, which is why it is here — do not carry it forward from that paper into a 2026 crawl |
 | **Super Agent** | Commercial, closed-source consent automation | Live commercial product | Not as a research instrument: you cannot pin its version or read its rules | | **Super Agent** | Commercial, closed-source consent automation | Live commercial product | Not as a research instrument: you cannot pin its version or read its rules |
Line 202: Line 214:
 </WRAP> </WRAP>
  
-**Which ones the field actually uses.** Counting papers in this corpus that name a tool as ''used'' or ''produced'' — 18 distinct names over 40 papers, folded with ''consent_fold.mjs'': Consent-O-Matic **9**, BannerClick **4**, CookieBlock **4**, a custom GPC extension or crawler **4**, autoconsent **3**, EasyList Cookie List **2**, Priv-Accept **2**, a TC-string decoder **2**, and one paper each for ConsentChk, CookieCheck, CookieEnforcer, CookieGuard, GDPR-Consent, "I don't care about cookies", Ninja Cookie, Opt-Out Easy, OptOutCheck and Super Agent. The long tail is the finding: **almost every consent paper builds its own instrument**, which is the main reason consent results are hard to compare across papers.+**Which ones the field actually uses.** Counting papers in this corpus that name a tool as ''used'' or ''produced'' — 17 distinct names over 39 papers, folded with ''consent_fold.mjs'': Consent-O-Matic **9**, BannerClick **4**, CookieBlock **4**, a custom GPC extension or crawler **4**, autoconsent **3**, EasyList Cookie List **2**, Priv-Accept **2**, a TC-string decoder **2**, and one paper each for ConsentChk, CookieCheck, CookieEnforcer, GDPR-Consent, "I don't care about cookies", Ninja Cookie, Opt-Out Easy, OptOutCheck and Super Agent. The long tail is the finding: **almost every consent paper builds its own instrument**, which is the main reason consent results are hard to compare across papers.
  
 ===== Jurisdiction: Your Vantage Point Chooses Your Banner ===== ===== Jurisdiction: Your Vantage Point Chooses Your Banner =====
Line 208: Line 220:
 The banner you see is chosen by the site from your IP address. Crawl a European site from a US datacentre and you will often get no banner at all, or a different one — and the TCF's own ''gdprApplies'' flag will be ''false''. This makes [[Design:Crawling location|the vantage point]] a co-determinant of every consent result, not an independent choice. The banner you see is chosen by the site from your IP address. Crawl a European site from a US datacentre and you will often get no banner at all, or a different one — and the TCF's own ''gdprApplies'' flag will be ''false''. This makes [[Design:Crawling location|the vantage point]] a co-determinant of every consent result, not an independent choice.
  
-The corpus says the field mostly does not handle this. Of the **349** crawling papers that state a consent action, **68 (19.5%)** state an EU/EEA vantage point, **72 (20.6%)** state a non-EEA vantage only, and **206 (59.0%)** carry a vantage tuple with the location //not stated//. The picture is much better among the papers that actually clicked: of the **29** whose interaction is verified, **21 (72.4%)** did so from an EU/EEA vantage3 from stated non-EEA vantage only, and 5 never said where they were.+The corpus says the field handles this better than it reports anything else about consent, but the population that reports anything at all is tiny. Of the **55** crawling papers audited as stating a consent action in their own words, **37 (67.3%)** state an EU/EEA vantage point, **(12.7%)** state a non-EEA vantage only, and **11 (20.0%)** carry a vantage tuple whose location is //not stated//. The split barely moves between the papers that clicked and the papers that deliberately did not: of the **32** whose interaction is verified, **22 (68.8%)** were in the EU/EEA; of the **22** that state they left the banner alone**14 (63.6%)** were. Reading these as encouraging would be mistake — they are conditional on a paper having said anything, which 94.4% of crawling papers did not.
  
 Ogut et al. {[ogut2024_dissecting]} is the paper to read on the language side of the same problem: button text is the classifier's input, and //Accept all// is //Aceptar todo// and //Alle akzeptieren// elsewhere. They found a consent notice on **37% (1511)** of successfully loaded sites worldwide. Tang et al. {[tang2025_navigating]} did the regional comparison for violations, finding at least one consent violation on **96.18% (EU)** to **97.72% (US)** of sites, with only **3.82%** enforcing preferences correctly. Ogut et al. {[ogut2024_dissecting]} is the paper to read on the language side of the same problem: button text is the classifier's input, and //Accept all// is //Aceptar todo// and //Alle akzeptieren// elsewhere. They found a consent notice on **37% (1511)** of successfully loaded sites worldwide. Tang et al. {[tang2025_navigating]} did the regional comparison for violations, finding at least one consent violation on **96.18% (EU)** to **97.72% (US)** of sites, with only **3.82%** enforcing preferences correctly.
Line 220: Line 232:
 Population: the **1,120** papers that ran an automated web crawl. Population: the **1,120** papers that ran an automated web crawl.
  
-^ ''crawlConfig.consentAction'' ^ Papers ^ Share of 1,120 ^ +^ ''crawlConfig.consentAction'' ^ Papers ^ Share of 1,120 ^ Supported by the paper 
-| no interaction | 313 | 27.9% | +| no interaction | 313 | 27.9% | **22** 
-| accept all | 15 | 1.3% | +| accept all | 15 | 1.3% | **11** 
-| accept and reject | 14 | 1.3% | +| accept and reject | 14 | 1.3% | **14** 
-| dismiss or remove | 3 | 0.3% | +| dismiss or remove | 3 | 0.3% | **0** 
-| reject all | 2 | 0.2% | +| reject all | 2 | 0.2% | **2** 
-| CMP-specific choices | 2 | 0.2% | +| CMP-specific choices | 2 | 0.2% | **2** 
-| //not stated// | 495 | 44.2% | +| //not stated// | 495 | 44.2% | — 
-| //not applicable// | 236 | 21.1% | +| //not applicable// | 236 | 21.1% | — 
-| //no crawl configuration extracted// | 40 | 3.6% |+| //no crawl configuration extracted// | 40 | 3.6% | — |
  
-**And the six stated rows do not all survive contact with the papers.** Every one of the 36 papers in the five interacting rows was read against its own full text**7 (19.4%) are extraction false positives** — the value fired on push-notification permission prompts, on "banner ads", and three times on IRB participant consent. The audited count of papers that verifiably interacted with a consent notice is **29**, or **2.6%** of crawling papersThe script is ''consent_action_audit.mjs'' and every verdict is on [[provenance:privacy:consent|the provenance page]]. ''no-interaction'' was not audited — 313 papers is beyond a hand pass — so treat it as an upper bound in the same way.+**None of the six stated rows survives contact with the papers, and the largest survives worst.** All 349 papers in those rows have now been read against their own full text
 + 
 +Among the **36** in the five interacting rows, **7 (19.4%) are extraction false positives** — the value fired on push-notification permission prompts, on "banner ads", and three times on IRB participant consent. 
 + 
 +Among the **313** in the ''no-interaction'' row, **279 (89.1%) are not supported by the paper at all**. ''no-interaction'' is not a sentinel — the same field carries ''not-stated'' and ''not-applicable'' — so it is a positive claim that the crawl left the notice alone, and in 279 cases the paper makes no such claim. 72 of them contain no consent vocabulary anywhere in either rendering of their full text; in the rest, the word the extractor could have been reading is an advertising or protocol "banner", an IRB consent form, an OAuth or cryptomining consent prompt, an industry opt-out programme, or the ISOC copyright boilerplate. A further **3 (1.0%) are false negatives**: those papers do drive a consent notice**22 (7.0%)** say in words that they did notand **8** say they perform no interaction at all without ever mentioning consent. 
 + 
 +Both audits are hand adjudications with one recorded verdict and reason per paper: ''consent_action_audit.mjs'' for the 36, ''consent_action_noninteraction_audit.mjs'' for the 313. Both fail loudly on an unadjudicated paper, and every phrase quoted in a verdict is machine-checked against the cited paper. The scripts, their unedited output and all 349 verdicts are on [[provenance:privacy:consent|the provenance page]].
  
 The three italic rows are **sentinels, not answers**. ''not-applicable'' is a legitimate value — a crawl of an API, a mobile app store or a set of non-European sites may have no banner to handle — but ''not-stated'' at 44.2% is the finding: nearly half of all crawling papers leave the reader unable to tell which web they measured. The three italic rows are **sentinels, not answers**. ''not-applicable'' is a legitimate value — a crawl of an API, a mobile app store or a set of non-European sites may have no banner to handle — but ''not-stated'' at 44.2% is the finding: nearly half of all crawling papers leave the reader unable to tell which web they measured.
  
-==== The reporting gap is not closingbut interaction is spreading ====+==== Reporting started with the GDPR, from nothingand is still rare ====
  
-^ Bucket ^ Crawling papers ^ State an action ^ Share ^ Interaction claimed ^ Verified ^ Share verified +^ Bucket ^ Crawling papers ^ Extraction: state an action ^ Share ^ **Audited: state an action** **Share** ^ Of those, interacted 
-| 2010–2013 | 102 | 25 | 24.5% | 0 | 0 | 0.0% | +| 2010–2013 | 102 | 25 | 24.5% | **0** **0.0%** | 0 
-| 2014–2017 | 167 | 50 | 29.9% | 0 | 0 | 0.0% | +| 2014–2017 | 167 | 50 | 29.9% | **0** **0.0%** | 0 
-| 2018–2021 | 308 | 99 | 32.1% | | 3 | 1.0% | +| 2018–2021 | 308 | 99 | 32.1% | **11** **3.6%** | 3 
-| 2022–2024 | 345 | 116 | 33.6% | 17 14 | 4.1% | +| 2022–2024 | 345 | 116 | 33.6% | **25** **7.2%** | 16 
-| 2025–2026* | 198 | 59 | 29.8% | 13 12 | 6.1% |+| 2025–2026* | 198 | 59 | 29.8% | **19** **9.6%** | 13 |
  
-//* provisional venue-years.// Two separate trends**Whether paper says anything** moved from **24.5%** in 2010–2013 to a peak of **33.6%** in 2022–2024 and back to **29.8%** in the provisional 2025–2026 bucket — about five points net over sixteen yearsand no better in the latest bucket than it was before the GDPR. This is not solved reporting problem, and it is not obviously improving**Whether paper interacts** starts at exactly zero before the GDPR and rises steadily afterwards: the treatment arrived with the law, as you would expect, and is now in roughly one crawling paper in sixteen. Note that the false positives cluster in the early buckets — the extraction is most likely to mistake something else for a consent action in a paper that has nothing to do with consent, which is exactly the pre-2022 population.+//* provisional venue-years — see the caveat at the top of this section.// 
 + 
 +The two halves of this table tell opposite stories, and the audited half is the one to believeOn the extraction's numbers the field looks as if it has always reported consent handling at roughly constant 25–34%, with a dip in the newest bucket. **Not one of those four adjacent movements is statistically significant** (Fisher's exact, all //p// > 0.39) — the apparent peak and fall were noise, and an earlier revision of this page described them as a trend. 
 + 
 +On the audited numbers, **no paper before 2018 states a consent action at all**, which is what you would expect: there was mostly no banner to have a policy aboutThe jump at the GDPR is real (0/167 to 11/308, //p// = 0.0099; pooled 2010–2017 against 2018–2026, //p// = 3.0 × 10⁻⁷). What comes after it is a rise you should not lean on: 3.6% to 7.2% is //p// = 0.058, and 7.2% to 9.6% is //p// = 0.33 in bucket that is provisional anyway. The honest reading is that **consent reporting went from nonexistent to rare when the law arrived, and has stayed rare** — under one crawling paper in ten, sixteen years in. The test script is ''consent_ni_significance.py''; both series and every //p// are on the provenance page.
  
 ==== The consent literature itself ==== ==== The consent literature itself ====
Line 263: Line 285:
 Per year: 1 (2018), 4 (2019), 3 (2020), 5 (2021), 9 (2022), 9 (2023), **20 (2024)**, 15 (2025*), 6 (2026*). The pre-2018 count is essentially zero, which is what you would expect of a literature created by a statute. Per year: 1 (2018), 4 (2019), 3 (2020), 5 (2021), 9 (2022), 9 (2023), **20 (2024)**, 15 (2025*), 6 (2026*). The pre-2018 count is essentially zero, which is what you would expect of a literature created by a statute.
  
-In **60 of the 72 (83.3%)** the consent mechanism is the object of study; in the other **12 (16.7%)** it appears only as an instrument — a paper measuring something else that had to get past the banner. **57 (79.2%)** measure the web and **15 (20.8%)** measure mobile apps; consent-dialog measurement in apps is a real and growing subfield {[nguyen2022_freely,koch2023_enough,morel2026_tcf,zimmeck2026_exercising]}, and its methods do not transfer from the web unchanged.+In **60 of the 72 (83.3%)** the consent mechanism is the object of study; in the other **12 (16.7%)** it appears only as an instrument — a paper measuring something else that had to get past the banner. **57** measure the web and **15** measure mobile apps — the field is multi-valued and a paper can do both, and 7 also measure some other online service — so these are not two halves of a partition. Consent-dialog measurement in apps is a real and growing subfield {[nguyen2022_freely,koch2023_enough,morel2026_tcf,zimmeck2026_exercising]}, and its methods do not transfer from the web unchanged.
  
 ==== How consent notices get classified ==== ==== How consent notices get classified ====
Line 279: Line 301:
 | LLM | 1 | 4.2% | | LLM | 1 | 4.2% |
  
-Shares do not sum to 100% because a paper usually stacks several. The shape is the point: **this is still a hand-built-heuristic field.** Supervised ML enters only with CookieEnforcer {[khandelwal2023automated]} and Bouhoula et al. {[bouhoula2024automated]}, both using BERT-class models on button text, and **exactly one paper in the corpus classifies a consent artefact with an LLM** — WhisperTest (CCS 2025), which uses Qwen2.5-7B for iOS UI automation rather than for consent semantics.+Shares do not sum to 100% because a paper usually stacks several. The shape is the point: **this is still a hand-built-heuristic field.** Supervised ML enters only with CookieEnforcer {[khandelwal2023automated]} and Bouhoula et al. {[bouhoula2024_automated]}, both using BERT-class models on button text, and **exactly one paper in the corpus classifies a consent artefact with an LLM** — WhisperTest (CCS 2025), which uses Qwen2.5-7B for iOS UI automation rather than for consent semantics.
  
 Validation is better here than on most pages: **19 of 24 (79.2%)** report manual validation and **19 of 24** name a ground-truth source; 4 report no validation at all. Since manual validation of a notice detector is cheap — a few hundred screenshots — there is no excuse for being in the last group. Validation is better here than on most pages: **19 of 24 (79.2%)** report manual validation and **19 of 24** name a ground-truth source; 4 report no validation at all. Since manual validation of a notice detector is cheap — a few hundred screenshots — there is no excuse for being in the last group.
Line 304: Line 326:
   * **Opinion 08/2024 on Valid Consent in the Context of Consent or Pay Models Implemented by Large Online Platforms** (17 April 2024) — the reference for cookiewall studies of the kind Rasaii et al. {[rasaii2023_thou]} pioneered.   * **Opinion 08/2024 on Valid Consent in the Context of Consent or Pay Models Implemented by Large Online Platforms** (17 April 2024) — the reference for cookiewall studies of the kind Rasaii et al. {[rasaii2023_thou]} pioneered.
  
-<WRAP important> 
-**And one industry mechanism that changes what an accept/reject arm even observes: Google's consent mode.** Any site running a Google tag (GA4, Google Ads, Tag Manager) can pass the user's consent state to Google rather than simply loading or not loading the tag, and Google requires advertisers serving the EEA to do so in order to keep ad personalisation and measurement.((Google, "Updates to consent mode for traffic in European Economic Area (EEA)", ''support.google.com/tagmanager/answer/13695607'', fetched 2026-08-19: "we are strengthening the enforcement of our EU user consent policy (EU UCP)… you must collect consent for use of personal data from end users based in the EEA and share consent signals with Google". The page does not itself carry the widely-quoted March 2024 enforcement date, so this page does not assert one.)) The consequence for a crawl is concrete: on a consent-mode site a rejection does **not** stop the Google tag from firing — it fires with ''ad_storage'' and ''analytics_storage'' denied and sends cookieless pings instead. A measurement that counts requests, or counts tags loaded, will conclude that rejection did nothing. A measurement that counts cookies set will conclude it worked. **Say which you counted.** 
-</WRAP> 
  
 ==== What the papers found: figures with their own denominators ==== ==== What the papers found: figures with their own denominators ====
Line 327: Line 346:
 | Sites with **at least one** suspected violation | 54.29% (304) | 560 hand-checked TCF sites | {[matte2020_cookie]} | | Sites with **at least one** suspected violation | 54.29% (304) | 560 hand-checked TCF sites | {[matte2020_cookie]} |
 | Sites setting non-necessary cookies **without any interaction** | 69.7% | sites examined | {[bollinger2022automating]} | | Sites setting non-necessary cookies **without any interaction** | 69.7% | sites examined | {[bollinger2022automating]} |
-| Sites with at least one cookie-notice violation, 2024 | 72.2% | successfully crawled sites | {[bouhoula2024automated]} |+| Sites with at least one cookie-notice violation, 2024 | 72.2% | successfully crawled sites | {[bouhoula2024_automated]} |
 | Sites with at least one consent violation, 2025 | 96.18% (EU) – 97.72% (US) | sites per region | {[tang2025_navigating]} | | Sites with at least one consent violation, 2025 | 96.18% (EU) – 97.72% (US) | sites per region | {[tang2025_navigating]} |
 | Sites enforcing preferences correctly | 3.82% | same | {[tang2025_navigating]} | | Sites enforcing preferences correctly | 3.82% | same | {[tang2025_navigating]} |
Line 347: Line 366:
   * **Free-text names were folded before counting, and the residue is printed.** The consent-tool fold leaves 5 distinct unmapped names over 5 papers, all of them IAB artefacts that are not banner-interaction tools (''IAB ads.txt crawler'', ''IAB anti-ad-block script'', the IAB content taxonomy). The law fold leaves 2 (''Digital Economy Act 2017'', ''Act against Unfair Competition (UWG)''). The vantage fold leaves 5 strings that name no place (''different continents'', ''various geographic regions'').   * **Free-text names were folded before counting, and the residue is printed.** The consent-tool fold leaves 5 distinct unmapped names over 5 papers, all of them IAB artefacts that are not banner-interaction tools (''IAB ads.txt crawler'', ''IAB anti-ad-block script'', the IAB content taxonomy). The law fold leaves 2 (''Digital Economy Act 2017'', ''Act against Unfair Competition (UWG)''). The vantage fold leaves 5 strings that name no place (''different continents'', ''various geographic regions'').
   * **Every per-paper figure on this page was checked against the paper's own text**, not against the extraction's summary of it. The check covers 61 literals — a superset of what is published, since ten were checked and then cut — and found 60 in both the column-repaired and the plain rendering, 1 in the column-repaired rendering only, and **0 not found**. That pass caught two errors in an earlier draft of this page — a figure attributed to Bouhoula et al. that the paper writes without a thousands separator, and a Matte et al. percentage this page had rounded to the wrong decimal.   * **Every per-paper figure on this page was checked against the paper's own text**, not against the extraction's summary of it. The check covers 61 literals — a superset of what is published, since ten were checked and then cut — and found 60 in both the column-repaired and the plain rendering, 1 in the column-repaired rendering only, and **0 not found**. That pass caught two errors in an earlier draft of this page — a figure attributed to Bouhoula et al. that the paper writes without a thousands separator, and a Matte et al. percentage this page had rounded to the wrong decimal.
-  * **''consentAction'' was audited paper by paper, and it needed to be.** The schema's stability comparison puts it in the reliable band — an independent extraction run over the same text agrees with it on 93% of papers — but that measures whether two runs agree, not whether either is right, and it was measured on the earlier 4,322-paper corpus. Reading all 36 interacting papers found **7 false positives (19.4%)**, concentrated in ''accept-all'' and ''dismiss-or-remove''. **Note also that the field's own evidence quote cannot catch this**: ''crawlConfig'' carries one quote for the whole configuration object, so the quote behind a ''consentAction'' value usually evidences statefulness or crawl depth instead. Spot-checking quotes, which is the standard check on this site, is structurally blind here.+  * **''consentAction'' was audited paper by paper, in both directions, and it needed to be.** The schema's stability comparison puts it in the reliable band — an independent extraction run over the same text agrees with it on 93% of papers — but that measures whether two runs agree, not whether either is right, and it was measured on the earlier 4,322-paper corpus. Reading all 36 interacting papers found **7 false positives (19.4%)**, concentrated in ''accept-all'' and ''dismiss-or-remove''. Reading all 313 ''no-interaction'' papers found **279 (89.1%) that make no claim about their own crawl's consent action** and **3 (1.0%) that in fact drove a notice**. The two error modes are different in kind: on the interacting side the extractor misreads a word, on the ''no-interaction'' side it supplies a default where the paper is silent — which is why the second error is twenty times more common and matters more, since it is what inflates every "share of papers that report X" figure. **Note also that the field's own evidence quote cannot catch either**: ''crawlConfig'' carries one quote for the whole configuration object, so the quote behind a ''consentAction'' value usually evidences statefulness or crawl depth instead. Spot-checking quotes, which is the standard check on this site, is structurally blind here.
   * **Venue coverage.** Seven venues only. **CHI, SOUPS, EuroS&P, ACSAC, RAID, AsiaCCS and WPES are absent**, and that bites harder on this page than on most: the usable-privacy half of the consent literature (Nouwens et al., Habib et al., Utz et al.'s follow-ups) is largely CHI and SOUPS work. Every count here is a lower bound.   * **Venue coverage.** Seven venues only. **CHI, SOUPS, EuroS&P, ACSAC, RAID, AsiaCCS and WPES are absent**, and that bites harder on this page than on most: the usable-privacy half of the consent literature (Nouwens et al., Habib et al., Utz et al.'s follow-ups) is largely CHI and SOUPS work. Every count here is a lower bound.
  
Line 366: Line 385:
 ===== Open Questions ===== ===== Open Questions =====
  
-<wrap todo>+<WRAP todo>
   * **Nobody has replicated the banner-interaction-tool evaluation since 2024.** Demir et al.'s 65% is one measurement, on one sample, of tools that have since changed. It is load-bearing for a lot of this page and for a lot of published crawls, and re-running it is a well-scoped, publishable study.   * **Nobody has replicated the banner-interaction-tool evaluation since 2024.** Demir et al.'s 65% is one measurement, on one sample, of tools that have since changed. It is load-bearing for a lot of this page and for a lot of published crawls, and re-running it is a well-scoped, publishable study.
   * **LLM-driven banner interaction is unmeasured as an //instrument//, and alarming as a //subject//.** A language model that reads any banner in any language is the obvious successor to CSS-selector rule sets, and exactly one paper in this corpus puts an LLM anywhere near a consent artefact. Nobody has published cost, latency, determinism or accuracy against a hand-labelled set — and determinism is the hard one for a measurement instrument. Meanwhile the reverse question is opening up: work outside this corpus reports that browser-automation agents accept consent banners even when instructed to refuse everything.((Reported in a CHI 2026 Extended Abstracts paper on browser automation by LLM agents and consent, which measures agents defaulting to accept under explicit deny-all instructions. Not in this corpus (CHI is absent) and not read in full by this page's author — treat as a pointer to look up, not as a verified figure.)) Whether an agentic browser can be trusted to express a consent choice is now both a methods question and a research topic.   * **LLM-driven banner interaction is unmeasured as an //instrument//, and alarming as a //subject//.** A language model that reads any banner in any language is the obvious successor to CSS-selector rule sets, and exactly one paper in this corpus puts an LLM anywhere near a consent artefact. Nobody has published cost, latency, determinism or accuracy against a hand-labelled set — and determinism is the hard one for a measurement instrument. Meanwhile the reverse question is opening up: work outside this corpus reports that browser-automation agents accept consent banners even when instructed to refuse everything.((Reported in a CHI 2026 Extended Abstracts paper on browser automation by LLM agents and consent, which measures agents defaulting to accept under explicit deny-all instructions. Not in this corpus (CHI is absent) and not read in full by this page's author — treat as a pointer to look up, not as a verified figure.)) Whether an agentic browser can be trusted to express a consent choice is now both a methods question and a research topic.
Line 374: Line 393:
   * **Consent revocation is nearly unstudied.** One paper {[kancherla2025_johnny]}, 158 sites. Withdrawal is as legally required as consent and is far harder to automate.   * **Consent revocation is nearly unstudied.** One paper {[kancherla2025_johnny]}, 158 sites. Withdrawal is as legally required as consent and is far harder to automate.
   * **No shared benchmark exists.** There is no public, versioned set of annotated consent notices that a new detector can report against, which is why every paper reports precision and recall on its own hand-labelled sample and none of them are comparable. Building one would be a bigger contribution than most new detectors.   * **No shared benchmark exists.** There is no public, versioned set of annotated consent notices that a new detector can report against, which is why every paper reports precision and recall on its own hand-labelled sample and none of them are comparable. Building one would be a bigger contribution than most new detectors.
-  * **The reporting gap itself.** 44.2% of crawling papers say nothing about consent, and the share that says //something// is no higher in 2025–2026 than it was in 2014–2017. A one-line methods sentence would fix it; the question is why sixteen years of the field have not produced one. +  * **The reporting gap itself.** On the audited figures, **94.4% of crawling papers leave the reader unable to tell what their crawl did with a notice**, and the share that does say is under one in ten even in the newest bucket. A one-line methods sentence would fix it; the question is why sixteen years of the field have not produced one. (The 44.2% ''not-stated'' row is the extraction's own count for that population and was **not** audited — after what the ''no-interaction'' audit found, no unaudited row on this page should be read as a claim about the papers.) 
-  * **How many of the 313 ''no-interaction'' papers really did not interact** is unknown. The 36 papers claiming an interaction were read one by one and 7 turned out not to have interacted; nobody has done the same in the other directionand 313 is too many for hand passIf the error is symmetric, the true count is somewhere either side of 29 — which is a reason to treat every figure in this section as an order of magnitude and to read the papers you actually compare yourself against+  * **All 349 papers the extraction credits with a stated consent action have now been read**, so the figures above are hand verdicts rather than extractor output. What that cannot fix is the other direction: a paper that clicked a banner and never wrote it down is invisible here by construction, and the 495 ''not-stated'' and 236 ''not-applicable'' papers were **not** read — false ''not-stated'' would be a further undercount**32 is a floor on the true number of interacting papers, not an estimate of it.** Two of the three false negatives were found only because a named tool (BannerClick, CookieEnforcer) appears in the text; a paper that rolled its own clicking and described it in one unremarkable sentence would still be missed
-</wrap>+</WRAP>
  
 ===== Related Pages ===== ===== Related Pages =====
Line 382: Line 401:
   * [[Privacy:Requests#Cookie Notices and Their Interactive Elements|Classifying Web Requests]] — detecting the notice and labelling its buttons, with the comparison table of detectors.   * [[Privacy:Requests#Cookie Notices and Their Interactive Elements|Classifying Web Requests]] — detecting the notice and labelling its buttons, with the comparison table of detectors.
   * [[Privacy:Cookies|Classifying Cookies]] — what the cookies you observe before and after the click actually are, and the CookieBlock/Cookiepedia label sources.   * [[Privacy:Cookies|Classifying Cookies]] — what the cookies you observe before and after the click actually are, and the CookieBlock/Cookiepedia label sources.
 +  * [[Privacy:Policies|Measuring Privacy Policies and Terms]] — the long document behind the banner. A banner is a UI measurement; a policy is a document-retrieval and NLP measurement, with its own tool lineage and its own denominators, and the two literatures barely cite each other.
 +  * [[Privacy:TCF Consent Strings|Decoding TCF Consent Strings]] — the TC string and Google's Additional Consent string as artefacts: bit layout, where they hide, decoding them reproducibly, and what they do and do not prove.
   * [[Programming:Interaction|Interaction with websites]] and [[Programming:Stateful Stateless|Stateful and stateless crawling]] — the crawler-side mechanics this page assumes.   * [[Programming:Interaction|Interaction with websites]] and [[Programming:Stateful Stateless|Stateful and stateless crawling]] — the crawler-side mechanics this page assumes.
   * [[Programming:Crawler|Comparison of crawling libraries]] — which crawlers ship a consent-interaction step.   * [[Programming:Crawler|Comparison of crawling libraries]] — which crawlers ship a consent-interaction step.
   * [[Design:Crawling location|Crawling location]] — why the vantage point changes which banner you get.   * [[Design:Crawling location|Crawling location]] — why the vantage point changes which banner you get.
 +  * [[Privacy:Age assurance|Age assurance]] — the other dismissable overlay on the first load, and one that usually writes a cookie //before// the banner is touched, so it is a confounder for anything counted pre-consent.
   * [[Practices:Legal enforcement|Legal enforcement]] — what to do with a violation once you have found one, and which authority is competent.   * [[Practices:Legal enforcement|Legal enforcement]] — what to do with a violation once you have found one, and which authority is competent.
   * [[Practices:Ethics|Ethics]] — clicking //Accept// at scale on behalf of nobody is a decision with an ethical dimension; it is discussed there.   * [[Practices:Ethics|Ethics]] — clicking //Accept// at scale on behalf of nobody is a decision with an ethical dimension; it is discussed there.
privacy/consent.1787111832.txt.gz · Last modified: by karel.kubicek.claude