| Both sides previous revisionPrevious revisionNext revision | Previous revision |
| privacy:consent [2026/09/04 17:29] – Citekey consolidation 2026-09-04: repoint duplicate bibliography keys (fouad2022my/boettger2025_regional/ahmad2026_ipfp/bouhoula2024automated/lerner2016internet) to the kept key; no prose or figure changed. Authored by Claude karel.kubicek.claude | privacy:consent [2026/09/16 11:20] (current) – Add the boundary with privacy:data_subject_rights to the GPC section: this page sets the signal, that page owns it as an opt-out right. Authored by Claude karel.kubicek.claude |
|---|
| |
| <WRAP important> | <WRAP important> |
| **The finding that should shape your methods section.** Of the **1,120** papers in [[literature:corpus|this corpus]] that ran an automated web crawl, **349 (31.2%)** say what they did about consent notices at all. Of those 349, **313 say they did not interact with the notice**. That leaves 36 that appear to have interacted with one — and reading all 36 against their own full text leaves **29 papers in the whole 2010–2026 corpus, across seven venues, that verifiably interacted with a consent notice.** That is **2.6% of all crawling papers**.((The other 7 are extraction false positives, hand-adjudicated on 2026-08-19 and listed one by one on [[provenance:privacy:consent|the provenance page]]. They are the reason the page reports 29 rather than 36; see [[#Almost nobody says what they did about consent]].)) | **The finding that should shape your methods section.** Of the **1,120** papers in [[literature:corpus|this corpus]] that ran an automated web crawl, **55 (4.9%)** say in their own words what they did about consent notices. **32** of those describe interacting with one; **22** state that they deliberately did not; 1 does both, in different arms.((Every one of the 349 papers the extraction credits with a stated consent action has now been read against its own full text — the 36 claiming an interaction on 2026-08-19, the 313 labelled //no-interaction// on 2026-09-05. **279 of those 313 give no evidence of their own crawl's consent action**: the label is the extractor's default, not the paper's claim. (Most of the 279 do use the word //consent// somewhere — about an IRB form, an OAuth screen, a banner ad — which is exactly how the extractor came to fill the field.) Full verdicts, one line per paper, on [[provenance:privacy:consent|the provenance page]]; see [[#Almost nobody says what they did about consent]].)) A further 8 papers say they perform no page interaction whatsoever without ever mentioning consent, which takes the lenient count to **63 (5.6%)**. |
| |
| Two consequences. First, the reviewer question "what did you do about the banner?" has no established answer to point at, so **state yours explicitly** — you will be in a minority that does. Second, when you compare your prevalence number against a published one, check what that paper did with the notice before you conclude anything from the difference. Excluding the 236 papers for which the question genuinely does not arise, **about three in five say nothing**, and the difference may be entirely the treatment. | That is the number to carry away: **for 1,057 of 1,120 crawling papers (94.4%) you cannot tell from the paper which web was measured** — the pre-consent one or the post-consent one. |
| | |
| | Two consequences. First, the reviewer question "what did you do about the banner?" has no established answer to point at, so **state yours explicitly** — you will be in a small minority that does. Second, when you compare your prevalence number against a published one, check what that paper did with the notice before you conclude anything from the difference. Excluding even the 236 papers for which the question genuinely does not arise, **more than nine in ten still say nothing**, and the difference between your number and theirs may be entirely the treatment. |
| </WRAP> | </WRAP> |
| |
| There are six things a crawl can do, and the corpus records which one each paper chose. They are not interchangeable and they are not on a scale — each answers a different question. | There are six things a crawl can do, and the corpus records which one each paper chose. They are not interchangeable and they are not on a scale — each answers a different question. |
| |
| ^ Action ^ What it measures ^ What it costs you ^ Papers stating it ^ Verified ^ | ^ Action ^ What it measures ^ What it costs you ^ Extraction ^ Verified ^ |
| | **no interaction** — load and leave the banner alone | The **pre-consent** web: what a site does before it has any legal basis. This is the right treatment for a violation study, because tracking before consent is the violation | Not what a user experiences. Any "how much tracking is there" number from this treatment is a floor, not a level | **313** | not audited | | | **no interaction** — load and leave the banner alone | The **pre-consent** web: what a site does before it has any legal basis. This is the right treatment for a violation study, because tracking before consent is the violation | Not what a user experiences. Any "how much tracking is there" number from this treatment is a floor, not a level | 313 | **22** | |
| | **accept all** | The **upper bound**: everything the site is prepared to do with permission. The right treatment for enumerating vendors, purposes and the full cookie set | Says nothing about compliance, and over-states what a typical user is exposed to | 15 | **11** | | | **accept all** | The **upper bound**: everything the site is prepared to do with permission. The right treatment for enumerating vendors, purposes and the full cookie set | Says nothing about compliance, and over-states what a typical user is exposed to | 15 | **11** | |
| | **reject all** | Whether refusal is honoured. Only meaningful when paired with another arm | On its own it is uninterpretable — you cannot tell "respects rejection" from "has no tracking anyway" | 2 | **2** | | | **reject all** | Whether refusal is honoured. Only meaningful when paired with another arm | On its own it is uninterpretable — you cannot tell "respects rejection" from "has no tracking anyway" | 2 | **2** | |
| | **dismiss or remove** — close the banner, or delete it from the DOM | Nothing about consent. Useful only to unblock a crawl whose real subject is something else | Removing the banner from the DOM is **not** a consent choice: no consent string is written, and the site may behave as it does pre-consent. Say "we removed the overlay", never "we declined" | 3 | **0** | | | **dismiss or remove** — close the banner, or delete it from the DOM | Nothing about consent. Useful only to unblock a crawl whose real subject is something else | Removing the banner from the DOM is **not** a consent choice: no consent string is written, and the site may behave as it does pre-consent. Say "we removed the overlay", never "we declined" | 3 | **0** | |
| |
| The //Verified// column comes from reading all 36 candidate papers' own full text (see [[#Almost nobody says what they did about consent|below]] and [[provenance:privacy:consent|the provenance page]]). The pattern in it is worth carrying away even if you never touch this corpus: **the enum is trustworthy exactly where the paper had to describe two arms**, and unreliable where a single ambiguous word in a methods paragraph could be misread. | The //Verified// column comes from reading each of these 349 papers' own full text (see [[#Almost nobody says what they did about consent|below]] and [[provenance:privacy:consent|the provenance page]]). Two patterns in it are worth carrying away even if you never touch this corpus. **The enum is trustworthy exactly where the paper had to describe two arms**, and unreliable where a single ambiguous word in a methods paragraph could be misread. And **//no interaction// is not a finding about the field, it is a finding about the extractor**: 279 of those 313 papers say nothing about what their own crawl did with a notice, and a further **3** turned out to have driven one after all — the notice-driving crawl of CookieEnforcer {[khandelwal2023automated]}, the BannerClick accept/reject arms of Lin et al. {[lin2024_browsing]}, and the three TCF consent modes of Morel et al. {[morel2026_tcf]}. Read a paper before you count it in either direction. |
| |
| <WRAP important> | <WRAP important> |
| **Two arms or no claim.** If your result is "sites track users who rejected", you need the rejection arm //and// a baseline. If your result is "sites track before consent", you need the no-interaction arm. A single accept-all crawl supports neither. Only **14 papers in this corpus** run both an accept and a reject arm — and that is the one value the audit above confirmed at 14 out of 14. It is why so many consent findings are hard to compare. | **Two arms or no claim.** If your result is "sites track users who rejected", you need the rejection arm //and// a baseline. If your result is "sites track before consent", you need the no-interaction arm. A single accept-all crawl supports neither. Only **16 papers in this corpus** run both an accept and a reject arm — the 14 the extraction labels ''accept-and-reject'', which the audit confirmed at 14 out of 14, plus two the audit recovered from the //no interaction// pile. It is why so many consent findings are hard to compare. |
| </WRAP> | </WRAP> |
| |
| |
| <WRAP important> | <WRAP important> |
| Whether the reload can carry the decision at all is a [[Programming:Stateful Stateless|stateful/stateless]] question. A stateless crawl that clears the profile between visits cannot observe post-consent behaviour on a **later** visit, because the consent cookie went with the profile. In this corpus only **115 crawling papers state both a consent action and a statefulness**, and **93 of those are the no-interaction case**. If you interact with banners, statefulness is not an independent choice — say which you used and why. | Whether the reload can carry the decision at all is a [[Programming:Stateful Stateless|stateful/stateless]] question. A stateless crawl that clears the profile between visits cannot observe post-consent behaviour on a **later** visit, because the consent cookie went with the profile. In this corpus only **32 crawling papers state, in their own words, both a consent action and a statefulness** — 20 that interacted and 12 that deliberately did not. (The extraction puts that figure at 115, but 93 of those are //no-interaction// labels the paper never makes; see below.) Of the 32, **18 are stateless**, which is the combination that most often makes the consent arm unobservable. If you interact with banners, statefulness is not an independent choice — say which you used and why. |
| </WRAP> | </WRAP> |
| |
| |
| DNT sent a ''DNT: 1'' request header and exposed ''navigator.doNotTrack''. It failed because nothing obliged anyone to honour it: Libert {[libert2018_automated]} found that only **7%** of privacy policies even contained the string "do not track", and of the ones that did, **64.80% explicitly said they did not honour it** against **19.46%** that committed to honouring it. Among 25 third-party data collectors, nine mentioned DNT and **none offered unqualified support**. | DNT sent a ''DNT: 1'' request header and exposed ''navigator.doNotTrack''. It failed because nothing obliged anyone to honour it: Libert {[libert2018_automated]} found that only **7%** of privacy policies even contained the string "do not track", and of the ones that did, **64.80% explicitly said they did not honour it** against **19.46%** that committed to honouring it. Among 25 third-party data collectors, nine mentioned DNT and **none offered unqualified support**. |
| | |
| | **Where this page stops on GPC.** This section is about **setting** the signal from a crawl and reading what the site recorded. [[Privacy:Data subject rights]] owns GPC as an //opt-out right//: which statutes make it binding, the US Privacy and GPP strings that carry the opt-out onward, what the corpus has measured about whether sites honour it, and the access and deletion requests that sit beside it. If your question is "does this site comply", start there; if it is "how do I make my browser say it", stay here. |
| |
| It is formally dead. The W3C Tracking Protection Working Group concluded its work and republished both specifications as **W3C Working Group Notes on 17 January 2019**, saying in the status section that "there has not been sufficient deployment of these extensions (as defined) to justify further advancement".((''w3.org/TR/tracking-dnt/'', //Tracking Preference Expression (DNT)//, W3C Working Group Note 17 January 2019. Fetched 2026-08-19.)) Browsers have since diverged rather than converged: **Safari** dropped DNT alongside ITP 2.1 in 2019, and **Firefox removed the checkbox in version 135 (4 February 2025)**, whose release notes point users at "Tell websites not to sell or share my data" — which is GPC.((Mozilla, //Firefox 135.0 release notes//: "The 'Do Not Track' checkbox has been removed from preferences. If you wish to ask websites to respect your privacy, you can use the 'Tell websites not to sell or share my data' setting instead. This option is built on top of the Global Privacy Control (GPC)." Fetched 2026-08-19.)) **Chrome still exposes a DNT toggle.** If your crawl runs a default Chrome profile you may be sending ''DNT'' without meaning to; if it runs a current Firefox, the setting you find in the UI is GPC, not DNT. Check what your browser actually sends rather than what you assume. | It is formally dead. The W3C Tracking Protection Working Group concluded its work and republished both specifications as **W3C Working Group Notes on 17 January 2019**, saying in the status section that "there has not been sufficient deployment of these extensions (as defined) to justify further advancement".((''w3.org/TR/tracking-dnt/'', //Tracking Preference Expression (DNT)//, W3C Working Group Note 17 January 2019. Fetched 2026-08-19.)) Browsers have since diverged rather than converged: **Safari** dropped DNT alongside ITP 2.1 in 2019, and **Firefox removed the checkbox in version 135 (4 February 2025)**, whose release notes point users at "Tell websites not to sell or share my data" — which is GPC.((Mozilla, //Firefox 135.0 release notes//: "The 'Do Not Track' checkbox has been removed from preferences. If you wish to ask websites to respect your privacy, you can use the 'Tell websites not to sell or share my data' setting instead. This option is built on top of the Global Privacy Control (GPC)." Fetched 2026-08-19.)) **Chrome still exposes a DNT toggle.** If your crawl runs a default Chrome profile you may be sending ''DNT'' without meaning to; if it runs a current Firefox, the setting you find in the UI is GPC, not DNT. Check what your browser actually sends rather than what you assume. |
| ^ Tool ^ What it actually does ^ State on 2026-08-19 ^ Use it when ^ | ^ Tool ^ What it actually does ^ State on 2026-08-19 ^ Use it when ^ |
| | **Consent-O-Matic** ([[https://github.com/cavi-au/Consent-O-Matic|cavi-au/Consent-O-Matic]]), from the team behind {[nouwens2020_dark]} | Per-CMP declarative rules. The only widely used tool that can express **purpose-level** choices rather than just accept-or-dismiss | **Alive.** ''rules/'' holds **204** rule files; last commit on ''master'' **2025-11-07**; latest release **v1.1.5**, 2025-06-17 | You need ''cmp-specific-choices'', or a reject that is a real reject. Coverage is bounded by the 204 rules — everything else is untouched | | | **Consent-O-Matic** ([[https://github.com/cavi-au/Consent-O-Matic|cavi-au/Consent-O-Matic]]), from the team behind {[nouwens2020_dark]} | Per-CMP declarative rules. The only widely used tool that can express **purpose-level** choices rather than just accept-or-dismiss | **Alive.** ''rules/'' holds **204** rule files; last commit on ''master'' **2025-11-07**; latest release **v1.1.5**, 2025-06-17 | You need ''cmp-specific-choices'', or a reject that is a real reject. Coverage is bounded by the 204 rules — everything else is untouched | |
| | **autoconsent** ([[https://github.com/duckduckgo/autoconsent|duckduckgo/autoconsent]]) | A library, not an extension: detects the CMP and drives it. Ships **571** auto-generated and **331** hand-authored site rules | **Alive and the most actively maintained of the set.** Release **v16.23.0** on 2026-08-18; ''main'' committed the same day | You are embedding consent handling in your own crawler. It is a library with a stable API, which is what you want. Note the repo says the reference extension build is deliberately not published to stores — the functionality ships inside DuckDuckGo's own browsers | | | **autoconsent** ([[https://github.com/duckduckgo/autoconsent|duckduckgo/autoconsent]]) | A library, not an extension: detects the CMP and drives it. Ships **567** auto-generated and **356** hand-authored site rules (counted on 2026-09-05) | **Alive and the most actively maintained of the set.** Release **v16.37.0** on 2026-09-05 — fourteen releases in the eighteen days since this row first recorded v16.23.0, which is the rate you are committing to if you pin it | You are embedding consent handling in your own crawler. It is a library with a stable API, which is what you want. Note the repo says the reference extension build is deliberately not published to stores — the functionality ships inside DuckDuckGo's own browsers | |
| | **BannerClick** ([[https://github.com/bannerclick/bannerclick|bannerclick/bannerclick]]) {[rasaii2023_thou]} | An **[[Programming:Crawler:OpenWPM|OpenWPM]] custom command**: detect the banner, then accept or reject it, with the detection and the interaction separable | **Alive.** Default branch ''bannerclick_v0.26.0'', last commit **2025-07-01**; a ''_pets25_artifact'' tag accompanies {[rasaii2025_crumbs]} | You are already on OpenWPM and want both arms. This is the lowest-friction path to an accept/reject design | | | **BannerClick** ([[https://github.com/bannerclick/bannerclick|bannerclick/bannerclick]]) {[rasaii2023_thou]} | An **[[Programming:Crawler:OpenWPM|OpenWPM]] custom command**: detect the banner, then accept or reject it, with the detection and the interaction separable | **Alive.** Default branch ''bannerclick_v0.26.0'', last commit **2025-07-01**; a ''_pets25_artifact'' tag accompanies {[rasaii2025_crumbs]} | You are already on OpenWPM and want both arms. This is the lowest-friction path to an accept/reject design | |
| | **Priv-Accept** ([[https://github.com/marty90/priv-accept|marty90/priv-accept]]) | Selenium plus a keyword heuristic. **Accept only** — there is no reject arm | **Stale.** Last commit on ''main'' **2022-04-13**. Not archived, but four years of Selenium and ChromeDriver drift stand between you and it | You want a cheap accept-all arm and are prepared to fix it. Two papers in this corpus still use it | | | **Priv-Accept** ([[https://github.com/marty90/priv-accept|marty90/priv-accept]]) | Selenium plus a keyword heuristic. **Accept only** — there is no reject arm | **Stale.** Last commit on ''main'' **2022-04-13**. Not archived, but four years of Selenium and ChromeDriver drift stand between you and it | You want a cheap accept-all arm and are prepared to fix it. Two papers in this corpus still use it | |
| | **CookieBlock** ([[https://github.com/dibollinger/CookieBlock|dibollinger/CookieBlock]]) {[bollinger2022automating]} | Not a banner tool. It **classifies cookies by purpose and deletes the ones you rejected** — the enforcement half, not the interaction half | **Stale, and a Manifest V2 extension.** Last commit **2023-12-08**; the crawler **2023-06-03**; the published AMO build dates from 2022 | You want purpose labels for observed cookies. See [[Privacy:Cookies|Classifying Cookies]]. Do not assume the shipped extension still loads in a current Chrome | | | **CookieBlock** ([[https://github.com/dibollinger/CookieBlock|dibollinger/CookieBlock]]) {[bollinger2022automating]} | Not a banner tool. It **classifies cookies by purpose and deletes the ones you rejected** — the enforcement half, not the interaction half | **Stale, and a Manifest V2 extension.** Last commit **2023-12-08**; the crawler **2023-06-03**; the published AMO build dates from 2022 | You want purpose labels for observed cookies. See [[Privacy:Cookies|Classifying Cookies]]. Do not assume the shipped extension still loads in a current Chrome | |
| | **"I don't care about cookies"** | **Hides** banners far more often than it answers them. Acquired by Avast | **Original stale** (published build last updated 2023-11). The maintained Manifest V3 successor is the community fork [[https://github.com/OhMyGuus/I-Still-Dont-Care-About-Cookies|OhMyGuus/I-Still-Dont-Care-About-Cookies]], last commit **2026-06-21** | Almost never, in research. Hiding a banner is ''dismiss-or-remove'', not consent — see the warning below | | | **"I don't care about cookies"** | **Hides** banners far more often than it answers them. Acquired by Avast | **Original stale** (the Firefox build on addons.mozilla.org is **v3.5.0**, last updated **2023-12-06**, checked 2026-09-05). The maintained Manifest V3 successor is the community fork [[https://github.com/OhMyGuus/I-Still-Dont-Care-About-Cookies|OhMyGuus/I-Still-Dont-Care-About-Cookies]], last commit **2026-06-21** | Almost never, in research. Hiding a banner is ''dismiss-or-remove'', not consent — see the warning below | |
| | **Ninja Cookie** | Rule-driven banner rejection | **Abandoned.** The project domain is parked, the GitLab repository has been silent since **2022-02**, and the Firefox listing is gone | Never. It appears in {[demir2024_bannertools]}, which is why it is here — do not carry it forward from that paper into a 2026 crawl | | | **Ninja Cookie** | Rule-driven banner rejection | **Abandoned.** The project domain is parked, the GitLab repository has been silent since **2022-02**, and the Firefox listing is gone | Never. It appears in {[demir2024_bannertools]}, which is why it is here — do not carry it forward from that paper into a 2026 crawl | |
| | **Super Agent** | Commercial, closed-source consent automation | Live commercial product | Not as a research instrument: you cannot pin its version or read its rules | | | **Super Agent** | Commercial, closed-source consent automation | Live commercial product | Not as a research instrument: you cannot pin its version or read its rules | |
| The banner you see is chosen by the site from your IP address. Crawl a European site from a US datacentre and you will often get no banner at all, or a different one — and the TCF's own ''gdprApplies'' flag will be ''false''. This makes [[Design:Crawling location|the vantage point]] a co-determinant of every consent result, not an independent choice. | The banner you see is chosen by the site from your IP address. Crawl a European site from a US datacentre and you will often get no banner at all, or a different one — and the TCF's own ''gdprApplies'' flag will be ''false''. This makes [[Design:Crawling location|the vantage point]] a co-determinant of every consent result, not an independent choice. |
| |
| The corpus says the field mostly does not handle this. Of the **349** crawling papers that state a consent action, **68 (19.5%)** state an EU/EEA vantage point, **71 (20.3%)** state a non-EEA vantage only, and **206 (59.0%)** carry a vantage tuple whose location is //not stated//. The remaining four are 3 papers with no vantage tuple at all and 1 that names a place the geo fold cannot resolve. The picture is much better among the papers that actually clicked: of the **29** whose interaction is verified, **21 (72.4%)** did so from an EU/EEA vantage, 3 from a stated non-EEA vantage only, and 5 never said where they were. | The corpus says the field handles this better than it reports anything else about consent, but the population that reports anything at all is tiny. Of the **55** crawling papers audited as stating a consent action in their own words, **37 (67.3%)** state an EU/EEA vantage point, **7 (12.7%)** state a non-EEA vantage only, and **11 (20.0%)** carry a vantage tuple whose location is //not stated//. The split barely moves between the papers that clicked and the papers that deliberately did not: of the **32** whose interaction is verified, **22 (68.8%)** were in the EU/EEA; of the **22** that state they left the banner alone, **14 (63.6%)** were. Reading these as encouraging would be a mistake — they are conditional on a paper having said anything, which 94.4% of crawling papers did not. |
| |
| Ogut et al. {[ogut2024_dissecting]} is the paper to read on the language side of the same problem: button text is the classifier's input, and //Accept all// is //Aceptar todo// and //Alle akzeptieren// elsewhere. They found a consent notice on **37% (1511)** of successfully loaded sites worldwide. Tang et al. {[tang2025_navigating]} did the regional comparison for violations, finding at least one consent violation on **96.18% (EU)** to **97.72% (US)** of sites, with only **3.82%** enforcing preferences correctly. | Ogut et al. {[ogut2024_dissecting]} is the paper to read on the language side of the same problem: button text is the classifier's input, and //Accept all// is //Aceptar todo// and //Alle akzeptieren// elsewhere. They found a consent notice on **37% (1511)** of successfully loaded sites worldwide. Tang et al. {[tang2025_navigating]} did the regional comparison for violations, finding at least one consent violation on **96.18% (EU)** to **97.72% (US)** of sites, with only **3.82%** enforcing preferences correctly. |
| Population: the **1,120** papers that ran an automated web crawl. | Population: the **1,120** papers that ran an automated web crawl. |
| |
| ^ ''crawlConfig.consentAction'' ^ Papers ^ Share of 1,120 ^ | ^ ''crawlConfig.consentAction'' ^ Papers ^ Share of 1,120 ^ Supported by the paper ^ |
| | no interaction | 313 | 27.9% | | | no interaction | 313 | 27.9% | **22** | |
| | accept all | 15 | 1.3% | | | accept all | 15 | 1.3% | **11** | |
| | accept and reject | 14 | 1.3% | | | accept and reject | 14 | 1.3% | **14** | |
| | dismiss or remove | 3 | 0.3% | | | dismiss or remove | 3 | 0.3% | **0** | |
| | reject all | 2 | 0.2% | | | reject all | 2 | 0.2% | **2** | |
| | CMP-specific choices | 2 | 0.2% | | | CMP-specific choices | 2 | 0.2% | **2** | |
| | //not stated// | 495 | 44.2% | | | //not stated// | 495 | 44.2% | — | |
| | //not applicable// | 236 | 21.1% | | | //not applicable// | 236 | 21.1% | — | |
| | //no crawl configuration extracted// | 40 | 3.6% | | | //no crawl configuration extracted// | 40 | 3.6% | — | |
| |
| **And the six stated rows do not all survive contact with the papers.** Every one of the 36 papers in the five interacting rows was read against its own full text; **7 (19.4%) are extraction false positives** — the value fired on push-notification permission prompts, on "banner ads", and three times on IRB participant consent. The audited count of papers that verifiably interacted with a consent notice is **29**, or **2.6%** of crawling papers. The script is ''consent_action_audit.mjs'' and every verdict is on [[provenance:privacy:consent|the provenance page]]. ''no-interaction'' was not audited — 313 papers is beyond a hand pass — so treat it as an upper bound in the same way. | **None of the six stated rows survives contact with the papers, and the largest survives worst.** All 349 papers in those rows have now been read against their own full text. |
| | |
| | Among the **36** in the five interacting rows, **7 (19.4%) are extraction false positives** — the value fired on push-notification permission prompts, on "banner ads", and three times on IRB participant consent. |
| | |
| | Among the **313** in the ''no-interaction'' row, **279 (89.1%) are not supported by the paper at all**. ''no-interaction'' is not a sentinel — the same field carries ''not-stated'' and ''not-applicable'' — so it is a positive claim that the crawl left the notice alone, and in 279 cases the paper makes no such claim. 72 of them contain no consent vocabulary anywhere in either rendering of their full text; in the rest, the word the extractor could have been reading is an advertising or protocol "banner", an IRB consent form, an OAuth or cryptomining consent prompt, an industry opt-out programme, or the ISOC copyright boilerplate. A further **3 (1.0%) are false negatives**: those papers do drive a consent notice. **22 (7.0%)** say in words that they did not, and **8** say they perform no interaction at all without ever mentioning consent. |
| | |
| | Both audits are hand adjudications with one recorded verdict and reason per paper: ''consent_action_audit.mjs'' for the 36, ''consent_action_noninteraction_audit.mjs'' for the 313. Both fail loudly on an unadjudicated paper, and every phrase quoted in a verdict is machine-checked against the cited paper. The scripts, their unedited output and all 349 verdicts are on [[provenance:privacy:consent|the provenance page]]. |
| |
| The three italic rows are **sentinels, not answers**. ''not-applicable'' is a legitimate value — a crawl of an API, a mobile app store or a set of non-European sites may have no banner to handle — but ''not-stated'' at 44.2% is the finding: nearly half of all crawling papers leave the reader unable to tell which web they measured. | The three italic rows are **sentinels, not answers**. ''not-applicable'' is a legitimate value — a crawl of an API, a mobile app store or a set of non-European sites may have no banner to handle — but ''not-stated'' at 44.2% is the finding: nearly half of all crawling papers leave the reader unable to tell which web they measured. |
| |
| ==== The reporting gap is not closing, but interaction is spreading ==== | ==== Reporting started with the GDPR, from nothing, and is still rare ==== |
| | |
| | ^ Bucket ^ Crawling papers ^ Extraction: state an action ^ Share ^ **Audited: state an action** ^ **Share** ^ Of those, interacted ^ |
| | | 2010–2013 | 102 | 25 | 24.5% | **0** | **0.0%** | 0 | |
| | | 2014–2017 | 167 | 50 | 29.9% | **0** | **0.0%** | 0 | |
| | | 2018–2021 | 308 | 99 | 32.1% | **11** | **3.6%** | 3 | |
| | | 2022–2024 | 345 | 116 | 33.6% | **25** | **7.2%** | 16 | |
| | | 2025–2026* | 198 | 59 | 29.8% | **19** | **9.6%** | 13 | |
| | |
| | //* provisional venue-years — see the caveat at the top of this section.// |
| |
| ^ Bucket ^ Crawling papers ^ State an action ^ Share ^ Interaction claimed ^ Verified ^ Share verified ^ | The two halves of this table tell opposite stories, and the audited half is the one to believe. On the extraction's numbers the field looks as if it has always reported consent handling at roughly a constant 25–34%, with a dip in the newest bucket. **Not one of those four adjacent movements is statistically significant** (Fisher's exact, all //p// > 0.39) — the apparent peak and fall were noise, and an earlier revision of this page described them as a trend. |
| | 2010–2013 | 102 | 25 | 24.5% | 0 | 0 | 0.0% | | |
| | 2014–2017 | 167 | 50 | 29.9% | 0 | 0 | 0.0% | | |
| | 2018–2021 | 308 | 99 | 32.1% | 6 | 3 | 1.0% | | |
| | 2022–2024 | 345 | 116 | 33.6% | 17 | 14 | 4.1% | | |
| | 2025–2026* | 198 | 59 | 29.8% | 13 | 12 | 6.1% | | |
| |
| //* provisional venue-years.// Two separate trends. **Whether a paper says anything** moved from **24.5%** in 2010–2013 to a peak of **33.6%** in 2022–2024 and back to **29.8%** in the provisional 2025–2026 bucket — about five points net over sixteen years, and no better in the latest bucket than it was before the GDPR. This is not a solved reporting problem, and it is not obviously improving. **Whether a paper interacts** starts at exactly zero before the GDPR and rises steadily afterwards: the treatment arrived with the law, as you would expect, and is now in roughly one crawling paper in sixteen. Note that the false positives cluster in the early buckets — the extraction is most likely to mistake something else for a consent action in a paper that has nothing to do with consent, which is exactly the pre-2022 population. | On the audited numbers, **no paper before 2018 states a consent action at all**, which is what you would expect: there was mostly no banner to have a policy about. The jump at the GDPR is real (0/167 to 11/308, //p// = 0.0099; pooled 2010–2017 against 2018–2026, //p// = 3.0 × 10⁻⁷). What comes after it is a rise you should not lean on: 3.6% to 7.2% is //p// = 0.058, and 7.2% to 9.6% is //p// = 0.33 in a bucket that is provisional anyway. The honest reading is that **consent reporting went from nonexistent to rare when the law arrived, and has stayed rare** — under one crawling paper in ten, sixteen years in. The test script is ''consent_ni_significance.py''; both series and every //p// are on the provenance page. |
| |
| ==== The consent literature itself ==== | ==== The consent literature itself ==== |
| * **Free-text names were folded before counting, and the residue is printed.** The consent-tool fold leaves 5 distinct unmapped names over 5 papers, all of them IAB artefacts that are not banner-interaction tools (''IAB ads.txt crawler'', ''IAB anti-ad-block script'', the IAB content taxonomy). The law fold leaves 2 (''Digital Economy Act 2017'', ''Act against Unfair Competition (UWG)''). The vantage fold leaves 5 strings that name no place (''different continents'', ''various geographic regions''). | * **Free-text names were folded before counting, and the residue is printed.** The consent-tool fold leaves 5 distinct unmapped names over 5 papers, all of them IAB artefacts that are not banner-interaction tools (''IAB ads.txt crawler'', ''IAB anti-ad-block script'', the IAB content taxonomy). The law fold leaves 2 (''Digital Economy Act 2017'', ''Act against Unfair Competition (UWG)''). The vantage fold leaves 5 strings that name no place (''different continents'', ''various geographic regions''). |
| * **Every per-paper figure on this page was checked against the paper's own text**, not against the extraction's summary of it. The check covers 61 literals — a superset of what is published, since ten were checked and then cut — and found 60 in both the column-repaired and the plain rendering, 1 in the column-repaired rendering only, and **0 not found**. That pass caught two errors in an earlier draft of this page — a figure attributed to Bouhoula et al. that the paper writes without a thousands separator, and a Matte et al. percentage this page had rounded to the wrong decimal. | * **Every per-paper figure on this page was checked against the paper's own text**, not against the extraction's summary of it. The check covers 61 literals — a superset of what is published, since ten were checked and then cut — and found 60 in both the column-repaired and the plain rendering, 1 in the column-repaired rendering only, and **0 not found**. That pass caught two errors in an earlier draft of this page — a figure attributed to Bouhoula et al. that the paper writes without a thousands separator, and a Matte et al. percentage this page had rounded to the wrong decimal. |
| * **''consentAction'' was audited paper by paper, and it needed to be.** The schema's stability comparison puts it in the reliable band — an independent extraction run over the same text agrees with it on 93% of papers — but that measures whether two runs agree, not whether either is right, and it was measured on the earlier 4,322-paper corpus. Reading all 36 interacting papers found **7 false positives (19.4%)**, concentrated in ''accept-all'' and ''dismiss-or-remove''. **Note also that the field's own evidence quote cannot catch this**: ''crawlConfig'' carries one quote for the whole configuration object, so the quote behind a ''consentAction'' value usually evidences statefulness or crawl depth instead. Spot-checking quotes, which is the standard check on this site, is structurally blind here. | * **''consentAction'' was audited paper by paper, in both directions, and it needed to be.** The schema's stability comparison puts it in the reliable band — an independent extraction run over the same text agrees with it on 93% of papers — but that measures whether two runs agree, not whether either is right, and it was measured on the earlier 4,322-paper corpus. Reading all 36 interacting papers found **7 false positives (19.4%)**, concentrated in ''accept-all'' and ''dismiss-or-remove''. Reading all 313 ''no-interaction'' papers found **279 (89.1%) that make no claim about their own crawl's consent action** and **3 (1.0%) that in fact drove a notice**. The two error modes are different in kind: on the interacting side the extractor misreads a word, on the ''no-interaction'' side it supplies a default where the paper is silent — which is why the second error is twenty times more common and matters more, since it is what inflates every "share of papers that report X" figure. **Note also that the field's own evidence quote cannot catch either**: ''crawlConfig'' carries one quote for the whole configuration object, so the quote behind a ''consentAction'' value usually evidences statefulness or crawl depth instead. Spot-checking quotes, which is the standard check on this site, is structurally blind here. |
| * **Venue coverage.** Seven venues only. **CHI, SOUPS, EuroS&P, ACSAC, RAID, AsiaCCS and WPES are absent**, and that bites harder on this page than on most: the usable-privacy half of the consent literature (Nouwens et al., Habib et al., Utz et al.'s follow-ups) is largely CHI and SOUPS work. Every count here is a lower bound. | * **Venue coverage.** Seven venues only. **CHI, SOUPS, EuroS&P, ACSAC, RAID, AsiaCCS and WPES are absent**, and that bites harder on this page than on most: the usable-privacy half of the consent literature (Nouwens et al., Habib et al., Utz et al.'s follow-ups) is largely CHI and SOUPS work. Every count here is a lower bound. |
| |
| * **Consent revocation is nearly unstudied.** One paper {[kancherla2025_johnny]}, 158 sites. Withdrawal is as legally required as consent and is far harder to automate. | * **Consent revocation is nearly unstudied.** One paper {[kancherla2025_johnny]}, 158 sites. Withdrawal is as legally required as consent and is far harder to automate. |
| * **No shared benchmark exists.** There is no public, versioned set of annotated consent notices that a new detector can report against, which is why every paper reports precision and recall on its own hand-labelled sample and none of them are comparable. Building one would be a bigger contribution than most new detectors. | * **No shared benchmark exists.** There is no public, versioned set of annotated consent notices that a new detector can report against, which is why every paper reports precision and recall on its own hand-labelled sample and none of them are comparable. Building one would be a bigger contribution than most new detectors. |
| * **The reporting gap itself.** 44.2% of crawling papers say nothing about consent, and the share that says //something// is no higher in 2025–2026 than it was in 2014–2017. A one-line methods sentence would fix it; the question is why sixteen years of the field have not produced one. | * **The reporting gap itself.** On the audited figures, **94.4% of crawling papers leave the reader unable to tell what their crawl did with a notice**, and the share that does say is under one in ten even in the newest bucket. A one-line methods sentence would fix it; the question is why sixteen years of the field have not produced one. (The 44.2% ''not-stated'' row is the extraction's own count for that population and was **not** audited — after what the ''no-interaction'' audit found, no unaudited row on this page should be read as a claim about the papers.) |
| * **How many of the 313 ''no-interaction'' papers really did not interact** is unknown. The 36 papers claiming an interaction were read one by one and 7 turned out not to have interacted; nobody has done the same in the other direction, and 313 is too many for a hand pass. If the error is symmetric, the true count is somewhere either side of 29 — which is a reason to treat every figure in this section as an order of magnitude and to read the papers you actually compare yourself against. | * **All 349 papers the extraction credits with a stated consent action have now been read**, so the figures above are hand verdicts rather than extractor output. What that cannot fix is the other direction: a paper that clicked a banner and never wrote it down is invisible here by construction, and the 495 ''not-stated'' and 236 ''not-applicable'' papers were **not** read — a false ''not-stated'' would be a further undercount. **32 is a floor on the true number of interacting papers, not an estimate of it.** Two of the three false negatives were found only because a named tool (BannerClick, CookieEnforcer) appears in the text; a paper that rolled its own clicking and described it in one unremarkable sentence would still be missed. |
| </WRAP> | </WRAP> |
| |
| * [[Privacy:Requests#Cookie Notices and Their Interactive Elements|Classifying Web Requests]] — detecting the notice and labelling its buttons, with the comparison table of detectors. | * [[Privacy:Requests#Cookie Notices and Their Interactive Elements|Classifying Web Requests]] — detecting the notice and labelling its buttons, with the comparison table of detectors. |
| * [[Privacy:Cookies|Classifying Cookies]] — what the cookies you observe before and after the click actually are, and the CookieBlock/Cookiepedia label sources. | * [[Privacy:Cookies|Classifying Cookies]] — what the cookies you observe before and after the click actually are, and the CookieBlock/Cookiepedia label sources. |
| | * [[Privacy:Policies|Measuring Privacy Policies and Terms]] — the long document behind the banner. A banner is a UI measurement; a policy is a document-retrieval and NLP measurement, with its own tool lineage and its own denominators, and the two literatures barely cite each other. |
| * [[Privacy:TCF Consent Strings|Decoding TCF Consent Strings]] — the TC string and Google's Additional Consent string as artefacts: bit layout, where they hide, decoding them reproducibly, and what they do and do not prove. | * [[Privacy:TCF Consent Strings|Decoding TCF Consent Strings]] — the TC string and Google's Additional Consent string as artefacts: bit layout, where they hide, decoding them reproducibly, and what they do and do not prove. |
| * [[Programming:Interaction|Interaction with websites]] and [[Programming:Stateful Stateless|Stateful and stateless crawling]] — the crawler-side mechanics this page assumes. | * [[Programming:Interaction|Interaction with websites]] and [[Programming:Stateful Stateless|Stateful and stateless crawling]] — the crawler-side mechanics this page assumes. |
| * [[Programming:Crawler|Comparison of crawling libraries]] — which crawlers ship a consent-interaction step. | * [[Programming:Crawler|Comparison of crawling libraries]] — which crawlers ship a consent-interaction step. |
| * [[Design:Crawling location|Crawling location]] — why the vantage point changes which banner you get. | * [[Design:Crawling location|Crawling location]] — why the vantage point changes which banner you get. |
| | * [[Privacy:Age assurance|Age assurance]] — the other dismissable overlay on the first load, and one that usually writes a cookie //before// the banner is touched, so it is a confounder for anything counted pre-consent. |
| * [[Practices:Legal enforcement|Legal enforcement]] — what to do with a violation once you have found one, and which authority is competent. | * [[Practices:Legal enforcement|Legal enforcement]] — what to do with a violation once you have found one, and which authority is competent. |
| * [[Practices:Ethics|Ethics]] — clicking //Accept// at scale on behalf of nobody is a decision with an ethical dimension; it is discussed there. | * [[Practices:Ethics|Ethics]] — clicking //Accept// at scale on behalf of nobody is a decision with an ethical dimension; it is discussed there. |