User Tools

Site Tools


privacy:darkpatterns

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
privacy:darkpatterns [2026/08/21 02:54] – Fix: separate Nayak et al. F1 metrics from the 11,118-site deployment corpus (reviewer finding). Authored by Claude karel.kubicek.claudeprivacy:darkpatterns [2026/08/21 02:57] (current) – Apply generic review: correct 5859->5855 denominator, unsupported 2025 agent-targeting claim, mislabelled share-vs-rate rows, method-table periods, cite sun2026 in the agent section. Authored by Claude karel.kubicek.claude
Line 4: Line 4:
  
   * **As the thing you measure.** If you are auditing consent notices, cancellation flows or data-rights portals, the asymmetry //is// the result. Counting it means committing to a taxonomy, an operationalisation per category, and a denominator — and the field's operationalisations differ enough that two papers reporting "interface interference" are often not measuring the same thing.   * **As the thing you measure.** If you are auditing consent notices, cancellation flows or data-rights portals, the asymmetry //is// the result. Counting it means committing to a taxonomy, an operationalisation per category, and a denominator — and the field's operationalisations differ enough that two papers reporting "interface interference" are often not measuring the same thing.
-  * **As a confounder in every other crawl.** A banner engineered to make refusal expensive is also engineered against your crawlerand since 2025 explicitly against LLM-driven agents. If your pipeline clicks "Accept" because the accept button was the only one it could find, the pattern has become part of your instrument.+  * **As a confounder in every other crawl.** A banner engineered to make refusal expensive is also, incidentally, engineered against your crawler — and two 2026 papers show LLM-driven agents are measurably susceptible to the same designs. If your pipeline clicks "Accept" because the accept button was the only one it could find, the pattern has become part of your instrument.
  
 <WRAP important> <WRAP important>
Line 14: Line 14:
 There is no single legal or academic definition, and this is the central methodological problem: a prevalence figure is uninterpretable without the taxonomy that produced it. There is no single legal or academic definition, and this is the central methodological problem: a prevalence figure is uninterpretable without the taxonomy that produced it.
  
-The current answer to "which taxonomy" is **Gray et al.'s ontology** {[gray2024_ontology]}, which harmonises ten prior regulatory and academic taxonomies into 65 pattern types across three levels of abstraction — //high-level// strategies, //meso-level// angles of attack, //low-level// concrete patterns. It is the right starting point in 2026 for two reasons. First, it is a mapping rather than a competitor: if your data was labelled with Mathur et al.'s categories {[mathur2019_scale]} or with the taxonomy embedded in a regulator's guidance, the ontology tells you where those labels sit. Second, its low/meso/high split is exactly the axis on which measurement papers talk past each other — an automated crawl detects low-level patterns (a pre-ticked box, a greyed-out button), while a user study measures high-level effects (did the design change the decision), and reporting one as the other is the commonest overclaim in this area.+The current answer to "which taxonomy" is **Gray et al.'s ontology** {[gray2024_ontology]}, which harmonises ten prior regulatory and academic taxonomies into 65 pattern types across three levels of abstraction — //high-level// strategies, //meso-level// angles of attack, //low-level// concrete patterns. It is the right starting point in 2026 for two reasons. First, it is a mapping rather than a competitor: if your data was labelled with Mathur et al.'s categories {[mathur2019_scale]} or with the taxonomy embedded in a regulator's guidance, the ontology tells you where those labels sit. Second, its low/meso/high split is exactly the axis on which measurement papers talk past each other — an automated crawl detects low-level patterns (a pre-ticked box, a greyed-out button), while a user study measures high-level effects (did the design change the decision), and reporting one as the other is a common overclaim in this area.
  
 Older taxonomies still worth knowing, and their status: Older taxonomies still worth knowing, and their status:
Line 23: Line 23:
 | Gray et al., ontology of dark patterns knowledge {[gray2024_ontology]} | 2024 | **Current default.** Use it, or state your mapping onto it. | | Gray et al., ontology of dark patterns knowledge {[gray2024_ontology]} | 2024 | **Current default.** Use it, or state your mapping onto it. |
  
-Legally, the position as of August 2026 is still fragmentation, and the page you are writing should not claim otherwise. The DSA prohibits dark patterns on online platforms in Article 25 and describes them in recital 67, but Article 25 explicitly steps aside where the UCPD or GDPR applies; the UCPD prohibits "misleading" and "aggressive" practices without using the term and without addressing digital interfaces in its Annex I; the GDPR does not name them at all, though consent-obtaining techniques can be assessed under it.((European Parliamentary Research Service, //Regulating dark patterns in the EU: Towards digital fairness//, PE 767.191, January 2025.)) A Digital Fairness Act intended to consolidate this was announced in the Commission's 2026 work programme for Q4 2026 and, as of this writing, **has not been tabled** — so do not cite it as law.((European Parliament Legislative Train Schedule, //Digital Fairness Act//, status "Announced", CWP indicative date Q4 2026; checked 2026-08-21.))+Legally, the position as of August 2026 is still fragmentation, and your paper should not claim otherwise. The DSA prohibits dark patterns on online platforms in Article 25 and describes them in recital 67, but Article 25 explicitly steps aside where the UCPD or GDPR applies; the UCPD prohibits "misleading" and "aggressive" practices without using the term and without addressing digital interfaces in its Annex I; the GDPR does not name them at all, though consent-obtaining techniques can be assessed under it.((European Parliamentary Research Service, //Regulating dark patterns in the EU: Towards digital fairness//, PE 767.191, January 2025.)) A Digital Fairness Act intended to consolidate this was announced in the Commission's 2026 work programme for Q4 2026 and, as of this writing, **has not been tabled** — so do not cite it as law.((European Parliament Legislative Train Schedule, //Digital Fairness Act//, status "Announced", CWP indicative date Q4 2026; checked 2026-08-21.))
  
 ===== What the S&P corpus has actually measured ===== ===== What the S&P corpus has actually measured =====
  
-Of the 5,859 papers with extracted full text, **48 discuss dark patterns substantively** (at least five occurrences of //dark pattern// / //deceptive design// / //deceptive pattern// / //manipulative design// in the full text, whitespace-collapsed); a further 83 mention the term only in passing or in related work. That 48 is the honest denominator for anything on this page. Its year distribution — 2 papers before 2020, 10 in 2020–2022, 18 in 2023–2024, 18 in 2025–2026 (both provisional venue-years) — says the topic arrived in security venues late and is growing.+Of the 5,855 papers whose ''paper.cols.txt'' the probe could read (4 of the 5,859 extracted papers have no readable full text and are silently counted as negatives here), **48 discuss dark patterns substantively** (at least five occurrences of //dark pattern// / //deceptive design// / //deceptive pattern// / //manipulative design// in the full text, whitespace-collapsed); a further 83 mention the term only in passing or in related work. That 48 is the honest denominator for anything on this page. Its year distribution — 2 papers before 2020, 10 in 2020–2022, 18 in 2023–2024, 18 in 2025–2026 (both provisional venue-years) — says the topic arrived in security venues late and is growing.
  
 Almost all of it is about **consent notices**. The measured figures below are quoted with each paper's own denominator; note how different the denominators are, which is why the percentages should never be averaged. Almost all of it is about **consent notices**. The measured figures below are quoted with each paper's own denominator; note how different the denominators are, which is why the percentages should never be averaged.
Line 39: Line 39:
 | Forced action (non-notice links unreachable before interacting) | 46.4% | 48,843 websites with a cookie notice | {[bouhoula2024_automated]} | | Forced action (non-notice links unreachable before interacting) | 46.4% | 48,843 websites with a cookie notice | {[bouhoula2024_automated]} |
 | Mobile consent dialogs violating at least one design requirement | 98.8% (429) | apps showing a //proper// consent dialog — itself only 11.9% of apps analysed | {[koch2023_enough]} | | Mobile consent dialogs violating at least one design requirement | 98.8% (429) | apps showing a //proper// consent dialog — itself only 11.9% of apps analysed | {[koch2023_enough]} |
-Deceptive patterns detected per site (VLM-based) | 49.02% (375 of 765) | accessible websites reached by the crawl | {[shi2025_shades]} | +Websites with at least one deceptive pattern detected (vision-model pipeline) | 49.02% (375 of 765) | accessible websites reached by the crawl | {[shi2025_shades]} | 
-Deceptive patterns in mobile apps (same pipeline) | 25.68% (246 of 958) | apps from which screenshots could be collected | {[shi2025_shades]} |+Apps with at least one deceptive pattern detected (same pipeline) | 25.68% (246 of 958) | apps from which screenshots could be collected | {[shi2025_shades]} |
  
 Two things to take from that table rather than from any single row. The first is that **the denominator does most of the work**: "67.8% show interface interference" is a statement about sites that offered a reject option, and the sites that offered none are excluded from it, so the figure understates the population-level asymmetry rather than overstating it. The second is that the effect-side evidence is thinner and comes from experiments, not crawls: a randomised banner study found refusal rates moving from 17% on a control banner to 34% with a highlighted decline button and 47% when consequences were spelled out {[bielova2024_designpatterns]}, and the effect partly persisted when participants later saw a neutral banner. Prevalence and effect are different claims; the corpus supports the first much better than the second. Two things to take from that table rather than from any single row. The first is that **the denominator does most of the work**: "67.8% show interface interference" is a statement about sites that offered a reject option, and the sites that offered none are excluded from it, so the figure understates the population-level asymmetry rather than overstating it. The second is that the effect-side evidence is thinner and comes from experiments, not crawls: a randomised banner study found refusal rates moving from 17% on a control banner to 34% with a highlighted decline button and 47% when consequences were spelled out {[bielova2024_designpatterns]}, and the effect partly persisted when participants later saw a neutral banner. Prevalence and effect are different claims; the corpus supports the first much better than the second.
Line 48: Line 48:
 Dating these matters, because the cheap method and the current method are not the same one. Dating these matters, because the cheap method and the current method are not the same one.
  
-^ Method ^ Period ^ Status ^ +^ Method ^ Years it is used in the corpus ^ Status ^ 
-| Manual audit of screenshots against a coded variable list | 2016–2020 | **Still the ground truth.** {[utz2019_informed]}, {[matte2020_cookie]}, {[toth2022_darkpatterns]}. Does not scale, and remains what everything else is validated against. |+| Manual audit of screenshots against a coded variable list | 2016–2022, and still used for validation after | **Still the ground truth.** {[utz2019_informed]}, {[matte2020_cookie]}, {[toth2022_darkpatterns]}. Does not scale, and remains what everything else is validated against. |
 | DOM and CSS heuristics: pre-ticked ''checked'' attributes, computed colour/contrast of button pairs, text-style comparison, reachability of links behind the overlay | 2022–2024 | **Current practice for low-level patterns at scale**, and the highest-precision option for the narrow set of patterns it covers — {[bouhoula2024_automated]} report a 0.0% false-positive rate for dark patterns against 500 hand-annotated sites, versus 9.4% for the harder privacy-violation judgements in the same pipeline. Cannot see anything requiring semantics. | | DOM and CSS heuristics: pre-ticked ''checked'' attributes, computed colour/contrast of button pairs, text-style comparison, reachability of links behind the overlay | 2022–2024 | **Current practice for low-level patterns at scale**, and the highest-precision option for the narrow set of patterns it covers — {[bouhoula2024_automated]} report a 0.0% false-positive rate for dark patterns against 500 hand-annotated sites, versus 9.4% for the harder privacy-violation judgements in the same pipeline. Cannot see anything requiring semantics. |
-| Regular expressions / keyword matching over dialog text | 2023 | **Superseded for anything but a first filter.** Used at scale for mobile dialogs {[koch2023_enough]}; brittle across languages, and the language question is on [[programming:multilingual_support|Multilingual support]]. |+| Regular expressions / keyword matching over dialog text | 2023 (one paper at scale) | **Superseded for anything but a first filter.** Used at scale for mobile dialogs {[koch2023_enough]}; brittle across languages, and the language question is on [[programming:multilingual_support|Multilingual support]]. |
 | Vision models over screenshots: object detection to localise UI elements, then a multimodal model to classify | 2025–2026 | **Current for the categories heuristics cannot reach.** {[shi2025_shades]} (YOLO-family detector plus a multimodal classifier, 65-type taxonomy); {[nayak2025_automatically]} report F1 0.93 for deceptive-pattern detection and F1 0.91 for UI-element localisation on 5,879 elements, then deploy the pipeline over a separate corpus of 11,118 websites (6,626 Tranco domains plus 4,492 Shopify stores) — note that the 11,118 is the deployment corpus, **not** the set the F1 was computed on. This is where the field is moving, and it is also the least externally validated: both are 2025 results in the provisional slice of the corpus. | | Vision models over screenshots: object detection to localise UI elements, then a multimodal model to classify | 2025–2026 | **Current for the categories heuristics cannot reach.** {[shi2025_shades]} (YOLO-family detector plus a multimodal classifier, 65-type taxonomy); {[nayak2025_automatically]} report F1 0.93 for deceptive-pattern detection and F1 0.91 for UI-element localisation on 5,879 elements, then deploy the pipeline over a separate corpus of 11,118 websites (6,626 Tranco domains plus 4,492 Shopify stores) — note that the 11,118 is the deployment corpus, **not** the set the F1 was computed on. This is where the field is moving, and it is also the least externally validated: both are 2025 results in the provisional slice of the corpus. |
 | LLM-driven agents that //traverse// a workflow and classify what they encounter | 2026 | **Emerging, and not yet a measurement instrument you should trust unsupervised.** {[sun2026_suitability]} audited CCPA data-rights workflows across 456 data brokers, verifying agent completion at 87% and 79% in two phases, and got category prevalence estimates spanning 15.2%–48.6% across eight categories — wide intervals that are the point of the paper, not a footnote to it. | | LLM-driven agents that //traverse// a workflow and classify what they encounter | 2026 | **Emerging, and not yet a measurement instrument you should trust unsupervised.** {[sun2026_suitability]} audited CCPA data-rights workflows across 456 data brokers, verifying agent completion at 87% and 79% in two phases, and got category prevalence estimates spanning 15.2%–48.6% across eight categories — wide intervals that are the point of the paper, not a footnote to it. |
Line 61: Line 61:
 This is the part a fresh PhD student is least likely to expect, and it is the newest material in the corpus. This is the part a fresh PhD student is least likely to expect, and it is the newest material in the corpus.
  
-If your crawler interacts with consent notices — see [[privacy:consent|Granting consent to websites]] and [[programming:interaction|Interaction with websites]] — then interface asymmetry is a systematic bias in your treatment assignment, not background noise. A pipeline that clicks the visually dominant button reproduces the operator's intended outcome and will report a consent rate that is a property of your heuristic. Two 2026 papers make this concrete for agentic crawlersLLM-based web agents followed the dark pattern rather than the task an average of **41%** of the time when a single pattern was present, with susceptibility varying strongly by category and by agentand prompt-level countermeasures recovering only part of it {[ersoy2026_investigating]}.+If your crawler interacts with consent notices — see [[privacy:consent|Granting consent to websites]] and [[programming:interaction|Interaction with websites]] — then interface asymmetry is a systematic bias in your treatment assignment, not background noise. A pipeline that clicks the visually dominant button reproduces the operator's intended outcome and will report a consent rate that is a property of your heuristic. Two 2026 papers make this concrete for agentic crawlersLLM-based web agents followed the dark pattern rather than the task an average of **41%** of the time when a single pattern was present, with susceptibility varying strongly by category and by agent — per-agent rates reach 72.3% — and prompt-level countermeasures recovering only part of it {[ersoy2026_investigating]}. And when an agent is the auditing instrument rather than the subject, its category estimates spread across 15.2%–48.6% depending on configuration {[sun2026_suitability]}, which is the same problem seen from the measurement side.
  
 Practical consequences for a crawl: Practical consequences for a crawl:
privacy/darkpatterns.txt · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki