| Both sides previous revisionPrevious revision | |
| provenance:design:algorithm_audits [2026/09/11 11:02] – Reviewer findings from four passes; correct the probe-2 residue (146->158) and the probe-7 narrative; add the recall-repair sweeps and the shared-denominator judgement call. Authored by Claude karel.kubicek.claude | provenance:design:algorithm_audits [2026/09/11 11:21] (current) – Generic review log (9 findings, all accepted); the null adjudication and its definition; verdicts for all 58 shortlisted papers; residue re-derived (157/391). Authored by Claude karel.kubicek.claude |
|---|
| ^ Question ^ Answer ^ | ^ Question ^ Answer ^ |
| | Does the wiki already cover it? | No. [[Programming:Stateful stateless]] owns the profile axis and lists personalisation as a phenomenon that needs a stateful design; [[Statistics:Hypothesis testing]] owns the test; [[Design:Platforms]] owns sock puppets as an **access route** and says in its own Open Questions that //"Search-engine and ads-ecosystem auditing has no page on this wiki"//. Nobody owned the experimental design. | | | Does the wiki already cover it? | No. [[Programming:Stateful stateless]] owns the profile axis and lists personalisation as a phenomenon that needs a stateful design; [[Statistics:Hypothesis testing]] owns the test; [[Design:Platforms]] owns sock puppets as an **access route** and says in its own Open Questions that //"Search-engine and ads-ecosystem auditing has no page on this wiki"//. Nobody owned the experimental design. | |
| | How large is the in-corpus population? | **32** papers under a written inclusion rule (below), of 5,859. Twelve further candidates were read and rejected. | | | How large is the in-corpus population? | **34** papers under a written inclusion rule (below), of 5,859. **36** further candidates were read and rejected. The figure was 32 until a generic review pass found that 26 shortlisted papers had been read and dropped without a written verdict, and that one of them satisfied the rule. | |
| | Is it growing or historical? | Growing: 11 papers in 2020–2023 and 10 in 2024–2026*, against 6 in 2010–2015. | | | Is it growing or historical? | Not historical. Per year: 1.00 (2010–2015), 1.25 (2016–2019), 3.00 (2020–2023), 3.67 (2024–2026*). //n// = 34 will not carry a trend claim and the page does not make one; what it says is that more of this work appeared in the last six years than in the first ten. | |
| | Could it be a section instead? | It could have been ~3 KB on [[Design:Automated measurements]]. It was not, for two reasons: (a) that page is a **routing** page between crawl / scan / app, and an audit is a fourth instrument that cuts across the crawl branch rather than sitting beside it; (b) the material that makes the page worth writing — control arms, carry-over, the noise floor — is design advice, not routing, and would have doubled the length of a page whose job is to be short. | | | Could it be a section instead? | It could have been ~3 KB on [[Design:Automated measurements]]. It was not, for two reasons: (a) that page is a **routing** page between crawl / scan / app, and an audit is a fourth instrument that cuts across the crawl branch rather than sitting beside it; (b) the material that makes the page worth writing — control arms, carry-over, the noise floor — is design advice, not routing, and would have doubled the length of a page whose job is to be short. | |
| | The counter-argument | The 32 is a **lower bound with a known bias** (see //The screening loss// below), so a reader could reasonably say the page is built on a population the corpus cannot see properly. That is stated on the page itself, in its own box, rather than buried here. | | | The counter-argument | The 32 is a **lower bound with a known bias** (see //The screening loss// below), so a reader could reasonably say the page is built on a population the corpus cannot see properly. That is stated on the page itself, in its own box, rather than buried here. | |
| |
| - **(T) Treatment.** It deliberately varies a property of the measuring identity or request — profile history, declared attribute, location, device, opt-out setting, ad creative — and holds the rest fixed. | - **(T) Treatment.** It deliberately varies a property of the measuring identity or request — profile history, declared attribute, location, device, opt-out setting, ad creative — and holds the rest fixed. |
| - **(O) Outcome.** What it measures is the platform's own response: ads served, results ranked, prices quoted, feed or recommendation contents, or an access decision. | - **(O) Outcome.** What it measures is the platform's own **discriminating** response to the identity it has built: ads served, results ranked, prices quoted, feed or recommendation contents, or the profile the platform reports back. **Not reachability.** A study whose outcome is whether you can reach the site at all — geoblocking, censorship, Tor-exit refusal — varies a **vantage point** rather than an identity and belongs to [[Design:Blocking and geodifference]]. This clause was tightened on 2026-09-11 after a review pass; it excludes no paper already in the set, and it is what keeps //403 Forbidden: A Global View of CDN Geoblocking// (IMC 2018) out. Without it, every geodifference study in the corpus would be an algorithm audit and the page would duplicate a neighbour. |
| - **(C) Contrast.** The result is a difference (or a bounded absence of difference) **between arms**, not a prevalence over a crawl of many sites. | - **(C) Contrast.** The result is a difference (or a bounded absence of difference) **between arms**, not a prevalence over a crawl of many sites. |
| |
| ===== The adjudication ===== | ===== The adjudication ===== |
| |
| 108 candidates scored ≥3 across the probes were shortlisted with their abstracts (''_aa_short.mjs''). The 58 scoring ≥4, plus 17 lower-scored papers flagged by hand, were passed through ''_aa_adj.mjs'', which prints every sentence in the paper matching an arm/treatment/control/persona pattern. Each verdict was then made by reading those sentences, and where they were ambiguous, by grepping the paper's methods section directly. | 108 candidates scored ≥3 across the probes were shortlisted with their abstracts (''_aa_short.mjs''). The 58 scoring ≥4 were passed through ''_aa_adj.mjs'', which prints every sentence in the paper matching an arm/treatment/control/persona pattern; a further 12 lower-scored or probe-7 papers were read directly. Each verdict was made by reading those sentences, and where they were ambiguous, by grepping the paper's methods section. |
| | |
| | **A generic review pass found the weak point here.** 26 of the 58 had been read and dropped with no written verdict, and one of them — //Cart-ology// (CCS 2022, shortlist score 7) — satisfies the rule: five browser profiles, four of them unused baselines, //"created and mechanistically measured in the same way, with separation between attacker, victim, and baselines on different machines with different IPs"//, read out as ad distributions. It had been rejected on the strength of ''_aa_adj.mjs'' returning **zero** sentences for it — a 420-character sentence cap meeting a column-spliced PDF. **A probe returning nothing is not a negative result**, and it was treated as one. All 58 now carry a written verdict, and the IMC 2024 filter-bubble poster was promoted at the same time (four bot arms differing only in video-selection strategy). The population moved from 32 to **34**. |
| |
| **Every verdict, with the sentence that settled it, is in the script output below** (''--list''). The script refuses to run if any entry lacks an adjudication note. | **Every verdict, with the sentence that settled it, is in the script output below** (''--list''). The script refuses to run if any entry lacks an adjudication note. |
| |
| * **''statistics.method''** is free text and ~20% stable run-to-run. It is folded to an alphanumeric skeleton (lowercase, non-alphanumerics stripped) and **paper-counted**, and it is published as a **ranking**, not as percentages. The fold does not merge synonyms: ''Holm-Bonferroni correction'' (3), ''Holm-Bonferroni'' (1) and ''Holm-Bonferroni method'' (1) are three rows in the raw output. An unfolded reading would publish "Holm–Bonferroni 3"; the true paper count for the Holm–Bonferroni family is **5 of 32**, and for any Bonferroni-family correction **7 of 32**. The report script now prints the hand-folded family counts alongside the raw skeleton ranking, so the page quotes a number the script produced rather than one assembled in prose. Both are in the output below. | * **''statistics.method''** is free text and ~20% stable run-to-run. It is folded to an alphanumeric skeleton (lowercase, non-alphanumerics stripped) and **paper-counted**, and it is published as a **ranking**, not as percentages. The fold does not merge synonyms: ''Holm-Bonferroni correction'' (3), ''Holm-Bonferroni'' (1) and ''Holm-Bonferroni method'' (1) are three rows in the raw output. An unfolded reading would publish "Holm–Bonferroni 3"; the true paper count for the Holm–Bonferroni family is **5 of 32**, and for any Bonferroni-family correction **7 of 32**. The report script now prints the hand-folded family counts alongside the raw skeleton ranking, so the page quotes a number the script produced rather than one assembled in prose. Both are in the output below. |
| * **The eight noise-baseline phrasings** are not a fold but a deliberately widened probe (''_aa_noise.mjs''): the narrow term "A/A test" returns 2 papers corpus-wide, so seven further phrasings were added. The count rose from 0 to 11 of 32. Residue is not applicable — the probe reports term presence, and the page states explicitly that a paper can build a null without any of the eight phrasings. | * **The eight noise-baseline phrasings** are not a fold but a deliberately widened probe (''_aa_noise.mjs''): the narrow term "A/A test" returns 2 papers corpus-wide, so seven further phrasings were added, taking the count from 0 to 11. **That probe is now published as a //vocabulary// finding only**, not as a practice rate — see the next section. |
| | |
| | **Residue of the candidate probes.** Probe 2 returned **190** candidates. **28** of the final 34 audits and **5** of the 36 rejections lie inside it — the rest were caught by probes 4–7 — leaving a residue of **157** papers dropped at title-and-abstract level without an individual note. The union across all seven probes is **461**, of which **70** were adjudicated in depth, so **391** rest on a title-level read. That is the honest residue of this page. |
| | |
| | ===== The null: why it is adjudicated and not counted ===== |
| | |
| | The page's central claim is that establishing a null is the skipped step. The first version of this log said that claim rested on //"a term count plus an eight-way concept probe, never as 'nobody does this'"//, and the page then went ahead and wrote **"only 11 of 32 establish a same-treatment baseline at all"** — which is a practice claim, from a term probe, three times over. A generic review pass caught the gap between the caveat and the sentences. |
| | |
| | It was right, and the probe is wrong **in both directions**: |
| | |
| | ^ Paper ^ Term that matched ^ What the sentence actually says ^ |
| | | {[robertson2018_auditing]} | ''noise floor'' | //"Hannak et al. found (1) evidence of general personalization above a noise floor"// — a description of somebody else's result. This paper pairs standard and incognito windows; it has no same-treatment pair | |
| | | {[iqbal2022_left]} | ''same treatment'' | //"whether the SFA of any given email service provided same treatment to similar emails from candidates of different political affiliations"// — the research question, not a baseline | |
| | | {[oh2022_cartology]} | **none of the eight** | four unused baseline profiles //"created and mechanistically measured in the same way"//, with the divergence of one read as the noise. The clearest A/A in the corpus, invisible to the probe | |
| | |
| | So the question was put to all 34 papers by hand, the same way the population was, with a stated definition and an evidence sentence per verdict, in ''NULLS'' in ''algorithm_audits_set.mjs''. **A paper establishes a null if it measures outcome variation under the same treatment**, by any of three routes: **A/A** (arms differing in nothing), **repeats** (one condition run several times with the spread reported), or **generated** (a null distribution built from the paper's own observations — permutation over arm labels, or a randomised baseline). |
| | |
| | Result: **10 of 34 (29.4%)**, against 11 of 32 from the probe. The two numbers are close and that is a coincidence, not corroboration: the sets differ by five papers. |
| |
| **Residue of the candidate probes.** Probe 2 returned **190** candidates. **27** of the final 32 audits and **5** of the 12 rejections lie inside it — the other 5 and 7 were caught by probes 4–7 — leaving a residue of **158** papers dropped at title-and-abstract level without an individual note. The union across all seven probes is **461**, of which **44** were adjudicated in depth, so **417** rest on a title-level read. That is the honest residue of this page. | What this is not: |
| | * **Not mechanical.** It is a judgement per paper. {[mai2025_more]} is counted in on //repeats// — six runs of each condition a week apart, with outliers cut on the country-specific standard deviation — and a second reader could reasonably say that is replication rather than a null. {[roongta2025_sheep]} is counted **out**, because its fresh control browser establishes a baseline set of ad exchanges for a **mediation** analysis rather than a same-treatment outcome spread. Both are arguable. The evidence sentence is published for each so the disagreement can be about a specific paper. |
| | * **Not a quality verdict.** Several of the 24 report effects far larger than any plausible noise floor. What they cannot do is show the reader that. |
| | * **Not re-measured after the population grew.** ''NULLS'' was adjudicated against the 34-paper set, including both papers added in the same pass. |
| |
| ===== Recall repair ===== | ===== Recall repair ===== |
| Two sweeps were run //after// the population was settled, because the probe-7 result above showed the scoring had silently discarded 9.4% of it. | Two sweeps were run //after// the population was settled, because the probe-7 result above showed the scoring had silently discarded 9.4% of it. |
| |
| **1. The dropped tail of probe 2.** ''scripts/_aa_residue.mjs'' takes the 158 probe-2 candidates that were never adjudicated and re-scores them with the apparatus-density probe that recovered the three late finds, at a lower threshold and requiring an outcome term. **34 of the 158 clear it. All 34 were read at title level; none is a differential platform audit.** They fall into three groups: user-study randomised trials where the treatment is applied to a //person// and the outcome is that person's behaviour (phishing warnings, Tor nudges, consent dialogs, personalised cookie banners), platform-side A/B tests run //with// the platform rather than against it (//How Intention Informed Recommendations Modulate Choices//, //Reducing Symbiosis Bias through Better A/B Tests//), and passing mentions. Output below. | **1. The dropped tail of probe 2.** ''scripts/_aa_residue.mjs'' takes the probe-2 candidates that were never adjudicated and re-scores them with the apparatus-density probe that recovered the three late finds, at a lower threshold and requiring an outcome term. **33 of the 157 clear it. All were read at title level; none is a differential platform audit.** They fall into three groups: user-study randomised trials where the treatment is applied to a //person// and the outcome is that person's behaviour (phishing warnings, Tor nudges, consent dialogs, personalised cookie banners), platform-side A/B tests run //with// the platform rather than against it (//How Intention Informed Recommendations Modulate Choices//, //Reducing Symbiosis Bias through Better A/B Tests//), and passing mentions. Output below. |
| |
| **2. The five zero-years.** The audit set has no papers in 2011, 2012, 2013, 2017 or 2021. ''_aa_gap.mjs'' requires ≥4 apparatus matches, which is too strict to prove a **negative** — it returns nothing at all for 2017, so a zero from it is uninformative. ''scripts/_aa_zeroyears.mjs'' drops the threshold to ≥2 and requires no outcome term, and refuses to run if any year it is given is not actually a zero-year. Over **1,002** papers in those five years it finds **2** with ≥2 apparatus terms: //An Army of Me// (TheWebConf 2017), which is an observational study of sockpuppets other people created, and //Spotting fake reviewer groups in consumer reviews// (TheWebConf 2012), which is not an audit either. **The zero-years are real in this corpus** — which is a different claim from their being real in the field, since 2013 is the year of Hannak et al. and the screen dropped it. | **2. The five zero-years.** The audit set has no papers in 2011, 2012, 2013, 2017 or 2021. ''_aa_gap.mjs'' requires ≥4 apparatus matches, which is too strict to prove a **negative** — it returns nothing at all for 2017, so a zero from it is uninformative. ''scripts/_aa_zeroyears.mjs'' drops the threshold to ≥2 and requires no outcome term, and refuses to run if any year it is given is not actually a zero-year. Over **1,002** papers in those five years it finds **2** with ≥2 apparatus terms: //An Army of Me// (TheWebConf 2017), which is an observational study of sockpuppets other people created, and //Spotting fake reviewer groups in consumer reviews// (TheWebConf 2012), which is not an audit either. **The zero-years are real in this corpus** — which is a different claim from their being real in the field, since 2013 is the year of Hannak et al. and the screen dropped it. |
| |
| * **The size of this literature.** The corpus cannot give it; see //The screening loss//. Closing it needs a pass over FAccT, EuroS&P and the IR venues, which is out of scope for a corpus-backed page. | * **The size of this literature.** The corpus cannot give it; see //The screening loss//. Closing it needs a pass over FAccT, EuroS&P and the IR venues, which is out of scope for a corpus-backed page. |
| * **Whether the 146 dropped probe-2 candidates contain any true audits.** They were read at title-and-abstract level only. A full-text read of all 146 would settle it and was not done. | * **Whether the papers dropped at title level contain any more true audits.** The union across all seven probes is 461; 70 were adjudicated in depth, so **391** rest on a title-and-abstract read. The probe-2 tail was re-swept (//Recall repair//) and a generic review pass found one more audit in the score-≥4 shortlist; both are reasons to expect the 391 is not empty. A full-text read of all of them would settle it and was not done. |
| * **Whether papers build a null without naming it.** The eight-term probe is a lower bound. Establishing the real rate needs 32 methods sections read for the //concept//, which is a different and slower exercise than the verdict read that produced the population. | * **Whether papers build a null without naming it.** The eight-term probe is a lower bound. Establishing the real rate needs 32 methods sections read for the //concept//, which is a different and slower exercise than the verdict read that produced the population. |
| * **Effect sizes.** Every audit measures a different outcome on a different platform with a different metric. Nothing is poolable, and the page publishes no cross-paper effect size deliberately. | * **Effect sizes.** Every audit measures a different outcome on a different platform with a different metric. Nothing is poolable, and the page publishes no cross-paper effect size deliberately. |
| ==== Generic (''fable'') ==== | ==== Generic (''fable'') ==== |
| |
| GENERICREVIEW | Handed the corrected text after the three focused passes. No checklist; asked for whatever they were not looking for. It returned nine findings, ranked. **All nine accepted**; the first three changed the page materially. |
| | |
| | ^ # ^ Finding ^ Verdict and what changed ^ |
| | | 1 | The central claim states a **practice** ("only 11 of 32 establish a same-treatment baseline", "the practice is missing too", and the heading "The null is missing from two thirds of them") on the strength of a **term count** — and the reviewer printed the matching sentence behind all 11 hits, showing the probe over-counts as well as under-counts | **accepted.** All 34 papers were read for the question and hand-adjudicated with an evidence sentence each: **10 of 34**. The term table stays, relabelled as a vocabulary finding. See [[#The null: why it is adjudicated and not counted]] | |
| | | 2 | The adjudication dropped candidates silently: **26 of the 58** score-≥4 shortlist papers had no verdict, //"every verdict is in the output"// was true only of the 44 that got one, and one of the 26 — //Cart-ology// — satisfies the rule | **accepted, and it was right about Cart-ology.** All 58 now carry a written verdict. Cart-ology and the IMC 2024 poster promoted; population **32 → 34**; rejections **12 → 36**. Root cause recorded above: ''_aa_adj.mjs'' returned zero sentences for Cart-ology and that silence was read as a negative | |
| | | 3 | **Nothing links to the page** except the two routing pages, and [[Design:Platforms]] still says twice that this topic //"has no page on this wiki"// — the exact sentence used to justify creating it | **accepted.** Back-links added from the neighbour pages, and the two stale sentences on [[Design:Platforms]] repointed | |
| | | 4 | //"Most audit papers supply the first three and skip the fourth"// contradicts the page's own 43.8% vantage-location figure | **accepted.** Replaced with the actual figures | |
| | | 5 | //"Twenty profiles per condition ... for the same total cost"// is false — same measurements, ten times the training, and the page says training is the expensive part | **accepted.** Rewritten to say what it actually costs and why the corpus is full of two-profile designs. The unsupported "audits are its most natural habitat" was cut too | |
| | | 6 | {[sun2026_suitability]} is in neither ''AUDITS'' nor ''REJECTED'' yet gets four mentions | **accepted.** Trimmed to one table row that says plainly it is not one of the 34, plus the link | |
| | | 7 | **"This method is not historical. It is growing."** is not what a 6/4/4/3-year window table with //n// = 34 from a biased screen supports | **accepted.** Replaced with papers-per-year (1.00, 1.25, 3.00, 3.67), the observation that the second window is barely above the first, and a narrower claim | |
| | | 8 | //"used continuously 2010 → 2026"// for personas is not script-derived; the first ''persona'' hit is 2016 and there are none in 2017–2019 | **accepted.** Corrected, with the note that {[guha2010_challenges]} does the same thing and spells it //personae// | |
| | | 9 | Smaller: **146** survived in one place after being corrected to 158 elsewhere; //"the two papers most cited"// with no citation count; //"never drive a browser"// said of studies that run in participants' browsers; //"Chrome 141"// stale in a sentence dated 2026-09 | **all accepted.** The residue figures were re-derived for the new sets (**157** and **391**); the "most cited" clause is gone with the rewritten section; the eight no-config papers are described accurately; Chrome pinned to **154**, fetched from the Chrome version-history API on 2026-09-11 | |
| | |
| | **What it found fine, and said so:** length (33 KB against 50–70 KB neighbours), no duplication of the four neighbour pages, the arms-versus-waves noise-floor distinction, the 26-vs-14 sock-puppet reconciliation, the //What this corpus cannot tell you// box, the re-based reporting table, and the //What to report// list with its one-sentence template. |
| | |
| | **One thing it flagged that was not a defect:** the literal token ''GENERICREVIEW'' was live on this page at the time of review. That was this section's own placeholder. |
| |
| ===== The report script ===== | ===== The report script ===== |
| // -- profile history, declared attribute, location, device, opt-out | // -- profile history, declared attribute, location, device, opt-out |
| // setting, ad creative -- and holds the rest fixed; | // setting, ad creative -- and holds the rest fixed; |
| // (O) the outcome it measures is the platform's own response: ads served, | // (O) the outcome it measures is the platform's own DISCRIMINATING response to |
| // results ranked, prices quoted, feed or recommendation contents, or an | // the identity it has built: ads served, results ranked, prices quoted, feed |
| // access/labelling decision; | // or recommendation contents, or the profile the platform reports back. |
| | // NOT reachability. A study whose outcome is whether you can reach the site |
| | // at all -- geoblocking, censorship, Tor-exit refusal -- varies a vantage |
| | // point rather than an identity and belongs to design:blocking_and_geodifference. |
| | // This clause was tightened on 2026-09-11; it excludes no paper already in |
| | // the set, and it is what keeps "403 Forbidden: A Global View of CDN |
| | // Geoblocking" (IMC 2018) out. |
| // (C) the result is a difference (or a bounded absence of difference) | // (C) the result is a difference (or a bounded absence of difference) |
| // BETWEEN arms, not a prevalence over a crawl of many sites. | // BETWEEN arms, not a prevalence over a crawl of many sites. |
| ['WWW/2020/stop-tracking-me-bro-differential-tracking-of-user-demographics-on-hyper-partisa', | ['WWW/2020/stop-tracking-me-bro-differential-tracking-of-user-demographics-on-hyper-partisa', |
| '"We create 9 carefully crafted personas representing different genders and age groups"'], | '"We create 9 carefully crafted personas representing different genders and age groups"'], |
| | ['CCS/2022/cart-ology-intercepting-targeted-advertising-via-ad-network-identity-entanglemen', |
| | 'five profiles -- two blank baselines, two account baselines, one entangled attack profile: "All profiles are created and mechanistically measured in the same way, with separation between attacker, victim, and baselines on different machines with different IPs"'], |
| ['IMC/2022/measurement-and-analysis-of-implied-identity-in-ad-delivery-optimization', | ['IMC/2022/measurement-and-analysis-of-implied-identity-in-ad-delivery-optimization', |
| '"We ran 200 versions of this ad at the same time, all from the same account and with the same budget"'], | '"We ran 200 versions of this ad at the same time, all from the same account and with the same budget"'], |
| ['PETS/2024/opted-out-yet-tracked-are-regulations-enough-to-protect-your-privacy', | ['PETS/2024/opted-out-yet-tracked-are-regulations-enough-to-protect-your-privacy', |
| '"We then visit the filtered websites with user interest personas" -- 16 interest personas, two regions'], | '"We then visit the filtered websites with user interest personas" -- 16 interest personas, two regions'], |
| | ['IMC/2024/poster-identifying-filter-bubble-based-on-feed-level-embedding-similarity-analys', |
| | 'four bot arms differing only in the video-selection strategy: "The bot selected a video having (a) the minimum JCC, (b) JCC larger than the minimum (random selection), (c) JCC larger than the mean, and (d) JCC larger than WCC"; outcome is the recommended feed'], |
| ['WWW/2024/tiktok-and-the-art-of-personalization-investigating-exploration-and-exploitation', | ['WWW/2024/tiktok-and-the-art-of-personalization-investigating-exploration-and-exploitation', |
| '"validate our results using a baseline generated from automated TikTok bots, as well as a randomized baseline"'], | '"validate our results using a baseline generated from automated TikTok bots, as well as a randomized baseline"'], |
| ['IEEE-SP/2023/collaborative-ad-transparency-promises-and-limitations', 'explicitly the non-persona alternative: "One method that does not use fake personas"'], | ['IEEE-SP/2023/collaborative-ad-transparency-promises-and-limitations', 'explicitly the non-persona alternative: "One method that does not use fake personas"'], |
| ['USENIX/2023/problematic-advertising-and-its-disparate-exposure-on-facebook', 'donated ad exposure from real users; disparity is observational'], | ['USENIX/2023/problematic-advertising-and-its-disparate-exposure-on-facebook', 'donated ad exposure from real users; disparity is observational'], |
| ['CCS/2022/cart-ology-intercepting-targeted-advertising-via-ad-network-identity-entanglemen', 'attack on ad-network identity linking; no persona arms'], | |
| ['CCS/2022/privacy-limitations-of-interest-based-advertising-on-the-web-a-post-mortem-empir', 'randomised control is a data permutation over a browsing panel, not a live-platform arm'], | ['CCS/2022/privacy-limitations-of-interest-based-advertising-on-the-web-a-post-mortem-empir', 'randomised control is a data permutation over a browsing panel, not a live-platform arm'], |
| ['IEEE-SP/2024/targeted-and-troublesome-tracking-and-advertising-on-childrens-websites', 'fresh profile per page visit; the paper itself says "Future work could extend our method to incorporate personas"'], | ['IEEE-SP/2024/targeted-and-troublesome-tracking-and-advertising-on-childrens-websites', 'fresh profile per page visit; the paper itself says "Future work could extend our method to incorporate personas"'], |
| ['WWW/2026/when-ads-become-profiles-uncovering-the-invisible-risk-of-web-advertising-at-sca', '"Random Control group" is a model ablation, not a platform arm'], | ['WWW/2026/when-ads-become-profiles-uncovering-the-invisible-risk-of-web-advertising-at-sca', '"Random Control group" is a model ablation, not a platform arm'], |
| ['IEEE-SP/2022/deployment-of-source-address-validation-by-network-operators-a-randomized-contro', 'an RCT, but the treatment is a notification to operators; outcome is not platform output'], | ['IEEE-SP/2022/deployment-of-source-address-validation-by-network-operators-a-randomized-contro', 'an RCT, but the treatment is a notification to operators; outcome is not platform output'], |
| | // The rest of the score->=4 shortlist. A generic review found these had been read |
| | // and dropped without a written verdict, and that one of the papers so dropped |
| | // (Cart-ology) in fact satisfied the rule. Every shortlisted paper now has a verdict. |
| | ['NDSS/2015/bloom-cookies-web-search-personalization-without-user-tracking', 'a privacy-preserving personalisation design evaluated on search logs; no arms'], |
| | ['CCS/2018/peeling-the-onions-user-experience-layer-examining-naturalistic-use-of-the-tor-b', 'a 19-participant UX study of Tor Browser; outcome is user experience'], |
| | ['USENIX/2023/strategies-and-vulnerabilities-of-participants-in-venezuelan-influence-operation', 'interviews with 19 influence-operation participants; the sockpuppets are theirs, not the authors'], |
| | ['CCS/2011/policy-auditing-over-incomplete-logs-theory-implementation-and-applications', 'compliance checking over audit logs; shares the word "audit" and nothing else'], |
| | ['WWW/2013/your-browsing-behavior-for-a-big-mac-economics-of-personal-information-online', '168 recruited participants valuing their own PII in an auction; a user study'], |
| | ['IEEE-SP/2015/effective-real-time-android-application-auditing', 'program analysis of Android apps; "auditing" means taint tracking'], |
| | ['USENIX/2017/exploring-user-perceptions-of-discrimination-in-online-targeted-advertising', 'a survey of user perceptions; the randomised factors are survey vignettes, not platform arms'], |
| | ['IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking', 'vantage points in 177 countries, but the outcome is REACHABILITY. Excluded by the tightened (O) clause; belongs to design:blocking_and_geodifference'], |
| | ['CCS/2019/the-art-and-craft-of-fraudulent-app-promotion-in-google-play', 'a study of app-store fraud workers who operate sockpuppets; observational'], |
| | ['WWW/2022/characterizing-detecting-and-predicting-online-ban-evasion', "Wikipedia's own labelled sockpuppet groups; observational, no arms"], |
| | ['WWW/2022/fairness-audit-of-machine-learning-models-with-confidential-computing', 'ML fairness auditing inside a TEE; no live platform'], |
| | ['IEEE-SP/2022/towards-automated-auditing-for-account-and-session-management-flaws-in-single-si', 'SSO implementation flaws; outcome is a security bug, not a discriminating response'], |
| | ['NDSS/2023/tactics-threats-targets-modeling-disinformation-and-its-mitigation', 'interviews with fact-checkers and analysts; sockpuppets are the subject, not the instrument'], |
| | ['WWW/2023/scoping-fairness-objectives-and-identifying-fairness-metrics-for-recommender-sys', "practitioner interviews about fairness metrics; no measurement of a platform"], |
| | ['IEEE-SP/2023/when-and-why-do-people-want-ad-targeting-explanations-evidence-from-a-four-week', 'a four-week field study of what people want from ad explanations; outcome is attitudes'], |
| | ['USENIX/2023/auditing-framework-apis-via-inferred-app-side-security-specifications', 'Android framework API access control; "auditing" is static analysis'], |
| | ['CCS/2024/curator-attack-when-blackbox-differential-privacy-auditing-loses-its-power', 'differential-privacy auditing; a different object with the same name'], |
| | ['USENIX/2024/efficient-privacy-auditing-in-federated-learning', 'membership inference against an FL model; not a platform'], |
| | ['USENIX/2024/what-do-you-want-from-theory-alone-experimenting-with-tight-auditing-of-differen', 'DP synthetic-data auditing; not a platform'], |
| | ['USENIX/2024/fledging-will-continue-until-privacy-improves-empirical-analysis-of-googles-priv', 'security analysis of the FLEDGE/Protected Audience API; the outcome is API behaviour and attack feasibility, not a served-ad contrast across arms'], |
| | ['NDSS/2025/exploring-user-perceptions-of-security-auditing-in-the-web3-ecosystem', 'user perceptions of smart-contract audits'], |
| | ['PETS/2025/privacy-perceptions-and-behaviors-towards-targeted-advertising-on-social-media-a', 'an n=412 cross-country survey; outcome is attitudes'], |
| | ['PETS/2026/audagent-automated-auditing-of-privacy-policy-compliance-in-ai-agents', 'policy-compliance checking of AI agents; no arms'], |
| | ['PETS/2026/privacy-in-theory-bugs-in-practice-grey-box-auditing-of-differential-privacy-lib', 'DP library auditing; not a platform'], |
| | ['WWW/2026/question-the-questions-auditing-representation-in-online-deliberative-processes', 'algorithm design for question selection in deliberative polls; not a measurement of a deployed platform'], |
| ]; | ]; |
| | |
| | // Does the paper measure outcome variation under the SAME treatment? That is the |
| | // step the page argues is skipped, and a term probe cannot answer it: the eight |
| | // phrasings in _aa_noise.mjs let in a "noise floor" that describes someone else's |
| | // paper, and miss Cart-ology's four identically-measured baseline profiles, which |
| | // match none of them. So it is hand-adjudicated here, like the population, with |
| | // the sentence that settled it. Three routes count: |
| | // (a) A/A -- two or more arms that differ in nothing; |
| | // (b) repeats -- the same condition run several times, with the spread reported; |
| | // (c) generated -- a null distribution built from the paper's own observations |
| | // (permutation over arm labels, or a randomised baseline). |
| | // Anything not listed here was read and found to state a difference without ever |
| | // measuring what a non-difference looks like. |
| | export const NULLS = { |
| | 'IMC/2010/challenges-in-measuring-online-advertising-systems': |
| | ['A/A', '"Even queries launched simultaneously from two identically configured clients on the same subnet can produce wildly different ads"; and "In this paper we perform all analysis relative to a control experiment"'], |
| | 'CCS/2015/sunlight-fine-grained-targeting-detection-at-scale-with-statistical-confidence': |
| | ['generated', 'exact/random permutation test against an explicit H0, on a held-out split of the profiles, with Benjamini-Yekutieli or Holm correction'], |
| | 'IMC/2015/location-location-location-the-impact-of-geolocation-on-web-search-personalizati': |
| | ['A/A', '"The red line compares two treatments at the baseline location (i.e., the experimental control), and thus shows the noise floor."'], |
| | 'PETS/2015/automated-experiments-on-ad-privacy-settings': |
| | ['generated', 'permutation test over the arm labels -- "The permutation test randomly permutes ... the control and experimental treatments" -- run over "blocks of nearly identical agents"'], |
| | 'WWW/2019/measuring-political-personalization-of-google-news-search': |
| | ['A/A', 'edit distances tested "among four identical browser profiles" as the reference for the trained-profile comparison'], |
| | 'CCS/2022/cart-ology-intercepting-targeted-advertising-via-ad-network-identity-entanglemen': |
| | ['A/A', 'four unused baseline profiles measured identically; one diverged and the authors read that as the noise: "Such non-determinism is expected, validating our strategy of deploying numerous baseline profiles and the use of normalization."'], |
| | 'WWW/2024/tiktok-and-the-art-of-personalization-investigating-exploration-and-exploitation': |
| | ['generated', '"a randomized baseline" against which the personalisation features are scored'], |
| | 'PETS/2025/more-and-scammier-ads-the-perils-of-youtubes-ad-privacy-settings': |
| | ['repeats', 'each condition run six times at least a week apart ("we repeat each experiment six times"), with outliers cut at standard deviations of the country-specific mean'], |
| | 'IMC/2025/from-voice-to-ads-auditing-commercial-smart-speakers-for-targeted-advertising-ba': |
| | ['A/A + generated', 'paired puppets differing in nothing, scheduled in parallel; plus a permutation test against a random dataset'], |
| | 'PETS/2026/ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis': |
| | ['A/A + generated', '"our baseline persona uses the same neutral persona type for control and treatment. For the baseline persona, the majority of measurement pairs produce insignificant tests" -- an A/A that behaved like one -- plus a 9,999-permutation test'], |
| | }; |
| </file> | </file> |
| |
| const byKey = new Map(papers.map((p) => [`${p.venue}/${p.year}/${p.slug}`, p])); | const byKey = new Map(papers.map((p) => [`${p.venue}/${p.year}/${p.slug}`, p])); |
| |
| import { AUDITS, REJECTED } from './algorithm_audits_set.mjs'; | import { AUDITS, REJECTED, NULLS } from './algorithm_audits_set.mjs'; |
| |
| // --------------------------------------------------------------------------- | // --------------------------------------------------------------------------- |
| // Pin them here so an accidental edit to the set fails the run instead of | // Pin them here so an accidental edit to the set fails the run instead of |
| // silently moving every figure on the page. | // silently moving every figure on the page. |
| if (audit.length !== 32) die(`audit population is ${audit.length}, the page says 32 -- update the page and this contract together`); | if (audit.length !== 34) die(`audit population is ${audit.length}, the page says 34 -- update the page and this contract together`); |
| if (REJECTED.length !== 12) die(`rejected set is ${REJECTED.length}, the page says 12 -- update the page and this contract together`); | if (REJECTED.length !== 36) die(`rejected set is ${REJECTED.length}, the page says 36 -- update the page and this contract together`); |
| if (AUDITS.some(([k]) => REJECTED.some(([r]) => r === k))) die('a key is in both AUDITS and REJECTED'); | if (AUDITS.some(([k]) => REJECTED.some(([r]) => r === k))) die('a key is in both AUDITS and REJECTED'); |
| for (const [key] of REJECTED) if (!byKey.get(key)) die(`rejected paper not in the extraction: ${key}`); | for (const [key] of REJECTED) if (!byKey.get(key)) die(`rejected paper not in the extraction: ${key}`); |
| |
| // --- by year ----------------------------------------------------------------- | // --- by year ----------------------------------------------------------------- |
| console.log('## Audit papers by year (denominator: the 32-paper audit set)'); | console.log('## Audit papers by year (denominator: the audit set)'); |
| const years = [...new Set(papers.map((p) => p.year))].sort(); | const years = [...new Set(papers.map((p) => p.year))].sort(); |
| let run = 0; | let run = 0; |
| console.log(` years with zero audit papers: ${gaps.join(', ')}`); | console.log(` years with zero audit papers: ${gaps.join(', ')}`); |
| console.log(''); | console.log(''); |
| | let wsum = 0; |
| for (const [lo, hi] of [[2010, 2015], [2016, 2019], [2020, 2023], [2024, 2026]]) { | for (const [lo, hi] of [[2010, 2015], [2016, 2019], [2020, 2023], [2024, 2026]]) { |
| const n = audit.filter((p) => p.year >= lo && p.year <= hi).length; | const n = audit.filter((p) => p.year >= lo && p.year <= hi).length; |
| console.log(` ${lo}-${hi}: ${n} of ${N} (${pct(n, N)})`); | wsum += n; |
| | const span = hi - lo + 1; |
| | console.log(` ${lo}-${hi} (${span} yr): ${n} of ${N} (${pct(n, N)}), ${(n / span).toFixed(2)} papers/year`); |
| } | } |
| | if (wsum !== N) die(`window buckets sum to ${wsum}, not ${N}`); |
| console.log(''); | console.log(''); |
| |
| ['states a vantage location', audit, vantageStated, papers.filter((p) => p.vantage.length > 0), vantageStated, 'measuredFrom (3,908)'], | ['states a vantage location', audit, vantageStated, papers.filter((p) => p.vantage.length > 0), vantageStated, 'measuredFrom (3,908)'], |
| ]; | ]; |
| console.log(` ${'indicator'.padEnd(34)} ${'audit (32)'.padStart(14)} baseline`); | console.log(` ${'indicator'.padEnd(34)} ${`audit (${N})`.padStart(14)} baseline`); |
| for (const [label, aSet, aFn, bSet, bFn, bName] of rows) { | for (const [label, aSet, aFn, bSet, bFn, bName] of rows) { |
| const a = aSet.filter(aFn).length; | const a = aSet.filter(aFn).length; |
| console.log(''); | console.log(''); |
| console.log('## The crawl-config rows again, on a SHARED denominator'); | console.log('## The crawl-config rows again, on a SHARED denominator'); |
| console.log(" The rows above compare 32 audits (8 of which have no crawl config) against"); | console.log(` The rows above compare ${N} audits (${N - audit.filter((p) => p.crawlConfig !== null).length} of which have no crawl config) against`); |
| console.log(' 1,120 crawled papers (40 of which have none). A field that can only be stated'); | console.log(' 1,120 crawled papers (40 of which have none). A field that can only be stated'); |
| console.log(' on a paper that HAS a config must be counted over papers that have one, or the'); | console.log(' on a paper that HAS a config must be counted over papers that have one, or the'); |
| for (const [k, [c, a]] of Object.entries(vocabCount)) | for (const [k, [c, a]] of Object.entries(vocabCount)) |
| console.log(` ${k.padEnd(32)} ${String(c).padStart(7)} ${String(a).padStart(5)} of ${N}`); | console.log(` ${k.padEnd(32)} ${String(c).padStart(7)} ${String(a).padStart(5)} of ${N}`); |
| | console.log(''); |
| | |
| | // --- the null, hand-adjudicated --- |
| | console.log('## Establishing a null, HAND-ADJUDICATED (not a term count)'); |
| | console.log(' Does the paper measure outcome variation under the SAME treatment?'); |
| | console.log(' Routes: A/A (arms differing in nothing), repeats (same condition run'); |
| | console.log(" several times, spread reported), generated (null built from the paper's"); |
| | console.log(' own data). Evidence sentence per paper in algorithm_audits_set.mjs.'); |
| | for (const k of Object.keys(NULLS)) if (!AUDITS.some(([a]) => a === k)) die(`NULLS key not in AUDITS: ${k}`); |
| | const byRoute = {}; |
| | for (const [k, [route]] of Object.entries(NULLS)) (byRoute[route] ||= []).push(k); |
| | for (const [route, ks] of Object.entries(byRoute)) console.log(` ${route.padEnd(16)} ${ks.length}`); |
| | const nNull = Object.keys(NULLS).length; |
| | console.log(` ${'TOTAL'.padEnd(16)} ${nNull} of ${N} (${pct(nNull, N)})`); |
| | console.log(` no null established: ${N - nNull} of ${N} (${pct(N - nNull, N)})`); |
| | for (const [k, [route, why]] of Object.entries(NULLS)) { |
| | const p = byKey.get(k); |
| | console.log(` ${p.year} ${p.venue.padEnd(8)} [${route}] ${p.title}`); |
| | console.log(` ${why}`); |
| | } |
| console.log(''); | console.log(''); |
| |
| crawled 1120 | crawled 1120 |
| inferential (any non-descriptive statistic) 1762 | inferential (any non-descriptive statistic) 1762 |
| AUDIT (hand-adjudicated, rule above) 32 | AUDIT (hand-adjudicated, rule above) 34 |
| adjudicated and rejected 12 | adjudicated and rejected 36 |
| |
| ## Audit papers by year (denominator: the 32-paper audit set) | ## Audit papers by year (denominator: the audit set) |
| 2010 1 # | 2010 1 # |
| 2011 0 | 2011 0 |
| 2020 2 ## | 2020 2 ## |
| 2021 0 | 2021 0 |
| 2022 5 ##### | 2022 6 ###### |
| 2023 4 #### | 2023 4 #### |
| 2024 2 ## | 2024 3 ### |
| 2025 5 ##### *provisional | 2025 5 ##### *provisional |
| 2026 3 ### *provisional | 2026 3 ### *provisional |
| total 32 | total 34 |
| years with zero audit papers: 2011, 2012, 2013, 2017, 2021 | years with zero audit papers: 2011, 2012, 2013, 2017, 2021 |
| |
| 2010-2015: 6 of 32 (18.8%) | 2010-2015 (6 yr): 6 of 34 (17.6%), 1.00 papers/year |
| 2016-2019: 5 of 32 (15.6%) | 2016-2019 (4 yr): 5 of 34 (14.7%), 1.25 papers/year |
| 2020-2023: 11 of 32 (34.4%) | 2020-2023 (4 yr): 12 of 34 (35.3%), 3.00 papers/year |
| 2024-2026: 10 of 32 (31.3%) | 2024-2026 (3 yr): 11 of 34 (32.4%), 3.67 papers/year |
| |
| ## Audit papers by venue, against that venue's own output | ## Audit papers by venue, against that venue's own output |
| CCS 3 of 990 0.3% | CCS 4 of 990 0.4% |
| IEEE-SP 1 of 767 0.1% | IEEE-SP 1 of 767 0.1% |
| IMC 7 of 638 1.1% | IMC 8 of 638 1.3% |
| NDSS 1 of 701 0.1% | NDSS 1 of 701 0.1% |
| PETS 10 of 510 2.0% | PETS 10 of 510 2.0% |
| |
| ## Platform measured (multi-valued; does not sum to N) | ## Platform measured (multi-valued; does not sum to N) |
| web 25 of 32 78.1% | web 27 of 34 79.4% |
| other-online-service 18 of 32 56.3% | other-online-service 19 of 34 55.9% |
| mobile 4 of 32 12.5% | mobile 4 of 34 11.8% |
| iot 3 of 32 9.4% | iot 3 of 34 8.8% |
| offline 0 of 32 0.0% | offline 0 of 34 0.0% |
| |
| ## What audit papers report, vs the comparable corpus baseline | ## What audit papers report, vs the comparable corpus baseline |
| Each row states both populations. The baseline is the population named, | Each row states both populations. The baseline is the population named, |
| not "all papers", so the two cells are comparable. | not "all papers", so the two cells are comparable. |
| indicator audit (32) baseline | indicator audit (34) baseline |
| runs a non-descriptive statistic 21/32 (65.6%) 1637/5118 (32.0%) of empirical (5,118) | runs a non-descriptive statistic 21/34 (61.8%) 1637/5118 (32.0%) of empirical (5,118) |
| states an ethics-review outcome 20/32 (62.5%) 1728/5118 (33.8%) of empirical (5,118) | states an ethics-review outcome 20/34 (58.8%) 1728/5118 (33.8%) of empirical (5,118) |
| states artifact availability 22/32 (68.8%) 2890/5118 (56.5%) of empirical (5,118) | states artifact availability 22/34 (64.7%) 2890/5118 (56.5%) of empirical (5,118) |
| has a crawlConfig at all 24/32 (75.0%) 1080/1120 (96.4%) of crawled (1,120) | has a crawlConfig at all 26/34 (76.5%) 1080/1120 (96.4%) of crawled (1,120) |
| states crawl statefulness 22/32 (68.8%) 219/1120 (19.6%) of crawled (1,120) | states crawl statefulness 24/34 (70.6%) 219/1120 (19.6%) of crawled (1,120) |
| states interaction depth 21/32 (65.6%) 841/1120 (75.1%) of crawled (1,120) | states interaction depth 22/34 (64.7%) 841/1120 (75.1%) of crawled (1,120) |
| states a vantage location 14/32 (43.8%) 1228/3908 (31.4%) of measuredFrom (3,908) | states a vantage location 15/34 (44.1%) 1228/3908 (31.4%) of measuredFrom (3,908) |
| |
| ## The crawl-config rows again, on a SHARED denominator | ## The crawl-config rows again, on a SHARED denominator |
| The rows above compare 32 audits (8 of which have no crawl config) against | The rows above compare 34 audits (8 of which have no crawl config) against |
| 1,120 crawled papers (40 of which have none). A field that can only be stated | 1,120 crawled papers (40 of which have none). A field that can only be stated |
| on a paper that HAS a config must be counted over papers that have one, or the | on a paper that HAS a config must be counted over papers that have one, or the |
| two cells are not comparable. Both framings are printed; the page uses this one. | two cells are not comparable. Both framings are printed; the page uses this one. |
| statefulness audit 22/24 ( 91.7%) crawled 219/1080 (20.3%) | statefulness audit 24/26 ( 92.3%) crawled 219/1080 (20.3%) |
| interactionDepth audit 21/24 ( 87.5%) crawled 841/1080 (77.9%) | interactionDepth audit 22/26 ( 84.6%) crawled 841/1080 (77.9%) |
| consentAction audit 11/24 ( 45.8%) crawled 349/1080 (32.3%) | consentAction audit 12/26 ( 46.2%) crawled 349/1080 (32.3%) |
| headless audit 5/24 ( 20.8%) crawled 140/1080 (13.0%) | headless audit 5/26 ( 19.2%) crawled 140/1080 (13.0%) |
| -- the 8 audit papers with no crawl config at all: | -- the 8 audit papers with no crawl config at all: |
| 2015 IMC Peeking Beneath the Hood of Uber. | 2015 IMC Peeking Beneath the Hood of Uber. |
| 1 Counts and percentages of advertisements | 1 Counts and percentages of advertisements |
| -- correction families, folded by hand (paper-counted; the skeleton fold does not merge these): | -- correction families, folded by hand (paper-counted; the skeleton fold does not merge these): |
| Holm-Bonferroni (any spelling) 5 of 32 | Holm-Bonferroni (any spelling) 5 of 34 |
| Bonferroni without Holm 3 of 32 | Bonferroni without Holm 3 of 34 |
| any Bonferroni-family 7 of 32 | any Bonferroni-family 7 of 34 |
| Benjamini-Hochberg/Yekutieli 1 of 32 | Benjamini-Hochberg/Yekutieli 1 of 34 |
| -- 11 of 32 audit papers report no non-descriptive statistic at all: | -- 13 of 34 audit papers report no non-descriptive statistic at all: |
| 2010 IMC Challenges in measuring online advertising systems. | 2010 IMC Challenges in measuring online advertising systems. |
| 2014 CCS Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising. | 2014 CCS Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising. |
| 2018 WWW AdBudgetKiller: Online Advertising Budget Draining Attack. | 2018 WWW AdBudgetKiller: Online Advertising Budget Draining Attack. |
| 2019 PETS Investigating sources of PII used in Facebook’s targeted advertising | 2019 PETS Investigating sources of PII used in Facebook’s targeted advertising |
| | 2022 CCS Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement. |
| 2022 NDSS HARPO: Learning to Subvert Online Behavioral Advertising | 2022 NDSS HARPO: Learning to Subvert Online Behavioral Advertising |
| 2023 CCS Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects. | 2023 CCS Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects. |
| 2023 PETS A Utility-Preserving Obfuscation Approach for YouTube Recommendations | 2023 PETS A Utility-Preserving Obfuscation Approach for YouTube Recommendations |
| | 2024 IMC Poster: Identifying Filter Bubble Based on Feed-Level Embedding Similarity Analysis. |
| 2025 PETS Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants | 2025 PETS Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants |
| 2026 IEEE-SP Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube. | 2026 IEEE-SP Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube. |
| full text present for 5855 of 5859 papers | full text present for 5855 of 5859 papers |
| term corpus in audit set | term corpus in audit set |
| sock puppet 26 5 of 32 | sock puppet 26 5 of 34 |
| persona 190 14 of 32 | persona 190 14 of 34 |
| control profile/account/persona 33 11 of 32 | control profile/account/persona 33 11 of 34 |
| treatment group/profile 66 5 of 32 | treatment group/profile 66 5 of 34 |
| trained profile 4 1 of 32 | trained profile 4 1 of 34 |
| A/A test 2 0 of 32 | A/A test 2 0 of 34 |
| noise floor 35 2 of 32 | noise floor 35 2 of 34 |
| price discrimination/steering 33 9 of 32 | price discrimination/steering 33 9 of 34 |
| | |
| | ## Establishing a null, HAND-ADJUDICATED (not a term count) |
| | Does the paper measure outcome variation under the SAME treatment? |
| | Routes: A/A (arms differing in nothing), repeats (same condition run |
| | several times, spread reported), generated (null built from the paper's |
| | own data). Evidence sentence per paper in algorithm_audits_set.mjs. |
| | A/A 4 |
| | generated 3 |
| | repeats 1 |
| | A/A + generated 2 |
| | TOTAL 10 of 34 (29.4%) |
| | no null established: 24 of 34 (70.6%) |
| | 2010 IMC [A/A] Challenges in measuring online advertising systems. |
| | "Even queries launched simultaneously from two identically configured clients on the same subnet can produce wildly different ads"; and "In this paper we perform all analysis relative to a control experiment" |
| | 2015 CCS [generated] Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence. |
| | exact/random permutation test against an explicit H0, on a held-out split of the profiles, with Benjamini-Yekutieli or Holm correction |
| | 2015 IMC [A/A] Location, Location, Location: The Impact of Geolocation on Web Search Personalization. |
| | "The red line compares two treatments at the baseline location (i.e., the experimental control), and thus shows the noise floor." |
| | 2015 PETS [generated] Automated Experiments on Ad Privacy Settings |
| | permutation test over the arm labels -- "The permutation test randomly permutes ... the control and experimental treatments" -- run over "blocks of nearly identical agents" |
| | 2019 WWW [A/A] Measuring Political Personalization of Google News Search. |
| | edit distances tested "among four identical browser profiles" as the reference for the trained-profile comparison |
| | 2022 CCS [A/A] Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement. |
| | four unused baseline profiles measured identically; one diverged and the authors read that as the noise: "Such non-determinism is expected, validating our strategy of deploying numerous baseline profiles and the use of normalization." |
| | 2024 WWW [generated] TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds. |
| | "a randomized baseline" against which the personalisation features are scored |
| | 2025 PETS [repeats] More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings |
| | each condition run six times at least a week apart ("we repeat each experiment six times"), with outliers cut at standard deviations of the country-specific mean |
| | 2025 IMC [A/A + generated] From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics. |
| | paired puppets differing in nothing, scheduled in parallel; plus a permutation test against a random dataset |
| | 2026 PETS [A/A + generated] Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores |
| | "our baseline persona uses the same neutral persona type for control and treatment. For the baseline persona, the majority of measurement pairs produce insignificant tests" -- an A/A that behaved like one -- plus a 9,999-permutation test |
| |
| ## Screening loss: audit-topical papers in the index but not in the extraction | ## Screening loss: audit-topical papers in the index but not in the extraction |
| crawled 1120 | crawled 1120 |
| inferential (any non-descriptive statistic) 1762 | inferential (any non-descriptive statistic) 1762 |
| AUDIT (hand-adjudicated, rule above) 32 | AUDIT (hand-adjudicated, rule above) 34 |
| adjudicated and rejected 12 | adjudicated and rejected 36 |
| |
| ## Audit papers by year (denominator: the 32-paper audit set) | ## Audit papers by year (denominator: the audit set) |
| 2010 1 # | 2010 1 # |
| 2011 0 | 2011 0 |
| 2020 2 ## | 2020 2 ## |
| 2021 0 | 2021 0 |
| 2022 5 ##### | 2022 6 ###### |
| 2023 4 #### | 2023 4 #### |
| 2024 2 ## | 2024 3 ### |
| 2025 5 ##### *provisional | 2025 5 ##### *provisional |
| 2026 3 ### *provisional | 2026 3 ### *provisional |
| total 32 | total 34 |
| years with zero audit papers: 2011, 2012, 2013, 2017, 2021 | years with zero audit papers: 2011, 2012, 2013, 2017, 2021 |
| |
| 2010-2015: 6 of 32 (18.8%) | 2010-2015: 6 of 34 (17.6%) |
| 2016-2019: 5 of 32 (15.6%) | 2016-2019: 5 of 34 (14.7%) |
| 2020-2023: 11 of 32 (34.4%) | 2020-2023: 12 of 34 (35.3%) |
| 2024-2026: 10 of 32 (31.3%) | 2024-2026: 11 of 34 (32.4%) |
| |
| ## Audit papers by venue, against that venue's own output | ## Audit papers by venue, against that venue's own output |
| CCS 3 of 990 0.3% | CCS 4 of 990 0.4% |
| IEEE-SP 1 of 767 0.1% | IEEE-SP 1 of 767 0.1% |
| IMC 7 of 638 1.1% | IMC 8 of 638 1.3% |
| NDSS 1 of 701 0.1% | NDSS 1 of 701 0.1% |
| PETS 10 of 510 2.0% | PETS 10 of 510 2.0% |
| |
| ## Platform measured (multi-valued; does not sum to N) | ## Platform measured (multi-valued; does not sum to N) |
| web 25 of 32 78.1% | web 27 of 34 79.4% |
| other-online-service 18 of 32 56.3% | other-online-service 19 of 34 55.9% |
| mobile 4 of 32 12.5% | mobile 4 of 34 11.8% |
| iot 3 of 32 9.4% | iot 3 of 34 8.8% |
| offline 0 of 32 0.0% | offline 0 of 34 0.0% |
| |
| ## What audit papers report, vs the comparable corpus baseline | ## What audit papers report, vs the comparable corpus baseline |
| Each row states both populations. The baseline is the population named, | Each row states both populations. The baseline is the population named, |
| not "all papers", so the two cells are comparable. | not "all papers", so the two cells are comparable. |
| indicator audit (32) baseline | indicator audit (34) baseline |
| runs a non-descriptive statistic 21/32 (65.6%) 1637/5118 (32.0%) of empirical (5,118) | runs a non-descriptive statistic 21/34 (61.8%) 1637/5118 (32.0%) of empirical (5,118) |
| states an ethics-review outcome 20/32 (62.5%) 1728/5118 (33.8%) of empirical (5,118) | states an ethics-review outcome 20/34 (58.8%) 1728/5118 (33.8%) of empirical (5,118) |
| states artifact availability 22/32 (68.8%) 2890/5118 (56.5%) of empirical (5,118) | states artifact availability 22/34 (64.7%) 2890/5118 (56.5%) of empirical (5,118) |
| has a crawlConfig at all 24/32 (75.0%) 1080/1120 (96.4%) of crawled (1,120) | has a crawlConfig at all 26/34 (76.5%) 1080/1120 (96.4%) of crawled (1,120) |
| states crawl statefulness 22/32 (68.8%) 219/1120 (19.6%) of crawled (1,120) | states crawl statefulness 24/34 (70.6%) 219/1120 (19.6%) of crawled (1,120) |
| states interaction depth 21/32 (65.6%) 841/1120 (75.1%) of crawled (1,120) | states interaction depth 22/34 (64.7%) 841/1120 (75.1%) of crawled (1,120) |
| states a vantage location 14/32 (43.8%) 1228/3908 (31.4%) of measuredFrom (3,908) | states a vantage location 15/34 (44.1%) 1228/3908 (31.4%) of measuredFrom (3,908) |
| |
| ## The crawl-config rows again, on a SHARED denominator | ## The crawl-config rows again, on a SHARED denominator |
| The rows above compare 32 audits (8 of which have no crawl config) against | The rows above compare 34 audits (8 of which have no crawl config) against |
| 1,120 crawled papers (40 of which have none). A field that can only be stated | 1,120 crawled papers (40 of which have none). A field that can only be stated |
| on a paper that HAS a config must be counted over papers that have one, or the | on a paper that HAS a config must be counted over papers that have one, or the |
| two cells are not comparable. Both framings are printed; the page uses this one. | two cells are not comparable. Both framings are printed; the page uses this one. |
| statefulness audit 22/24 ( 91.7%) crawled 219/1080 (20.3%) | statefulness audit 24/26 ( 92.3%) crawled 219/1080 (20.3%) |
| interactionDepth audit 21/24 ( 87.5%) crawled 841/1080 (77.9%) | interactionDepth audit 22/26 ( 84.6%) crawled 841/1080 (77.9%) |
| consentAction audit 11/24 ( 45.8%) crawled 349/1080 (32.3%) | consentAction audit 12/26 ( 46.2%) crawled 349/1080 (32.3%) |
| headless audit 5/24 ( 20.8%) crawled 140/1080 (13.0%) | headless audit 5/26 ( 19.2%) crawled 140/1080 (13.0%) |
| -- the 8 audit papers with no crawl config at all: | -- the 8 audit papers with no crawl config at all: |
| 2015 IMC Peeking Beneath the Hood of Uber. | 2015 IMC Peeking Beneath the Hood of Uber. |
| 1 Counts and percentages of advertisements | 1 Counts and percentages of advertisements |
| -- correction families, folded by hand (paper-counted; the skeleton fold does not merge these): | -- correction families, folded by hand (paper-counted; the skeleton fold does not merge these): |
| Holm-Bonferroni (any spelling) 5 of 32 | Holm-Bonferroni (any spelling) 5 of 34 |
| Bonferroni without Holm 3 of 32 | Bonferroni without Holm 3 of 34 |
| any Bonferroni-family 7 of 32 | any Bonferroni-family 7 of 34 |
| Benjamini-Hochberg/Yekutieli 1 of 32 | Benjamini-Hochberg/Yekutieli 1 of 34 |
| -- 11 of 32 audit papers report no non-descriptive statistic at all: | -- 13 of 34 audit papers report no non-descriptive statistic at all: |
| 2010 IMC Challenges in measuring online advertising systems. | 2010 IMC Challenges in measuring online advertising systems. |
| 2014 CCS Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising. | 2014 CCS Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising. |
| 2018 WWW AdBudgetKiller: Online Advertising Budget Draining Attack. | 2018 WWW AdBudgetKiller: Online Advertising Budget Draining Attack. |
| 2019 PETS Investigating sources of PII used in Facebook’s targeted advertising | 2019 PETS Investigating sources of PII used in Facebook’s targeted advertising |
| | 2022 CCS Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement. |
| 2022 NDSS HARPO: Learning to Subvert Online Behavioral Advertising | 2022 NDSS HARPO: Learning to Subvert Online Behavioral Advertising |
| 2023 CCS Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects. | 2023 CCS Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects. |
| 2023 PETS A Utility-Preserving Obfuscation Approach for YouTube Recommendations | 2023 PETS A Utility-Preserving Obfuscation Approach for YouTube Recommendations |
| | 2024 IMC Poster: Identifying Filter Bubble Based on Feed-Level Embedding Similarity Analysis. |
| 2025 PETS Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants | 2025 PETS Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants |
| 2026 IEEE-SP Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube. | 2026 IEEE-SP Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube. |
| full text present for 5855 of 5859 papers | full text present for 5855 of 5859 papers |
| term corpus in audit set | term corpus in audit set |
| sock puppet 26 5 of 32 | sock puppet 26 5 of 34 |
| persona 190 14 of 32 | persona 190 14 of 34 |
| control profile/account/persona 33 11 of 32 | control profile/account/persona 33 11 of 34 |
| treatment group/profile 66 5 of 32 | treatment group/profile 66 5 of 34 |
| trained profile 4 1 of 32 | trained profile 4 1 of 34 |
| A/A test 2 0 of 32 | A/A test 2 0 of 34 |
| noise floor 35 2 of 32 | noise floor 35 2 of 34 |
| price discrimination/steering 33 9 of 32 | price discrimination/steering 33 9 of 34 |
| | |
| | ## Establishing a null, HAND-ADJUDICATED (not a term count) |
| | Does the paper measure outcome variation under the SAME treatment? |
| | Routes: A/A (arms differing in nothing), repeats (same condition run |
| | several times, spread reported), generated (null built from the paper's |
| | own data). Evidence sentence per paper in algorithm_audits_set.mjs. |
| | A/A 4 |
| | generated 3 |
| | repeats 1 |
| | A/A + generated 2 |
| | TOTAL 10 of 34 (29.4%) |
| | no null established: 24 of 34 (70.6%) |
| | 2010 IMC [A/A] Challenges in measuring online advertising systems. |
| | "Even queries launched simultaneously from two identically configured clients on the same subnet can produce wildly different ads"; and "In this paper we perform all analysis relative to a control experiment" |
| | 2015 CCS [generated] Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence. |
| | exact/random permutation test against an explicit H0, on a held-out split of the profiles, with Benjamini-Yekutieli or Holm correction |
| | 2015 IMC [A/A] Location, Location, Location: The Impact of Geolocation on Web Search Personalization. |
| | "The red line compares two treatments at the baseline location (i.e., the experimental control), and thus shows the noise floor." |
| | 2015 PETS [generated] Automated Experiments on Ad Privacy Settings |
| | permutation test over the arm labels -- "The permutation test randomly permutes ... the control and experimental treatments" -- run over "blocks of nearly identical agents" |
| | 2019 WWW [A/A] Measuring Political Personalization of Google News Search. |
| | edit distances tested "among four identical browser profiles" as the reference for the trained-profile comparison |
| | 2022 CCS [A/A] Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement. |
| | four unused baseline profiles measured identically; one diverged and the authors read that as the noise: "Such non-determinism is expected, validating our strategy of deploying numerous baseline profiles and the use of normalization." |
| | 2024 WWW [generated] TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds. |
| | "a randomized baseline" against which the personalisation features are scored |
| | 2025 PETS [repeats] More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings |
| | each condition run six times at least a week apart ("we repeat each experiment six times"), with outliers cut at standard deviations of the country-specific mean |
| | 2025 IMC [A/A + generated] From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics. |
| | paired puppets differing in nothing, scheduled in parallel; plus a permutation test against a random dataset |
| | 2026 PETS [A/A + generated] Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores |
| | "our baseline persona uses the same neutral persona type for control and treatment. For the baseline persona, the majority of measurement pairs produce insignificant tests" -- an A/A that behaved like one -- plus a 9,999-permutation test |
| |
| ## Screening loss: audit-topical papers in the index but not in the extraction | ## Screening loss: audit-topical papers in the index but not in the extraction |
| WWW/2020/stop-tracking-me-bro-differential-tracking-of-user-demographics-on-hyper-partisa | WWW/2020/stop-tracking-me-bro-differential-tracking-of-user-demographics-on-hyper-partisa |
| IN: "We create 9 carefully crafted personas representing different genders and age groups" | IN: "We create 9 carefully crafted personas representing different genders and age groups" |
| | 2022 CCS Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement. |
| | CCS/2022/cart-ology-intercepting-targeted-advertising-via-ad-network-identity-entanglemen |
| | IN: five profiles -- two blank baselines, two account baselines, one entangled attack profile: "All profiles are created and mechanistically measured in the same way, with separation between attacker, victim, and baselines on different machines with different IPs" |
| 2022 IMC Measurement and analysis of implied identity in ad delivery optimization. | 2022 IMC Measurement and analysis of implied identity in ad delivery optimization. |
| IMC/2022/measurement-and-analysis-of-implied-identity-in-ad-delivery-optimization | IMC/2022/measurement-and-analysis-of-implied-identity-in-ad-delivery-optimization |
| PETS/2024/opted-out-yet-tracked-are-regulations-enough-to-protect-your-privacy | PETS/2024/opted-out-yet-tracked-are-regulations-enough-to-protect-your-privacy |
| IN: "We then visit the filtered websites with user interest personas" -- 16 interest personas, two regions | IN: "We then visit the filtered websites with user interest personas" -- 16 interest personas, two regions |
| | 2024 IMC Poster: Identifying Filter Bubble Based on Feed-Level Embedding Similarity Analysis. |
| | IMC/2024/poster-identifying-filter-bubble-based-on-feed-level-embedding-similarity-analys |
| | IN: four bot arms differing only in the video-selection strategy: "The bot selected a video having (a) the minimum JCC, (b) JCC larger than the minimum (random selection), (c) JCC larger than the mean, and (d) JCC larger than WCC"; outcome is the recommended feed |
| 2024 WWW TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds. | 2024 WWW TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds. |
| WWW/2024/tiktok-and-the-art-of-personalization-investigating-exploration-and-exploitation | WWW/2024/tiktok-and-the-art-of-personalization-investigating-exploration-and-exploitation |
| 2023 USENIX Problematic Advertising and its Disparate Exposure on Facebook | 2023 USENIX Problematic Advertising and its Disparate Exposure on Facebook |
| OUT: donated ad exposure from real users; disparity is observational | OUT: donated ad exposure from real users; disparity is observational |
| 2022 CCS Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement. | |
| OUT: attack on ad-network identity linking; no persona arms | |
| 2022 CCS Privacy Limitations of Interest-based Advertising on The Web: A Post-mortem Empirical Analysis of Google's FLoC. | 2022 CCS Privacy Limitations of Interest-based Advertising on The Web: A Post-mortem Empirical Analysis of Google's FLoC. |
| OUT: randomised control is a data permutation over a browsing panel, not a live-platform arm | OUT: randomised control is a data permutation over a browsing panel, not a live-platform arm |
| 2022 IEEE-SP Deployment of Source Address Validation by Network Operators: A Randomized Control Trial. | 2022 IEEE-SP Deployment of Source Address Validation by Network Operators: A Randomized Control Trial. |
| OUT: an RCT, but the treatment is a notification to operators; outcome is not platform output | OUT: an RCT, but the treatment is a notification to operators; outcome is not platform output |
| | 2015 NDSS Bloom Cookies: Web Search Personalization without User Tracking |
| | OUT: a privacy-preserving personalisation design evaluated on search logs; no arms |
| | 2018 CCS Peeling the Onion's User Experience Layer: Examining Naturalistic Use of the Tor Browser. |
| | OUT: a 19-participant UX study of Tor Browser; outcome is user experience |
| | 2023 USENIX Strategies and Vulnerabilities of Participants in Venezuelan Influence Operations |
| | OUT: interviews with 19 influence-operation participants; the sockpuppets are theirs, not the authors |
| | 2011 CCS Policy auditing over incomplete logs: theory, implementation and applications. |
| | OUT: compliance checking over audit logs; shares the word "audit" and nothing else |
| | 2013 WWW Your browsing behavior for a big mac: economics of personal information online. |
| | OUT: 168 recruited participants valuing their own PII in an auction; a user study |
| | 2015 IEEE-SP Effective Real-Time Android Application Auditing. |
| | OUT: program analysis of Android apps; "auditing" means taint tracking |
| | 2017 USENIX Exploring User Perceptions of Discrimination in Online Targeted Advertising |
| | OUT: a survey of user perceptions; the randomised factors are survey vignettes, not platform arms |
| | 2018 IMC 403 Forbidden: A Global View of CDN Geoblocking. |
| | OUT: vantage points in 177 countries, but the outcome is REACHABILITY. Excluded by the tightened (O) clause; belongs to design:blocking_and_geodifference |
| | 2019 CCS The Art and Craft of Fraudulent App Promotion in Google Play. |
| | OUT: a study of app-store fraud workers who operate sockpuppets; observational |
| | 2022 WWW Characterizing, Detecting, and Predicting Online Ban Evasion. |
| | OUT: Wikipedia's own labelled sockpuppet groups; observational, no arms |
| | 2022 WWW Fairness Audit of Machine Learning Models with Confidential Computing. |
| | OUT: ML fairness auditing inside a TEE; no live platform |
| | 2022 IEEE-SP Towards Automated Auditing for Account and Session Management Flaws in Single Sign-On Deployments. |
| | OUT: SSO implementation flaws; outcome is a security bug, not a discriminating response |
| | 2023 NDSS Tactics, Threats & Targets: Modeling Disinformation and its Mitigation |
| | OUT: interviews with fact-checkers and analysts; sockpuppets are the subject, not the instrument |
| | 2023 WWW Scoping Fairness Objectives and Identifying Fairness Metrics for Recommender Systems: The Practitioners' Perspective. |
| | OUT: practitioner interviews about fairness metrics; no measurement of a platform |
| | 2023 IEEE-SP When and Why Do People Want Ad Targeting Explanations? Evidence from a Four-Week, Mixed-Methods Field Study. |
| | OUT: a four-week field study of what people want from ad explanations; outcome is attitudes |
| | 2023 USENIX Auditing Framework APIs via Inferred App-side Security Specifications |
| | OUT: Android framework API access control; "auditing" is static analysis |
| | 2024 CCS Curator Attack: When Blackbox Differential Privacy Auditing Loses Its Power. |
| | OUT: differential-privacy auditing; a different object with the same name |
| | 2024 USENIX Efficient Privacy Auditing in Federated Learning |
| | OUT: membership inference against an FL model; not a platform |
| | 2024 USENIX "What do you want from theory alone?" Experimenting with Tight Auditing of Differentially Private Synthetic Data Generation |
| | OUT: DP synthetic-data auditing; not a platform |
| | 2024 USENIX Fledging Will Continue Until Privacy Improves: Empirical Analysis of Google's Privacy-Preserving Targeted Advertising |
| | OUT: security analysis of the FLEDGE/Protected Audience API; the outcome is API behaviour and attack feasibility, not a served-ad contrast across arms |
| | 2025 NDSS Exploring User Perceptions of Security Auditing in the Web3 Ecosystem |
| | OUT: user perceptions of smart-contract audits |
| | 2025 PETS Privacy Perceptions and Behaviors Towards Targeted Advertising on Social Media: A Cross-Country Study on the Effect of Culture and Religion |
| | OUT: an n=412 cross-country survey; outcome is attitudes |
| | 2026 PETS AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents |
| | OUT: policy-compliance checking of AI agents; no arms |
| | 2026 PETS Privacy in Theory, Bugs in Practice: Grey-Box Auditing of Differential Privacy Libraries |
| | OUT: DP library auditing; not a platform |
| | 2026 WWW Question the Questions: Auditing Representation in Online Deliberative Processes. |
| | OUT: algorithm design for question selection in deliberative polls; not a measurement of a deployed platform |
| |
| ## Screening loss in full | ## Screening loss in full |
| |
| <file text _aa_noise-output.txt> | <file text _aa_noise-output.txt> |
| AUDIT keys parsed from report script: 32 | AUDIT keys parsed from report script: 34 |
| term corpus audit | term corpus audit |
| A/A test 2 0 | A/A test 2 0 |
| two arms with the same treatment 13 1 | two arms with the same treatment 13 1 |
| |
| audit papers with >=1 noise-baseline term: 11 of 32 | audit papers with >=1 noise-baseline term: 11 of 34 |
| CCS/2015/sunlight-fine-grained-targeting-detection-at-scale-with-statistical-confidence | CCS/2015/sunlight-fine-grained-targeting-detection-at-scale-with-statistical-confidence |
| permutation / randomisation test | permutation / randomisation test |
| PETS/2020/inferring-tracker-advertiser-relationships-in-the-online-advertising-ecosystem-u | PETS/2020/inferring-tracker-advertiser-relationships-in-the-online-advertising-ecosystem-u |
| WWW/2020/stop-tracking-me-bro-differential-tracking-of-user-demographics-on-hyper-partisa | WWW/2020/stop-tracking-me-bro-differential-tracking-of-user-demographics-on-hyper-partisa |
| | CCS/2022/cart-ology-intercepting-targeted-advertising-via-ad-network-identity-entanglemen |
| IMC/2022/measurement-and-analysis-of-implied-identity-in-ad-delivery-optimization | IMC/2022/measurement-and-analysis-of-implied-identity-in-ad-delivery-optimization |
| IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising-a-field-measurement | IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising-a-field-measurement |
| WWW/2023/a-method-to-assess-and-explain-disparate-impact-in-online-retailing | WWW/2023/a-method-to-assess-and-explain-disparate-impact-in-online-retailing |
| PETS/2024/opted-out-yet-tracked-are-regulations-enough-to-protect-your-privacy | PETS/2024/opted-out-yet-tracked-are-regulations-enough-to-protect-your-privacy |
| | IMC/2024/poster-identifying-filter-bubble-based-on-feed-level-embedding-similarity-analys |
| PETS/2025/echoes-of-privacy-uncovering-the-profiling-practices-of-voice-assistants | PETS/2025/echoes-of-privacy-uncovering-the-profiling-practices-of-voice-assistants |
| PETS/2025/sheeps-clothing-wolfish-intent-automated-detection-and-evaluation-of-problematic | PETS/2025/sheeps-clothing-wolfish-intent-automated-detection-and-evaluation-of-problematic |
| |
| <file text _aa_residue-output.txt> | <file text _aa_residue-output.txt> |
| loose=190 adjudicated=32 residue=158 residue_with_apparatus_and_outcome=34 | loose=190 adjudicated=33 residue=157 residue_with_apparatus_and_outcome=33 |
| 25 2021 CCS The Effect of Google Search on Software Security: Unobtrusive Security Interventions via Content Re-ranking. | 25 2021 CCS The Effect of Google Search on Software Security: Unobtrusive Security Interventions via Content Re-ranking. |
| 14 2020 WWW Finding a Choice in a Haystack: Automatic Extraction of Opt-Out Statements from Privacy Policy Text. | 14 2020 WWW Finding a Choice in a Haystack: Automatic Extraction of Opt-Out Statements from Privacy Policy Text. |
| 12 2025 IEEE-SP Restricting the Link: Effects of Focused Attention and Time Delay on Phishing Warning Effectiveness. | 12 2025 IEEE-SP Restricting the Link: Effects of Focused Attention and Time Delay on Phishing Warning Effectiveness. |
| 10 2023 USENIX Strategies and Vulnerabilities of Participants in Venezuelan Influence Operations | |
| 9 2022 PETS Increasing Adoption of Tor Browser Using Informational and Planning Nudges | 9 2022 PETS Increasing Adoption of Tor Browser Using Informational and Planning Nudges |
| 5 2019 WWW How Intention Informed Recommendations Modulate Choices: A Field Study of Spoken Word Content. | 5 2019 WWW How Intention Informed Recommendations Modulate Choices: A Field Study of Spoken Word Content. |