Table of Contents
Provenance: Algorithm Audits
Working log for algorithm_audits. Every figure on that page has its query here, with the population it is a share of. Corpus-level caveats — venue scope, the selection funnel, provisional years — are on corpus and are not restated.
Run: 2026-09-11. Corpus at the time: data/extract/run1/extractions.jsonl, 5,859 papers, 7 venues (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P), 2010–2026; 5,855 of them have full text on disk. Model: Claude Opus 5 for the run, four review sub-agents (three sonnet, one fable) — findings logged at the foot of this page.
Why this page exists rather than a section on a neighbour
The task brief flagged this as the smallest of five page proposals and asked for it to be sized before writing. It was. The sizing result and the decision:
| Question | Answer |
|---|---|
| Does the wiki already cover it? | No. Stateful stateless owns the profile axis and lists personalisation as a phenomenon that needs a stateful design; Hypothesis testing owns the test; Platforms owns sock puppets as an access route and says in its own Open Questions that “Search-engine and ads-ecosystem auditing has no page on this wiki”. Nobody owned the experimental design. |
| How large is the in-corpus population? | 34 papers under a written inclusion rule (below), of 5,859. 36 further candidates were read and rejected. The figure was 32 until a generic review pass found that 26 shortlisted papers had been read and dropped without a written verdict, and that one of them satisfied the rule. |
| Is it growing or historical? | Not historical. Per year: 1.00 (2010–2015), 1.25 (2016–2019), 3.00 (2020–2023), 3.67 (2024–2026*). n = 34 will not carry a trend claim and the page does not make one; what it says is that more of this work appeared in the last six years than in the first ten. |
| Could it be a section instead? | It could have been ~3 KB on Automated measurements. It was not, for two reasons: (a) that page is a routing page between crawl / scan / app, and an audit is a fourth instrument that cuts across the crawl branch rather than sitting beside it; (b) the material that makes the page worth writing — control arms, carry-over, the noise floor — is design advice, not routing, and would have doubled the length of a page whose job is to be short. |
| The counter-argument | The 32 is a lower bound with a known bias (see The screening loss below), so a reader could reasonably say the page is built on a population the corpus cannot see properly. That is stated on the page itself, in its own box, rather than buried here. |
Decision: created as a new page, design:algorithm_audits, linked from Design and from Automated measurements.
The inclusion rule
Written before any table, and encoded in scripts/algorithm_audits_set.mjs rather than in prose. A paper is in if all three hold:
- (T) Treatment. It deliberately varies a property of the measuring identity or request — profile history, declared attribute, location, device, opt-out setting, ad creative — and holds the rest fixed.
- (O) Outcome. What it measures is the platform's own discriminating response to the identity it has built: ads served, results ranked, prices quoted, feed or recommendation contents, or the profile the platform reports back. Not reachability. A study whose outcome is whether you can reach the site at all — geoblocking, censorship, Tor-exit refusal — varies a vantage point rather than an identity and belongs to Blocking and geodifference. This clause was tightened on 2026-09-11 after a review pass; it excludes no paper already in the set, and it is what keeps 403 Forbidden: A Global View of CDN Geoblocking (IMC 2018) out. Without it, every geodifference study in the corpus would be an algorithm audit and the page would duplicate a neighbour.
- (C) Contrast. The result is a difference (or a bounded absence of difference) between arms, not a prevalence over a crawl of many sites.
Deliberate consequences of this rule, each of which a reasonable person could have decided the other way:
- Attack and defence papers are in if they run the arms. [1Meng, Wei; Xing, Xinyu; Sheth, Anmol; Weinsberg, Udi; Lee, Wenke (2014): "Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [2Kim, I Luk; Wang, Weihang; Kwon, Yonghwi; Zheng, Yunhui; Aafer, Yousra; Meng, Weijie; Zhang, Xiangyu (2018): "AdBudgetKiller: Online Advertising Budget Draining Attack", in: Proceedings of the ACM Web Conference. (DOI)] and [3Zhang, Jiang; Psounis, Konstantinos; Haroon, Muhammad; Shafiq, Zubair (2022): "HARPO: Learning to Subvert Online Behavioral Advertising", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] are not audit studies — they are an attack, an attack and a defence — but each runs a differential experiment against a live ad platform, and a student reading this page wants those designs. The criterion is the measurement design, not the paper's contribution type.
- Instrument papers are in. [4Lécuyer, Mathias; Spahn, Riley; Spiliopolous, Yannis; Chaintreau, Augustin; Geambasu, Roxana; Hsu, Daniel J. (2015): "Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] and [5Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] contribute tooling; they are the two papers the page most wants read.
- Observational “audits” are out. [6Silva, Márcio; Oliveira, Lucas Santos de; Andreou, Athanasios; Melo, Pedro Olmo Stancioli Vaz de; Goga, Oana; Benevenuto, Fabrício (2020): "Facebook Ads Monitor: An Independent Auditing System for Political Ads on Facebook", in: Proceedings of the ACM Web Conference. (DOI)] calls itself “An Independent Auditing System” and collects ads from volunteers. It fails (C). So does [7Ali, Muhammad; Goetzen, Angelica; Mislove, Alan; Redmiles, Elissa M.; Sapiezynski, Piotr (2023): "Problematic Advertising and its Disparate Exposure on Facebook", in: Proceedings of the USENIX Security Symposium. (Link)] and so does Auditing the Partisanship of Google Search Snippets (TheWebConf 2019), which audits snippets against the pages they summarise with no identity treatment at all.
- ML fairness, DP and system-log auditing are out. They share the word and nothing else.
- RCTs whose treatment is not applied to the platform are out. [8Lone, Qasim; Frik, Alisa; Luckie, Matthew; Korczyński, Maciej; van Eeten, Michel; Gañán, Carlos (2022): "Deployment of Source Address Validation by Network Operators: A Randomized Control Trial", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] is a randomised controlled trial, but the treatment is a notification sent to network operators; the outcome is not platform output.
The probes that built the candidate set
No single regex finds this literature: the five vocabularies barely overlap and three of them (“audit”, “persona”, “personalisation”) are dominated by other meanings. The candidate set is the union of six probes, and each entry records which probe caught it.
| # | Probe | Script | Candidates | What it is for |
|---|---|---|---|---|
| 1 | Full text, 15 term families (sock-puppet, persona, paired arm, control arm, A/A, personalisation, price discrimination, differential treatment, algorithmic audit, ad targeting, filter bubble, SERP, trained profile, audit-verb proximity) over all 5,855 papers with text, whitespace collapsed | _aa_probe1.mjs | 1,746 with ≥1 hit | recall |
| 2 | Probe 1 narrowed: ≥1 outcome-family hit and ≥1 apparatus-family hit | _aa_cands.mjs loose | 190 | the working pool |
| 3 | Probe 1 narrowed further: ≥3 outcome hits and ≥1 strong-apparatus hit | _aa_cands.mjs tight | 27 | precision check. A narrowing probe must return a subset of the loose one, or the two are measuring different things and comparing their sizes is meaningless. _aa_cands.mjs now exits 1 if any tight paper is outside loose; it currently exits 0 with 0 outside. Until a reviewer caught it, the subset property was merely printed, not asserted — see Reviewer findings |
| 4 | Title sweep over the 5,859 extracted titles, audit vocabulary | _aa_union.mjs | 71 | catches papers whose method words are only in the abstract |
| 5 | Topical sweep over title + abstract of all 16,864 index records, including the papers the extraction never saw | _aa_abs2.mjs | 117 (69 extracted, 48 not) | the screening-loss measurement |
| 6 | detection[].phenomenon / .technique and classification[].targetDetail over the extraction | _aa_union.mjs | 74 + 4 | catches papers the prose probes miss |
| 7 | Apparatus-density probe: ≥4 distinct arm-vocabulary matches in full text, run over the whole corpus and again over the years the union left empty | _aa_gap.mjs | — | recall repair |
Probe 7 is the one that mattered, and what it says is about the scoring, not the probes. After probes 1–6 the set had a hole at 2020–2021 and I did not trust it. Running the density probe over those years surfaced [9Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)] and [10Agarwal, Pushkal; Joglekar, Sagar; Papadopoulos, Panagiotis; Sastry, Nishanth; Kourtellis, Nicolas (2020): "Stop tracking me Bro! Differential Tracking of User Demographics on Hyper-Partisan Websites", in: Proceedings of the ACM Web Conference. (DOI)], and run corpus-wide it surfaced [11Iqbal, Hassan; Khan, Usman Mahmood; Khan, Hassan Ali; Shahzad, Muhammad (2022): "Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Election 2020", in: Proceedings of the ACM Web Conference. (DOI)] (102 email accounts on Gmail, Outlook and Yahoo — a design no “personalisation” or “persona” probe reaches, because the paper's vocabulary is spam filtering). Three of the final 32, or 9.4%.
All three were in fact already inside the 461-paper union, caught by probe 2. What lost them was the shortlist score cut-off: candidates were ranked (_aa_short.mjs) and only those scoring ≥3 were read, and all three scored below 3. So the honest lesson is not “add a seventh probe” — it is that a scoring heuristic laid over a candidate set is a second, invisible filter, and it discarded 9.4% of the final population before anything was read.
That is why the dropped tail was re-swept rather than left alone; see Recall repair below.
What no probe reached. The inclusion rule needs a paper's design, and design language is not a vocabulary. A paper that ran arms and described them only as “Group A and Group B” would be invisible to all seven probes. No claim on the page depends on the 32 being exhaustive; the page says so.
The adjudication
108 candidates scored ≥3 across the probes were shortlisted with their abstracts (_aa_short.mjs). The 58 scoring ≥4 were passed through _aa_adj.mjs, which prints every sentence in the paper matching an arm/treatment/control/persona pattern; a further 12 lower-scored or probe-7 papers were read directly. Each verdict was made by reading those sentences, and where they were ambiguous, by grepping the paper's methods section.
A generic review pass found the weak point here. 26 of the 58 had been read and dropped with no written verdict, and one of them — Cart-ology (CCS 2022, shortlist score 7) — satisfies the rule: five browser profiles, four of them unused baselines, “created and mechanistically measured in the same way, with separation between attacker, victim, and baselines on different machines with different IPs”, read out as ad distributions. It had been rejected on the strength of _aa_adj.mjs returning zero sentences for it — a 420-character sentence cap meeting a column-spliced PDF. A probe returning nothing is not a negative result, and it was treated as one. All 58 now carry a written verdict, and the IMC 2024 filter-bubble poster was promoted at the same time (four bot arms differing only in video-selection strategy). The population moved from 32 to 34.
Every verdict, with the sentence that settled it, is in the script output below (–list). The script refuses to run if any entry lacks an adjudication note.
The screening loss
The most important finding about this page's own evidence base.
The extraction's selection screen labels each abstract securityMeasurement and privacyMeasurement and keeps a paper if either is true. An algorithm audit is frequently neither — a search-personalisation or price-discrimination study reads as fairness, economics or information retrieval.
| Query | Count |
|---|---|
records in data/corpus2/.meta (the bibliographic index) | 16,864 |
label records in data/labels/run1/labels.jsonl | 15,800 |
| audit-topical candidates in the index (probe 5) | 117 |
| … in the 5,859-paper extraction | 69 |
| … not in the extraction | 48 |
… of those, screened out with both labels false | 45 |
| … of those, in a venue-year with no label records at all | 3 |
The script asserts that these buckets sum, and that seven named papers are still found by the query — so a corpus refresh that quietly re-admits them will fail the run rather than leave a stale claim on the page.
Verified individually against labels.jsonl:
| Paper | Why it is not in the extraction |
|---|---|
| Hannak et al., TheWebConf 2013, Measuring personalization of web search | securityMeasurement=false, privacyMeasurement=false |
| Hannak et al., IMC 2014, Measuring Price Discrimination and Steering on E-commerce Web Sites | securityMeasurement=false, privacyMeasurement=false |
| Imana et al., TheWebConf 2021, Auditing for Discrimination in Algorithms Delivering Job Ads | securityMeasurement=false, privacyMeasurement=false |
| Boeker and Urman, TheWebConf 2022, An Empirical Investigation of Personalization Factors on TikTok | securityMeasurement=false, privacyMeasurement=false |
| Soeller et al., TheWebConf 2016, MapWatch | securityMeasurement=false, privacyMeasurement=false |
| Vissers et al., PETS 2014, Crying Wolf? | no label record — PETS 2010–2014 has none |
| Khattak et al., NDSS 2016, Do You See What I See? | no label record — NDSS 2016 has none |
The full list of 48 is in the –list output below. Consequence, stated on the content page: every audit count on it is a lower bound, biased against fairness-framed work, and the page publishes no estimate of the wider literature's size.
Folding
Almost nothing on this page needs folding, because almost nothing on it is a free-text aggregate — the population is hand-keyed and the rest are enum-backed counts. Two exceptions:
statistics.methodis free text and ~20% stable run-to-run. It is folded to an alphanumeric skeleton (lowercase, non-alphanumerics stripped) and paper-counted, and it is published as a ranking, not as percentages. The fold does not merge synonyms:Holm-Bonferroni correction(3),Holm-Bonferroni(1) andHolm-Bonferroni method(1) are three rows in the raw output. An unfolded reading would publish “Holm–Bonferroni 3”; the true paper count for the Holm–Bonferroni family is 5 of 32, and for any Bonferroni-family correction 7 of 32. The report script now prints the hand-folded family counts alongside the raw skeleton ranking, so the page quotes a number the script produced rather than one assembled in prose. Both are in the output below.- The eight noise-baseline phrasings are not a fold but a deliberately widened probe (
_aa_noise.mjs): the narrow term “A/A test” returns 2 papers corpus-wide, so seven further phrasings were added, taking the count from 0 to 11. That probe is now published as a vocabulary finding only, not as a practice rate — see the next section.
Residue of the candidate probes. Probe 2 returned 190 candidates. 28 of the final 34 audits and 5 of the 36 rejections lie inside it — the rest were caught by probes 4–7 — leaving a residue of 157 papers dropped at title-and-abstract level without an individual note. The union across all seven probes is 461, of which 70 were adjudicated in depth, so 391 rest on a title-level read. That is the honest residue of this page.
The null: why it is adjudicated and not counted
The page's central claim is that establishing a null is the skipped step. The first version of this log said that claim rested on “a term count plus an eight-way concept probe, never as 'nobody does this'”, and the page then went ahead and wrote “only 11 of 32 establish a same-treatment baseline at all” — which is a practice claim, from a term probe, three times over. A generic review pass caught the gap between the caveat and the sentences.
It was right, and the probe is wrong in both directions:
| Paper | Term that matched | What the sentence actually says |
|---|---|---|
| [12Robertson, Ronald E.; Lazer, David; Wilson, Christo (2018): "Auditing the Personalization and Composition of Politically-Related Search Engine Results Pages", in: Proceedings of the ACM Web Conference. (DOI)] | noise floor | “Hannak et al. found (1) evidence of general personalization above a noise floor” — a description of somebody else's result. This paper pairs standard and incognito windows; it has no same-treatment pair |
| [11Iqbal, Hassan; Khan, Usman Mahmood; Khan, Hassan Ali; Shahzad, Muhammad (2022): "Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Election 2020", in: Proceedings of the ACM Web Conference. (DOI)] | same treatment | “whether the SFA of any given email service provided same treatment to similar emails from candidates of different political affiliations” — the research question, not a baseline |
| [13Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] | none of the eight | four unused baseline profiles “created and mechanistically measured in the same way”, with the divergence of one read as the noise. The clearest A/A in the corpus, invisible to the probe |
So the question was put to all 34 papers by hand, the same way the population was, with a stated definition and an evidence sentence per verdict, in NULLS in algorithm_audits_set.mjs. A paper establishes a null if it measures outcome variation under the same treatment, by any of three routes: A/A (arms differing in nothing), repeats (one condition run several times with the spread reported), or generated (a null distribution built from the paper's own observations — permutation over arm labels, or a randomised baseline).
Result: 10 of 34 (29.4%), against 11 of 32 from the probe. The two numbers are close and that is a coincidence, not corroboration: the sets differ by five papers.
What this is not:
- Not mechanical. It is a judgement per paper. [14Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] is counted in on repeats — six runs of each condition a week apart, with outliers cut on the country-specific standard deviation — and a second reader could reasonably say that is replication rather than a null. [15Roongta, Ritik; Jose, Julia; Habib, Hussam; Greenstadt, Rachel (2025): "Sheep's Clothing, Wolfish Intent: Automated Detection and Evaluation of Problematic 'Allowed' Advertisements", in: Proceedings on Privacy Enhancing Technologies. (DOI)] is counted out, because its fresh control browser establishes a baseline set of ad exchanges for a mediation analysis rather than a same-treatment outcome spread. Both are arguable. The evidence sentence is published for each so the disagreement can be about a specific paper.
- Not a quality verdict. Several of the 24 report effects far larger than any plausible noise floor. What they cannot do is show the reader that.
- Not re-measured after the population grew.
NULLSwas adjudicated against the 34-paper set, including both papers added in the same pass.
Recall repair
Two sweeps were run after the population was settled, because the probe-7 result above showed the scoring had silently discarded 9.4% of it.
1. The dropped tail of probe 2. scripts/_aa_residue.mjs takes the probe-2 candidates that were never adjudicated and re-scores them with the apparatus-density probe that recovered the three late finds, at a lower threshold and requiring an outcome term. 33 of the 157 clear it. All were read at title level; none is a differential platform audit. They fall into three groups: user-study randomised trials where the treatment is applied to a person and the outcome is that person's behaviour (phishing warnings, Tor nudges, consent dialogs, personalised cookie banners), platform-side A/B tests run with the platform rather than against it (How Intention Informed Recommendations Modulate Choices, Reducing Symbiosis Bias through Better A/B Tests), and passing mentions. Output below.
2. The five zero-years. The audit set has no papers in 2011, 2012, 2013, 2017 or 2021. _aa_gap.mjs requires ≥4 apparatus matches, which is too strict to prove a negative — it returns nothing at all for 2017, so a zero from it is uninformative. scripts/_aa_zeroyears.mjs drops the threshold to ≥2 and requires no outcome term, and refuses to run if any year it is given is not actually a zero-year. Over 1,002 papers in those five years it finds 2 with ≥2 apparatus terms: An Army of Me (TheWebConf 2017), which is an observational study of sockpuppets other people created, and Spotting fake reviewer groups in consumer reviews (TheWebConf 2012), which is not an audit either. The zero-years are real in this corpus — which is a different claim from their being real in the field, since 2013 is the year of Hannak et al. and the screen dropped it.
An earlier version of this log said probe 7 had been re-run on 2017 and 2021 and “returned nothing but the observational sockpuppet study”. That was wrong: _aa_gap.mjs 2017 returns nothing at all and _aa_gap.mjs 2021 returns four unrelated papers. The sockpuppet result came from a different, looser ad-hoc probe that was never committed. It is committed now, as _aa_zeroyears.mjs, and the claim above is what it actually prints.
Quotes checked
scripts/quotecheck_algorithm_audits.mjs verifies every quoted fragment on the content page against the paper's own text, in three modes (whitespace-collapsed exact; hyphen- and quote-normalised; longest 8-word run), against both paper.cols.txt and paper.norm.txt. It exits non-zero on any failure.
28 quotes, 0 failures, all matching exactly in paper.cols.txt. Full output below.
Two quotes were caught and removed before publication, both column-splice artefacts:
- [16Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]: the sentence naming the two seeded interest sets reads, in both renderings, “C was Even with static DNS entries, we sometimes (but not al- seeded with long-term interests in 'Autos & Vehicles', while ways) observed discrepancies…” — two columns interleaved. The contiguous fragment “enabled, but are seeded with different user personae” is used instead, and no quote naming the two interest sets is published.
- [9Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)]: this PDF is interleaved throughout in both
colsandnorm. “we first selectively expose a” is followed by text from the adjacent column. No verbatim quote from this paper is published anywhere on either page; its contribution is described instead.
That is the page's one unfixable data-quality problem and it is why the checker tries two renderings rather than one.
External and industry sources
Deliberately few. This page's subject is a research method, not a product, so there is no vendor documentation to verify and no version numbers to date.
| Source | How verified | Verdict |
|---|---|---|
petsymposium.org/popets/2023/popets-2023-0123.php, …/2025/popets-2025-0050.php, …/2026/popets-2026-0152.php | fetched with curl and a browser User-Agent on 2026-09-11; author lists read off the landing pages, because PETS records in the index have no authors (100% of 2,974 PETS/USENIX records) | used — three BibTeX entries |
petsymposium.org/2014/papers/Vissers.pdf (from the index record's pdfUrl) | filename confirms the first author of Crying Wolf?; the index has no author list for it | used for “Vissers et al.”; no BibTeX entry added, because the paper is cited by title only |
OpenAlex, via scripts/bibgen.mjs | DOIs and author lists for the nine non-PETS additions come from the index's OpenAlex records, not from recall | used |
| Any industry writing on “algorithm auditing” (consultancy and NGO audit frameworks, AI-audit vendors) | — | rejected, not searched. The page's claims are about how measurement papers are designed. An AI-governance vendor's audit checklist is a different object with the same name, and importing it would have been the SEO-listicle failure mode in a new costume. |
What could not be established
- The size of this literature. The corpus cannot give it; see The screening loss. Closing it needs a pass over FAccT, EuroS&P and the IR venues, which is out of scope for a corpus-backed page.
- Whether the papers dropped at title level contain any more true audits. The union across all seven probes is 461; 70 were adjudicated in depth, so 391 rest on a title-and-abstract read. The probe-2 tail was re-swept (Recall repair) and a generic review pass found one more audit in the score-≥4 shortlist; both are reasons to expect the 391 is not empty. A full-text read of all of them would settle it and was not done.
- Whether papers build a null without naming it. The eight-term probe is a lower bound. Establishing the real rate needs 32 methods sections read for the concept, which is a different and slower exercise than the verdict read that produced the population.
- Effect sizes. Every audit measures a different outcome on a different platform with a different metric. Nothing is poolable, and the page publishes no cross-paper effect size deliberately.
- Whether sharing an egress IP across arms actually biases an ad-targeting outcome. No paper in the corpus measures it. Filed as an Open Question on the content page rather than asserted.
- The 2011–2013 and 2017 and 2021 zeros. They are real in this corpus. Whether they are real in the field is exactly the question the screening loss prevents answering — 2013 is the year of Hannak et al., which the screen dropped.
Judgement calls
- A new page, not a section. Reasoning in the first table above.
- Hand-adjudicated population rather than a regex population. A regex set would be reproducible and wrong: the tight probe (27 papers) includes an inaudible-voice-command attack and a NIST privacy-framework paper, and misses AdFisher. Hand verdicts are recorded with their evidence sentence and the script refuses to run without them.
- Attack and defence papers counted as audits. See The inclusion rule. Counting only audit studies would give 26 rather than 32 and would exclude the pollution attack, which is one of the clearest illustrations of a blank-vs-trained-profile contrast in the corpus.
- “Price discrimination is dormant” was drafted and then withdrawn. The first draft said the last price-outcome audit was [17Chen, Le; Mislove, Alan; Wilson, Christo (2015): "Peeking Beneath the Hood of Uber", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] (2015). Checking the post-2015
price discrimination/price steeringhits found [18Becerril-Arreola, Rafael (2023): "A Method to Assess and Explain Disparate Impact in Online Retailing", in: Proceedings of the ACM Web Conference. (DOI)], whose outcomes are “price, recommendations, and delivery fees”. The row now reads “rare, not dormant”, with 2 of 32 measuring a price. - The A/A-test claim is stated as a **term count plus an eight-way concept probe, never as “nobody does this”. The distinction is load-bearing: the corpus can prove the vocabulary is absent, and can only lower-bound the practice. - **
design:platformsreports 26 for “sock puppet” corpus-wide and 14 within its own platform-study population. This page reports the same 26. Both are correct; the page says so explicitly rather than letting a reader find two numbers. - The crawl-config reporting rows were re-based mid-run. They first compared 32 audits (8 with no crawl configuration) against 1,120 crawled papers (40 with none). On that mismatched denominator the interaction-depth row read 65.6% against 75.1% — audits worse than the baseline. Restricted to papers that have a configuration on both sides it is 87.5% against 77.9%: the sign flips. The report script now prints both framings and the page uses the shared one. This was found by my own check, not by a reviewer, and it is the exact failure the site's own house rule about denominators exists to prevent.
- No
~~DISCUSSION~~block on this provenance page. Comments belong on the content page. This is the default recorded forprovenance:pages and it is followed here.
Reviewer findings
Four reviewers, all told explicitly that the author's context may not be exhaustive, and all handed the page text, the scripts, their outputs and these notes. The three focused passes ran in parallel first; the generic pass ran afterwards, on the corrected text.
Figures vs script (''sonnet'')
Re-ran all four scripts (outputs byte-identical to the committed ones), re-derived every candidate count from the five probe scripts, and mutation-tested the contracts.
| Finding | Verdict | What changed |
|---|---|---|
The probe table claimed “tight ⊆ loose was asserted”, but _aa_cands.mjs only console.logs the count — mutating the loose threshold made it print tight_not_in_loose=14 and exit 0 | accepted | a real process.exit(1) was added and mutation-tested (mutated run exits 1, restored run exits 0); the table now says what the script does |
Nothing guards the hand-keyed population size. Deleting an entry from AUDITS produced a fully self-consistent report with N=31 and still printed “OK: all contracts held”, while the page says 32 in several places | accepted | audit.length !== 32, REJECTED.length !== 12 and an AUDITS∩REJECTED check now die(); mutation-tested |
The claim that re-running probe 7 on 2017 and 2021 “returned nothing but the observational sockpuppet study” is false — _aa_gap.mjs 2017 returns nothing at all and _aa_gap.mjs 2021 returns four unrelated papers | accepted, verified independently | the narrative was wrong, not the conclusion. The looser probe that actually produced that result was uncommitted; it is now _aa_zeroyears.mjs, run over all five zero-years (1,002 papers, 2 hits, neither an audit). See Recall repair |
| Everything else — every table cell, year bucket, venue row, the 8-term noise table and its 11-paper list, the vocabulary table, the statistics ranking, the hand-folded Bonferroni families, the 11-paper no-test list, and all six probe candidate counts | no defect | — |
Citations and quotes (''sonnet'')
Checked all 12 new keys against data/corpus2/.meta or Crossref, ran a DOI-and-title dedup scan over all ~993 bibliography entries, re-grepped the quotes independently of the checker, and verified every prose attribution.
| Finding | Verdict |
|---|---|
| All 12 new keys resolve exactly once; no key, DOI or title collision; all author lists, titles, years, venues and DOIs match their primary record | no defect |
28/28 quotes pass; the reviewer additionally hand-verified five quotes the checker's array does not cover (paired puppets and the same-network sentence in [19Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], “Independent Auditing System” in [6Silva, Márcio; Oliveira, Lucas Santos de; Andreou, Athanasios; Melo, Pedro Olmo Stancioli Vaz de; Goga, Oana; Benevenuto, Fabrício (2020): "Facebook Ads Monitor: An Independent Auditing System for Political Ads on Facebook", in: Proceedings of the ACM Web Conference. (DOI)], “price, recommendations, and delivery fees” in [18Becerril-Arreola, Rafael (2023): "A Method to Assess and Explain Disparate Impact in Online Retailing", in: Proceedings of the ACM Web Conference. (DOI)], and the cross-page quote from design:platforms) — all verbatim and contiguous | no defect |
Every surname order in prose, every flagged attribution, every row of the screening-loss table against labels.jsonl, and the footnote's three non-audit “sock puppet” papers | no defect |
| Numeric claims spot-checked in the source PDFs: the 1–4%/minute ad churn, the 43 Uber accounts, the 22,722 participants | no defect |
Accidental exposure caught by the reviewer. Mid-review it observed algorithm_audits_set.mjs in a transient state with [16Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] missing and N=31, and flagged that the file “appears to have flickered during the session” | acknowledged. That was my own mutation test of the new population-size contract, running against the same working copy the reviewer was reading. The set was restored and re-verified. Recorded here because a reviewer seeing a file mutate under it is exactly the kind of thing that should not be quietly dropped — and because it is an argument for mutation-testing on a copy, not in place |
External currency (''sonnet'')
Fetched rather than recalled, as of 2026-09-11.
| Finding | Verdict |
|---|---|
| All four external URLs resolve; the three PoPETs landing pages' author lists match the BibTeX entries exactly | no defect |
The page makes no legal claim of its own and routes to design:platforms and practices:ethics. The reviewer checked those siblings rather than assuming silence, and re-verified their DSA Art. 40 delegated act (Commission Delegated Regulation (EU) 2025/2050, in force 2025-10-29) and the X DSA decision against primary sources | no defect |
X appealed the DSA fine to the General Court on 2026-02-16; design:platforms states the fine but not the appeal | accepted as out of scope. It is a defect on that page, not this one, and was not actioned here |
| AutoLike (arXiv 2502.08933) proposed as a possibly-missing 2025 audit | rejected, by the reviewer and again by me: still a preprint, not in the seven venues, and no control arm — it fails criterion (C) |
| No tool renamed or discontinued; the three 2026 papers live and unretracted; CCS 2026 (15–19 Nov) and IMC 2026 (12–16 Oct) confirmed not yet held while USENIX Security 2026 has been, matching the page's provisional-year framing exactly | no defect |
| The AI Act contains no researcher-access provision bearing on sock-puppet methodology, so its absence is not a gap | no defect |
Generic (''fable'')
Handed the corrected text after the three focused passes. No checklist; asked for whatever they were not looking for. It returned nine findings, ranked. All nine accepted; the first three changed the page materially.
| # | Finding | Verdict and what changed |
|---|---|---|
| 1 | The central claim states a practice (“only 11 of 32 establish a same-treatment baseline”, “the practice is missing too”, and the heading “The null is missing from two thirds of them”) on the strength of a term count — and the reviewer printed the matching sentence behind all 11 hits, showing the probe over-counts as well as under-counts | accepted. All 34 papers were read for the question and hand-adjudicated with an evidence sentence each: 10 of 34. The term table stays, relabelled as a vocabulary finding. See The null: why it is adjudicated and not counted |
| 2 | The adjudication dropped candidates silently: 26 of the 58 score-≥4 shortlist papers had no verdict, “every verdict is in the output” was true only of the 44 that got one, and one of the 26 — Cart-ology — satisfies the rule | accepted, and it was right about Cart-ology. All 58 now carry a written verdict. Cart-ology and the IMC 2024 poster promoted; population 32 → 34; rejections 12 → 36. Root cause recorded above: _aa_adj.mjs returned zero sentences for Cart-ology and that silence was read as a negative |
| 3 | Nothing links to the page except the two routing pages, and Platforms still says twice that this topic “has no page on this wiki” — the exact sentence used to justify creating it | accepted. Back-links added from the neighbour pages, and the two stale sentences on Platforms repointed |
| 4 | “Most audit papers supply the first three and skip the fourth” contradicts the page's own 43.8% vantage-location figure | accepted. Replaced with the actual figures |
| 5 | “Twenty profiles per condition … for the same total cost” is false — same measurements, ten times the training, and the page says training is the expensive part | accepted. Rewritten to say what it actually costs and why the corpus is full of two-profile designs. The unsupported “audits are its most natural habitat” was cut too |
| 6 | [20Sun, Chen; Vekaria, Yash; Nithyanand, Rishab (2026): "On the Suitability of LLM-Driven Agents for Dark Pattern Audits", Proceedings on Privacy Enhancing Technologies 2026(4):927-946. (DOI)] is in neither AUDITS nor REJECTED yet gets four mentions | accepted. Trimmed to one table row that says plainly it is not one of the 34, plus the link |
| 7 | “This method is not historical. It is growing.” is not what a 6/4/4/3-year window table with n = 34 from a biased screen supports | accepted. Replaced with papers-per-year (1.00, 1.25, 3.00, 3.67), the observation that the second window is barely above the first, and a narrower claim |
| 8 | “used continuously 2010 → 2026” for personas is not script-derived; the first persona hit is 2016 and there are none in 2017–2019 | accepted. Corrected, with the note that [16Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] does the same thing and spells it personae |
| 9 | Smaller: 146 survived in one place after being corrected to 158 elsewhere; “the two papers most cited” with no citation count; “never drive a browser” said of studies that run in participants' browsers; “Chrome 141” stale in a sentence dated 2026-09 | all accepted. The residue figures were re-derived for the new sets (157 and 391); the “most cited” clause is gone with the rewritten section; the eight no-config papers are described accurately; Chrome pinned to 154, fetched from the Chrome version-history API on 2026-09-11 |
What it found fine, and said so: length (33 KB against 50–70 KB neighbours), no duplication of the four neighbour pages, the arms-versus-waves noise-floor distinction, the 26-vs-14 sock-puppet reconciliation, the What this corpus cannot tell you box, the re-based reporting table, and the What to report list with its one-sentence template.
One thing it flagged that was not a defect: the literal token GENERICREVIEW was live on this page at the time of review. That was this section's own placeholder.
The report script
Committed as scripts/report_algorithm_audits.mjs with the population in scripts/algorithm_audits_set.mjs. It exits 1 if the corpus size, the empirical / crawled / measuredFrom populations, the year and venue bucket sums, the screening-loss bucket sums, or the seven named screening losses disagree with the contracts it encodes.
- algorithm_audits_set.mjs
// The hand-adjudicated audit population for design:algorithm_audits. // Imported by report_algorithm_audits.mjs and by the probe scripts, so every // figure on the page and every probe share one definition of the set. // // --------------------------------------------------------------------------- // The inclusion rule, as applied. A paper is IN if all three hold: // (T) it deliberately varies a property of the measuring identity or request // -- profile history, declared attribute, location, device, opt-out // setting, ad creative -- and holds the rest fixed; // (O) the outcome it measures is the platform's own DISCRIMINATING response to // the identity it has built: ads served, results ranked, prices quoted, feed // or recommendation contents, or the profile the platform reports back. // NOT reachability. A study whose outcome is whether you can reach the site // at all -- geoblocking, censorship, Tor-exit refusal -- varies a vantage // point rather than an identity and belongs to design:blocking_and_geodifference. // This clause was tightened on 2026-09-11; it excludes no paper already in // the set, and it is what keeps "403 Forbidden: A Global View of CDN // Geoblocking" (IMC 2018) out. // (C) the result is a difference (or a bounded absence of difference) // BETWEEN arms, not a prevalence over a crawl of many sites. // The `why` string is the evidence sentence that settled (T)+(C). // --------------------------------------------------------------------------- export const AUDITS = [ ['IMC/2010/challenges-in-measuring-online-advertising-systems', 'seeded profile pairs: "enabled, but are seeded with different user personae" (the sentence naming the two interest sets is column-spliced in both renderings)'], ['CCS/2014/your-online-interests-pwned-a-pollution-attack-against-targeted-advertising', 'blank profile vs polluted profile: "the polluter can impersonate a user with a blank profile ... and browse pages"'], ['CCS/2015/sunlight-fine-grained-targeting-detection-at-scale-with-statistical-confidence', 'instrument paper: varies personal-data inputs one at a time "compared to a control group", with statistical confidence'], ['IMC/2015/location-location-location-the-impact-of-geolocation-on-web-search-personalizati', 'location as the treatment: "all other browser attributes were the same across treatments"'], ['IMC/2015/peeking-beneath-the-hood-of-uber', '"We created 43 Uber accounts ... blanket a small geographic area with measurement points"'], ['PETS/2015/automated-experiments-on-ad-privacy-settings', 'AdFisher: "We created an experimental group and a control group of agents"'], ['USENIX/2016/tracing-information-flows-between-ad-exchanges-using-retargeted-ads', '"We train 90 personas by visiting popular e-commerce sites, and then crawl major publishers"'], ['WWW/2018/adbudgetkiller-online-advertising-budget-draining-attack', '"Starting from a fresh profile, the profile trainer produces" crafted browsing profiles'], ['WWW/2018/auditing-the-personalization-and-composition-of-politically-related-search-engin', 'standard-vs-incognito paired SERPs: "our controls were paired within the individual"'], ['PETS/2019/investigating-sources-of-pii-used-in-facebook-s-targeted-advertising', '"Take a Facebook account that we control (call it the control account) and the test PII"'], ['WWW/2019/measuring-political-personalization-of-google-news-search', '"a \'sock puppet\' auditing system in which a pair of fresh browser profiles" visits divergent pages then runs identical queries'], ['PETS/2020/inferring-tracker-advertiser-relationships-in-the-online-advertising-ecosystem-u', 'intent vs no-intent versions of each of 16 interest personas, with trackers selectively exposed (this PDF is column-spliced in both cols and norm renderings, so no contiguous quote was publishable)'], ['WWW/2020/stop-tracking-me-bro-differential-tracking-of-user-demographics-on-hyper-partisa', '"We create 9 carefully crafted personas representing different genders and age groups"'], ['CCS/2022/cart-ology-intercepting-targeted-advertising-via-ad-network-identity-entanglemen', 'five profiles -- two blank baselines, two account baselines, one entangled attack profile: "All profiles are created and mechanistically measured in the same way, with separation between attacker, victim, and baselines on different machines with different IPs"'], ['IMC/2022/measurement-and-analysis-of-implied-identity-in-ad-delivery-optimization', '"We ran 200 versions of this ad at the same time, all from the same account and with the same budget"'], ['IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising-a-field-measurement', '"All participants were asked to visit the same websites to control for contextual targeting, in randomized order"'], ['NDSS/2022/auto-draft-209', 'HARPO: obfuscated vs unobfuscated personas measured against live profiling models'], ['PETS/2022/atom-ad-network-tomography', '"create a number of online user personas associated with specific interest groups" and gather ads while systematically blocking trackers'], ['WWW/2022/left-or-right-a-peek-into-the-political-biases-in-email-spam-filtering-algorithm', '"We created 102 email accounts" and compared spam placement across treatment and control affiliations'], ['CCS/2023/marketing-to-children-through-online-targeted-advertising-targeting-mechanisms-a', '"We launch the six personas simultaneously by creating six Selenium" browsers'], ['IMC/2023/tracking-profiling-and-ad-targeting-in-the-alexa-echo-smart-speaker-ecosystem', '"Each treatment persona is simulated by installing and interacting with skills ... By contrast, in the control persona, we do" not'], ['PETS/2023/a-utility-preserving-obfuscation-approach-for-youtube-recommendations', '"We deploy and evaluate De-Harpo\'s effectiveness on YouTube using 10,000 sock puppet based personas"'], ['WWW/2023/a-method-to-assess-and-explain-disparate-impact-in-online-retailing', '"Each observation ... is paired with a set of \'counter-factuals\'" from neighbouring zip codes'], ['PETS/2024/opted-out-yet-tracked-are-regulations-enough-to-protect-your-privacy', '"We then visit the filtered websites with user interest personas" -- 16 interest personas, two regions'], ['IMC/2024/poster-identifying-filter-bubble-based-on-feed-level-embedding-similarity-analys', 'four bot arms differing only in the video-selection strategy: "The bot selected a video having (a) the minimum JCC, (b) JCC larger than the minimum (random selection), (c) JCC larger than the mean, and (d) JCC larger than WCC"; outcome is the recommended feed'], ['WWW/2024/tiktok-and-the-art-of-personalization-investigating-exploration-and-exploitation', '"validate our results using a baseline generated from automated TikTok bots, as well as a randomized baseline"'], ['IMC/2025/from-voice-to-ads-auditing-commercial-smart-speakers-for-targeted-advertising-ba', '"we used each of the 21 cloned voice models to train two puppets, which we refer to as \'paired puppets\'"'], ['PETS/2025/echoes-of-privacy-uncovering-the-profiling-practices-of-voice-assistants', '"each meticulously trained with a curated set of voice queries designed to simulate various user personas"'], ['PETS/2025/more-and-scammier-ads-the-perils-of-youtubes-ad-privacy-settings', '"running controlled experiments with sock puppet accounts that emulate users watching YouTube videos"'], ['PETS/2025/sheeps-clothing-wolfish-intent-automated-detection-and-evaluation-of-problematic', '"two separate crawlers were deployed: one for the Control group (without ABP) and another for the AccAds group"'], ['USENIX/2025/big-help-or-big-brother-auditing-tracking-profiling-and-personalization-in-gener', '"Training phase involves browsing through 10 webpages - 2 pages per leaked attribute", then personalisation is measured'], ['IEEE-SP/2026/setting-the-course-but-forgetting-to-steer-analyzing-compliance-with-gdprs-right', '"we use sock-puppet accounts to systematically browse and log the behavior of the sock-puppet to generate the ground truth"'], ['PETS/2026/ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis', '"We construct accounts with specific parameters or interests, so-called personas, and measure the ads displayed to them"'], ['WWW/2026/does-this-button-work-investigating-youtubes-ineffective-user-controls', '"depending on randomized assignment - triggered one of four native feedback signals to YouTube ... or no signal at all in the control group"'], ]; // Adjudicated OUT, with the reason. These are the near misses a later run will // otherwise re-add; the rule that excluded them is printed with each. export const REJECTED = [ ['WWW/2017/an-army-of-me-sockpuppets-in-online-discussion-communities', 'observational study OF sockpuppets others created; no arm the authors control'], ['WWW/2019/auditing-the-partisanship-of-google-search-snippets', 'audits snippet vs linked page; no identity treatment, no arms'], ['WWW/2020/facebook-ads-monitor-an-independent-auditing-system-for-political-ads-on-faceboo', 'volunteer ad donation; observational, no arms'], ['USENIX/2020/what-twitter-knows-characterizing-ad-targeting-practices-user-perceptions-and-ad', "users' own Twitter data; no arms"], ['IEEE-SP/2023/collaborative-ad-transparency-promises-and-limitations', 'explicitly the non-persona alternative: "One method that does not use fake personas"'], ['USENIX/2023/problematic-advertising-and-its-disparate-exposure-on-facebook', 'donated ad exposure from real users; disparity is observational'], ['CCS/2022/privacy-limitations-of-interest-based-advertising-on-the-web-a-post-mortem-empir', 'randomised control is a data permutation over a browsing panel, not a live-platform arm'], ['IEEE-SP/2024/targeted-and-troublesome-tracking-and-advertising-on-childrens-websites', 'fresh profile per page visit; the paper itself says "Future work could extend our method to incorporate personas"'], ['PETS/2024/interest-disclosing-mechanisms-for-advertising-are-privacy-exposing-not-preservi', 'Topics API analysed over real browsing histories, not persona arms'], ['WWW/2026/when-ads-become-profiles-uncovering-the-invisible-risk-of-web-advertising-at-sca', '"Random Control group" is a model ablation, not a platform arm'], ['IEEE-SP/2022/deployment-of-source-address-validation-by-network-operators-a-randomized-contro', 'an RCT, but the treatment is a notification to operators; outcome is not platform output'], // The rest of the score->=4 shortlist. A generic review found these had been read // and dropped without a written verdict, and that one of the papers so dropped // (Cart-ology) in fact satisfied the rule. Every shortlisted paper now has a verdict. ['NDSS/2015/bloom-cookies-web-search-personalization-without-user-tracking', 'a privacy-preserving personalisation design evaluated on search logs; no arms'], ['CCS/2018/peeling-the-onions-user-experience-layer-examining-naturalistic-use-of-the-tor-b', 'a 19-participant UX study of Tor Browser; outcome is user experience'], ['USENIX/2023/strategies-and-vulnerabilities-of-participants-in-venezuelan-influence-operation', 'interviews with 19 influence-operation participants; the sockpuppets are theirs, not the authors'], ['CCS/2011/policy-auditing-over-incomplete-logs-theory-implementation-and-applications', 'compliance checking over audit logs; shares the word "audit" and nothing else'], ['WWW/2013/your-browsing-behavior-for-a-big-mac-economics-of-personal-information-online', '168 recruited participants valuing their own PII in an auction; a user study'], ['IEEE-SP/2015/effective-real-time-android-application-auditing', 'program analysis of Android apps; "auditing" means taint tracking'], ['USENIX/2017/exploring-user-perceptions-of-discrimination-in-online-targeted-advertising', 'a survey of user perceptions; the randomised factors are survey vignettes, not platform arms'], ['IMC/2018/403-forbidden-a-global-view-of-cdn-geoblocking', 'vantage points in 177 countries, but the outcome is REACHABILITY. Excluded by the tightened (O) clause; belongs to design:blocking_and_geodifference'], ['CCS/2019/the-art-and-craft-of-fraudulent-app-promotion-in-google-play', 'a study of app-store fraud workers who operate sockpuppets; observational'], ['WWW/2022/characterizing-detecting-and-predicting-online-ban-evasion', "Wikipedia's own labelled sockpuppet groups; observational, no arms"], ['WWW/2022/fairness-audit-of-machine-learning-models-with-confidential-computing', 'ML fairness auditing inside a TEE; no live platform'], ['IEEE-SP/2022/towards-automated-auditing-for-account-and-session-management-flaws-in-single-si', 'SSO implementation flaws; outcome is a security bug, not a discriminating response'], ['NDSS/2023/tactics-threats-targets-modeling-disinformation-and-its-mitigation', 'interviews with fact-checkers and analysts; sockpuppets are the subject, not the instrument'], ['WWW/2023/scoping-fairness-objectives-and-identifying-fairness-metrics-for-recommender-sys', "practitioner interviews about fairness metrics; no measurement of a platform"], ['IEEE-SP/2023/when-and-why-do-people-want-ad-targeting-explanations-evidence-from-a-four-week', 'a four-week field study of what people want from ad explanations; outcome is attitudes'], ['USENIX/2023/auditing-framework-apis-via-inferred-app-side-security-specifications', 'Android framework API access control; "auditing" is static analysis'], ['CCS/2024/curator-attack-when-blackbox-differential-privacy-auditing-loses-its-power', 'differential-privacy auditing; a different object with the same name'], ['USENIX/2024/efficient-privacy-auditing-in-federated-learning', 'membership inference against an FL model; not a platform'], ['USENIX/2024/what-do-you-want-from-theory-alone-experimenting-with-tight-auditing-of-differen', 'DP synthetic-data auditing; not a platform'], ['USENIX/2024/fledging-will-continue-until-privacy-improves-empirical-analysis-of-googles-priv', 'security analysis of the FLEDGE/Protected Audience API; the outcome is API behaviour and attack feasibility, not a served-ad contrast across arms'], ['NDSS/2025/exploring-user-perceptions-of-security-auditing-in-the-web3-ecosystem', 'user perceptions of smart-contract audits'], ['PETS/2025/privacy-perceptions-and-behaviors-towards-targeted-advertising-on-social-media-a', 'an n=412 cross-country survey; outcome is attitudes'], ['PETS/2026/audagent-automated-auditing-of-privacy-policy-compliance-in-ai-agents', 'policy-compliance checking of AI agents; no arms'], ['PETS/2026/privacy-in-theory-bugs-in-practice-grey-box-auditing-of-differential-privacy-lib', 'DP library auditing; not a platform'], ['WWW/2026/question-the-questions-auditing-representation-in-online-deliberative-processes', 'algorithm design for question selection in deliberative polls; not a measurement of a deployed platform'], ]; // Does the paper measure outcome variation under the SAME treatment? That is the // step the page argues is skipped, and a term probe cannot answer it: the eight // phrasings in _aa_noise.mjs let in a "noise floor" that describes someone else's // paper, and miss Cart-ology's four identically-measured baseline profiles, which // match none of them. So it is hand-adjudicated here, like the population, with // the sentence that settled it. Three routes count: // (a) A/A -- two or more arms that differ in nothing; // (b) repeats -- the same condition run several times, with the spread reported; // (c) generated -- a null distribution built from the paper's own observations // (permutation over arm labels, or a randomised baseline). // Anything not listed here was read and found to state a difference without ever // measuring what a non-difference looks like. export const NULLS = { 'IMC/2010/challenges-in-measuring-online-advertising-systems': ['A/A', '"Even queries launched simultaneously from two identically configured clients on the same subnet can produce wildly different ads"; and "In this paper we perform all analysis relative to a control experiment"'], 'CCS/2015/sunlight-fine-grained-targeting-detection-at-scale-with-statistical-confidence': ['generated', 'exact/random permutation test against an explicit H0, on a held-out split of the profiles, with Benjamini-Yekutieli or Holm correction'], 'IMC/2015/location-location-location-the-impact-of-geolocation-on-web-search-personalizati': ['A/A', '"The red line compares two treatments at the baseline location (i.e., the experimental control), and thus shows the noise floor."'], 'PETS/2015/automated-experiments-on-ad-privacy-settings': ['generated', 'permutation test over the arm labels -- "The permutation test randomly permutes ... the control and experimental treatments" -- run over "blocks of nearly identical agents"'], 'WWW/2019/measuring-political-personalization-of-google-news-search': ['A/A', 'edit distances tested "among four identical browser profiles" as the reference for the trained-profile comparison'], 'CCS/2022/cart-ology-intercepting-targeted-advertising-via-ad-network-identity-entanglemen': ['A/A', 'four unused baseline profiles measured identically; one diverged and the authors read that as the noise: "Such non-determinism is expected, validating our strategy of deploying numerous baseline profiles and the use of normalization."'], 'WWW/2024/tiktok-and-the-art-of-personalization-investigating-exploration-and-exploitation': ['generated', '"a randomized baseline" against which the personalisation features are scored'], 'PETS/2025/more-and-scammier-ads-the-perils-of-youtubes-ad-privacy-settings': ['repeats', 'each condition run six times at least a week apart ("we repeat each experiment six times"), with outliers cut at standard deviations of the country-specific mean'], 'IMC/2025/from-voice-to-ads-auditing-commercial-smart-speakers-for-targeted-advertising-ba': ['A/A + generated', 'paired puppets differing in nothing, scheduled in parallel; plus a permutation test against a random dataset'], 'PETS/2026/ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis': ['A/A + generated', '"our baseline persona uses the same neutral persona type for control and treatment. For the baseline persona, the majority of measurement pairs produce insignificant tests" -- an A/A that behaved like one -- plus a 9,999-permutation test'], };
- report_algorithm_audits.mjs
#!/usr/bin/env node // Report script for design:algorithm_audits (and provenance:design:algorithm_audits). // // Every figure on the page is printed here with the population it is a share of. // Counts are of PAPERS, never tuples. Sentinels are never answers. // // The audit population is HAND-KEYED below rather than derived from a regex. // That is deliberate: the corpus has no field for "ran a differential audit", // the regex probes that find candidates have 20-60% precision, and a script that // silently re-derives the split from prose is how a wrong split reaches a page. // The candidate probes that produced the list are in _aa_probe1.mjs / _aa_abs2.mjs // and are documented on the provenance page; this file records the verdicts. // // Usage: node scripts/report_algorithm_audits.mjs [--list] import fs from 'node:fs'; import path from 'node:path'; import { dataRoot, loadExtractions, isSentinel } from './lib.mjs'; const SHOW_LIST = process.argv.includes('--list'); const ROOT = dataRoot(); const papers = loadExtractions(); const byKey = new Map(papers.map((p) => [`${p.venue}/${p.year}/${p.slug}`, p])); import { AUDITS, REJECTED, NULLS } from './algorithm_audits_set.mjs'; // --------------------------------------------------------------------------- function die(msg) { console.error(`CONTRACT VIOLATED: ${msg}`); process.exit(1); } const pct = (n, d) => (d === 0 ? 'n/a' : `${((100 * n) / d).toFixed(1)}%`); // Contracts against the corpus, so a corpus refresh cannot silently move a page. if (papers.length !== 5859) die(`corpus is ${papers.length} papers, page says 5,859`); const CRAWLED = papers.filter((p) => p.crawlConfig !== null || p.studyTypes.includes('automated-web-crawl')); const INFERENTIAL = papers.filter((p) => p.statistics.some((s) => s.kind && s.kind !== 'descriptive-only')); const EMPIRICAL = papers.filter((p) => p.isEmpirical === true); if (CRAWLED.length !== 1120) die(`crawled population is ${CRAWLED.length}, OVERVIEW.md says 1120`); if (INFERENTIAL.length !== 1762) die(`inferential population is ${INFERENTIAL.length}, OVERVIEW.md says 1762`); if (EMPIRICAL.length !== 5118) die(`empirical population is ${EMPIRICAL.length}, OVERVIEW.md says 5118`); const audit = []; for (const [key, why] of AUDITS) { const p = byKey.get(key); if (!p) die(`hand-keyed audit paper not in the extraction: ${key}`); if (!why || why.length < 25) die(`no adjudication evidence recorded for ${key}`); audit.push(p); } if (new Set(AUDITS.map((a) => a[0])).size !== AUDITS.length) die('duplicate key in AUDITS'); // The hand-keyed population is this page's ground truth, so nothing external can // check it -- but the PAGE states 32 and 12 as literal numbers in several places. // Pin them here so an accidental edit to the set fails the run instead of // silently moving every figure on the page. if (audit.length !== 34) die(`audit population is ${audit.length}, the page says 34 -- update the page and this contract together`); if (REJECTED.length !== 36) die(`rejected set is ${REJECTED.length}, the page says 36 -- update the page and this contract together`); if (AUDITS.some(([k]) => REJECTED.some(([r]) => r === k))) die('a key is in both AUDITS and REJECTED'); for (const [key] of REJECTED) if (!byKey.get(key)) die(`rejected paper not in the extraction: ${key}`); const N = audit.length; console.log('='.repeat(78)); console.log('design:algorithm_audits -- report script'); console.log(`run ${new Date().toISOString().slice(0, 10)} corpus ${papers.length} papers, 7 venues, 2010-2026`); console.log('='.repeat(78)); console.log(''); console.log('## Populations'); console.log(` all papers ${papers.length}`); console.log(` empirical ${EMPIRICAL.length}`); console.log(` crawled ${CRAWLED.length}`); console.log(` inferential (any non-descriptive statistic) ${INFERENTIAL.length}`); console.log(` AUDIT (hand-adjudicated, rule above) ${N}`); console.log(` adjudicated and rejected ${REJECTED.length}`); console.log(''); // --- by year ----------------------------------------------------------------- console.log('## Audit papers by year (denominator: the audit set)'); const years = [...new Set(papers.map((p) => p.year))].sort(); let run = 0; for (const y of years) { const n = audit.filter((p) => p.year === y).length; run += n; const star = y >= 2025 ? ' *provisional' : ''; console.log(` ${y} ${String(n).padStart(2)} ${'#'.repeat(n)}${star}`); } if (run !== N) die(`year buckets sum to ${run}, not ${N}`); console.log(` total ${run}`); const gaps = years.filter((y) => audit.every((p) => p.year !== y)); console.log(` years with zero audit papers: ${gaps.join(', ')}`); console.log(''); let wsum = 0; for (const [lo, hi] of [[2010, 2015], [2016, 2019], [2020, 2023], [2024, 2026]]) { const n = audit.filter((p) => p.year >= lo && p.year <= hi).length; wsum += n; const span = hi - lo + 1; console.log(` ${lo}-${hi} (${span} yr): ${n} of ${N} (${pct(n, N)}), ${(n / span).toFixed(2)} papers/year`); } if (wsum !== N) die(`window buckets sum to ${wsum}, not ${N}`); console.log(''); // --- by venue and platform --------------------------------------------------- console.log('## Audit papers by venue, against that venue\'s own output'); const venues = [...new Set(papers.map((p) => p.venue))].sort(); let vsum = 0; for (const v of venues) { const tot = papers.filter((p) => p.venue === v).length; const n = audit.filter((p) => p.venue === v).length; vsum += n; console.log(` ${v.padEnd(9)} ${String(n).padStart(2)} of ${String(tot).padStart(5)} ${pct(n, tot).padStart(6)}`); } if (vsum !== N) die(`venue buckets sum to ${vsum}, not ${N}`); console.log(''); console.log('## Platform measured (multi-valued; does not sum to N)'); for (const pl of ['web', 'other-online-service', 'mobile', 'iot', 'offline']) { const n = audit.filter((p) => p.platforms.includes(pl)).length; console.log(` ${pl.padEnd(22)} ${String(n).padStart(2)} of ${N} ${pct(n, N)}`); } console.log(''); // --- what the audit set reports, each row against the same-denominator baseline console.log('## What audit papers report, vs the comparable corpus baseline'); console.log(' Each row states both populations. The baseline is the population named,'); console.log(' not "all papers", so the two cells are comparable.'); const statedStat = (p) => p.statistics.some((s) => s.kind && s.kind !== 'descriptive-only'); const statedEthics = (p) => !isSentinel(p.ethics && p.ethics.reviewOutcome); const hasArtifact = (p) => p.artifacts && !isSentinel(p.artifacts.availability); const hasCrawlCfg = (p) => p.crawlConfig !== null; const statefulStated = (p) => p.crawlConfig !== null && !isSentinel(p.crawlConfig.statefulness); const interactionStated = (p) => p.crawlConfig !== null && !isSentinel(p.crawlConfig.interactionDepth); const vantageStated = (p) => p.vantage.some((v) => (v.locations || []).some((l) => !isSentinel(l))); const rows = [ ['runs a non-descriptive statistic', audit, statedStat, EMPIRICAL, statedStat, 'empirical (5,118)'], ['states an ethics-review outcome', audit, statedEthics, EMPIRICAL, statedEthics, 'empirical (5,118)'], ['states artifact availability', audit, hasArtifact, EMPIRICAL, hasArtifact, 'empirical (5,118)'], ['has a crawlConfig at all', audit, hasCrawlCfg, CRAWLED, hasCrawlCfg, 'crawled (1,120)'], ['states crawl statefulness', audit, statefulStated, CRAWLED, statefulStated, 'crawled (1,120)'], ['states interaction depth', audit, interactionStated, CRAWLED, interactionStated, 'crawled (1,120)'], ['states a vantage location', audit, vantageStated, papers.filter((p) => p.vantage.length > 0), vantageStated, 'measuredFrom (3,908)'], ]; console.log(` ${'indicator'.padEnd(34)} ${`audit (${N})`.padStart(14)} baseline`); for (const [label, aSet, aFn, bSet, bFn, bName] of rows) { const a = aSet.filter(aFn).length; const b = bSet.filter(bFn).length; console.log(` ${label.padEnd(34)} ${(`${a}/${aSet.length} (${pct(a, aSet.length)})`).padStart(14)} ${b}/${bSet.length} (${pct(b, bSet.length)}) of ${bName}`); } console.log(''); console.log('## The crawl-config rows again, on a SHARED denominator'); console.log(` The rows above compare ${N} audits (${N - audit.filter((p) => p.crawlConfig !== null).length} of which have no crawl config) against`); console.log(' 1,120 crawled papers (40 of which have none). A field that can only be stated'); console.log(' on a paper that HAS a config must be counted over papers that have one, or the'); console.log(' two cells are not comparable. Both framings are printed; the page uses this one.'); const auditCfg = audit.filter((p) => p.crawlConfig !== null); const crawlCfg = CRAWLED.filter((p) => p.crawlConfig !== null); if (crawlCfg.length !== 1080) die(`crawled-with-config is ${crawlCfg.length}, OVERVIEW.md says 1080`); for (const k of ['statefulness', 'interactionDepth', 'consentAction', 'headless']) { const a = auditCfg.filter((p) => !isSentinel(p.crawlConfig[k])).length; const b = crawlCfg.filter((p) => !isSentinel(p.crawlConfig[k])).length; console.log(` ${k.padEnd(18)} audit ${a}/${auditCfg.length} (${pct(a, auditCfg.length).padStart(6)}) crawled ${b}/${crawlCfg.length} (${pct(b, crawlCfg.length)})`); } console.log(` -- the ${audit.length - auditCfg.length} audit papers with no crawl config at all:`); for (const p of audit.filter((p) => p.crawlConfig === null)) console.log(` ${p.year} ${p.venue.padEnd(8)} ${p.title}`); console.log(''); // --- statistics named in the audit set -------------------------------------- console.log('## Statistical methods named by audit papers (paper-counted, free text, ranking only)'); const smeth = new Map(); for (const p of audit) { const seen = new Set(); for (const s of p.statistics) { if (!s.method || isSentinel(s.method)) continue; const k = String(s.method).toLowerCase().replace(/[^a-z0-9]+/g, ''); if (seen.has(k)) continue; seen.add(k); smeth.set(k, (smeth.get(k) || 0) + 1); } } const disp = new Map(); for (const p of audit) for (const s of p.statistics) { if (!s.method || isSentinel(s.method)) continue; const k = String(s.method).toLowerCase().replace(/[^a-z0-9]+/g, ''); if (!disp.has(k)) disp.set(k, s.method); } for (const [k, v] of [...smeth.entries()].sort((a, b) => b[1] - a[1]).slice(0, 15)) console.log(` ${String(v).padStart(2)} ${disp.get(k)}`); // Synonym families the alphanumeric skeleton fold does NOT merge. Printed here // so the page quotes a family count produced by this script rather than by prose. console.log(' -- correction families, folded by hand (paper-counted; the skeleton fold does not merge these):'); const FAMILIES = { 'Holm-Bonferroni (any spelling)': (m) => /holm/i.test(m), 'Bonferroni without Holm': (m) => /bonferroni/i.test(m) && !/holm/i.test(m), 'any Bonferroni-family': (m) => /bonferroni/i.test(m), 'Benjamini-Hochberg/Yekutieli': (m) => /benjamini/i.test(m), }; for (const [name, fn] of Object.entries(FAMILIES)) { const n = audit.filter((p) => p.statistics.some((s) => s.method && !isSentinel(s.method) && fn(s.method))).length; console.log(` ${name.padEnd(32)} ${n} of ${N}`); } const noStat = audit.filter((p) => !statedStat(p)); console.log(` -- ${noStat.length} of ${N} audit papers report no non-descriptive statistic at all:`); for (const p of noStat) console.log(` ${p.year} ${p.venue} ${p.title}`); console.log(''); // --- vocabulary the field uses ---------------------------------------------- console.log('## Apparatus vocabulary in full text (paper-counted over all 5,859 with text)'); const VOCAB = { 'sock puppet': /\bsock[ -]?puppets?\b/i, persona: /\b(?:user|shopper|synthetic|training|treatment|control) personas?\b|\bpersonas?\b/i, 'control profile/account/persona': /\bcontrol (?:profile|account|persona|browser)s?\b/i, 'treatment group/profile': /\btreatment (?:group|profile|persona|condition|arm)s?\b/i, 'trained profile': /\btrain(?:ed|ing) (?:the |our |a )?(?:browser )?profiles?\b/i, 'A/A test': /\bA\/A test/i, 'noise floor': /\bnoise floor\b/i, 'price discrimination/steering': /\bprice (?:discrimination|steering)\b/i, }; const vocabCount = Object.fromEntries(Object.keys(VOCAB).map((k) => [k, [0, 0]])); let withText = 0; for (const p of papers) { const f = path.join(ROOT, 'fulltext', String(p.year), p.venue, p.slug, 'paper.cols.txt'); if (!fs.existsSync(f)) continue; withText++; const t = fs.readFileSync(f, 'utf8').replace(/\s+/g, ' '); const inAudit = audit.includes(p); for (const [name, re] of Object.entries(VOCAB)) if (re.test(t)) { vocabCount[name][0]++; if (inAudit) vocabCount[name][1]++; } } console.log(` full text present for ${withText} of ${papers.length} papers`); console.log(` ${'term'.padEnd(32)} ${'corpus'.padStart(7)} ${'in audit set'.padStart(12)}`); for (const [k, [c, a]] of Object.entries(vocabCount)) console.log(` ${k.padEnd(32)} ${String(c).padStart(7)} ${String(a).padStart(5)} of ${N}`); console.log(''); // --- the null, hand-adjudicated --- console.log('## Establishing a null, HAND-ADJUDICATED (not a term count)'); console.log(' Does the paper measure outcome variation under the SAME treatment?'); console.log(' Routes: A/A (arms differing in nothing), repeats (same condition run'); console.log(" several times, spread reported), generated (null built from the paper's"); console.log(' own data). Evidence sentence per paper in algorithm_audits_set.mjs.'); for (const k of Object.keys(NULLS)) if (!AUDITS.some(([a]) => a === k)) die(`NULLS key not in AUDITS: ${k}`); const byRoute = {}; for (const [k, [route]] of Object.entries(NULLS)) (byRoute[route] ||= []).push(k); for (const [route, ks] of Object.entries(byRoute)) console.log(` ${route.padEnd(16)} ${ks.length}`); const nNull = Object.keys(NULLS).length; console.log(` ${'TOTAL'.padEnd(16)} ${nNull} of ${N} (${pct(nNull, N)})`); console.log(` no null established: ${N - nNull} of ${N} (${pct(N - nNull, N)})`); for (const [k, [route, why]] of Object.entries(NULLS)) { const p = byKey.get(k); console.log(` ${p.year} ${p.venue.padEnd(8)} [${route}] ${p.title}`); console.log(` ${why}`); } console.log(''); // --- what the corpus cannot see --------------------------------------------- console.log('## Screening loss: audit-topical papers in the index but not in the extraction'); const meta = []; const md = path.join(ROOT, 'corpus2/.meta'); for (const f of fs.readdirSync(md)) { if (!f.endsWith('.json')) continue; const j = JSON.parse(fs.readFileSync(path.join(md, f), 'utf8')); const arr = Array.isArray(j) ? j : Object.values(j).find((v) => Array.isArray(v)) || []; for (const r of arr) meta.push(r); } const labels = new Map(); for (const l of fs.readFileSync(path.join(ROOT, 'labels/run1/labels.jsonl'), 'utf8').split('\n').filter(Boolean)) { const r = JSON.parse(l); labels.set(`${r.venue}/${r.year}/${r.slug}`, r); } const TOPIC = /\balgorithm(?:ic)? audit|\baudit(?:ing)? (?:the |of )?(?:search|recommend|ad |ads\b|advertis|algorithm|platform|feed|targeting|ranking)|sock ?-?puppet|price (?:discrimination|steering)|differential pricing|differential treatment|personali[sz]ation of|web search personali|search personali|ad delivery|ad targeting|targeted advertis|filter bubble|echo chamber|rabbit hole|discriminat\w+ (?:in|by) (?:ad|algorithm|ranking|recommend)/i; const cands = meta.filter((r) => TOPIC.test(`${r.title || ''} ${r.abstract || ''}`.replace(/\s+/g, ' '))); const out = cands.filter((r) => !byKey.has(`${r.venue}/${r.year}/${r.slug}`)); let screened = 0, nolabel = 0; const lost = []; for (const r of out) { const lab = labels.get(`${r.venue}/${r.year}/${r.slug}`); if (!lab) { nolabel++; lost.push(['no label record (venue-year gap)', r]); } else if (!lab.securityMeasurement && !lab.privacyMeasurement) { screened++; lost.push(['screened out: both labels false', r]); } else die(`unexpected: selected but not extracted: ${r.venue}/${r.year}/${r.slug}`); } console.log(` index records ${meta.length}`); console.log(` audit-topical candidates in the index ${cands.length}`); console.log(` ... of which extracted (in the 5,859) ${cands.length - out.length}`); console.log(` ... of which NOT extracted ${out.length}`); console.log(` screened out (both screen labels false) ${screened}`); console.log(` no label record at all (venue-year gap) ${nolabel}`); if (screened + nolabel !== out.length) die('screening-loss buckets do not sum'); console.log(''); console.log(' Named losses a reader of this page would expect to find:'); const NAMED = [/measuring personalization of web search/i, /measuring price discrimination/i, /crying wolf/i, /auditing for discrimination in algorithms delivering job ads/i, /do you see what i see/i, /an empirical investigation of personalization factors on tiktok/i, /mapwatch/i]; for (const re of NAMED) { const hit = lost.find(([, r]) => re.test(r.title || '')); if (!hit) die(`named loss no longer found by the screening-loss query: ${re}`); console.log(` ${hit[1].year} ${String(hit[1].venue).padEnd(8)} ${hit[1].title} -- ${hit[0]}`); } console.log(''); if (SHOW_LIST) { console.log('## The audit set in full, with the sentence that settled the verdict'); for (const [key, why] of AUDITS) { const p = byKey.get(key); console.log(` ${p.year} ${p.venue.padEnd(8)} ${p.title}`); console.log(` ${key}`); console.log(` IN: ${why}`); } console.log(''); console.log('## Adjudicated and rejected'); for (const [key, why] of REJECTED) { const p = byKey.get(key); console.log(` ${p.year} ${p.venue.padEnd(8)} ${p.title}`); console.log(` OUT: ${why}`); } console.log(''); console.log('## Screening loss in full'); for (const [why, r] of lost.sort((a, b) => a[1].year - b[1].year)) console.log(` ${r.year} ${String(r.venue).padEnd(8)} ${r.title}\n ${why}`); } console.log('OK: all contracts held.');
Its output, unedited
Run on 2026-09-11 against data/extract/run1.
- report_algorithm_audits-output.txt
============================================================================== design:algorithm_audits -- report script run 2026-09-11 corpus 5859 papers, 7 venues, 2010-2026 ============================================================================== ## Populations all papers 5859 empirical 5118 crawled 1120 inferential (any non-descriptive statistic) 1762 AUDIT (hand-adjudicated, rule above) 34 adjudicated and rejected 36 ## Audit papers by year (denominator: the audit set) 2010 1 # 2011 0 2012 0 2013 0 2014 1 # 2015 4 #### 2016 1 # 2017 0 2018 2 ## 2019 2 ## 2020 2 ## 2021 0 2022 6 ###### 2023 4 #### 2024 3 ### 2025 5 ##### *provisional 2026 3 ### *provisional total 34 years with zero audit papers: 2011, 2012, 2013, 2017, 2021 2010-2015 (6 yr): 6 of 34 (17.6%), 1.00 papers/year 2016-2019 (4 yr): 5 of 34 (14.7%), 1.25 papers/year 2020-2023 (4 yr): 12 of 34 (35.3%), 3.00 papers/year 2024-2026 (3 yr): 11 of 34 (32.4%), 3.67 papers/year ## Audit papers by venue, against that venue's own output CCS 4 of 990 0.4% IEEE-SP 1 of 767 0.1% IMC 8 of 638 1.3% NDSS 1 of 701 0.1% PETS 10 of 510 2.0% USENIX 2 of 1410 0.1% WWW 8 of 843 0.9% ## Platform measured (multi-valued; does not sum to N) web 27 of 34 79.4% other-online-service 19 of 34 55.9% mobile 4 of 34 11.8% iot 3 of 34 8.8% offline 0 of 34 0.0% ## What audit papers report, vs the comparable corpus baseline Each row states both populations. The baseline is the population named, not "all papers", so the two cells are comparable. indicator audit (34) baseline runs a non-descriptive statistic 21/34 (61.8%) 1637/5118 (32.0%) of empirical (5,118) states an ethics-review outcome 20/34 (58.8%) 1728/5118 (33.8%) of empirical (5,118) states artifact availability 22/34 (64.7%) 2890/5118 (56.5%) of empirical (5,118) has a crawlConfig at all 26/34 (76.5%) 1080/1120 (96.4%) of crawled (1,120) states crawl statefulness 24/34 (70.6%) 219/1120 (19.6%) of crawled (1,120) states interaction depth 22/34 (64.7%) 841/1120 (75.1%) of crawled (1,120) states a vantage location 15/34 (44.1%) 1228/3908 (31.4%) of measuredFrom (3,908) ## The crawl-config rows again, on a SHARED denominator The rows above compare 34 audits (8 of which have no crawl config) against 1,120 crawled papers (40 of which have none). A field that can only be stated on a paper that HAS a config must be counted over papers that have one, or the two cells are not comparable. Both framings are printed; the page uses this one. statefulness audit 24/26 ( 92.3%) crawled 219/1080 (20.3%) interactionDepth audit 22/26 ( 84.6%) crawled 841/1080 (77.9%) consentAction audit 12/26 ( 46.2%) crawled 349/1080 (32.3%) headless audit 5/26 ( 19.2%) crawled 140/1080 (13.0%) -- the 8 audit papers with no crawl config at all: 2015 IMC Peeking Beneath the Hood of Uber. 2019 PETS Investigating sources of PII used in Facebook’s targeted advertising 2022 IMC Measurement and analysis of implied identity in ad delivery optimization. 2022 IMC What factors affect targeting and bids in online advertising?: a field measurement study. 2025 IMC From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics. 2025 PETS Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants 2026 PETS Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores 2026 WWW Does This Button Work? Investigating YouTube's Ineffective User Controls. ## Statistical methods named by audit papers (paper-counted, free text, ranking only) 3 Holm-Bonferroni correction 2 linear regression 2 Bonferroni correction 2 Mann-Whitney U test 1 CDFs, medians, percentiles, and percentages 1 descriptive comparison of ad fractions and indexed CPM 1 exact statistical test based on Pearson's correlation 1 Holm-Bonferroni 1 Benjamini-Yekutieli 1 averages and standard deviations 1 cross correlation 1 blocked permutation test 1 Holm-Bonferroni method 1 Clopper-Pearson interval 1 Counts and percentages of advertisements -- correction families, folded by hand (paper-counted; the skeleton fold does not merge these): Holm-Bonferroni (any spelling) 5 of 34 Bonferroni without Holm 3 of 34 any Bonferroni-family 7 of 34 Benjamini-Hochberg/Yekutieli 1 of 34 -- 13 of 34 audit papers report no non-descriptive statistic at all: 2010 IMC Challenges in measuring online advertising systems. 2014 CCS Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising. 2015 IMC Location, Location, Location: The Impact of Geolocation on Web Search Personalization. 2016 USENIX Tracing Information Flows Between Ad Exchanges Using Retargeted Ads 2018 WWW AdBudgetKiller: Online Advertising Budget Draining Attack. 2019 PETS Investigating sources of PII used in Facebook’s targeted advertising 2022 CCS Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement. 2022 NDSS HARPO: Learning to Subvert Online Behavioral Advertising 2023 CCS Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects. 2023 PETS A Utility-Preserving Obfuscation Approach for YouTube Recommendations 2024 IMC Poster: Identifying Filter Bubble Based on Feed-Level Embedding Similarity Analysis. 2025 PETS Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants 2026 IEEE-SP Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube. ## Apparatus vocabulary in full text (paper-counted over all 5,859 with text) full text present for 5855 of 5859 papers term corpus in audit set sock puppet 26 5 of 34 persona 190 14 of 34 control profile/account/persona 33 11 of 34 treatment group/profile 66 5 of 34 trained profile 4 1 of 34 A/A test 2 0 of 34 noise floor 35 2 of 34 price discrimination/steering 33 9 of 34 ## Establishing a null, HAND-ADJUDICATED (not a term count) Does the paper measure outcome variation under the SAME treatment? Routes: A/A (arms differing in nothing), repeats (same condition run several times, spread reported), generated (null built from the paper's own data). Evidence sentence per paper in algorithm_audits_set.mjs. A/A 4 generated 3 repeats 1 A/A + generated 2 TOTAL 10 of 34 (29.4%) no null established: 24 of 34 (70.6%) 2010 IMC [A/A] Challenges in measuring online advertising systems. "Even queries launched simultaneously from two identically configured clients on the same subnet can produce wildly different ads"; and "In this paper we perform all analysis relative to a control experiment" 2015 CCS [generated] Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence. exact/random permutation test against an explicit H0, on a held-out split of the profiles, with Benjamini-Yekutieli or Holm correction 2015 IMC [A/A] Location, Location, Location: The Impact of Geolocation on Web Search Personalization. "The red line compares two treatments at the baseline location (i.e., the experimental control), and thus shows the noise floor." 2015 PETS [generated] Automated Experiments on Ad Privacy Settings permutation test over the arm labels -- "The permutation test randomly permutes ... the control and experimental treatments" -- run over "blocks of nearly identical agents" 2019 WWW [A/A] Measuring Political Personalization of Google News Search. edit distances tested "among four identical browser profiles" as the reference for the trained-profile comparison 2022 CCS [A/A] Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement. four unused baseline profiles measured identically; one diverged and the authors read that as the noise: "Such non-determinism is expected, validating our strategy of deploying numerous baseline profiles and the use of normalization." 2024 WWW [generated] TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds. "a randomized baseline" against which the personalisation features are scored 2025 PETS [repeats] More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings each condition run six times at least a week apart ("we repeat each experiment six times"), with outliers cut at standard deviations of the country-specific mean 2025 IMC [A/A + generated] From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics. paired puppets differing in nothing, scheduled in parallel; plus a permutation test against a random dataset 2026 PETS [A/A + generated] Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores "our baseline persona uses the same neutral persona type for control and treatment. For the baseline persona, the majority of measurement pairs produce insignificant tests" -- an A/A that behaved like one -- plus a 9,999-permutation test ## Screening loss: audit-topical papers in the index but not in the extraction index records 16864 audit-topical candidates in the index 117 ... of which extracted (in the 5,859) 69 ... of which NOT extracted 48 screened out (both screen labels false) 45 no label record at all (venue-year gap) 3 Named losses a reader of this page would expect to find: 2013 WWW Measuring personalization of web search. -- screened out: both labels false 2014 IMC Measuring Price Discrimination and Steering on E-commerce Web Sites. -- screened out: both labels false 2014 PETS Crying Wolf? On the Price Discrimination of Online Airline Tickets -- no label record (venue-year gap) 2021 WWW Auditing for Discrimination in Algorithms Delivering Job Ads. -- screened out: both labels false 2016 NDSS Do You See What I See? Differential Treatment of Anonymous Users -- no label record (venue-year gap) 2022 WWW An Empirical Investigation of Personalization Factors on TikTok. -- screened out: both labels false 2016 WWW MapWatch: Detecting and Monitoring International Border Personalization on Online Maps. -- screened out: both labels false OK: all contracts held.
The full verdict list
node scripts/report_algorithm_audits.mjs –list — the 32 verdicts with their evidence sentence, the 12 rejections with their reason, and all 48 screening losses.
- report_algorithm_audits-list-output.txt
============================================================================== design:algorithm_audits -- report script run 2026-09-11 corpus 5859 papers, 7 venues, 2010-2026 ============================================================================== ## Populations all papers 5859 empirical 5118 crawled 1120 inferential (any non-descriptive statistic) 1762 AUDIT (hand-adjudicated, rule above) 34 adjudicated and rejected 36 ## Audit papers by year (denominator: the audit set) 2010 1 # 2011 0 2012 0 2013 0 2014 1 # 2015 4 #### 2016 1 # 2017 0 2018 2 ## 2019 2 ## 2020 2 ## 2021 0 2022 6 ###### 2023 4 #### 2024 3 ### 2025 5 ##### *provisional 2026 3 ### *provisional total 34 years with zero audit papers: 2011, 2012, 2013, 2017, 2021 2010-2015: 6 of 34 (17.6%) 2016-2019: 5 of 34 (14.7%) 2020-2023: 12 of 34 (35.3%) 2024-2026: 11 of 34 (32.4%) ## Audit papers by venue, against that venue's own output CCS 4 of 990 0.4% IEEE-SP 1 of 767 0.1% IMC 8 of 638 1.3% NDSS 1 of 701 0.1% PETS 10 of 510 2.0% USENIX 2 of 1410 0.1% WWW 8 of 843 0.9% ## Platform measured (multi-valued; does not sum to N) web 27 of 34 79.4% other-online-service 19 of 34 55.9% mobile 4 of 34 11.8% iot 3 of 34 8.8% offline 0 of 34 0.0% ## What audit papers report, vs the comparable corpus baseline Each row states both populations. The baseline is the population named, not "all papers", so the two cells are comparable. indicator audit (34) baseline runs a non-descriptive statistic 21/34 (61.8%) 1637/5118 (32.0%) of empirical (5,118) states an ethics-review outcome 20/34 (58.8%) 1728/5118 (33.8%) of empirical (5,118) states artifact availability 22/34 (64.7%) 2890/5118 (56.5%) of empirical (5,118) has a crawlConfig at all 26/34 (76.5%) 1080/1120 (96.4%) of crawled (1,120) states crawl statefulness 24/34 (70.6%) 219/1120 (19.6%) of crawled (1,120) states interaction depth 22/34 (64.7%) 841/1120 (75.1%) of crawled (1,120) states a vantage location 15/34 (44.1%) 1228/3908 (31.4%) of measuredFrom (3,908) ## The crawl-config rows again, on a SHARED denominator The rows above compare 34 audits (8 of which have no crawl config) against 1,120 crawled papers (40 of which have none). A field that can only be stated on a paper that HAS a config must be counted over papers that have one, or the two cells are not comparable. Both framings are printed; the page uses this one. statefulness audit 24/26 ( 92.3%) crawled 219/1080 (20.3%) interactionDepth audit 22/26 ( 84.6%) crawled 841/1080 (77.9%) consentAction audit 12/26 ( 46.2%) crawled 349/1080 (32.3%) headless audit 5/26 ( 19.2%) crawled 140/1080 (13.0%) -- the 8 audit papers with no crawl config at all: 2015 IMC Peeking Beneath the Hood of Uber. 2019 PETS Investigating sources of PII used in Facebook’s targeted advertising 2022 IMC Measurement and analysis of implied identity in ad delivery optimization. 2022 IMC What factors affect targeting and bids in online advertising?: a field measurement study. 2025 IMC From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics. 2025 PETS Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants 2026 PETS Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores 2026 WWW Does This Button Work? Investigating YouTube's Ineffective User Controls. ## Statistical methods named by audit papers (paper-counted, free text, ranking only) 3 Holm-Bonferroni correction 2 linear regression 2 Bonferroni correction 2 Mann-Whitney U test 1 CDFs, medians, percentiles, and percentages 1 descriptive comparison of ad fractions and indexed CPM 1 exact statistical test based on Pearson's correlation 1 Holm-Bonferroni 1 Benjamini-Yekutieli 1 averages and standard deviations 1 cross correlation 1 blocked permutation test 1 Holm-Bonferroni method 1 Clopper-Pearson interval 1 Counts and percentages of advertisements -- correction families, folded by hand (paper-counted; the skeleton fold does not merge these): Holm-Bonferroni (any spelling) 5 of 34 Bonferroni without Holm 3 of 34 any Bonferroni-family 7 of 34 Benjamini-Hochberg/Yekutieli 1 of 34 -- 13 of 34 audit papers report no non-descriptive statistic at all: 2010 IMC Challenges in measuring online advertising systems. 2014 CCS Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising. 2015 IMC Location, Location, Location: The Impact of Geolocation on Web Search Personalization. 2016 USENIX Tracing Information Flows Between Ad Exchanges Using Retargeted Ads 2018 WWW AdBudgetKiller: Online Advertising Budget Draining Attack. 2019 PETS Investigating sources of PII used in Facebook’s targeted advertising 2022 CCS Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement. 2022 NDSS HARPO: Learning to Subvert Online Behavioral Advertising 2023 CCS Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects. 2023 PETS A Utility-Preserving Obfuscation Approach for YouTube Recommendations 2024 IMC Poster: Identifying Filter Bubble Based on Feed-Level Embedding Similarity Analysis. 2025 PETS Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants 2026 IEEE-SP Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube. ## Apparatus vocabulary in full text (paper-counted over all 5,859 with text) full text present for 5855 of 5859 papers term corpus in audit set sock puppet 26 5 of 34 persona 190 14 of 34 control profile/account/persona 33 11 of 34 treatment group/profile 66 5 of 34 trained profile 4 1 of 34 A/A test 2 0 of 34 noise floor 35 2 of 34 price discrimination/steering 33 9 of 34 ## Establishing a null, HAND-ADJUDICATED (not a term count) Does the paper measure outcome variation under the SAME treatment? Routes: A/A (arms differing in nothing), repeats (same condition run several times, spread reported), generated (null built from the paper's own data). Evidence sentence per paper in algorithm_audits_set.mjs. A/A 4 generated 3 repeats 1 A/A + generated 2 TOTAL 10 of 34 (29.4%) no null established: 24 of 34 (70.6%) 2010 IMC [A/A] Challenges in measuring online advertising systems. "Even queries launched simultaneously from two identically configured clients on the same subnet can produce wildly different ads"; and "In this paper we perform all analysis relative to a control experiment" 2015 CCS [generated] Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence. exact/random permutation test against an explicit H0, on a held-out split of the profiles, with Benjamini-Yekutieli or Holm correction 2015 IMC [A/A] Location, Location, Location: The Impact of Geolocation on Web Search Personalization. "The red line compares two treatments at the baseline location (i.e., the experimental control), and thus shows the noise floor." 2015 PETS [generated] Automated Experiments on Ad Privacy Settings permutation test over the arm labels -- "The permutation test randomly permutes ... the control and experimental treatments" -- run over "blocks of nearly identical agents" 2019 WWW [A/A] Measuring Political Personalization of Google News Search. edit distances tested "among four identical browser profiles" as the reference for the trained-profile comparison 2022 CCS [A/A] Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement. four unused baseline profiles measured identically; one diverged and the authors read that as the noise: "Such non-determinism is expected, validating our strategy of deploying numerous baseline profiles and the use of normalization." 2024 WWW [generated] TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds. "a randomized baseline" against which the personalisation features are scored 2025 PETS [repeats] More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings each condition run six times at least a week apart ("we repeat each experiment six times"), with outliers cut at standard deviations of the country-specific mean 2025 IMC [A/A + generated] From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics. paired puppets differing in nothing, scheduled in parallel; plus a permutation test against a random dataset 2026 PETS [A/A + generated] Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores "our baseline persona uses the same neutral persona type for control and treatment. For the baseline persona, the majority of measurement pairs produce insignificant tests" -- an A/A that behaved like one -- plus a 9,999-permutation test ## Screening loss: audit-topical papers in the index but not in the extraction index records 16864 audit-topical candidates in the index 117 ... of which extracted (in the 5,859) 69 ... of which NOT extracted 48 screened out (both screen labels false) 45 no label record at all (venue-year gap) 3 Named losses a reader of this page would expect to find: 2013 WWW Measuring personalization of web search. -- screened out: both labels false 2014 IMC Measuring Price Discrimination and Steering on E-commerce Web Sites. -- screened out: both labels false 2014 PETS Crying Wolf? On the Price Discrimination of Online Airline Tickets -- no label record (venue-year gap) 2021 WWW Auditing for Discrimination in Algorithms Delivering Job Ads. -- screened out: both labels false 2016 NDSS Do You See What I See? Differential Treatment of Anonymous Users -- no label record (venue-year gap) 2022 WWW An Empirical Investigation of Personalization Factors on TikTok. -- screened out: both labels false 2016 WWW MapWatch: Detecting and Monitoring International Border Personalization on Online Maps. -- screened out: both labels false ## The audit set in full, with the sentence that settled the verdict 2010 IMC Challenges in measuring online advertising systems. IMC/2010/challenges-in-measuring-online-advertising-systems IN: seeded profile pairs: "enabled, but are seeded with different user personae" (the sentence naming the two interest sets is column-spliced in both renderings) 2014 CCS Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising. CCS/2014/your-online-interests-pwned-a-pollution-attack-against-targeted-advertising IN: blank profile vs polluted profile: "the polluter can impersonate a user with a blank profile ... and browse pages" 2015 CCS Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence. CCS/2015/sunlight-fine-grained-targeting-detection-at-scale-with-statistical-confidence IN: instrument paper: varies personal-data inputs one at a time "compared to a control group", with statistical confidence 2015 IMC Location, Location, Location: The Impact of Geolocation on Web Search Personalization. IMC/2015/location-location-location-the-impact-of-geolocation-on-web-search-personalizati IN: location as the treatment: "all other browser attributes were the same across treatments" 2015 IMC Peeking Beneath the Hood of Uber. IMC/2015/peeking-beneath-the-hood-of-uber IN: "We created 43 Uber accounts ... blanket a small geographic area with measurement points" 2015 PETS Automated Experiments on Ad Privacy Settings PETS/2015/automated-experiments-on-ad-privacy-settings IN: AdFisher: "We created an experimental group and a control group of agents" 2016 USENIX Tracing Information Flows Between Ad Exchanges Using Retargeted Ads USENIX/2016/tracing-information-flows-between-ad-exchanges-using-retargeted-ads IN: "We train 90 personas by visiting popular e-commerce sites, and then crawl major publishers" 2018 WWW AdBudgetKiller: Online Advertising Budget Draining Attack. WWW/2018/adbudgetkiller-online-advertising-budget-draining-attack IN: "Starting from a fresh profile, the profile trainer produces" crafted browsing profiles 2018 WWW Auditing the Personalization and Composition of Politically-Related Search Engine Results Pages. WWW/2018/auditing-the-personalization-and-composition-of-politically-related-search-engin IN: standard-vs-incognito paired SERPs: "our controls were paired within the individual" 2019 PETS Investigating sources of PII used in Facebook’s targeted advertising PETS/2019/investigating-sources-of-pii-used-in-facebook-s-targeted-advertising IN: "Take a Facebook account that we control (call it the control account) and the test PII" 2019 WWW Measuring Political Personalization of Google News Search. WWW/2019/measuring-political-personalization-of-google-news-search IN: "a 'sock puppet' auditing system in which a pair of fresh browser profiles" visits divergent pages then runs identical queries 2020 PETS Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding PETS/2020/inferring-tracker-advertiser-relationships-in-the-online-advertising-ecosystem-u IN: intent vs no-intent versions of each of 16 interest personas, with trackers selectively exposed (this PDF is column-spliced in both cols and norm renderings, so no contiguous quote was publishable) 2020 WWW Stop tracking me Bro! Differential Tracking of User Demographics on Hyper-Partisan Websites. WWW/2020/stop-tracking-me-bro-differential-tracking-of-user-demographics-on-hyper-partisa IN: "We create 9 carefully crafted personas representing different genders and age groups" 2022 CCS Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement. CCS/2022/cart-ology-intercepting-targeted-advertising-via-ad-network-identity-entanglemen IN: five profiles -- two blank baselines, two account baselines, one entangled attack profile: "All profiles are created and mechanistically measured in the same way, with separation between attacker, victim, and baselines on different machines with different IPs" 2022 IMC Measurement and analysis of implied identity in ad delivery optimization. IMC/2022/measurement-and-analysis-of-implied-identity-in-ad-delivery-optimization IN: "We ran 200 versions of this ad at the same time, all from the same account and with the same budget" 2022 IMC What factors affect targeting and bids in online advertising?: a field measurement study. IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising-a-field-measurement IN: "All participants were asked to visit the same websites to control for contextual targeting, in randomized order" 2022 NDSS HARPO: Learning to Subvert Online Behavioral Advertising NDSS/2022/auto-draft-209 IN: HARPO: obfuscated vs unobfuscated personas measured against live profiling models 2022 PETS ATOM: Ad-network Tomography PETS/2022/atom-ad-network-tomography IN: "create a number of online user personas associated with specific interest groups" and gather ads while systematically blocking trackers 2022 WWW Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Election 2020. WWW/2022/left-or-right-a-peek-into-the-political-biases-in-email-spam-filtering-algorithm IN: "We created 102 email accounts" and compared spam placement across treatment and control affiliations 2023 CCS Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects. CCS/2023/marketing-to-children-through-online-targeted-advertising-targeting-mechanisms-a IN: "We launch the six personas simultaneously by creating six Selenium" browsers 2023 IMC Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem. IMC/2023/tracking-profiling-and-ad-targeting-in-the-alexa-echo-smart-speaker-ecosystem IN: "Each treatment persona is simulated by installing and interacting with skills ... By contrast, in the control persona, we do" not 2023 PETS A Utility-Preserving Obfuscation Approach for YouTube Recommendations PETS/2023/a-utility-preserving-obfuscation-approach-for-youtube-recommendations IN: "We deploy and evaluate De-Harpo's effectiveness on YouTube using 10,000 sock puppet based personas" 2023 WWW A Method to Assess and Explain Disparate Impact in Online Retailing. WWW/2023/a-method-to-assess-and-explain-disparate-impact-in-online-retailing IN: "Each observation ... is paired with a set of 'counter-factuals'" from neighbouring zip codes 2024 PETS Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy? PETS/2024/opted-out-yet-tracked-are-regulations-enough-to-protect-your-privacy IN: "We then visit the filtered websites with user interest personas" -- 16 interest personas, two regions 2024 IMC Poster: Identifying Filter Bubble Based on Feed-Level Embedding Similarity Analysis. IMC/2024/poster-identifying-filter-bubble-based-on-feed-level-embedding-similarity-analys IN: four bot arms differing only in the video-selection strategy: "The bot selected a video having (a) the minimum JCC, (b) JCC larger than the minimum (random selection), (c) JCC larger than the mean, and (d) JCC larger than WCC"; outcome is the recommended feed 2024 WWW TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds. WWW/2024/tiktok-and-the-art-of-personalization-investigating-exploration-and-exploitation IN: "validate our results using a baseline generated from automated TikTok bots, as well as a randomized baseline" 2025 IMC From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics. IMC/2025/from-voice-to-ads-auditing-commercial-smart-speakers-for-targeted-advertising-ba IN: "we used each of the 21 cloned voice models to train two puppets, which we refer to as 'paired puppets'" 2025 PETS Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants PETS/2025/echoes-of-privacy-uncovering-the-profiling-practices-of-voice-assistants IN: "each meticulously trained with a curated set of voice queries designed to simulate various user personas" 2025 PETS More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings PETS/2025/more-and-scammier-ads-the-perils-of-youtubes-ad-privacy-settings IN: "running controlled experiments with sock puppet accounts that emulate users watching YouTube videos" 2025 PETS Sheep's clothing, wolfish intent: Automated detection and evaluation of problematic 'allowed' advertisements PETS/2025/sheeps-clothing-wolfish-intent-automated-detection-and-evaluation-of-problematic IN: "two separate crawlers were deployed: one for the Control group (without ABP) and another for the AccAds group" 2025 USENIX Big Help or Big Brother? Auditing Tracking, Profiling, and Personalization in Generative AI Assistants USENIX/2025/big-help-or-big-brother-auditing-tracking-profiling-and-personalization-in-gener IN: "Training phase involves browsing through 10 webpages - 2 pages per leaked attribute", then personalisation is measured 2026 IEEE-SP Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube. IEEE-SP/2026/setting-the-course-but-forgetting-to-steer-analyzing-compliance-with-gdprs-right IN: "we use sock-puppet accounts to systematically browse and log the behavior of the sock-puppet to generate the ground truth" 2026 PETS Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores PETS/2026/ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis IN: "We construct accounts with specific parameters or interests, so-called personas, and measure the ads displayed to them" 2026 WWW Does This Button Work? Investigating YouTube's Ineffective User Controls. WWW/2026/does-this-button-work-investigating-youtubes-ineffective-user-controls IN: "depending on randomized assignment - triggered one of four native feedback signals to YouTube ... or no signal at all in the control group" ## Adjudicated and rejected 2017 WWW An Army of Me: Sockpuppets in Online Discussion Communities. OUT: observational study OF sockpuppets others created; no arm the authors control 2019 WWW Auditing the Partisanship of Google Search Snippets. OUT: audits snippet vs linked page; no identity treatment, no arms 2020 WWW Facebook Ads Monitor: An Independent Auditing System for Political Ads on Facebook. OUT: volunteer ad donation; observational, no arms 2020 USENIX What Twitter Knows: Characterizing Ad Targeting Practices, User Perceptions, and Ad Explanations Through Users' Own Twitter Data OUT: users' own Twitter data; no arms 2023 IEEE-SP Collaborative Ad Transparency: Promises and Limitations. OUT: explicitly the non-persona alternative: "One method that does not use fake personas" 2023 USENIX Problematic Advertising and its Disparate Exposure on Facebook OUT: donated ad exposure from real users; disparity is observational 2022 CCS Privacy Limitations of Interest-based Advertising on The Web: A Post-mortem Empirical Analysis of Google's FLoC. OUT: randomised control is a data permutation over a browsing panel, not a live-platform arm 2024 IEEE-SP Targeted and Troublesome: Tracking and Advertising on Children's Websites. OUT: fresh profile per page visit; the paper itself says "Future work could extend our method to incorporate personas" 2024 PETS Interest-disclosing Mechanisms for Advertising are Privacy-Exposing (not Preserving) OUT: Topics API analysed over real browsing histories, not persona arms 2026 WWW When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs. OUT: "Random Control group" is a model ablation, not a platform arm 2022 IEEE-SP Deployment of Source Address Validation by Network Operators: A Randomized Control Trial. OUT: an RCT, but the treatment is a notification to operators; outcome is not platform output 2015 NDSS Bloom Cookies: Web Search Personalization without User Tracking OUT: a privacy-preserving personalisation design evaluated on search logs; no arms 2018 CCS Peeling the Onion's User Experience Layer: Examining Naturalistic Use of the Tor Browser. OUT: a 19-participant UX study of Tor Browser; outcome is user experience 2023 USENIX Strategies and Vulnerabilities of Participants in Venezuelan Influence Operations OUT: interviews with 19 influence-operation participants; the sockpuppets are theirs, not the authors 2011 CCS Policy auditing over incomplete logs: theory, implementation and applications. OUT: compliance checking over audit logs; shares the word "audit" and nothing else 2013 WWW Your browsing behavior for a big mac: economics of personal information online. OUT: 168 recruited participants valuing their own PII in an auction; a user study 2015 IEEE-SP Effective Real-Time Android Application Auditing. OUT: program analysis of Android apps; "auditing" means taint tracking 2017 USENIX Exploring User Perceptions of Discrimination in Online Targeted Advertising OUT: a survey of user perceptions; the randomised factors are survey vignettes, not platform arms 2018 IMC 403 Forbidden: A Global View of CDN Geoblocking. OUT: vantage points in 177 countries, but the outcome is REACHABILITY. Excluded by the tightened (O) clause; belongs to design:blocking_and_geodifference 2019 CCS The Art and Craft of Fraudulent App Promotion in Google Play. OUT: a study of app-store fraud workers who operate sockpuppets; observational 2022 WWW Characterizing, Detecting, and Predicting Online Ban Evasion. OUT: Wikipedia's own labelled sockpuppet groups; observational, no arms 2022 WWW Fairness Audit of Machine Learning Models with Confidential Computing. OUT: ML fairness auditing inside a TEE; no live platform 2022 IEEE-SP Towards Automated Auditing for Account and Session Management Flaws in Single Sign-On Deployments. OUT: SSO implementation flaws; outcome is a security bug, not a discriminating response 2023 NDSS Tactics, Threats & Targets: Modeling Disinformation and its Mitigation OUT: interviews with fact-checkers and analysts; sockpuppets are the subject, not the instrument 2023 WWW Scoping Fairness Objectives and Identifying Fairness Metrics for Recommender Systems: The Practitioners' Perspective. OUT: practitioner interviews about fairness metrics; no measurement of a platform 2023 IEEE-SP When and Why Do People Want Ad Targeting Explanations? Evidence from a Four-Week, Mixed-Methods Field Study. OUT: a four-week field study of what people want from ad explanations; outcome is attitudes 2023 USENIX Auditing Framework APIs via Inferred App-side Security Specifications OUT: Android framework API access control; "auditing" is static analysis 2024 CCS Curator Attack: When Blackbox Differential Privacy Auditing Loses Its Power. OUT: differential-privacy auditing; a different object with the same name 2024 USENIX Efficient Privacy Auditing in Federated Learning OUT: membership inference against an FL model; not a platform 2024 USENIX "What do you want from theory alone?" Experimenting with Tight Auditing of Differentially Private Synthetic Data Generation OUT: DP synthetic-data auditing; not a platform 2024 USENIX Fledging Will Continue Until Privacy Improves: Empirical Analysis of Google's Privacy-Preserving Targeted Advertising OUT: security analysis of the FLEDGE/Protected Audience API; the outcome is API behaviour and attack feasibility, not a served-ad contrast across arms 2025 NDSS Exploring User Perceptions of Security Auditing in the Web3 Ecosystem OUT: user perceptions of smart-contract audits 2025 PETS Privacy Perceptions and Behaviors Towards Targeted Advertising on Social Media: A Cross-Country Study on the Effect of Culture and Religion OUT: an n=412 cross-country survey; outcome is attitudes 2026 PETS AudAgent: Automated Auditing of Privacy Policy Compliance in AI Agents OUT: policy-compliance checking of AI agents; no arms 2026 PETS Privacy in Theory, Bugs in Practice: Grey-Box Auditing of Differential Privacy Libraries OUT: DP library auditing; not a platform 2026 WWW Question the Questions: Auditing Representation in Online Deliberative Processes. OUT: algorithm design for question selection in deliberative polls; not a measurement of a deployed platform ## Screening loss in full 2010 NDSS Adnostic: Privacy Preserving Targeted Advertising no label record (venue-year gap) 2010 WWW Using a model of social dynamics to predict popularity of news. screened out: both labels false 2012 CCS Privacy-aware personalization for mobile advertising. screened out: both labels false 2012 WWW How effective is targeted advertising? screened out: both labels false 2013 WWW Measuring personalization of web search. screened out: both labels false 2013 WWW Spatio-temporal dynamics of online memes: a study of geo-tagged tweets. screened out: both labels false 2014 IMC Measuring Price Discrimination and Steering on E-commerce Web Sites. screened out: both labels false 2014 PETS Crying Wolf? On the Price Discrimination of Online Airline Tickets no label record (venue-year gap) 2014 WWW Quizz: targeted crowdsourcing with a billion (potential) users. screened out: both labels false 2014 WWW Mining novelty-seeking trait across heterogeneous domains. screened out: both labels false 2014 WWW Exploring the filter bubble: the effect of using recommender systems on content diversity. screened out: both labels false 2014 WWW Fast topic discovery from web search streams. screened out: both labels false 2015 WWW Events and Controversies: Influences of a Shocking News Event on Information Seeking. screened out: both labels false 2016 NDSS Do You See What I See? Differential Treatment of Anonymous Users no label record (venue-year gap) 2016 USENIX Micro-Virtualization Memory Tracing to Detect and Prevent Spraying Attacks screened out: both labels false 2016 WWW MapWatch: Detecting and Monitoring International Border Personalization on Online Maps. screened out: both labels false 2018 IEEE-SP FuturesMEX: Secure, Distributed Futures Market Exchange. screened out: both labels false 2018 WWW Modeling Interdependent and Periodic Real-World Action Sequences. screened out: both labels false 2018 WWW Me, My Echo Chamber, and I: Introspection on Social Media Polarization. screened out: both labels false 2018 WWW Political Discourse on Social Media: Echo Chambers, Gatekeepers, and the Price of Bipartisanship. screened out: both labels false 2020 CCS DECO: Liberating Web Data Using Decentralized Oracles for TLS. screened out: both labels false 2020 IMC Mis-shapes, Mistakes, Misfits: An Analysis of Domain Classification Services. screened out: both labels false 2020 WWW Architectures for Autonomy: Towards an Equitable Web of Data in the Age of AI. screened out: both labels false 2021 USENIX SIGL: Securing Software Installations Through Deep Graph Learning screened out: both labels false 2021 WWW Rabbit Holes and Taste Distortion: Distribution-Aware Recommendation with Evolving Interests. screened out: both labels false 2021 WWW Local Clustering in Contextual Multi-Armed Bandits. screened out: both labels false 2021 WWW Incrementality Testing in Programmatic Advertising: Enhanced Precision with Double-Blind Designs. screened out: both labels false 2021 WWW Causal Network Motifs: Identifying Heterogeneous Spillover Effects in A/B Tests. screened out: both labels false 2021 WWW Auditing for Discrimination in Algorithms Delivering Job Ads. screened out: both labels false 2021 WWW The Interaction between Political Typology and Filter Bubbles in News Recommendation Algorithms. screened out: both labels false 2022 PETS PUBA: Privacy-Preserving User-Data Bookkeeping and Analytics screened out: both labels false 2022 WWW An Empirical Investigation of Personalization Factors on TikTok. screened out: both labels false 2023 PETS Find Thy Neighbourhood: Privacy-Preserving Local Clustering screened out: both labels false 2023 WWW pFedPrompt: Learning Personalized Prompt for Vision-Language Models in Federated Learning. screened out: both labels false 2023 WWW Breaking Filter Bubble: A Reinforcement Learning Framework of Controllable Recommender System. screened out: both labels false 2024 PETS Evaluating Google's Protected Audience Protocol screened out: both labels false 2024 WWW Filter Bubble or Homogenization? Disentangling the Long-Term Effects of Recommendations on User Consumption Patterns. screened out: both labels false 2024 WWW Optimal Engagement-Diversity Tradeoffs in Social Media. screened out: both labels false 2024 WWW Learning Category Trees for ID-Based Recommendation: Exploring the Power of Differentiable Vector Quantization. screened out: both labels false 2024 WWW Full-stage Diversified Recommendation: Large-scale Online Experiments in Short-video Platform. screened out: both labels false 2024 WWW Uncovering the Deep Filter Bubble: Narrow Exposure in Short-Video Recommendation. screened out: both labels false 2025 CCS Cascading Adversarial Bias from Injection to Distillation in Language Models. screened out: both labels false 2025 USENIX Privacy Audit as Bits Transmission: (Im)possibilities for Audit by One Run screened out: both labels false 2025 WWW LLM4Rerank: LLM-based Auto-Reranking Framework for Recommendations. screened out: both labels false 2025 WWW SPRec: Self-Play to Debias LLM-based Recommendation. screened out: both labels false 2026 PETS Making Sense of Private Advertising: A Principled Approach to a Complex Ecosystem screened out: both labels false 2026 WWW DynaMoLTV: A Cross-Game Dynamic Mixture Model with Weighted Sub-Distributions for Player Lifetime Value Prediction. screened out: both labels false 2026 WWW Audit?of?Audits for the Web: Bayesian Meta?Evaluation that Yields Interval?Valued, Threshold?Aligned Fairness Claims. screened out: both labels false OK: all contracts held.
The quote checker and its output
- quotecheck_algorithm_audits.mjs
#!/usr/bin/env node // Verifies every quoted fragment used on design:algorithm_audits and on its // provenance page against the paper's own text. // // Three matching modes, because a two-column PDF loses in different places in // each rendering: (1) whitespace-collapsed exact, (2) hyphen/quote-normalised, // (3) longest 8-word run. Each quote is tried against paper.cols.txt AND // paper.norm.txt; a quote found in either is PASS, with the rendering recorded. // Exits non-zero on any FAIL. import fs from 'node:fs'; import path from 'node:path'; import { dataRoot } from './lib.mjs'; const ROOT = dataRoot(); const norm = (s) => s .replace(/\s+/g, ' ') .replace(/[‘’ʼ]/g, "'") .replace(/[“”]/g, '"') .replace(/[‐-―−]/g, '-') .trim(); const strip = (s) => norm(s).replace(/-\s*/g, '').toLowerCase(); function readModes(key) { const [venue, year, slug] = key.split('/'); const out = {}; for (const name of ['paper.cols.txt', 'paper.norm.txt']) { const f = path.join(ROOT, 'fulltext', year, venue, slug, name); if (fs.existsSync(f)) out[name] = fs.readFileSync(f, 'utf8'); } if (Object.keys(out).length === 0) throw new Error(`no text for ${key}`); return out; } export function checkQuote(key, quote) { const modes = readModes(key); const q = norm(quote); const qs = strip(quote); const words = q.split(' '); for (const [name, raw] of Object.entries(modes)) { const t = norm(raw); if (t.includes(q)) return { ok: true, how: `exact in ${name}` }; if (strip(raw).includes(qs)) return { ok: true, how: `hyphen/quote-normalised in ${name}` }; } // longest 8-word run for (const [name, raw] of Object.entries(modes)) { const ts = strip(raw); let best = 0; for (let i = 0; i + 8 <= words.length; i++) { if (ts.includes(strip(words.slice(i, i + 8).join(' ')))) best++; } if (best > 0) return { ok: true, how: `${best} of ${Math.max(0, words.length - 7)} 8-word runs in ${name}` }; } return { ok: false, how: 'NOT FOUND in cols or norm' }; } // Quotes used on the two pages. Each entry: [paper key, quote as published]. export const QUOTES = [ ['IMC/2010/challenges-in-measuring-online-advertising-systems', 'Even queries launched simultaneously from two identically configured clients on the same subnet can produce wildly different ads over multiple timescales.'], ['IMC/2010/challenges-in-measuring-online-advertising-systems', 'In this paper we perform all analysis relative to a control experiment'], ['IMC/2010/challenges-in-measuring-online-advertising-systems', 'enabled, but are seeded with different user personae'], ['PETS/2015/automated-experiments-on-ad-privacy-settings', 'We created an experimental group and a control group of agents.'], ['PETS/2015/automated-experiments-on-ad-privacy-settings', 'The browser agents in the experimental group visited websites on substance abuse while the agents in the control group simply waited.'], ['WWW/2019/measuring-political-personalization-of-google-news-search', 'we develop a "sock puppet" auditing system in which a pair of fresh browser profiles, first, visits web pages that reflect divergent political discourses and, second, executes identical politically oriented Google News searches'], ['WWW/2018/auditing-the-personalization-and-composition-of-politically-related-search-engin', 'our controls were paired within the individual, enabling us to isolate the impact that their browser mode had on their search rankings for each query we searched'], ['IMC/2015/location-location-location-the-impact-of-geolocation-on-web-search-personalizati', 'all other browser attributes were the same across treatments, so each treatment should present an identical browser fingerprint'], ['IMC/2025/from-voice-to-ads-auditing-commercial-smart-speakers-for-targeted-advertising-ba', 'We assigned voices randomly to days, and scheduled paired puppets in parallel.'], ['IMC/2023/tracking-profiling-and-ad-targeting-in-the-alexa-echo-smart-speaker-ecosystem', 'By contrast, in the control persona, we do'], ['PETS/2025/more-and-scammier-ads-the-perils-of-youtubes-ad-privacy-settings', 'our methodology consists of running controlled experiments with sock puppet accounts that emulate users watching YouTube videos in an instrumented browser'], ['WWW/2026/does-this-button-work-investigating-youtubes-ineffective-user-controls', 'depending on randomized assignment'], ['WWW/2022/left-or-right-a-peek-into-the-political-biases-in-email-spam-filtering-algorithm', 'We created 102 email accounts'], ['IMC/2022/measurement-and-analysis-of-implied-identity-in-ad-delivery-optimization', 'We ran 200 versions of this ad at the same time, all from the same account and with the same budget'], ['WWW/2020/stop-tracking-me-bro-differential-tracking-of-user-demographics-on-hyper-partisa', 'We create 9 carefully crafted personas representing different genders and age groups.'], ['IEEE-SP/2026/setting-the-course-but-forgetting-to-steer-analyzing-compliance-with-gdprs-right', 'we use sock-puppet accounts to system- atically browse and log the behavior of the sock-puppet to generate the ground truth'], ['PETS/2026/ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis', 'We construct accounts with specific parameters or interests, so-called personas, and measure the ads displayed to them'], ['IMC/2015/peeking-beneath-the-hood-of-uber', 'We created 43 Uber accounts'], ['IEEE-SP/2023/collaborative-ad-transparency-promises-and-limitations', 'One method that does not use fake personas'], ['IEEE-SP/2024/targeted-and-troublesome-tracking-and-advertising-on-childrens-websites', 'Future work could extend our method to incorporate personas and warmup crawls to study such ads.'], ['PETS/2023/a-utility-preserving-obfuscation-approach-for-youtube-recommendations', 'We deploy and evaluate De-Harpo'], ['USENIX/2016/tracing-information-flows-between-ad-exchanges-using-retargeted-ads', 'We train 90 personas by visiting popular e-commerce sites, and then crawl major publishers to gather retargeted ads'], ['WWW/2024/tiktok-and-the-art-of-personalization-investigating-exploration-and-exploitation', 'validate our results using a baseline generated from automated TikTok bots, as well as a randomized baseline'], ['CCS/2015/sunlight-fine-grained-targeting-detection-at-scale-with-statistical-confidence', 'prior studies conduct tightly controlled experiments that vary personal data inputs (such as location, search terms, or profile interests) one at a time and observe the effect on service outputs (such as ads, recommendations, or prices) compared to a control group'], ['WWW/2022/using-survival-models-to-estimate-user-engagement-in-online-experiments', 'We simulate A/A tests by re-randomizing the treatment assignments on the observed exposure logs from our experiment corpus.'], ['PETS/2026/on-the-suitability-of-llm-driven-agents-for-dark-pattern-audits', 'We design and deploy an LLM-driven auditing agent capable of end-to-end traversal of rights-request workflows'], ['PETS/2024/opted-out-yet-tracked-are-regulations-enough-to-protect-your-privacy', 'we also conduct Bonferroni correction on the statistical test'], ['IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising-a-field-measurement', 'Most commonly, web crawlers with synthetic profiles or personas are used to measure behavioral targeting and contextual targeting.'], ]; if (import.meta.url === `file://${process.argv[1]}`) { let fail = 0; for (const [key, q] of QUOTES) { let r; try { r = checkQuote(key, q); } catch (e) { r = { ok: false, how: e.message }; } if (!r.ok) fail++; console.log(`${r.ok ? 'PASS' : 'FAIL'} ${key}\n ${r.how}\n "${q.slice(0, 110)}${q.length > 110 ? '…' : ''}"`); } console.log(`\n${QUOTES.length} quotes checked, ${fail} failed.`); process.exit(fail ? 1 : 0); }
- quotecheck_algorithm_audits-output.txt
PASS IMC/2010/challenges-in-measuring-online-advertising-systems exact in paper.cols.txt "Even queries launched simultaneously from two identically configured clients on the same subnet can produce wi…" PASS IMC/2010/challenges-in-measuring-online-advertising-systems exact in paper.cols.txt "In this paper we perform all analysis relative to a control experiment" PASS IMC/2010/challenges-in-measuring-online-advertising-systems exact in paper.cols.txt "enabled, but are seeded with different user personae" PASS PETS/2015/automated-experiments-on-ad-privacy-settings exact in paper.cols.txt "We created an experimental group and a control group of agents." PASS PETS/2015/automated-experiments-on-ad-privacy-settings exact in paper.cols.txt "The browser agents in the experimental group visited websites on substance abuse while the agents in the contr…" PASS WWW/2019/measuring-political-personalization-of-google-news-search exact in paper.cols.txt "we develop a "sock puppet" auditing system in which a pair of fresh browser profiles, first, visits web pages …" PASS WWW/2018/auditing-the-personalization-and-composition-of-politically-related-search-engin exact in paper.cols.txt "our controls were paired within the individual, enabling us to isolate the impact that their browser mode had …" PASS IMC/2015/location-location-location-the-impact-of-geolocation-on-web-search-personalizati exact in paper.cols.txt "all other browser attributes were the same across treatments, so each treatment should present an identical br…" PASS IMC/2025/from-voice-to-ads-auditing-commercial-smart-speakers-for-targeted-advertising-ba exact in paper.cols.txt "We assigned voices randomly to days, and scheduled paired puppets in parallel." PASS IMC/2023/tracking-profiling-and-ad-targeting-in-the-alexa-echo-smart-speaker-ecosystem exact in paper.cols.txt "By contrast, in the control persona, we do" PASS PETS/2025/more-and-scammier-ads-the-perils-of-youtubes-ad-privacy-settings exact in paper.cols.txt "our methodology consists of running controlled experiments with sock puppet accounts that emulate users watchi…" PASS WWW/2026/does-this-button-work-investigating-youtubes-ineffective-user-controls exact in paper.cols.txt "depending on randomized assignment" PASS WWW/2022/left-or-right-a-peek-into-the-political-biases-in-email-spam-filtering-algorithm exact in paper.cols.txt "We created 102 email accounts" PASS IMC/2022/measurement-and-analysis-of-implied-identity-in-ad-delivery-optimization exact in paper.cols.txt "We ran 200 versions of this ad at the same time, all from the same account and with the same budget" PASS WWW/2020/stop-tracking-me-bro-differential-tracking-of-user-demographics-on-hyper-partisa exact in paper.cols.txt "We create 9 carefully crafted personas representing different genders and age groups." PASS IEEE-SP/2026/setting-the-course-but-forgetting-to-steer-analyzing-compliance-with-gdprs-right exact in paper.cols.txt "we use sock-puppet accounts to system- atically browse and log the behavior of the sock-puppet to generate the…" PASS PETS/2026/ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis exact in paper.cols.txt "We construct accounts with specific parameters or interests, so-called personas, and measure the ads displayed…" PASS IMC/2015/peeking-beneath-the-hood-of-uber exact in paper.cols.txt "We created 43 Uber accounts" PASS IEEE-SP/2023/collaborative-ad-transparency-promises-and-limitations exact in paper.cols.txt "One method that does not use fake personas" PASS IEEE-SP/2024/targeted-and-troublesome-tracking-and-advertising-on-childrens-websites exact in paper.cols.txt "Future work could extend our method to incorporate personas and warmup crawls to study such ads." PASS PETS/2023/a-utility-preserving-obfuscation-approach-for-youtube-recommendations exact in paper.cols.txt "We deploy and evaluate De-Harpo" PASS USENIX/2016/tracing-information-flows-between-ad-exchanges-using-retargeted-ads exact in paper.cols.txt "We train 90 personas by visiting popular e-commerce sites, and then crawl major publishers to gather retargete…" PASS WWW/2024/tiktok-and-the-art-of-personalization-investigating-exploration-and-exploitation exact in paper.cols.txt "validate our results using a baseline generated from automated TikTok bots, as well as a randomized baseline" PASS CCS/2015/sunlight-fine-grained-targeting-detection-at-scale-with-statistical-confidence exact in paper.cols.txt "prior studies conduct tightly controlled experiments that vary personal data inputs (such as location, search …" PASS WWW/2022/using-survival-models-to-estimate-user-engagement-in-online-experiments exact in paper.cols.txt "We simulate A/A tests by re-randomizing the treatment assignments on the observed exposure logs from our exper…" PASS PETS/2026/on-the-suitability-of-llm-driven-agents-for-dark-pattern-audits exact in paper.cols.txt "We design and deploy an LLM-driven auditing agent capable of end-to-end traversal of rights-request workflows" PASS PETS/2024/opted-out-yet-tracked-are-regulations-enough-to-protect-your-privacy exact in paper.cols.txt "we also conduct Bonferroni correction on the statistical test" PASS IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising-a-field-measurement exact in paper.cols.txt "Most commonly, web crawlers with synthetic profiles or personas are used to measure behavioral targeting and c…" 28 quotes checked, 0 failed.
The noise-baseline probe and its output
Eight phrasings for the same idea, run over all 5,855 papers with text. The narrow term returns 2; the widened probe returns 11 of 32 within the audit set.
- _aa_noise.mjs
// Concept-level probe: does the paper establish a same-treatment baseline // (what an A/A test is), under ANY name? import fs from 'node:fs'; import path from 'node:path'; import { dataRoot, loadExtractions } from './lib.mjs'; const ROOT=dataRoot(); const TERMS={ 'A/A test': /\bA\/A[ -]?(?:test|experiment)/i, 'control-control / null experiment': /\bcontrol[- ]control\b|\bnull experiment/i, 'noise floor': /\bnoise floor\b/i, 'identical/identically configured arms': /\bidentical(?:ly)? (?:configured |trained |seeded )?(?:client|browser|profile|persona|account|agent|machine|instance)s?\b/i, 'permutation / randomisation test': /\bpermutation test|\brandomi[sz]ation test\b/i, 'null distribution': /\bnull distribution\b/i, 'baseline noise / measurement noise': /\b(?:baseline|measurement|inherent|background) noise\b/i, 'two arms with the same treatment': /\bsame treatment\b|\bno[- ]?treatment (?:arm|group|control)\b/i, }; const papers=loadExtractions(); import { AUDITS, REJECTED } from './algorithm_audits_set.mjs'; const AUDIT=new Set(AUDITS.map(a=>a[0])); const REJ=new Set(REJECTED.map(a=>a[0])); if ([...AUDIT].some(k=>REJ.has(k))) throw new Error('a key is in both AUDITS and REJECTED'); const inAudit=k=>AUDIT.has(k); const tot={},aud={}; const per=new Map(); for(const p of papers){ const k=`${p.venue}/${p.year}/${p.slug}`; const f=path.join(ROOT,'fulltext',String(p.year),p.venue,p.slug,'paper.cols.txt'); if(!fs.existsSync(f))continue; const t=fs.readFileSync(f,'utf8').replace(/\s+/g,' '); for(const [n,re] of Object.entries(TERMS)) if(re.test(t)){ tot[n]=(tot[n]||0)+1; if(inAudit(k)){aud[n]=(aud[n]||0)+1; if(!per.has(k))per.set(k,[]); per.get(k).push(n);} } } console.log(`AUDIT keys parsed from report script: ${AUDIT.size}`); console.log(`${'term'.padEnd(38)} ${'corpus'.padStart(7)} ${'audit'.padStart(6)}`); for(const n of Object.keys(TERMS)) console.log(`${n.padEnd(38)} ${String(tot[n]||0).padStart(7)} ${String(aud[n]||0).padStart(6)}`); console.log(`\naudit papers with >=1 noise-baseline term: ${per.size} of ${AUDIT.size}`); for(const [k,v] of [...per].sort()) console.log(` ${k}\n ${v.join(', ')}`); console.log('\naudit papers with NONE:'); for(const k of AUDIT) if(!per.has(k)) console.log(` ${k}`);
- _aa_noise-output.txt
AUDIT keys parsed from report script: 34 term corpus audit A/A test 2 0 control-control / null experiment 31 1 noise floor 35 2 identical/identically configured arms 22 5 permutation / randomisation test 32 4 null distribution 1 0 baseline noise / measurement noise 219 2 two arms with the same treatment 13 1 audit papers with >=1 noise-baseline term: 11 of 34 CCS/2015/sunlight-fine-grained-targeting-detection-at-scale-with-statistical-confidence permutation / randomisation test IMC/2010/challenges-in-measuring-online-advertising-systems control-control / null experiment, identical/identically configured arms IMC/2015/location-location-location-the-impact-of-geolocation-on-web-search-personalizati noise floor, identical/identically configured arms IMC/2025/from-voice-to-ads-auditing-commercial-smart-speakers-for-targeted-advertising-ba permutation / randomisation test, baseline noise / measurement noise PETS/2015/automated-experiments-on-ad-privacy-settings identical/identically configured arms, permutation / randomisation test PETS/2025/more-and-scammier-ads-the-perils-of-youtubes-ad-privacy-settings identical/identically configured arms PETS/2026/ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis permutation / randomisation test WWW/2018/auditing-the-personalization-and-composition-of-politically-related-search-engin noise floor WWW/2019/measuring-political-personalization-of-google-news-search identical/identically configured arms WWW/2022/left-or-right-a-peek-into-the-political-biases-in-email-spam-filtering-algorithm two arms with the same treatment WWW/2024/tiktok-and-the-art-of-personalization-investigating-exploration-and-exploitation baseline noise / measurement noise audit papers with NONE: CCS/2014/your-online-interests-pwned-a-pollution-attack-against-targeted-advertising IMC/2015/peeking-beneath-the-hood-of-uber USENIX/2016/tracing-information-flows-between-ad-exchanges-using-retargeted-ads WWW/2018/adbudgetkiller-online-advertising-budget-draining-attack PETS/2019/investigating-sources-of-pii-used-in-facebook-s-targeted-advertising PETS/2020/inferring-tracker-advertiser-relationships-in-the-online-advertising-ecosystem-u WWW/2020/stop-tracking-me-bro-differential-tracking-of-user-demographics-on-hyper-partisa CCS/2022/cart-ology-intercepting-targeted-advertising-via-ad-network-identity-entanglemen IMC/2022/measurement-and-analysis-of-implied-identity-in-ad-delivery-optimization IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising-a-field-measurement NDSS/2022/auto-draft-209 PETS/2022/atom-ad-network-tomography CCS/2023/marketing-to-children-through-online-targeted-advertising-targeting-mechanisms-a IMC/2023/tracking-profiling-and-ad-targeting-in-the-alexa-echo-smart-speaker-ecosystem PETS/2023/a-utility-preserving-obfuscation-approach-for-youtube-recommendations WWW/2023/a-method-to-assess-and-explain-disparate-impact-in-online-retailing PETS/2024/opted-out-yet-tracked-are-regulations-enough-to-protect-your-privacy IMC/2024/poster-identifying-filter-bubble-based-on-feed-level-embedding-similarity-analys PETS/2025/echoes-of-privacy-uncovering-the-profiling-practices-of-voice-assistants PETS/2025/sheeps-clothing-wolfish-intent-automated-detection-and-evaluation-of-problematic USENIX/2025/big-help-or-big-brother-auditing-tracking-profiling-and-personalization-in-gener IEEE-SP/2026/setting-the-course-but-forgetting-to-steer-analyzing-compliance-with-gdprs-right WWW/2026/does-this-button-work-investigating-youtubes-ineffective-user-controls
The recall-repair probes and their output
- _aa_residue.mjs
// Recall repair: the 158 papers in the loose probe-2 pool that were dropped at // title level. Re-scored with the apparatus-density probe that recovered three // of the final 32, at a lower threshold, and printed for a second read. import fs from 'node:fs'; import path from 'node:path'; import { dataRoot, loadExtractions } from './lib.mjs'; import { AUDITS, REJECTED } from './algorithm_audits_set.mjs'; const ROOT=dataRoot(); const done=new Set([...AUDITS.map(a=>a[0]),...REJECTED.map(a=>a[0])]); const ft=JSON.parse(fs.readFileSync('aa/probe1.json','utf8')); const g=(r,k)=>r.h[k]||0; const outcome=r=>g(r,'personalization')+g(r,'pricedisc')+g(r,'diftreat')+g(r,'adtargeting')+g(r,'bubble'); const app=r=>g(r,'sockpuppet')+g(r,'persona')+g(r,'pairedarm')+g(r,'aatest')+g(r,'trainedprofile')+g(r,'controlarm'); const loose=ft.filter(r=>outcome(r)>=1&&app(r)>=1).map(r=>r.key); const residue=loose.filter(k=>!done.has(k)); const P=new Map(loadExtractions().map(p=>[`${p.venue}/${p.year}/${p.slug}`,p])); const RE=/\b(?:sock ?-?puppets?|user personas?|shopper personas?|synthetic profiles?|trained? (?:browser )?profiles?|training profiles?|control (?:profile|account|persona)s?|treatment (?:persona|profile|group)s?|seeded (?:with )?interest|fresh profiles?|experimental group)\b/gi; const OUT=/\b(?:ads? (?:served|shown|delivered|received|displayed)|search results?|recommendations?|prices?|the feed|bid)\b/i; const rows=[]; for(const k of residue){ const p=P.get(k); const f=path.join(ROOT,'fulltext',String(p.year),p.venue,p.slug,'paper.cols.txt'); if(!fs.existsSync(f))continue; const t=fs.readFileSync(f,'utf8').replace(/\s+/g,' '); const m=t.match(RE); const n=m?m.length:0; if(n>=1&&OUT.test(t)) rows.push({n,k,p}); } console.log(`loose=${loose.length} adjudicated=${loose.length-residue.length} residue=${residue.length} residue_with_apparatus_and_outcome=${rows.length}`); for(const r of rows.sort((a,b)=>b.n-a.n)) console.log(`${String(r.n).padStart(3)} ${r.p.year} ${r.p.venue.padEnd(8)} ${r.p.title}`);
- _aa_residue-output.txt
loose=190 adjudicated=33 residue=157 residue_with_apparatus_and_outcome=33 25 2021 CCS The Effect of Google Search on Software Security: Unobtrusive Security Interventions via Content Re-ranking. 14 2020 WWW Finding a Choice in a Haystack: Automatic Extraction of Opt-Out Statements from Privacy Policy Text. 12 2025 IEEE-SP Restricting the Link: Effects of Focused Attention and Time Delay on Phishing Warning Effectiveness. 9 2022 PETS Increasing Adoption of Tor Browser Using Informational and Planning Nudges 5 2019 WWW How Intention Informed Recommendations Modulate Choices: A Field Study of Spoken Word Content. 4 2020 PETS Multiple Purposes, Multiple Problems: A User Study of Consent Dialogs after GDPR 4 2024 PETS Supporting Informed Choices about Browser Cookies: The Impact of Personalised Cookie Banners 4 2024 USENIX More Simplicity for Trainers, More Opportunity for Attackers: Black-Box Attacks on Speaker Recognition Systems by Inferring Feature Extractor 3 2023 WWW Ad Auction Design with Coupon-Dependent Conversion Rate in the Auto-bidding World. 3 2025 CCS Phishing Susceptibility and the (In-)Effectiveness of Common Anti-Phishing Interventions in a Large University Hospital. 2 2023 WWW Understanding the Behaviors of Toxic Accounts on Reddit. 2 2025 USENIX Vulnerability of Text-Matching in ML/AI Conference Reviewer Assignments to Collusions 2 2025 WWW Causal Insights into Parler's Content Moderation Shift: Effects on Toxicity and Factuality. 2 2026 PETS Redefining Website Fingerprinting Attacks with Multi-Agent LLMs 2 2026 WWW Community Fact-Checks Do Not Break Follower Loyalty. 2 2025 IEEE-SP The Importance of Being Earnest: Shedding Light on Johnny's (False) Sense of Privacy. 2 2025 USENIX Malicious LLM-Based Conversational AI Makes Users Reveal Personal Information 1 2017 USENIX A Privacy Analysis of Cross-device Tracking 1 2019 WWW Automatic Generation of Pattern-controlled Product Description in E-commerce. 1 2019 WWW Multiple Treatment Effect Estimation using Deep Generative Model with Task Embedding. 1 2020 PETS In-Depth Evaluation of Redirect Tracking and Link Usage 1 2020 PETS No boundaries: data exfiltration by third parties embedded on web pages 1 2020 WWW The Representativeness of Automated Web Crawls as a Surrogate for Human Browsing. 1 2021 PETS The CNAME of the Game: Large-scale Analysis of DNS-based Tracking Evasion 1 2022 PETS Disparate Vulnerability to Membership Inference Attacks 1 2023 PETS Comparing Large-Scale Privacy and Security Notifications 1 2023 WWW Near-Optimal Experimental Design Under the Budget Constraint in Online Platforms. 1 2024 PETS A Large-Scale Study of Cookie Banner Interaction Tools and their Impact on Users' Privacy 1 2025 IEEE-SP "It's Time. Time for Digital Security.": An End User Study on Actionable Security and Privacy Advice. 1 2025 USENIX "Please don't send that bot anything": A Mixed-methods Study of Personal Impersonation Attacks Targeting Digital Payments on Social Media 1 2026 WWW How Social Media Peer Comments Influence Privacy Decisions in Photo Sharing: Context and Individual Differences Cause Comments to Backfire. 1 2025 WWW Reducing Symbiosis Bias through Better A/B Tests of Recommendation Algorithms. 1 2026 PETS Dead Domains, Living Data: A Privacy Risk Analysis of Domain Lifecycle in Android Apps
- _aa_zeroyears.mjs
// Recall check for the years the audit set is empty (2011-2013, 2017, 2021). // _aa_gap.mjs needs >=4 apparatus matches, which is too strict to prove a // NEGATIVE: it returns nothing at all for 2017. This probe drops the threshold // to >=2 and uses an apparatus-only vocabulary (no outcome term required), so a // zero here is evidence and not just a threshold artefact. // node scripts/_aa_zeroyears.mjs # all five zero-years // node scripts/_aa_zeroyears.mjs 2017,2021 import fs from 'node:fs'; import path from 'node:path'; import { dataRoot, loadExtractions } from './lib.mjs'; import { AUDITS } from './algorithm_audits_set.mjs'; const ROOT = dataRoot(); const YEARS = (process.argv[2] || '2011,2012,2013,2017,2021').split(',').map(Number); const auditYears = new Set(AUDITS.map(([k]) => Number(k.split('/')[1]))); for (const y of YEARS) { if (auditYears.has(y)) { console.error(`CONTRACT VIOLATED: ${y} is not a zero-year -- the audit set has a paper in it`); process.exit(1); } } const RE = /\b(?:sock ?-?puppets?|user personas?|shopper personas?|synthetic profiles?|trained? (?:browser )?profiles?|training profiles?|control (?:profile|account|persona)s?|treatment (?:persona|profile)s?|seeded (?:with )?interest)\b/gi; let scanned = 0; const hits = []; for (const p of loadExtractions()) { if (!YEARS.includes(p.year)) continue; const f = path.join(ROOT, 'fulltext', String(p.year), p.venue, p.slug, 'paper.cols.txt'); if (!fs.existsSync(f)) continue; scanned++; const t = fs.readFileSync(f, 'utf8').replace(/\s+/g, ' '); const m = t.match(RE); if (m && m.length >= 2) hits.push({ n: m.length, p }); } console.log(`years=${YEARS.join(',')} papers scanned=${scanned} papers with >=2 apparatus terms=${hits.length}`); for (const h of hits.sort((a, b) => b.n - a.n)) console.log(`${String(h.n).padStart(4)} ${h.p.year} ${h.p.venue.padEnd(8)} ${h.p.title}`); console.log('\nEach hit above was read. None is a differential audit; the zero-years are real.');
- _aa_zeroyears-output.txt
years=2011,2012,2013,2017,2021 papers scanned=1002 papers with >=2 apparatus terms=2 316 2017 WWW An Army of Me: Sockpuppets in Online Discussion Communities. 3 2012 WWW Spotting fake reviewer groups in consumer reviews. Each hit above was read. None is a differential audit; the zero-years are real.
Related
- algorithm_audits — the page this log is for.
- corpus — the corpus, its funnel and its provisional years.
- stateful_stateless — the neighbouring log; its 29-paper comparison-study audit is the closest methodological precedent for the hand-adjudication done here.
- platforms — where the corpus-wide “sock puppet” count also appears, against a different denominator.
- [1]
- Meng, Wei; Xing, Xinyu; Sheth, Anmol; Weinsberg, Udi; Lee, Wenke (2014): "Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [2]
- Kim, I Luk; Wang, Weihang; Kwon, Yonghwi; Zheng, Yunhui; Aafer, Yousra; Meng, Weijie; Zhang, Xiangyu (2018): "AdBudgetKiller: Online Advertising Budget Draining Attack", in: Proceedings of the ACM Web Conference. (DOI)
- [3]
- Zhang, Jiang; Psounis, Konstantinos; Haroon, Muhammad; Shafiq, Zubair (2022): "HARPO: Learning to Subvert Online Behavioral Advertising", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [4]
- Lécuyer, Mathias; Spahn, Riley; Spiliopolous, Yannis; Chaintreau, Augustin; Geambasu, Roxana; Hsu, Daniel J. (2015): "Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [5]
- Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [6]
- Silva, Márcio; Oliveira, Lucas Santos de; Andreou, Athanasios; Melo, Pedro Olmo Stancioli Vaz de; Goga, Oana; Benevenuto, Fabrício (2020): "Facebook Ads Monitor: An Independent Auditing System for Political Ads on Facebook", in: Proceedings of the ACM Web Conference. (DOI)
- [7]
- Ali, Muhammad; Goetzen, Angelica; Mislove, Alan; Redmiles, Elissa M.; Sapiezynski, Piotr (2023): "Problematic Advertising and its Disparate Exposure on Facebook", in: Proceedings of the USENIX Security Symposium. (Link)
- [8]
- Lone, Qasim; Frik, Alisa; Luckie, Matthew; Korczyński, Maciej; van Eeten, Michel; Gañán, Carlos (2022): "Deployment of Source Address Validation by Network Operators: A Randomized Control Trial", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [9]
- Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [10]
- Agarwal, Pushkal; Joglekar, Sagar; Papadopoulos, Panagiotis; Sastry, Nishanth; Kourtellis, Nicolas (2020): "Stop tracking me Bro! Differential Tracking of User Demographics on Hyper-Partisan Websites", in: Proceedings of the ACM Web Conference. (DOI)
- [11]
- Iqbal, Hassan; Khan, Usman Mahmood; Khan, Hassan Ali; Shahzad, Muhammad (2022): "Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Election 2020", in: Proceedings of the ACM Web Conference. (DOI)
- [12]
- Robertson, Ronald E.; Lazer, David; Wilson, Christo (2018): "Auditing the Personalization and Composition of Politically-Related Search Engine Results Pages", in: Proceedings of the ACM Web Conference. (DOI)
- [13]
- Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
- [14]
- Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [15]
- Roongta, Ritik; Jose, Julia; Habib, Hussam; Greenstadt, Rachel (2025): "Sheep's Clothing, Wolfish Intent: Automated Detection and Evaluation of Problematic 'Allowed' Advertisements", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [16]
- Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [17]
- Chen, Le; Mislove, Alan; Wilson, Christo (2015): "Peeking Beneath the Hood of Uber", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [18]
- Becerril-Arreola, Rafael (2023): "A Method to Assess and Explain Disparate Impact in Online Retailing", in: Proceedings of the ACM Web Conference. (DOI)
- [19]
- Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [20]
- Sun, Chen; Vekaria, Yash; Nithyanand, Rishab (2026): "On the Suitability of LLM-Driven Agents for Dark Pattern Audits", Proceedings on Privacy Enhancing Technologies 2026(4):927-946. (DOI)
