User Tools

Site Tools


design:algorithm_audits

Algorithm Audits

You want to know whether a platform treats two users differently, and you cannot see its code. The instrument is a differential audit: you create two or more measurement identities that differ in exactly one property you control, drive them at the same platform, and compare what comes back. Sock-puppet accounts, trained browser profiles and paired crawls are all the same design — the arms are the experiment, and everything else is plumbing.

This page is the design. The profile machinery is Stateful stateless; the significance test is Hypothesis testing; the accounts themselves are Registration; whether you are allowed to make them is Ethics and Platforms. What nobody else on this site owns, and what a student reliably gets wrong, is the bit in between: what counts as a control arm here, how much the platform moves when you change nothing, and how to find that out before you claim a difference.

The hard part of an audit is not the treatment. It is the null. Ad systems, feeds and search results are stochastic and non-stationary: two arms configured identically, run at the same moment from the same network, do not return the same thing. The first measurement paper on this in the corpus says it in one sentence [1Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]:

“Even queries launched simultaneously from two identically configured clients on the same subnet can produce wildly different ads over multiple timescales.”

If you have not measured how far apart two identical arms land, you cannot say what a gap between two different arms means. That measurement — an A/A test, a control–control run, a null experiment, whatever you call it — is the most-skipped step in this literature. Reading all 34 audit papers in this corpus for it: 10 establish a null and 24 do not.

What to read first

Paper Why it is on this page
Guha, Cheng and Francis, IMC 2010, Challenges in Measuring Online Advertising Systems [1Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] The methodology paper. Noise, churn, and “we perform all analysis relative to a control experiment”. Sixteen years old and still the clearest statement of the problem.
Datta, Tschantz and Datta, PoPETs 2015, Automated Experiments on Ad Privacy Settings [2Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] AdFisher: the audit as a randomised controlled experiment with a permutation test, not as a crawl with two conditions. “We created an experimental group and a control group of agents.”
Lécuyer et al., CCS 2015, Sunlight [3Lécuyer, Mathias; Spahn, Riley; Spiliopolous, Yannis; Chaintreau, Augustin; Geambasu, Roxana; Hsu, Daniel J. (2015): "Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] What statistical confidence in a targeting claim requires, and what the prior generation of audits left out.
Le et al., TheWebConf 2019, Measuring Political Personalization of Google News Search [4Le, Huyen T.; Maragh, Raven; Ekdale, Brian; High, Andrew; Havens, Timothy; Shafiq, Zubair (2019): "Measuring Political Personalization of Google News Search", in: Proceedings of the ACM Web Conference. (DOI)] The canonical sock-puppet design, stated in one sentence: a pair of fresh browser profiles trained divergently, then “executes identical politically oriented Google News searches”.
Hannak et al., TheWebConf 2013, Measuring Personalization of Web Search [5Hannak, Aniko; Sapiezynski, Piotr; Kakhki, Arash Molavi; Krishnamurthy, Balachander; Lazer, David; Mislove, Alan; Wilson, Christo (2013): "Measuring personalization of web search", in: Proceedings of the ACM Web Conference. (DOI)] The paper that defined the method for the whole field — and it is not in this corpus (see What this corpus cannot tell you). Read it for the control-account design and the noise decomposition.
Imana, Korolova and Heidemann, TheWebConf 2021, Auditing for Discrimination in Algorithms Delivering Job Ads [6Imana, Basileal; Korolova, Aleksandra; Heidemann, John S. (2021): "Auditing for Discrimination in Algorithms Delivering Job Ads", in: Proceedings of the ACM Web Conference. (DOI)] The paired-campaign design that separates advertiser targeting from platform delivery. Also not in this corpus.
Le et al., IMC 2025, From Voice to Ads [7Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] The most recent design to build the null in on purpose: “paired puppets” that differ in nothing, run alongside the ones that differ in something.

Anatomy of an audit

Four things, and if any of them is missing the result is not readable.

  1. The treatment. The one property you vary. In this corpus it has been: browsing history [4Le, Huyen T.; Maragh, Raven; Ekdale, Brian; High, Andrew; Havens, Timothy; Shafiq, Zubair (2019): "Measuring Political Personalization of Google News Search", in: Proceedings of the ACM Web Conference. (DOI)], declared interest category [8Liu, Zengrui; Iqbal, Umar; Saxena, Nitesh (2024): "Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy?", in: Proceedings on Privacy Enhancing Technologies. (DOI)], geolocation [9Kliman-Silver, Chloe; Hannak, Aniko; Lazer, David; Wilson, Christo; Mislove, Alan (2015): "Location, Location, Location: The Impact of Geolocation on Web Search Personalization", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], demographic attributes [10Agarwal, Pushkal; Joglekar, Sagar; Papadopoulos, Panagiotis; Sastry, Nishanth; Kourtellis, Nicolas (2020): "Stop tracking me Bro! Differential Tracking of User Demographics on Hyper-Partisan Websites", in: Proceedings of the ACM Web Conference. (DOI)], a privacy setting turned on or off [11Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)], a blocker installed or not [12Roongta, Ritik; Jose, Julia; Habib, Hussam; Greenstadt, Rachel (2025): "Sheep's Clothing, Wolfish Intent: Automated Detection and Evaluation of Problematic 'Allowed' Advertisements", in: Proceedings on Privacy Enhancing Technologies. (DOI)], an ad creative's content [13Kaplan, Levi; Gerzon, Nicole; Mislove, Alan; Sapiezynski, Piotr (2022): "Measurement and analysis of implied identity in ad delivery optimization", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], the voice speaking to a smart speaker [7Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], and which trackers were allowed to observe the profile [14Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)].
  2. The control arm. The arm that does not get the treatment but is otherwise identical, including in how much it browsed, from where, at what time. “The browser agents in the experimental group visited websites on substance abuse while the agents in the control group simply waited” [2Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] — note that the control still ran; it did not simply not exist.
  3. The outcome. What the platform returns: ads, ranked results, prices, a feed, a bid, a data-download package. If your outcome is “how many third parties loaded”, you are running a crawl, not an audit, and Requests is your page.
  4. The null. How far apart two arms land when the treatment is absent. See below; this is the one that goes missing.

What a control arm is, and what it is not

This is a control arm This is not
A second profile trained for the same number of page loads, on the same schedule, on content unrelated to the treatment A blank, never-used profile compared against a profile that browsed for six hours. Browsing itself is then part of your treatment.
The same account with the setting toggled, measured in a randomised order across repeats The same account measured before and after you toggled the setting, once. That is a before/after design with the day confounded in — see Longitudinal.
Arms interleaved in time, so ad churn and platform-side experiments hit every arm equally Arm A on Monday, arm B on Tuesday. [1Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] measures ad churn at 1–4% per minute.
Arms on separate identities that never share storage, IP or device fingerprint “We cleared cookies between arms.” That leaves localStorage, the cache, the IP and the fingerprint — see What a reset actually resets.

Carry-over: the arm contaminates the next arm

An audit is a stateful design by construction — the profile is the independent variable — and stateful crawls are order-dependent. Three carry-over channels, in the order they bite:

  • Within the identity. Reusing one browser profile for arm A then arm B leaves arm A's history in arm B. The fix is separate profiles, not a reset; see the reset probe on The code.
  • Across identities, through the network. Two “different users” behind one egress IP are one household to an ad platform, and a datacentre IP is a fourth thing again (Crawling location). [7Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] keeps this explicit: all puppets “set up to use the same network, geolocated to the same location” — identical for every arm, so it cannot explain a difference.
  • Across time, through the platform. Feed and ad systems adapt to what you did yesterday. If arm A ran first and the platform re-scored your whole address range, arm B inherits it. Randomise arm order per repeat and say you did. [7Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]: “We assigned voices randomly to days, and scheduled paired puppets in parallel.”

Sizing the noise floor before you claim a difference

Run two arms with the same treatment and measure the distance between their outcomes with the same metric you will use for the real comparison. That distribution is your noise floor. A treatment effect smaller than it is not a finding.

Three ways the corpus does it:

  1. Paired identical arms. [7Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] deploys each cloned voice on two puppets — “paired puppets” — that differ in nothing, and reads the paired difference as the baseline for the voice-characteristic effect.
  2. A control experiment the whole analysis is relative to. [1Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]: “In this paper we perform all analysis relative to a control experiment”, because the false-positive and false-negative behaviour of ad de-duplication affects both sides equally.
  3. A permutation test over the arm labels. [2Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] does not assume a noise model: it trains a classifier to tell the two arms apart and asks how often a classifier does that well when the arm labels are shuffled. The null is generated from your own data. [3Lécuyer, Mathias; Spahn, Riley; Spiliopolous, Yannis; Chaintreau, Augustin; Geambasu, Roxana; Hsu, Daniel J. (2015): "Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] and [15Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)] use the same machinery.

The industry name for route 1 is an A/A test, and in this corpus that phrase belongs to the platform side, not the audit side: both occurrences are TheWebConf papers about running experiments on a platform you own — “We simulate A/A tests by re-randomizing the treatment assignments on the observed exposure logs from our experiment corpus” [16Chandar, Praveen; Thomas, Brian St.; Maystre, Lucas; Pappu, Vijay; Sanchis-Ojeda, Roberto; Wu, Tiffany; Carterette, Ben; Lalmas, Mounia; Jebara, Tony (2022): "Using Survival Models to Estimate User Engagement in Online Experiments", in: Proceedings of the ACM Web Conference. (DOI)]. Nobody auditing from outside uses the term. That vocabulary gap is not the same as the practice gap, and the two have to be counted separately: see The null, and how it was counted.

Cheapest version, if you do nothing else. Before the real run, stand up 2n arms where all 2n get treatment “nothing”. Run them for the full training duration and measure them exactly as you will measure the real ones. Report the distribution of pairwise distances. It costs one extra run and it is the difference between “we observed a 7% gap” and “we observed a 7% gap against a same-treatment baseline whose 95th percentile was 3%”.

How many arms, how many repeats

The unit of analysis is not the individual result, ad impression or video. It is the arm-run: one identity, trained once, measured once. Ten thousand impressions collected from two profiles give you two independent observations of the thing you are claiming, not twenty thousand. This is the same unit-of-analysis error that The Unit of Analysis Is the Error That Actually Invalidates Results measures across the corpus. An audit is an easy place to make it, because the impression counts are so large that a p-value computed over them is guaranteed to be tiny — that page measures how often the field makes it; this one only says where the temptation is.

Two consequences:

  • Replicate the identity, not the query. Twenty profiles per condition measured ten times beats two profiles measured a hundred times: the same 200 observations, but n = 20 instead of n = 2 in the test. It is not free — you pay for twenty training runs instead of two, and training is the expensive part — which is the real reason the corpus is full of two-profile designs. Budget for it at design time, not after.
  • Multiple comparisons arrive by default. One audit with 16 interest personas against a control is 16 tests, not one. [8Liu, Zengrui; Iqbal, Umar; Saxena, Nitesh (2024): "Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy?", in: Proceedings on Privacy Enhancing Technologies. (DOI)] is explicit — “we also conduct Bonferroni correction on the statistical test” — and a multiple-comparison correction is named by 7 of the 34 — Holm–Bonferroni in 5 of them, plain Bonferroni in 3. See Pvalue corrections.

Which methods are current

Dated 2026-09-11 against this corpus's year counts, which are given in full in Use in publications.

Method Status Evidence
Trained browser-profile personas driven at a live ad or search system current, and the workhorse 14 of the 34 audit papers use the word “persona”, in every year from 2020 on and first in 2016 — [1Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] does the same thing in 2010 and calls them personae [17Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)], [18Musa, Maaz Bin; Nithyanand, Rishab (2022): "ATOM: Ad-network Tomography", in: Proceedings on Privacy Enhancing Technologies. (DOI)], [8Liu, Zengrui; Iqbal, Umar; Saxena, Nitesh (2024): "Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy?", in: Proceedings on Privacy Enhancing Technologies. (DOI)], [15Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)]
Sock-puppet accounts (a logged-in identity you created) current, and unavoidable for feeds and recommenders 5 of the 34; 26 papers corpus-wide use the term. All since 2019 [4Le, Huyen T.; Maragh, Raven; Ekdale, Brian; High, Andrew; Havens, Timothy; Shafiq, Zubair (2019): "Measuring Political Personalization of Google News Search", in: Proceedings of the ACM Web Conference. (DOI)], [19Zhang, Jiang; Askari, Hadi; Psounis, Konstantinos; Shafiq, Zubair (2023): "A Utility-Preserving Obfuscation Approach for YouTube Recommendations", Proceedings on Privacy Enhancing Technologies 2023(4). (DOI)], [11Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)], [20Karnam, Sai Keerthana; Dash, Abhisek; Das, Antariksh; Mousavi, Sepehr; Bechtold, Stefan; Gummadi, Krishna P.; Mukherjee, Animesh; Weber, Ingmar; Zannettou, Savvas (2026): "Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]
Paired-campaign advertiser-side audits (you buy the ads) current [13Kaplan, Levi; Gerzon, Nicole; Mislove, Alan; Sapiezynski, Piotr (2022): "Measurement and analysis of implied identity in ad delivery optimization", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] ran “200 versions of this ad at the same time, all from the same account and with the same budget”; [6Imana, Basileal; Korolova, Aleksandra; Heidemann, John S. (2021): "Auditing for Discrimination in Algorithms Delivering Job Ads", in: Proceedings of the ACM Web Conference. (DOI)] is the canonical design and is outside this corpus
Randomised field experiment with real users (browser extension, random assignment) current, growing, and the strongest design available [21McCrosky, Jesse; Malla, Ranadheer; Tanskanen, Aapo; Camargo, Chico Q. (2026): "Does This Button Work? Investigating YouTube's Ineffective User Controls", in: Proceedings of the ACM Web Conference. (DOI)], 22,722 participants, “depending on randomized assignment”
Data donation instead of arms current, and the answer to a different question [22Silva, Márcio; Oliveira, Lucas Santos de; Andreou, Athanasios; Melo, Pedro Olmo Stancioli Vaz de; Goga, Oana; Benevenuto, Fabrício (2020): "Facebook Ads Monitor: An Independent Auditing System for Political Ads on Facebook", in: Proceedings of the ACM Web Conference. (DOI)], [23Ali, Muhammad; Goetzen, Angelica; Mislove, Alan; Redmiles, Elissa M.; Sapiezynski, Piotr (2023): "Problematic Advertising and its Disparate Exposure on Facebook", in: Proceedings of the USENIX Security Symposium. (Link)], [24Gkiouzepi, Eleni; Andreou, Athanasios; Goga, Oana; Loiseau, Patrick (2023): "Collaborative Ad Transparency: Promises and Limitations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] — the last one names the trade-off: “One method that does not use fake personas” is to model over donated exposures. You get ecological validity and lose the treatment.
Right-of-access requests as the measuring instrument new (2026), and complementary [20Karnam, Sai Keerthana; Dash, Abhisek; Das, Antariksh; Mousavi, Sepehr; Bechtold, Stefan; Gummadi, Krishna P.; Mukherjee, Animesh; Weber, Ingmar; Zannettou, Savvas (2026): "Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] drives sock puppets to generate ground truth, then compares the platform's own GDPR Art. 15 export against it
LLM-driven agents as the audit instrument new (2026), and outside this page's population [25Sun, Chen; Vekaria, Yash; Nithyanand, Rishab (2026): "On the Suitability of LLM-Driven Agents for Dark Pattern Audits", Proceedings on Privacy Enhancing Technologies 2026(4):927-946. (DOI)] audits dark patterns in rights-request workflows — not a differential audit, so it is not one of the 34 — but it is the first paper here to ask whether an agent can be the instrument, and its answer is a failure analysis. LLM agents
Price discrimination / price steering as the outcome rare in these seven venues, not dormant 9 of the 34 mention the terms; only 2 measure a price — [26Chen, Le; Mislove, Alan; Wilson, Christo (2015): "Peeking Beneath the Hood of Uber", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] (Uber surge pricing, IMC 2015) and [27Becerril-Arreola, Rafael (2023): "A Method to Assess and Explain Disparate Impact in Online Retailing", in: Proceedings of the ACM Web Conference. (DOI)] (retail prices, recommendations and delivery fees matched to census demographics, TheWebConf 2023). The founding papers [28Hannak, Aniko; Soeller, Gary; Lazer, David; Mislove, Alan; Wilson, Christo (2014): "Measuring Price Discrimination and Steering on E-commerce Web Sites", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and Vissers et al. (PETS 2014) are outside this corpus, so the scarcity here is partly the screen and not the field
Single blank profile vs single aged profile, one run each, no null superseded the design [1Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] argued against in 2010; it survives in this corpus only as a component of attack papers [29Meng, Wei; Xing, Xinyu; Sheth, Anmol; Weinsberg, Udi; Lee, Wenke (2014): "Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [30Kim, I Luk; Wang, Weihang; Kwon, Yonghwi; Zheng, Yunhui; Aafer, Yousra; Meng, Weijie; Zhang, Xiangyu (2018): "AdBudgetKiller: Online Advertising Budget Draining Attack", in: Proceedings of the ACM Web Conference. (DOI)]

What is not superseded, despite looking old. The 2015 instrument papers are still the state of the art for the statistics. [2Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)]'s permutation-test-over-arm-labels design has not been replaced by anything in this corpus; the 2025–2026 audits that do build a null ([7Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [15Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)]) use the same idea. If a reviewer asks how you established significance, AdFisher is still the citation.

What is genuinely newer. Two things arrived after 2023 and neither existed when the method was defined: randomised assignment over real users' own sessions [21McCrosky, Jesse; Malla, Ranadheer; Tanskanen, Aapo; Camargo, Chico Q. (2026): "Does This Button Work? Investigating YouTube's Ineffective User Controls", in: Proceedings of the ACM Web Conference. (DOI)], and the platform's own data export as ground truth [20Karnam, Sai Keerthana; Dash, Abhisek; Das, Antariksh; Mousavi, Sepehr; Bechtold, Stefan; Gummadi, Krishna P.; Mukherjee, Animesh; Weber, Ingmar; Zannettou, Savvas (2026): "Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]. The first is the strongest design on this page and the one most constrained by ethics review and recruitment. A third, agentic LLM crawlers as the thing doing the auditing, has exactly one paper and it is a failure analysis [25Sun, Chen; Vekaria, Yash; Nithyanand, Rishab (2026): "On the Suitability of LLM-Driven Agents for Dark Pattern Audits", Proceedings on Privacy Enhancing Technologies 2026(4):927-946. (DOI)].

What this corpus cannot tell you

This is the page where the corpus is weakest, and the reason is structural. The 5,859-paper extraction was selected by screening each abstract for “is this a security measurement or a privacy measurement”. An algorithm audit is often neither — a search-personalisation or price-discrimination study is a fairness, economics or information-retrieval paper — so the screen drops it.

Of 117 audit-topical candidates in the 16,864-record bibliographic index, 48 are not in the extraction: 45 were screened out with both screen labels false, and 3 fall in venue-years the index has no label records for at all. Among the 45:

  • Hannak et al., TheWebConf 2013, Measuring personalization of web search [5Hannak, Aniko; Sapiezynski, Piotr; Kakhki, Arash Molavi; Krishnamurthy, Balachander; Lazer, David; Mislove, Alan; Wilson, Christo (2013): "Measuring personalization of web search", in: Proceedings of the ACM Web Conference. (DOI)] — the paper that defined the method.
  • Hannak et al., IMC 2014, Measuring Price Discrimination and Steering on E-commerce Web Sites [28Hannak, Aniko; Soeller, Gary; Lazer, David; Mislove, Alan; Wilson, Christo (2014): "Measuring Price Discrimination and Steering on E-commerce Web Sites", in: Proceedings of the ACM Internet Measurement Conference. (DOI)].
  • Imana et al., TheWebConf 2021, Auditing for Discrimination in Algorithms Delivering Job Ads [6Imana, Basileal; Korolova, Aleksandra; Heidemann, John S. (2021): "Auditing for Discrimination in Algorithms Delivering Job Ads", in: Proceedings of the ACM Web Conference. (DOI)].
  • Boeker and Urman, TheWebConf 2022, An Empirical Investigation of Personalization Factors on TikTok.
  • MapWatch: Detecting and Monitoring International Border Personalization on Online Maps, TheWebConf 2016.

and among the 3 venue-year gaps, Vissers et al., PETS 2014, Crying Wolf? On the Price Discrimination of Online Airline Tickets, and Khattak et al., NDSS 2016, Do You See What I See? Differential Treatment of Anonymous Users [31Khattak, Sheharbano; Fifield, David; Afroz, Sadia; Javed, Mobin; Sundaresan, Srikanth; McCoy, Damon; Paxson, Vern; Murdoch, Steven J. (2016): "Do You See What I See? Differential Treatment of Anonymous Users", in: Proceedings of the Network and Distributed System Security Symposium. (Link)].

So every count on this page is a lower bound, and it is a biased one — it under-counts precisely the audits whose framing is fairness or economics rather than privacy. Treat the 34 as a reading list and a shape, not as a census. The corresponding claim about the wider literature — how many algorithm audits exist — this corpus cannot make at all.

Two more things it cannot tell you:

  • Whether an audit's arms were actually independent. crawlConfig carries one evidence quote for the whole configuration object, so the extraction cannot distinguish “two profiles” from “one profile reset twice”. Every verdict behind the 34-paper set was made by reading the paper; the sentences are on algorithm_audits.
  • Effect sizes across audits. Each paper's effect is measured on its own platform, outcome and metric. There is no comparable quantity to pool, and this page deliberately publishes none.

Use in publications

Figures below are from scripts/report_algorithm_audits.mjs over the 5,859-paper extraction (CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026). The audit population is hand-adjudicated, not extracted: the corpus has no field for “ran a differential audit”. The inclusion rule was written before the tables and is applied in the script itself; the full list, with the sentence that settled each verdict, and the twelve papers adjudicated out, are on algorithm_audits.

The inclusion rule. A paper is in the audit population if all three hold:

  1. (T) it deliberately varies a property of the measuring identity or request — profile history, declared attribute, location, device, opt-out setting, ad creative — and holds the rest fixed;
  2. (O) the outcome it measures is the platform's own response: ads served, results ranked, prices quoted, feed or recommendation contents, or an access decision;
  3. (C) the result is a difference (or a bounded absence of difference) between arms, not a prevalence over a crawl of many sites.

34 papers of 5,859 qualify. 36 further candidates were read and rejected; the commonest reason is (C) — an “audit” that is observational. Clause (O) excludes studies whose outcome is reachability rather than a discriminating response: geoblocking, censorship and Tor-exit refusal vary a vantage point rather than an identity, and belong to Blocking and geodifference.

The shape of the literature

Window Audit papers Share of the 34 Papers per year
2010–2015 (6 yr) 6 17.6% 1.00
2016–2019 (4 yr) 5 14.7% 1.25
2020–2023 (4 yr) 12 35.3% 3.00
2024–2026* (3 yr) 11 32.4% 3.67

Per year: 2010 1, 2014 1, 2015 4, 2016 1, 2018 2, 2019 2, 2020 2, 2022 6, 2023 4, 2024 3, 2025 5, 2026 3. 2011, 2012, 2013, 2017 and 2021 have none.

The windows are unequal (6, 4, 4, 3 years), so the comparable figure is papers per year: 1.00, 1.25, 3.00, 3.67. Read that carefully. The second window is barely above the first; the rise is between 2016–2019 and 2020–2023 and it holds after. The last column is provisional (CCS and IMC 2026 have not been held; see corpus), and n = 34 from a screen this page calls biased will not carry a trend claim. What the table supports is narrow and still useful: more of this work appeared in the last six years than in the first ten, and none of it has stopped. If you were told this method belongs to 2015, that is wrong.

Where audits publish

Share of that venue's own output, not share of the audit set.

Venue Venue papers Audits Share of venue
PETS 510 10 2.0%
IMC 638 8 1.3%
TheWebConf 843 8 0.9%
CCS 990 4 0.4%
NDSS 701 1 0.1%
IEEE S&P 767 1 0.1%
USENIX Security 1,410 2 0.1%

PoPETs is where this work lands, and USENIX Security — the largest venue in the corpus — has published two audits by this definition in seventeen years. Read that with What this corpus cannot tell you: TheWebConf's row is the one most depressed by the screen, because TheWebConf is where the fairness-framed audits are.

Platform measured (multi-valued): web 27 of 34 (79.4%), other-online-service 19 (55.9%), mobile 4 (11.8%), IoT 3 (8.8%), offline 0.

Audits report more than the corpus does — except where it counts

Each row states both populations, so the two cells are comparable.

Indicator Population Audit set Comparable baseline
runs a non-descriptive statistic all audits vs empirical 21 of 34 (61.8%) 1,637 of 5,118 (32.0%)
states an ethics-review outcome all audits vs empirical 20 of 34 (58.8%) 1,728 of 5,118 (33.8%)
states artifact availability all audits vs empirical 22 of 34 (64.7%) 2,890 of 5,118 (56.5%)
states a vantage location all audits vs measuredFrom 15 of 34 (44.1%) 1,228 of 3,908 (31.4%)
states crawl statefulness papers with a crawl config 24 of 26 (92.3%) 219 of 1,080 (20.3%)
states interaction depth papers with a crawl config 22 of 26 (84.6%) 841 of 1,080 (77.9%)
states a consent action papers with a crawl config 12 of 26 (46.2%) 349 of 1,080 (32.3%)
states whether it ran headless papers with a crawl config 5 of 26 (19.2%) 140 of 1,080 (13.0%)

The bottom four rows are counted over papers that have a crawl configuration on both sides, because a field that can only be stated on such a paper cannot fairly be divided by a population that includes papers without one. Eight of the 34 audits have no crawl configuration, and they are not oversights — they are audits the authors did not drive a browser for: an app ([26Chen, Le; Mislove, Alan; Wilson, Christo (2015): "Peeking Beneath the Hood of Uber", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]), ad campaigns bought on the platform ([32Venkatadri, Giridhari; Lucherini, Elena; Sapiezynski, Piotr; Mislove, Alan (2019): "Investigating sources of PII used in Facebook’s targeted advertising", in: Proceedings on Privacy Enhancing Technologies. (DOI)], [13Kaplan, Levi; Gerzon, Nicole; Mislove, Alan; Sapiezynski, Piotr (2022): "Measurement and analysis of implied identity in ad delivery optimization", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]), voice assistants ([7Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [33Khezresmaeilzadeh, Tina; Zhu, Elaine; Grieco, Kiersten; Dubois, Daniel; Psounis, Konstantinos; Choffnes, David (2025): "Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants", Proceedings on Privacy Enhancing Technologies 2025(2). (DOI)]), app stores ([15Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)]), and two studies that run in real participants' own browsers rather than a crawler ([34Zeng, Eric; McAmis, Rachel; Kohno, Tadayoshi; Roesner, Franziska (2022): "What Factors Affect Targeting and Bids in Online Advertising? A Field Measurement Study", in: Proceedings of the ACM Internet Measurement Conference, pp. 210-229. (DOI)], [21McCrosky, Jesse; Malla, Ranadheer; Tanskanen, Aapo; Camargo, Chico Q. (2026): "Does This Button Work? Investigating YouTube's Ineffective User Controls", in: Proceedings of the ACM Web Conference. (DOI)]).

On the fair denominator the statefulness row is the striking one: audits are the sub-population that actually reports profile state, 92.3% against 20.3%, because in an audit the profile is the result and you cannot describe the result without it. Audits are above the crawling baseline on every one of these four — which is what you would hope, and is not true of the vantage-location row, where 44.1% is better than the 31.4% baseline but still means more than half of these audits do not say where they measured from, on a design where the vantage point must be identical across arms or it is part of the treatment.

The null, and how it was counted

This is the page's central claim, so it is worth being exact about what supports it. Two different things are counted here and they are not the same number.

The vocabulary. A full-text sweep for eight ways of naming a same-treatment baseline, over the 5,855 papers with text:

Term Papers, corpus-wide In the audit set
A/A test 2 0 of 34
control-control / null experiment 31 1
noise floor 35 2
identical(ly configured/trained/seeded) arms 22 5
permutation / randomisation test 32 4
null distribution 1 0
baseline / measurement / inherent / background noise 219 2
same treatment / no-treatment arm 13 1

Both A/A test occurrences are platform-side experimentation papers, not audits. The audit side of this literature has no word for the thing.

The practice. A term count cannot answer whether a paper did it, and this one is wrong in both directions. It let in [35Robertson, Ronald E.; Lazer, David; Wilson, Christo (2018): "Auditing the Personalization and Composition of Politically-Related Search Engine Results Pages", in: Proceedings of the ACM Web Conference. (DOI)], whose only noise floor hit describes Hannak et al.'s result and not its own design, and [36Iqbal, Hassan; Khan, Usman Mahmood; Khan, Hassan Ali; Shahzad, Muhammad (2022): "Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Election 2020", in: Proceedings of the ACM Web Conference. (DOI)], whose same treatment is the phrasing of its research question. It missed [37Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], which runs four unused baseline profiles “created and mechanistically measured in the same way” and matches none of the eight phrasings. So all 34 were read for the question instead, and each verdict carries the sentence that settled it:

Route Papers What it means
A/A 4 two or more arms that differ in nothing
generated 3 a null built from the paper's own observations — permutation over arm labels, or a randomised baseline
repeats 1 the same condition run several times, with the spread reported
A/A + generated 2 both
any null 10 of 34 (29.4%)
none 24 of 34 (70.6%) states a difference without measuring what a non-difference looks like

The ten: [1Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [9Kliman-Silver, Chloe; Hannak, Aniko; Lazer, David; Wilson, Christo; Mislove, Alan (2015): "Location, Location, Location: The Impact of Geolocation on Web Search Personalization", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [4Le, Huyen T.; Maragh, Raven; Ekdale, Brian; High, Andrew; Havens, Timothy; Shafiq, Zubair (2019): "Measuring Political Personalization of Google News Search", in: Proceedings of the ACM Web Conference. (DOI)] and [37Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] (A/A), [3Lécuyer, Mathias; Spahn, Riley; Spiliopolous, Yannis; Chaintreau, Augustin; Geambasu, Roxana; Hsu, Daniel J. (2015): "Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] and [2Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] and [38Vombatkere, Karan; Mousavi, Sepehr; Zannettou, Savvas; Roesner, Franziska; Gummadi, Krishna P. (2024): "TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds", in: Proceedings of the ACM Web Conference. (DOI)] (generated), [11Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] (repeats), [7Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [15Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)] (both).

Two things this does not say. It is a judgement per paper, not a mechanical count, so a second reader would move one or two of them — the definition and every evidence sentence are on algorithm_audits so you can disagree with a specific verdict rather than with the total. And a paper with no null is not thereby wrong: several of the 24 report effects far larger than any plausible noise floor. What they cannot do is tell the reader that.

And a third of them run no inferential test

11 of the 34 report no statistic beyond descriptives, by the extraction's statistics.kind field: [1Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [29Meng, Wei; Xing, Xinyu; Sheth, Anmol; Weinsberg, Udi; Lee, Wenke (2014): "Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [9Kliman-Silver, Chloe; Hannak, Aniko; Lazer, David; Wilson, Christo; Mislove, Alan (2015): "Location, Location, Location: The Impact of Geolocation on Web Search Personalization", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [17Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)], [30Kim, I Luk; Wang, Weihang; Kwon, Yonghwi; Zheng, Yunhui; Aafer, Yousra; Meng, Weijie; Zhang, Xiangyu (2018): "AdBudgetKiller: Online Advertising Budget Draining Attack", in: Proceedings of the ACM Web Conference. (DOI)], [32Venkatadri, Giridhari; Lucherini, Elena; Sapiezynski, Piotr; Mislove, Alan (2019): "Investigating sources of PII used in Facebook’s targeted advertising", in: Proceedings on Privacy Enhancing Technologies. (DOI)], [39Zhang, Jiang; Psounis, Konstantinos; Haroon, Muhammad; Shafiq, Zubair (2022): "HARPO: Learning to Subvert Online Behavioral Advertising", in: Proceedings of the Network and Distributed System Security Symposium. (Link)], [40Medjkoune, Tinhinane; Goga, Oana; Senechal, Juliette (2023): "Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [19Zhang, Jiang; Askari, Hadi; Psounis, Konstantinos; Shafiq, Zubair (2023): "A Utility-Preserving Obfuscation Approach for YouTube Recommendations", Proceedings on Privacy Enhancing Technologies 2023(4). (DOI)], [33Khezresmaeilzadeh, Tina; Zhu, Elaine; Grieco, Kiersten; Dubois, Daniel; Psounis, Konstantinos; Choffnes, David (2025): "Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants", Proceedings on Privacy Enhancing Technologies 2025(2). (DOI)] and [20Karnam, Sai Keerthana; Dash, Abhisek; Das, Antariksh; Mousavi, Sepehr; Bechtold, Stefan; Gummadi, Krishna P.; Mukherjee, Animesh; Weber, Ingmar; Zannettou, Savvas (2026): "Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]. Several of those are attack or system papers where the audit is the evaluation rather than the claim, and [9Kliman-Silver, Chloe; Hannak, Aniko; Lazer, David; Wilson, Christo; Mislove, Alan (2015): "Location, Location, Location: The Impact of Geolocation on Web Search Personalization", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] builds a noise floor without a named test — so this is a count of reported tests, not a verdict on any paper. The methods that are named, paper-counted over the set: a Bonferroni-family correction 7 (Holm–Bonferroni 5, plain Bonferroni 3; one paper names both), linear regression 2, Mann–Whitney U 2, then a long tail of one each (blocked permutation test, Benjamini–Yekutieli, Clopper–Pearson interval, exact test on Pearson's correlation, cross-correlation). The 7 is a hand-folded family count printed by the report script; the automatic skeleton fold leaves Holm-Bonferroni correction (3), Holm-Bonferroni (1) and Holm-Bonferroni method (1) as three separate rows, which is why an unfolded reading of this field would have said 3. Free-text method names are ~20% stable run-to-run, so read the tail as a ranking. See Hypothesis testing.

Term in full text Papers, corpus-wide In the audit set
“persona” (any use) 190 14 of 34
“sock puppet” / “sockpuppet” 26 5 of 34
“control profile/account/persona/browser” 33 11 of 34
“treatment group/profile/persona/condition/arm” 66 5 of 34
“price discrimination” / “price steering” 33 9 of 34

These five vocabularies barely overlap, and none of them is a reliable handle on its own: the “sock puppet” 26 include an observational study of sockpuppets in discussion forums, a study of app-store fraud workers, and interviews with influence-operation participants, none of which is an audit. 1) If you search for one term you will miss most of the field — which is exactly what happened here, and the probes that recovered the rest are listed on the provenance page. Platforms publishes the same 26 against its own, smaller platform-study population, where it is 14; the two numbers are the same query over different denominators.

Methodology and limitations of these figures

  • Populations. empirical = 5,118; crawled = a crawl-configuration record exists or studyTypes includes automated-web-crawl = 1,120; measuredFrom = has a vantage tuple = 3,908. The audit population of 34 is hand-adjudicated against a written rule and is not an extraction field.
  • It is a lower bound, and a biased one. See What this corpus cannot tell you: 48 audit-topical papers in the index are outside the extraction, 45 of them because the selection screen looks for security or privacy and an audit is often neither.
  • Counts are of papers, never tuples. Sentinels (not-stated, none-mentioned) are never counted as answers.
  • Free-text fields are rankings. statistics.method agrees with itself on ~20% of exact strings run-to-run; the method table above is a ranking. Enum-backed rows (statefulness 98%, ethics review 68%) carry percentages.
  • The term sweep is a regex over prose. It gives a lower bound on the practice, not a measurement of it.
  • Seven venues. EuroS&P, ACSAC, RAID, AsiaCCS, CHI, SOUPS, FAccT and WWW's fairness track siblings are absent. FAccT's absence matters more here than on any other page on this site.
  • 2025–2026 are provisional. Do not read the last bucket as a completed period. corpus.
  • Full query log, the 34 verdicts and the 36 rejections, the probes and their residue, quote checks and reviewer findings: algorithm_audits.

What to report

A reviewer who knows this literature will look for these, in roughly this order. The corpus says the field is good at the first item and bad at the rest: 24 of the 26 audits that drive a browser state their profile state, but only 14 of 34 say where they measured from and only 10 of 34 establish a null.

  1. The arms, as a table: how many identities, what each one's treatment was, and what the control arm's treatment was. “We created personas” is not a design.
  2. What was held fixed across arms: egress network and geolocation, device/browser and version, time window, number and order of page loads or interactions, logged-in state.
  3. Training and measurement, separately: how long each identity was trained, how many observations were collected per arm afterwards, and whether measurement reset state or continued to accumulate (Stateful stateless).
  4. The null: the same-treatment baseline, its metric, and its distribution. If you did not run one, say so in the limitations rather than leaving the reader to infer it.
  5. The unit of analysis in the test, in words: “n = 20 profile-runs per condition”, not “n = 41,000 impressions”.
  6. The multiple-comparison correction, if you have more than one condition or more than one outcome (Pvalue corrections).
  7. Repeats and order randomisation, with the seed.
  8. The ethics story: fake accounts on someone else's service is the highest-risk access route on Platforms. Review outcome, what the accounts did, and what you did with them afterwards (Ethics).
  9. Attrition: identities that got banned, rate-limited or challenged, and whether losing them is correlated with the treatment. A treatment arm that gets banned more often is a finding, not a nuisance.

One sentence that does most of it: “Twelve Chrome 154 profiles (six treatment, six control) ran from one residential vantage in Germany, each trained on 200 page loads over four hours in randomised order, then measured on the same 50 publisher pages; six further control–control pairs established a same-treatment baseline, and all tests are over profile-runs (n = 6 per condition) with Holm–Bonferroni across the four outcomes.”

Open Questions

  • Nobody has published the noise floor for the platforms people audit most. [41Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)] did this for tracking measurement across days (Your noise floor); there is no equivalent for “how far apart do two identical YouTube sock puppets land”. It would be a short, highly citable paper and every audit after it would use it.
  • How much does the arms-are-not-independent problem actually cost? No paper in this corpus measures what sharing an egress IP across arms does to an ad-targeting outcome. It is assumed to matter; the size is unknown.
  • The corpus cannot size this method. The screening loss in What this corpus cannot tell you is measurable but not fixable from here. A pass over FAccT, EuroS&P and the IR venues would give the real denominator.
  • LLM-agent auditors are one paper old. [25Sun, Chen; Vekaria, Yash; Nithyanand, Rishab (2026): "On the Suitability of LLM-Driven Agents for Dark Pattern Audits", Proceedings on Privacy Enhancing Technologies 2026(4):927-946. (DOI)] characterises where they fail on dark-pattern audits. Whether an agentic crawler can hold an arm's confounders fixed — the whole game in this design — is untested.
  • Is a donated-data audit ever a substitute for arms? [24Gkiouzepi, Eleni; Andreou, Athanasios; Goga, Oana; Loiseau, Patrick (2023): "Collaborative Ad Transparency: Promises and Limitations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] argues the trade-off explicitly; nobody has run both designs on the same question and compared the answers.
  • Stateful stateless — the profile: what a reset resets, order dependence, and why an audit is a stateful design.
  • Hypothesis testing — the test, and the unit-of-analysis error this design invites.
  • Pvalue corrections — you have more conditions than you think.
  • Platforms — sock puppets as an access route, the ToS and ban risk, and the platforms the field actually measures.
  • Registration — making the accounts an audit needs.
  • Crawling location — the vantage point, which must be identical across arms or it is part of your treatment.
  • Longitudinal — the other noise floor: what moves between waves rather than between arms.
  • Ethics — creating accounts on someone else's service.
  • LLM agents — the 2026 instrument, and what it has been shown to get wrong.
  • Automated measurements — where this sits among crawl, scan and app analysis.
  • corpus — venue scope, the selection funnel that costs this page most, provisional years.
[1]
Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[2]
Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[3]
Lécuyer, Mathias; Spahn, Riley; Spiliopolous, Yannis; Chaintreau, Augustin; Geambasu, Roxana; Hsu, Daniel J. (2015): "Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[4]
Le, Huyen T.; Maragh, Raven; Ekdale, Brian; High, Andrew; Havens, Timothy; Shafiq, Zubair (2019): "Measuring Political Personalization of Google News Search", in: Proceedings of the ACM Web Conference. (DOI)
[5]
Hannak, Aniko; Sapiezynski, Piotr; Kakhki, Arash Molavi; Krishnamurthy, Balachander; Lazer, David; Mislove, Alan; Wilson, Christo (2013): "Measuring personalization of web search", in: Proceedings of the ACM Web Conference. (DOI)
[6]
Imana, Basileal; Korolova, Aleksandra; Heidemann, John S. (2021): "Auditing for Discrimination in Algorithms Delivering Job Ads", in: Proceedings of the ACM Web Conference. (DOI)
[7]
Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[8]
Liu, Zengrui; Iqbal, Umar; Saxena, Nitesh (2024): "Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy?", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[9]
Kliman-Silver, Chloe; Hannak, Aniko; Lazer, David; Wilson, Christo; Mislove, Alan (2015): "Location, Location, Location: The Impact of Geolocation on Web Search Personalization", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[10]
Agarwal, Pushkal; Joglekar, Sagar; Papadopoulos, Panagiotis; Sastry, Nishanth; Kourtellis, Nicolas (2020): "Stop tracking me Bro! Differential Tracking of User Demographics on Hyper-Partisan Websites", in: Proceedings of the ACM Web Conference. (DOI)
[11]
Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[12]
Roongta, Ritik; Jose, Julia; Habib, Hussam; Greenstadt, Rachel (2025): "Sheep's Clothing, Wolfish Intent: Automated Detection and Evaluation of Problematic 'Allowed' Advertisements", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[13]
Kaplan, Levi; Gerzon, Nicole; Mislove, Alan; Sapiezynski, Piotr (2022): "Measurement and analysis of implied identity in ad delivery optimization", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[14]
Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[15]
Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)
[16]
Chandar, Praveen; Thomas, Brian St.; Maystre, Lucas; Pappu, Vijay; Sanchis-Ojeda, Roberto; Wu, Tiffany; Carterette, Ben; Lalmas, Mounia; Jebara, Tony (2022): "Using Survival Models to Estimate User Engagement in Online Experiments", in: Proceedings of the ACM Web Conference. (DOI)
[17]
Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)
[18]
Musa, Maaz Bin; Nithyanand, Rishab (2022): "ATOM: Ad-network Tomography", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[19]
Zhang, Jiang; Askari, Hadi; Psounis, Konstantinos; Shafiq, Zubair (2023): "A Utility-Preserving Obfuscation Approach for YouTube Recommendations", Proceedings on Privacy Enhancing Technologies 2023(4). (DOI)
[20]
Karnam, Sai Keerthana; Dash, Abhisek; Das, Antariksh; Mousavi, Sepehr; Bechtold, Stefan; Gummadi, Krishna P.; Mukherjee, Animesh; Weber, Ingmar; Zannettou, Savvas (2026): "Setting the Course, but Forgetting to Steer: Analyzing Compliance with GDPR's Right of Access to Data by Instagram, TikTok, and Youtube", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[21]
McCrosky, Jesse; Malla, Ranadheer; Tanskanen, Aapo; Camargo, Chico Q. (2026): "Does This Button Work? Investigating YouTube's Ineffective User Controls", in: Proceedings of the ACM Web Conference. (DOI)
[22]
Silva, Márcio; Oliveira, Lucas Santos de; Andreou, Athanasios; Melo, Pedro Olmo Stancioli Vaz de; Goga, Oana; Benevenuto, Fabrício (2020): "Facebook Ads Monitor: An Independent Auditing System for Political Ads on Facebook", in: Proceedings of the ACM Web Conference. (DOI)
[23]
Ali, Muhammad; Goetzen, Angelica; Mislove, Alan; Redmiles, Elissa M.; Sapiezynski, Piotr (2023): "Problematic Advertising and its Disparate Exposure on Facebook", in: Proceedings of the USENIX Security Symposium. (Link)
[24]
Gkiouzepi, Eleni; Andreou, Athanasios; Goga, Oana; Loiseau, Patrick (2023): "Collaborative Ad Transparency: Promises and Limitations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[25]
Sun, Chen; Vekaria, Yash; Nithyanand, Rishab (2026): "On the Suitability of LLM-Driven Agents for Dark Pattern Audits", Proceedings on Privacy Enhancing Technologies 2026(4):927-946. (DOI)
[26]
Chen, Le; Mislove, Alan; Wilson, Christo (2015): "Peeking Beneath the Hood of Uber", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[27]
Becerril-Arreola, Rafael (2023): "A Method to Assess and Explain Disparate Impact in Online Retailing", in: Proceedings of the ACM Web Conference. (DOI)
[28]
Hannak, Aniko; Soeller, Gary; Lazer, David; Mislove, Alan; Wilson, Christo (2014): "Measuring Price Discrimination and Steering on E-commerce Web Sites", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[29]
Meng, Wei; Xing, Xinyu; Sheth, Anmol; Weinsberg, Udi; Lee, Wenke (2014): "Your Online Interests: Pwned! A Pollution Attack Against Targeted Advertising", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[30]
Kim, I Luk; Wang, Weihang; Kwon, Yonghwi; Zheng, Yunhui; Aafer, Yousra; Meng, Weijie; Zhang, Xiangyu (2018): "AdBudgetKiller: Online Advertising Budget Draining Attack", in: Proceedings of the ACM Web Conference. (DOI)
[31]
Khattak, Sheharbano; Fifield, David; Afroz, Sadia; Javed, Mobin; Sundaresan, Srikanth; McCoy, Damon; Paxson, Vern; Murdoch, Steven J. (2016): "Do You See What I See? Differential Treatment of Anonymous Users", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[32]
Venkatadri, Giridhari; Lucherini, Elena; Sapiezynski, Piotr; Mislove, Alan (2019): "Investigating sources of PII used in Facebook’s targeted advertising", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[33]
Khezresmaeilzadeh, Tina; Zhu, Elaine; Grieco, Kiersten; Dubois, Daniel; Psounis, Konstantinos; Choffnes, David (2025): "Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants", Proceedings on Privacy Enhancing Technologies 2025(2). (DOI)
[34]
Zeng, Eric; McAmis, Rachel; Kohno, Tadayoshi; Roesner, Franziska (2022): "What Factors Affect Targeting and Bids in Online Advertising? A Field Measurement Study", in: Proceedings of the ACM Internet Measurement Conference, pp. 210-229. (DOI)
[35]
Robertson, Ronald E.; Lazer, David; Wilson, Christo (2018): "Auditing the Personalization and Composition of Politically-Related Search Engine Results Pages", in: Proceedings of the ACM Web Conference. (DOI)
[36]
Iqbal, Hassan; Khan, Usman Mahmood; Khan, Hassan Ali; Shahzad, Muhammad (2022): "Left or Right: A Peek into the Political Biases in Email Spam Filtering Algorithms During US Election 2020", in: Proceedings of the ACM Web Conference. (DOI)
[37]
Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[38]
Vombatkere, Karan; Mousavi, Sepehr; Zannettou, Savvas; Roesner, Franziska; Gummadi, Krishna P. (2024): "TikTok and the Art of Personalization: Investigating Exploration and Exploitation on Social Media Feeds", in: Proceedings of the ACM Web Conference. (DOI)
[39]
Zhang, Jiang; Psounis, Konstantinos; Haroon, Muhammad; Shafiq, Zubair (2022): "HARPO: Learning to Subvert Online Behavioral Advertising", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[40]
Medjkoune, Tinhinane; Goga, Oana; Senechal, Juliette (2023): "Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[41]
Demir, Nurullah; Große-Kampmann, Matteo; Urban, Tobias; Wressnegger, Christian; Holz, Thorsten; Pohlmann, Norbert (2022): "Reproducibility and Replicability of Web Measurement Studies", in: Proceedings of the ACM Web Conference. (DOI)
1)
Kumar et al., TheWebConf 2017, An Army of Me: Sockpuppets in Online Discussion Communities; Rahman et al., CCS 2019, The Art and Craft of Fraudulent App Promotion in Google Play; Recabarren and Carbunar, USENIX Security 2023, Strategies and Vulnerabilities of Participants in Venezuelan Influence Operations.
You could leave a comment if you were logged in.
design/algorithm_audits.txt · Last modified: by karel.kubicek.claude