User Tools

Site Tools


privacy:advertising

Measuring Ad Targeting and Bidding

You want to know whether an ad was targeted — shown to this identity because of something the platform knew or inferred about it — and why, and what an advertiser paid to put it there. None of that is visible in the ad. What a measurement can observe is a short list of proxies: an ad that follows a visit it could only have followed if data moved; a difference between what two trained identities are shown; a bid or a clearing price; the platform's own explanation; the ads real people were shown, donated after the fact; the delivery statistics of an ad you bought yourself; and the creatives and landing pages a crawler collects. Each one proves a different, narrower thing than “targeted”, and each has its own denominator. This page is about choosing among them.

It sits between five pages that each own a piece of the mechanism, and it points to them rather than repeating them:

  • Ads txt owns the supply chain and header bidding / Prebid as an observable — how to find the auction on a page, read the bids, and why a stateless crawler's bid is a price for nobody. Here, bids are only one of seven signals of targeting.
  • Algorithm audits owns the paired-profile design: control arms, the null, carry-over between arms, how many repeats. Every persona result below inherits its rules.
  • Ad archives owns the platforms' ad repositories and their error rates.
  • Cookie syncing owns identifier exchange between parties; retargeting is how the papers here see it, not how they measure it.
  • Privacy Sandbox owns Topics and Protected Audience, which were built to replace third-party-cookie targeting and are being withdrawn.

Four things to know before you design anything.

  1. Pick the observable by what it can prove. Retargeting proves a data flow; a persona difference proves the platform used the trained attribute; a bid proves someone valued this persona differently; an explanation proves only what the platform chose to say. The table in The observables, and what each can prove is the page in one screen.
  2. The platform's explanation is not ground truth. In 79 campaigns that targeted two or three attributes at once, Facebook's “Why am I seeing this?” never showed more than one [1Andreou, Athanasios; Venkatadri, Giridhari; Goga, Oana; Gummadi, Krishna P.; Loiseau, Patrick; Mislove, Alan (2018): "Investigating Ad Transparency Mechanisms in Social Media: A Case Study of Facebook's Explanations", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)]; and in 2026, 98.9% of YouTube's explanation texts under the EU's Digital Services Act named only the main targeting form [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] — the explanations are coarse before you even ask whether they are true. Validate an explanation against a treatment you controlled, or do not use it as a label.
  3. The denominator is ads, impressions, bids, explanations, campaigns or personas — never “sites”. The same crawl is 24,961,698 impressions or 19,543 unique ads [3Lécuyer, Mathias; Spahn, Riley; Spiliopolous, Yannis; Chaintreau, Augustin; Geambasu, Roxana; Hsu, Daniel J. (2015): "Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]; say which, and say how you decided two impressions were the same ad. Of the 29 papers here that collect ad creatives with a crawler or app driver, only 9 are found saying how they identified a unique ad — a probe-scoped floor (Use in Publications).
  4. The instrument moved under the literature. Third-party cookies stayed in Chrome and the Privacy Sandbox ad APIs are being removed (Privacy Sandbox); RTB clearing prices went encrypted; advertiser-side targeting options narrowed; the DSA made per-ad explanations mandatory in the EU. The dated status of every method is in Which methods are current.

What to Read First

Paper Why
Guha, Cheng and Francis, IMC 2010, Challenges in Measuring Online Advertising Systems [4Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] The methodology paper for everything on this page: ad churn, how many reloads you need before you have seen an ad set, and — still the only measurement of it here — how badly each way of deciding “these two impressions are the same ad” miscounts.
Bashir et al., USENIX Security 2016, Tracing Information Flows Between Ad Exchanges Using Retargeted Ads [5Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)] Retargeting as a controlled stimulus, end to end: train shopping personas, crawl publishers, find the retargets among 31,850 labelled ad images, and read which exchange-to-exchange flow each one implies.
Datta, Tschantz and Datta, PoPETs 2015, Automated Experiments on Ad Privacy Settings [6Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] The randomised persona experiment, and the first statistically rigorous evidence that the platform's ad-settings page omits what it uses (the paper credits an earlier, manual study).
Papadopoulos et al., IMC 2017, If you are not paying for it, you are the product [7Papadopoulos, Panagiotis; Rodríguez, Pablo Rodríguez; Kourtellis, Nicolas; Laoutaris, Nikolaos (2017): "If you are not paying for it, you are the product: how much do advertisers pay to reach you?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] What an advertiser paid, from clearing-price notifications in real users' traffic — and the encrypted-price problem that makes this method historical.
Andreou et al., NDSS 2018, Investigating Ad Transparency Mechanisms in Social Media [1Andreou, Athanasios; Venkatadri, Giridhari; Goga, Oana; Gummadi, Krishna P.; Loiseau, Patrick; Mislove, Alan (2018): "Investigating Ad Transparency Mechanisms in Social Media: A Case Study of Facebook's Explanations", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)] How to test an explanation: buy the campaign whose targeting you know, and compare. Not in the corpus behind this site (its venue-year has no screening labels); read it anyway.
Zeng et al., IMC 2022, What factors affect targeting and bids in online advertising? [8Zeng, Eric; McAmis, Rachel; Kohno, Tadayoshi; Roesner, Franziska (2022): "What Factors Affect Targeting and Bids in Online Advertising? A Field Measurement Study", in: Proceedings of the ACM Internet Measurement Conference, pp. 210-229. (DOI)] The one study in this corpus of real people's ads and bids at once: 286 participants' own browsers, with their demographics, beside the header-bidding price each ad cleared at.
Medjkoune et al., CCS 2023, Marketing to Children Through Online Targeted Advertising [9Medjkoune, Tinhinane; Goga, Oana; Senechal, Juliette (2023): "Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] All three recent tools at once: trained profiles, the platform's explanation, and the authors' own campaigns used to check what the explanation means.
Mai et al., PoPETs 2025, More and Scammier Ads [10Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] A current persona design with explicit teardown, randomised settings and deduplicated creatives, and a finding a student will not expect: turning ad personalisation off made the ads worse.

The observables, and what each can prove

Seventy-one papers in the corpus observe ads or ad-auction outcomes and use them to ask a targeting, price, delivery or content question (papers that observe ads to study fraud, blocking or tracking are not counted); how they were found is in Use in Publications. The papers column counts them (a paper can use several observables).

Observable What it can prove What it cannot Papers, first–latest
Retargeting as a stimulus — visit a product page, then look for its ad elsewhere that a data flow happened between the site you visited and the exchange that served the ad how the data moved (cookie match, server-side, first-party ID); anything about ads you did not stimulate 3, 2016–2022
Personas vs control — trained browser profiles or accounts, differing in one attribute that the platform used the trained attribute, if the difference beats a same-treatment baseline why; and what real users see — a persona is a caricature of one 19, 2010–2026
Header-bidding bids read from the page that some bidder valued this persona differently, including bidders who were never sent the attribute; the rendered winning bid is the price of that impression where the auction is first-price what losing bidders would have paid; anything about server-side auctions — Ads txt 7, 2019–2025
RTB clearing-price notifications in real users' traffic what an advertiser paid, for the cleartext share the encrypted share, which is the more expensive one 3, 2017–2020
The platform's explanation / ad-preference page / advertiser list what the platform says it used what it actually used — found wanting in every check below 14, 2015–2026
Donated ad data — extension panels, platform data exports what real people were shown, with the platform's stated targeting a controlled treatment; the panel is who volunteered 11, 2018–2026
Advertiser side — buy the ad, read reach estimates and delivery how the platform delivers an ad whose targeting you set other advertisers' targeting; exact audience sizes (obfuscated) 10, 2017–2024
Creatives, screenshots and landing pages collected by a crawler or app driver what was served, where, from which vantage targeting, unless paired with personas or explanations 29, 2010–2026
Ad-transparency archives what the platform put in its repository — Ad archives ads it left out 7, 2020–2026

The first four answer whether an ad was targeted; the explanation and donation rows are the only ones that claim to answer why, and both rest on the platform's own account of the targeting. The next sections take them in that order.

Retargeting as a controlled stimulus

Retargeting is the one targeting signal a crawler can plant on purpose: visit a merchant, then see whether that merchant's ad appears on an unrelated publisher, and through which exchange. [5Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)] trained shopping personas on 738 e-commerce sites, crawled 150 publishers, and had the ad images labelled: of 31,850 images, 7,563 were behaviourally targeted and 5,102 were retargets, carrying 35,448 inclusion chains. Because each retarget implies that the serving exchange learned of the visit, the chains become evidence of which exchanges share data — and the result is a measurement of the syncing literature's own instrument: heuristic cookie-matching detection missed 31% of exchange pairs that the retargets proved were sharing. Two later papers use the same stimulus as an attack surface rather than a measurement: [11Kim, I Luk; Wang, Weihang; Kwon, Yonghwi; Zheng, Yunhui; Aafer, Yousra; Meng, Weijie; Zhang, Xiangyu (2018): "AdBudgetKiller: Online Advertising Budget Draining Attack", in: Proceedings of the ACM Web Conference. (DOI)] reverse-engineered the retargeting strategy of 254 of 291 advertisers, and [12Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] made a victim's shopping appear in the attacker's ads, visible roughly 15 hours after the attack — a useful number for how long a retargeting stimulus takes to show up.

What it cannot do: tell you which mechanism carried the visit. A retarget served today may have come through a third-party cookie, a hashed-email match or a server-side conversion API; [5Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)] reads the chain to classify it, and a chain is only visible for what the browser fetched. And it depends on the browser: the stimulus relies on the ad system recognising the same browser on two sites, which Safari and Firefox now block by default for third-party cookies and Chrome does not — see Browser protection. A retargeting result is a result about the browser you ran.

How common it is from the user's side has two measurements here, of different kinds: the platform's own label on a donated panel — 12% of Facebook ads were part of retargeting in [13Andreou, Athanasios; Silva, Márcio; Benevenuto, Fabrício; Goga, Oana; Loiseau, Patrick; Mislove, Alan (2019): "Measuring the Facebook Advertising Ecosystem", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)] (see below) — and self-report: participants in [8Zeng, Eric; McAmis, Rachel; Kohno, Tadayoshi; Roesner, Franziska (2022): "What Factors Affect Targeting and Bids in Online Advertising? A Field Measurement Study", in: Proceedings of the ACM Internet Measurement Conference, pp. 210-229. (DOI)] said they had visited the advertiser before for 18.3% of ads, and those ads cleared about $1.07 CPM higher.

Personas, training time and carry-over

The design — control arms, the same-treatment null, randomised order, the unit of analysis — is Algorithm audits, and nothing here overrides it. What is specific to ads is how long a persona takes to become targetable, and how dirty it gets.

What the platform needed How long, measured Source
Google voice-assistant profiling labels to appear 18.0 ± 4.1 days [14Khezresmaeilzadeh, Tina; Zhu, Elaine; Grieco, Kiersten; Dubois, Daniel; Psounis, Konstantinos; Choffnes, David (2025): "Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants", Proceedings on Privacy Enhancing Technologies 2025(2). (DOI)]
Alexa profiling labels 7.7 ± 1.3 days [14Khezresmaeilzadeh, Tina; Zhu, Elaine; Grieco, Kiersten; Dubois, Daniel; Psounis, Konstantinos; Choffnes, David (2025): "Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants", Proceedings on Privacy Enhancing Technologies 2025(2). (DOI)]
A phone number added for two-factor authentication to become targetable on Facebook 22 days [15Venkatadri, Giridhari; Lucherini, Elena; Sapiezynski, Piotr; Mislove, Alan (2019): "Investigating sources of PII used in Facebook’s targeted advertising", in: Proceedings on Privacy Enhancing Technologies. (DOI)]
An email address (and a phone number) given for login alerts 17 days [15Venkatadri, Giridhari; Lucherini, Elena; Sapiezynski, Piotr; Mislove, Alan (2019): "Investigating sources of PII used in Facebook’s targeted advertising", in: Proceedings on Privacy Enhancing Technologies. (DOI)]
A victim's shopping to appear through an entangled retargeting identity ~15 hours [12Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]
An ad set on one search page to be seen completely ~10 reloads before churn overtakes new ads [4Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]
The number of unique ads a persona is shown to stop growing 40 training sites per interest group [16Musa, Maaz Bin; Nithyanand, Rishab (2022): "ATOM: Ad-network Tomography", in: Proceedings on Privacy Enhancing Technologies. (DOI)]

Budget training in days, not page loads, and measure your own convergence as [16Musa, Maaz Bin; Nithyanand, Rishab (2022): "ATOM: Ad-network Tomography", in: Proceedings on Privacy Enhancing Technologies. (DOI)] did, because none of these transfers across platforms.

Carry-over inside a persona is the default, not an accident. Profiles pick up interests they were not trained on: in [17Barford, Paul; Canadi, Igor; Krushevskaja, Darja; Ma, Qiang; Muthukrishnan, S. (2014): "Adscape: Harvesting and Analyzing Online Display Ads", in: Proceedings of the ACM Web Conference. (DOI)], 60% of profiles gained more than nine new interests after 50 site visits, against an average of eight at the start. Every visit your crawler makes to measure a persona also trains it; either measure on pages whose topic you have modelled, or reset between measurements and accept the training cost again. [10Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] deleted each account's activity history between watch sequences and waited at least 12 hours “to avoid bot detection and for the activity deletion to take effect”.

Bids and clearing prices as a targeting signal

How to read header-bidding bids off a page, and why a clean crawler's median bid (2019) sits two orders of magnitude below real users' median winning bid (2021) — a comparison across two papers that no paper has run side by side — is on Observing header bidding. For a targeting question the bid has one property no other observable has: it reacts to data the page never sent. Personas that opted out under the GDPR and CCPA still drew higher bids than a control — and in the paper's table of bids from advertisers it had not leaked the interests to, all personas but one still bid higher on average, which it reads as server-side sharing [18Liu, Zengrui; Iqbal, Umar; Saxena, Nitesh (2024): "Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy?", in: Proceedings on Privacy Enhancing Technologies. (DOI)]; smart-speaker interaction raised web bids to as much as 30× the vanilla persona's mean [19Iqbal, Umar; Bahrami, Pouneh Nikkhah; Trimananda, Rahmadi; Cui, Hao; Gamero-Garrido, Alexander; Dubois, Daniel J.; Choffnes, David R.; Markopoulou, Athina; Roesner, Franziska; Shafiq, Zubair (2023): "Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]; changing only the browser fingerprint changed the trend, mean, median and maximum of the bids [20Liu, Zengrui; Dani, Jimmy; Cao, Yinzhi; Wu, Shujiang; Saxena, Nitesh (2025): "The First Early Evidence of the Use of Browser Fingerprinting for Online Tracking", in: Proceedings of the ACM Web Conference. (DOI)]. A bid is therefore the cheapest evidence of server-side data use you can collect from a client.

What it does not prove is that an ad was targeted: a higher bid is a valuation. Two cautions that come from the bid papers themselves. Zero bids are a large share of the data — 22% of all bids in [21Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)], which first says its data cannot tell misconfiguration from intent and then leans towards intentional underbidding for several large bidders — either way, a mean over non-zero bids is a different statistic. And the treatment effect is confounded by placement and the calendar: the ad slot [19Iqbal, Umar; Bahrami, Pouneh Nikkhah; Trimananda, Rahmadi; Cui, Hao; Gamero-Garrido, Alexander; Dubois, Daniel J.; Choffnes, David R.; Markopoulou, Athina; Roesner, Franziska; Shafiq, Zubair (2023): "Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], day of week [19Iqbal, Umar; Bahrami, Pouneh Nikkhah; Trimananda, Rahmadi; Cui, Hao; Gamero-Garrido, Alexander; Dubois, Daniel J.; Choffnes, David R.; Markopoulou, Athina; Roesner, Franziska; Shafiq, Zubair (2023): "Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and the pre-Christmas surge [8Zeng, Eric; McAmis, Rachel; Kohno, Tadayoshi; Roesner, Franziska (2022): "What Factors Affect Targeting and Bids in Online Advertising? A Field Measurement Study", in: Proceedings of the ACM Internet Measurement Conference, pp. 210-229. (DOI)] all move bids; a control persona run on the same day from the same network, on the same slots, is the minimum.

All bids are a valuation; the rendered winning bid is a price. Where the final auction is first-price — Google Ad Manager moved to unified first-price auctions in 2019 1) — the winning header bid that rendered is what the impression cleared at, which is how [8Zeng, Eric; McAmis, Rachel; Kohno, Tadayoshi; Roesner, Franziska (2022): "What Factors Affect Targeting and Bids in Online Advertising? A Field Measurement Study", in: Proceedings of the ACM Internet Measurement Conference, pp. 210-229. (DOI)] uses its 7,117 rendered winners (Observing header bidding). For platform ads, the ad archives publish spend as coarse ranges for the ads they cover (Ad archives). Outside client-side header bidding, the clearing price is mostly hidden. The one line of work in this corpus that read it — exchanges' win-notification URLs in real users' traffic — found that encrypted prices ran ~1.7× higher than cleartext ones and that the median user cost advertisers ~25 CPM over a year [7Papadopoulos, Panagiotis; Rodríguez, Pablo Rodríguez; Kourtellis, Nicolas; Laoutaris, Nikolaos (2017): "If you are not paying for it, you are the product: how much do advertisers pay to reach you?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]; a companion study put the median advertiser at 0.00071 € per delivered ad [22Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2018): "The Cost of Digital Advertisement: Comparing User and Advertiser Views", in: Proceedings of the ACM Web Conference. (DOI)], and [23Agarwal, Pushkal; Joglekar, Sagar; Papadopoulos, Panagiotis; Sastry, Nishanth; Kourtellis, Nicolas (2020): "Stop tracking me Bro! Differential Tracking of User Demographics on Hyper-Partisan Websites", in: Proceedings of the ACM Web Conference. (DOI)] used the same observable to price the ads its personas were shown on right- vs left-leaning hyper-partisan sites at a median $0.667 vs $0.561 CPM. The method needs cleartext price macros in traffic you are allowed to see, and the encrypted share is not a random sample of it — the paper had to model the encrypted prices. Its precursor, Olejnik et al.'s Selling Off Privacy at Auction [24Olejnik, Lukasz; Tran, Minh-Dung; Castelluccia, Claude (2014): "Selling Off Privacy at Auction", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)], is not in the corpus at all (no abstract in the index, so it was never screened). No paper in this page's population has used win-notification prices since 2020.

Explanations and donated data as ground truth, and how unreliable the explanations are

The platform's “Why am I seeing this ad?”, its ad-preference page and its list of advertisers who uploaded your data are the only observables that claim to state the targeting. They are the platform's account, and every controlled check here — three of them — found the explanation under-reporting, while the self-report checks found the inferred interests wrong from the user's side:

Explanation surface Checked against Result Source
Facebook ad explanation 79 own campaigns targeting two or three attributes only one attribute was ever shown [1Andreou, Athanasios; Venkatadri, Giridhari; Goga, Oana; Gummadi, Krishna P.; Loiseau, Patrick; Mislove, Alan (2018): "Investigating Ad Transparency Mechanisms in Social Media: A Case Study of Facebook's Explanations", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)]
Google Ad Settings a persona that visited substance-abuse sites visiting them “changed the ads shown but not the settings page”: the experimental group saw a rehab advertiser's ads 3,309 times, the control never [6Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)]
YouTube explanations, children's content 3,221 ads collected by six trained profiles, with seven own campaigns run to establish what a placement explanation means 76% of ads on one children's list carried a non-personalised explanation; 25% of ads on children-focused videos had a profiling explanation [9Medjkoune, Tinhinane; Goga, Oana; Senechal, Juliette (2023): "Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]
YouTube explanations under the DSA the explanation texts themselves, from real users' notices collected by the Who Targets Me extension 98.9% cite only the main targeting form — a granularity measurement, not a check against known targeting [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)]
Ad-preference interest lists (Facebook, Google, eXelate; BlueKai was collected but not rated) the participant's own judgement participants were strongly interested in only 27% of listed interests [25Bashir, Muhammad Ahmad; Farooq, Umar; Shahid, Maryam; Zaffar, Muhammad Fareed; Wilson, Christo (2019): "Quantity vs. Quality: Evaluating User Interest Profiles Using Ad Preference Managers", in: Proceedings of the Network and Distributed System Security Symposium. (Link)]
Interests assigned by Google, Facebook, Twitter, LinkedIn the participant's own judgement fewer than half relevant on every platform [26Caravaca, Francisco; González-Cabañas, José; Cuevas, Ángel; Cuevas, Rubén (2024): "Overprofiling Analysis on Major Internet Players", in: Proceedings on Privacy Enhancing Technologies. (DOI)]
Google Ad Settings interests participants assigned to their own activity only 21% of participant-assigned interests appeared in Ad Settings [27Farke, Florian M.; Balash, David G.; Golla, Maximilian; Aviv, Adam J. (2024): "How Does Connecting Online Activities to Advertising Inferences Impact Privacy Perceptions?", in: Proceedings on Privacy Enhancing Technologies. (DOI)]

Read the direction of each check. The first three are controlled — the researchers knew the targeting because they set it or trained it — and they show explanations under-report. The fourth measures how specific the explanations are, not whether they are true. The last three are self-report — they show inferred interests are often wrong from the user's point of view, which is a different claim. Neither lets you treat an explanation as a label. If an explanation is your outcome, validate it on a campaign you bought, as [1Andreou, Athanasios; Venkatadri, Giridhari; Goga, Oana; Gummadi, Krishna P.; Loiseau, Patrick; Mislove, Alan (2018): "Investigating Ad Transparency Mechanisms in Social Media: A Case Study of Facebook's Explanations", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)] and [9Medjkoune, Tinhinane; Goga, Oana; Senechal, Juliette (2023): "Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] did.

Donated data gets closer to real users: participants install an extension that records the ads and explanations they see ([13Andreou, Athanasios; Silva, Márcio; Benevenuto, Fabrício; Goga, Oana; Loiseau, Patrick; Mislove, Alan (2019): "Measuring the Facebook Advertising Ecosystem", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)], 622 users; [28Ali, Muhammad; Goetzen, Angelica; Mislove, Alan; Redmiles, Elissa M.; Sapiezynski, Piotr (2023): "Problematic Advertising and its Disparate Exposure on Facebook", in: Proceedings of the USENIX Security Symposium. (Link)]; [29Gkiouzepi, Eleni; Andreou, Athanasios; Goga, Oana; Loiseau, Patrick (2023): "Collaborative Ad Transparency: Promises and Limitations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]), or upload the platform's own data export — Twitter's contained 240,651 targeted ads and 30 targeting types for 231 participants [30Wei, Miranda; Stamos, Madison; Veys, Sophie; Reitinger, Nathan; Goodman, Justin; Herman, Margot; Filipczuk, Dorota; Weinshel, Ben; Mazurek, Michelle L.; Ur, Blase (2020): "What Twitter Knows: Characterizing Ad Targeting Practices, User Perceptions, and Ad Explanations Through Users' Own Twitter Data", in: Proceedings of the USENIX Security Symposium. (Link)]. From a panel like this, [13Andreou, Athanasios; Silva, Márcio; Benevenuto, Fabrício; Goga, Oana; Loiseau, Patrick; Mislove, Alan (2019): "Measuring the Facebook Advertising Ecosystem", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)] measured that 17% of Facebook ads used lookalike audiences and 12% retargeting — figures no crawler could produce. The cost is the loss of the treatment, and the panel: whoever installs a transparency extension is a sample of people who install transparency extensions. [29Gkiouzepi, Eleni; Andreou, Athanasios; Goga, Oana; Loiseau, Patrick (2023): "Collaborative Ad Transparency: Promises and Limitations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] tried to infer an advertiser's targeting formula back from a panel and got it right in 65% of experiments only where ten or more monitored users received the ad. Data exports as an instrument in their own right — including what they omit — are Data subject rights.

A per-impression explanation is not the same instrument as an archive. The DSA made both mandatory for very large platforms in the EU; Four Different Instruments Get Called "Ad Transparency" separates them, and the repositories' own error rates are there.

Buying the ad

The advertiser-side route inverts the problem: instead of guessing another advertiser's targeting from what you were shown, set the targeting yourself and read what the platform does with it. [31Kaplan, Levi; Gerzon, Nicole; Mislove, Alan; Sapiezynski, Piotr (2022): "Measurement and analysis of implied identity in ad delivery optimization", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] ran paired campaigns that differed only in the implied identity of the person in the image and found that images of Black people were delivered to audiences that were 73.8% Black against 56.3% for images of white people — a delivery skew no advertiser asked for. [32González-Cabañas, José; Cuevas, Ángel; Cuevas, Rubén; López-Fernández, Juan; García, David (2021): "Unique on Facebook: Formulation and Evidence of (Nano)targeting Individual Users with non-PII Data", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] used the same route to show that an individual can be targeted alone: 9 of 21 campaigns built from one user's interests reached exactly that user. [33El fraihi, Asmaa; Amieur, Nardjes; Rudametkin, Walter; Goga, Oana (2024): "Client-side and Server-side Tracking on Meta: Effectiveness and Accuracy", Proceedings on Privacy Enhancing Technologies 2024(3):431-445. (DOI)] used campaign reach as the oracle for a tracking question: Meta's server-side Conversions API matched 34% to 51% of website visitors to profiles. Two constraints set the design. The platform's audience-size estimates are deliberately noised — uniform noise between 0 and 20 before rounding, when [15Venkatadri, Giridhari; Lucherini, Elena; Sapiezynski, Piotr; Mislove, Alan (2019): "Investigating sources of PII used in Facebook’s targeted advertising", in: Proceedings on Privacy Enhancing Technologies. (DOI)] measured it — so a reach estimate is a bounded measurement, not a count; and the targeting options have been narrowed (see Which methods are current). Buying ads costs money, puts content in front of real people, and binds you to the platform's advertiser terms: Ethics.

Ad-content collection from a crawl, and its dedup

Most papers that collect ads do so to ask what the ads are — malicious (about 1% of 673,596 ads collected in 2014 [34Zarras, Apostolis; Kapravelos, Alexandros; Stringhini, Gianluca; Holz, Thorsten; Kruegel, Christopher; Vigna, Giovanni (2014): "The Dark Alleys of Madison Avenue: Understanding Malicious Advertisements", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]), political, predatory, inappropriate for children — rather than whom they target: 20 of the 29 crawl-collecting papers ask only that. The collection step is shared, and so are its three decisions.

Finding the ad. Filter lists are the usual detector — 14 of the 29 mention EasyList — and they miss: the crawler in [35Moti, Zahra; Senol, Asuman; Bostani, Hamid; Zuiderveen Borgesius, Frederik J.; Moonsamy, Veelasha; Mathur, Arunesh; Acar, Gunes (2024): "Targeted and Troublesome: Tracking and Advertising on Children's Websites", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] detected ads correctly in 85% of a hand-checked sample, captured a non-ad in 7.5% and a blank frame in 7.5%. Report the equivalent for yours. The mechanics of the lists are Filter lists.

Deciding when two impressions are the same ad. This is the step that sets your denominator, and the only paper here that measured its error is from 2010: on two sets of search text ads (fashion queries, and dress queries only), identifying an ad by its redirect URL, its display URL, title plus display URL, or title plus summary produced false-negative rates from 10% to 69% and false positives from 0% to 12% — and a false negative “has the effect of over-counting the number of unique ads” [4Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. The methods used since, where papers say:

Identity rule Used by
perceptual / average image hash on the creative or a screenshot [36Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] (dHash, then clustering), [37Yeung, Christina; Iqbal, Umar; O'Neil, Yekaterina Tsipenyuk; Kohno, Tadayoshi; Roesner, Franziska (2023): "Online Advertising in Ukraine and Russia During the 2022 Russian Invasion", in: Proceedings of the ACM Web Conference. (DOI)], [38Yeung, Christina; Kohno, Tadayoshi; Roesner, Franziska (2024): "Analyzing the (In)Accessibility of Online Advertisements", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] (plus the accessibility tree), [39Roongta, Ritik; Jose, Julia; Habib, Hussam; Greenstadt, Rachel (2025): "Sheep's Clothing, Wolfish Intent: Automated Detection and Evaluation of Problematic 'Allowed' Advertisements", in: Proceedings on Privacy Enhancing Technologies. (DOI)] (perceptual hashes grouped with Faiss)
visual similarity of landing pages rather than their bytes, then entity matching [40Chen, Gong; Meng, Wei; Copeland, John A. (2019): "Revisiting Mobile Advertising Threats with MAdLife", in: Proceedings of the ACM Web Conference. (DOI)]
MinHash locality-sensitive hashing over the ad's text (Jaccard similarity > 0.5), within groups sharing a landing-page domain [41Zeng, Eric; Wei, Miranda; Gregersen, Theo; Kohno, Tadayoshi; Roesner, Franziska (2021): "Polls, Clickbait, and Commemorative \$2 Bills: Problematic Political Advertising on News and Media Websites Around the 2020 U.S. Elections", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]
URL with and without its parameters [42Bashir, Muhammad Ahmad; Arshad, Sajjad; Wilson, Christo (2016): "Recommended For You: A First Look at Content Recommendation Networks", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] — 94% of ad URLs appeared on a single publisher, 85% after stripping parameters
the platform's own ID (video ID) [10Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)]

The ratio between the two units is large and varies with the platform, which is why it has to be reported:

Collection Impressions / ads collected Unique ads Source
Gmail ads, 2014 24,961,698 impressions 19,543 [3Lécuyer, Mathias; Spahn, Riley; Spiliopolous, Yannis; Chaintreau, Augustin; Geambasu, Roxana; Hsu, Daniel J. (2015): "Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)]
Display ads, 180 sites, 340 personas 875,209 impressions 175,495 distinct [17Barford, Paul; Canadi, Igor; Krushevskaja, Darja; Ma, Qiang; Muthukrishnan, S. (2014): "Adscape: Harvesting and Analyzing Online Display Ads", in: Proceedings of the ACM Web Conference. (DOI)]
News and media sites, six US cities, 2020–2021 1,402,245 ads 169,751 [41Zeng, Eric; Wei, Miranda; Gregersen, Theo; Kohno, Tadayoshi; Roesner, Franziska (2021): "Polls, Clickbait, and Commemorative \$2 Bills: Problematic Political Advertising on News and Media Websites Around the 2020 U.S. Elections", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]
YouTube pre-roll, 2024–2025 38,851 impressions 10,628 [10Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)]
Conflict-related ads, Ukraine / Russia / US, 2022 4,319 appearances 1,197 [37Yeung, Christina; Iqbal, Umar; O'Neil, Yekaterina Tsipenyuk; Kohno, Tadayoshi; Roesner, Franziska (2023): "Online Advertising in Ukraine and Russia During the 2022 Russian Invasion", in: Proceedings of the ACM Web Conference. (DOI)]

A share computed over impressions weights a creative by how often it was served; a share over unique ads weights every creative once. [10Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] reports its predatory-ad rate over impressions and gives both counts; do the same.

Where the crawler stands. Ads differ by vantage, and ad networks treat automation and datacentre addresses differently. [36Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] found that “many ads, including SEACMA ads, were tailored to the kind of operating system or browser being used”, and that two ad networks answered requests from the institution, Tor exit nodes and Amazon AWS ranges differently from residential ones — they “seemed to never serve SEACMA ads” there. A crawler on cloud IPs measures the ads served to crawlers on cloud IPs. See Crawling location and Crawler detection.

Classifying the content. Hand coding dominates, and LLMs arrive in 2024: seven of the 71 papers name an LLM, all 2024–2026 — as a second coder beside humans ([43Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], GPT-3.5), as a zero-shot classifier validated against experts ([39Roongta, Ritik; Jose, Julia; Habib, Hussam; Greenstadt, Rachel (2025): "Sheep's Clothing, Wolfish Intent: Automated Detection and Evaluation of Problematic 'Allowed' Advertisements", in: Proceedings on Privacy Enhancing Technologies. (DOI)]: GPT-4o-mini “attains an inter-annotator agreement of 0.74 across all categories”), as an extractor of audio ads from speech transcripts ([44Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], GPT-4o over Whisper output), and as the thing whose inference is the finding. The last case carries a warning: [45Chen, Baiyu; Tag, Benjamin; Xue, Hao; Angus, Daniel; Salim, Flora (2026): "When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs", in: Proceedings of the ACM Web Conference. (DOI)] infers users' attributes from the ads they were shown, and at the session level its Gemini model reached 59.13% accuracy on gender against 58.80% for a random-control baseline in the same table; per user over time the plain model reached 74.88% and a variant given extra cultural context 76.38%, against 70.39% for the random control. The baseline row is what makes the number readable — publish yours.

Pick the unit before you pick the method

The same word — “ads” — is at least seven different units in this literature, and none of them is a site.

Unit A published figure in that unit Its denominator, stated
Impression 184,218 ad impressions from 6,347 advertisers [46Ballard, Cameron; Goldstein, Ian; Mehta, Pulak; Smothers, Genesis; Take, Kejsi; Zhong, Victoria; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2022): "Conspiracy Brokers: Understanding the Monetization of YouTube Conspiracy Theories", in: Proceedings of the ACM Web Conference. (DOI)] YouTube videos from conspiracy and mainstream recommendation seeds
Unique ad (creative) 5,102 retargets [5Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)] 31,850 labelled ad images after four filters
Page ads on 36% of pages, targeted ads on 27% [35Moti, Zahra; Senol, Asuman; Bostani, Hamid; Zuiderveen Borgesius, Frederik J.; Moonsamy, Veelasha; Mathur, Arunesh; Acar, Gunes (2024): "Targeted and Troublesome: Tracking and Advertising on Children's Websites", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] pages crawled on 2,004 child-directed sites, seven crawls
Explanation 98.9% cite only the main form [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] YouTube explanation texts among 48,511 notices on four platforms
Bid 22% are zero [21Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)] all bids received by the paper's personas
Campaign 9 of 21 nanotargeting campaigns reached the one intended user [32González-Cabañas, José; Cuevas, Ángel; Cuevas, Rubén; López-Fernández, Juan; García, David (2021): "Unique on Facebook: Formulation and Evidence of (Nano)targeting Individual Users with non-PII Data", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] the authors' own Facebook campaigns
Persona / arm-run 42 cloned-voice puppets [44Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] the treatment arms; the n of the test — see How many arms, how many repeats
Participant 231 participants' ad files [30Wei, Miranda; Stamos, Madison; Veys, Sophie; Reitinger, Nathan; Goodman, Justin; Herman, Margot; Filipczuk, Dorota; Weinshel, Ben; Mazurek, Michelle L.; Ur, Blase (2020): "What Twitter Knows: Characterizing Ad Targeting Practices, User Perceptions, and Ad Explanations Through Users' Own Twitter Data", in: Proceedings of the USENIX Security Symposium. (Link)] people who uploaded their data and completed the study

A site count appears in almost every one of these papers as the sampling frame, not as the unit of the result. In [35Moti, Zahra; Senol, Asuman; Bostani, Hamid; Zuiderveen Borgesius, Frederik J.; Moonsamy, Veelasha; Mathur, Arunesh; Acar, Gunes (2024): "Targeted and Troublesome: Tracking and Advertising on Children's Websites", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], 36% of pages carried an ad and 73% of ads with a disclosure page (40,281 of them) were targeted: both correct, and not on one axis.

Which methods are current

Dated 2026-09-23. A ranking of what the literature did is a fact about the literature; a status below is an argument from the named papers, the year counts in Use in Publications and the platform changes cited. The 2025–2026 slice is provisional (CCS and IMC 2026 not held; IEEE S&P and TheWebConf 2026 incompletely indexed), and it is also where the newest methods appear, so a “current” label resting on it rests on the thinnest years in the corpus. Year counts per method are over n ≤ 29 papers and carry no trend.

Method Papers per bucket: 2010–2013 / 2014–2017 / 2018–2021 / 2022–2024 / 2025–2026* Status in 2026 Why
Trained personas vs a control 1 / 6 / 2 / 5 / 5 current, the workhorse Every platform change below changes what a persona sees, not whether the design works. Design rules: Algorithm audits
Crawl-collected creatives and landing pages 3 / 6 / 6 / 7 / 7 current The collection step is stable; what changed is classification (LLMs, 2024–) and that dedup is still rarely reported
Header-bidding bids as the dependent variable 0 / 0 / 2 / 4 / 1 current, with a shrinking window Sees only client-side auctions; the client/server split has not been re-measured since 2019 (Ads txt). Protected Audience, the on-device successor, is being withdrawn (Privacy Sandbox)
Platform explanations and ad-preference pages 0 / 1 / 6 / 3 / 4 current, never as ground truth The DSA requires every online platform to give, with each ad, “meaningful information … about the main parameters used to determine the recipient” — very large platforms from August 2023, all from 17 February 2024 2). Every check on this page found explanations incomplete. The interfaces also moved: Google rolled out My Ad Center for signed-in users on 2022-10-20, keeping Ad Settings for signed-out ones 3), so a 2015–2022 Ad Settings result describes an interface a signed-in persona no longer sees
Donated ad data (extension panels, data exports) 0 / 0 / 5 / 4 / 2 current, at the platform's tolerance The only route to real people's ads. On 2021-08-04 Meta disabled the accounts and platform access of NYU's Ad Observatory project 4); its Ad Observer extension is still distributed, and the Australian Ad Observatory describes itself as active 5). For the method in general see [47Angus, Daniel; Obeid, Abdul Karim; Burgess, Jean; Parker, Christine; Andrejevic, Mark; Carah, Nicholas; Tan, Xue Ying (2024): "Enabling Online Advertising Transparency through Data Donation Methods", Computational Communication Research 6(2). (DOI)]
Advertiser-side campaigns, reach estimates 0 / 1 / 5 / 4 / 0 current, narrowed Reach estimates are deliberately noised (uniform noise 0–20, then rounding, in 2019 [15Venkatadri, Giridhari; Lucherini, Elena; Sapiezynski, Piotr; Mislove, Alan (2019): "Investigating sources of PII used in Facebook’s targeted advertising", in: Proceedings on Privacy Enhancing Technologies. (DOI)]). From 2022-01-19 Meta removed detailed-targeting options referencing health, race or ethnicity, political affiliation, religion or sexual orientation 6) — the category of option [48González-Cabañas, José; Cuevas, Ángel; Cuevas, Rubén (2018): "Unveiling and Quantifying Facebook Exploitation of Sensitive Personal Data for Advertising Purposes", in: Proceedings of the USENIX Security Symposium. (Link)] showed in use — and in the EU the DSA forbids profiling-based ads on special categories of data and to known minors (Articles 26(3) and 28(2)). None of the 2025–2026 papers uses this route
Retargeting as a stimulus 0 / 1 / 1 / 1 / 0 current; browser-dependent, and unmeasured outside Chromium Needs the ad system to recognise one browser on two sites. The cookie route is closed by default outside Chrome; the hashed-email and server-side routes are not, and no paper here has run the stimulus across browsers. Chrome kept third-party cookies (2025-04-22); Safari has blocked them by default since 2020 and Firefox has partitioned them by default since 2022 7); details on Browser protection
RTB win-notification clearing prices 0 / 1 / 2 / 0 / 0 historical Needs cleartext price macros in real users' traffic; the encrypted share was already the expensive share in 2015 [7Papadopoulos, Panagiotis; Rodríguez, Pablo Rodríguez; Kourtellis, Nicolas; Laoutaris, Nikolaos (2017): "If you are not paying for it, you are the product: how much do advertisers pay to reach you?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], and the largest exchange documents its winning-price macro as encrypted for the buyer 8). No use since 2020
LLM classification of ad content or explanations 0 / 0 / 0 / 1 / 6 new (2024–) Seven papers name an LLM. Two validate it against human coders ([43Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [39Roongta, Ritik; Jose, Julia; Habib, Hussam; Greenstadt, Rachel (2025): "Sheep's Clothing, Wolfish Intent: Automated Detection and Evaluation of Problematic 'Allowed' Advertisements", in: Proceedings on Privacy Enhancing Technologies. (DOI)]), one uses it to extract ads from transcripts ([44Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]), one makes its inference the finding, against a random control ([45Chen, Baiyu; Tag, Benjamin; Xue, Hao; Angus, Daniel; Salim, Flora (2026): "When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs", in: Proceedings of the ACM Web Conference. (DOI)]). Treat it as a coder that needs an agreement statistic (Interrater agreement)
Privacy Sandbox auctions and Topics as the targeting observable — superseded before it was current Privacy Sandbox

What is not superseded, despite its age. Two 2014–2015 statistical designs for turning ad observations into a targeting claim — differential correlation across shadow accounts [49Lécuyer, Mathias; Ducoffe, Guillaume; Lan, Francis; Papancea, Andrei; Petsios, Theofilos; Spahn, Riley; Chaintreau, Augustin; Geambasu, Roxana (2014): "XRay: Enhancing the Web’s Transparency with Differential Correlation", in: Proceedings of the USENIX Security Symposium. (Link)] and its successor with statistical confidence [3Lécuyer, Mathias; Spahn, Riley; Spiliopolous, Yannis; Chaintreau, Augustin; Geambasu, Roxana; Hsu, Daniel J. (2015): "Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)] — have not been replaced in this corpus, and the randomised persona experiment of [6Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] is still the citation for significance (Which methods are current). The 2010 dedup-error table [4Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] has not been redone in this corpus.

What the 2025–2026 slice adds is not a new observable but new platforms — voice assistants ([44Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], [14Khezresmaeilzadeh, Tina; Zhu, Elaine; Grieco, Kiersten; Dubois, Daniel; Psounis, Konstantinos; Choffnes, David (2025): "Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants", Proceedings on Privacy Enhancing Technologies 2025(2). (DOI)], [50Sabir, Aafaq; B., Abhinaya S.; Ahmed, Dilawer; Das, Anupam (2025): "Analyzing Ad Prevalence, Characteristics, and Compliance in Alexa Skills", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]), app stores [51Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)], mini-games — and LLMs as coders. Nothing in it replaces personas, explanations or crawled creatives.

Crawl configuration that decides what you see

  • Consent and jurisdiction. Since TCF v2.2 (2023-05-16), consent is the only legal basis a TCF vendor can declare for building an ad profile and for selecting personalised ads 9), so a “no-interaction” crawl from a European vantage measures the non-personalised fallback of the vendors that honour the signal. State the consent action and the vantage together; Consent and TCF consent strings cover how. Among the 41 papers here with a crawl configuration, 16 do not state a consent action and 15 state that they did not interact.
  • Logged in or not, and which ad option the account chose. Platform ads (YouTube, Facebook, app stores) are targeted on the account; web display ads on cookies and identifiers. Say which identity you trained. An EU Facebook or Instagram account has chosen between personalised ads, a paid no-ads subscription (from November 2023) and, since 2024-11-12, “less personalized ads” that use context plus “a minimal set of data points including a person's age, location, gender, and how a person engages with ads” 10). The choice changed again: after the Commission found Meta in breach of the Digital Markets Act, Meta undertook to present new options — fully personalised ads, or sharing less data for “more limited personalised advertising” — to EU users in January 2026 11). Whichever option the account is on is a treatment variable; record it and its date. Only 6 of those 41 crawling papers created accounts or logged in; of the 36 papers here that measure a platform service (other than the open web), 18 have no crawl configuration at all.
  • Vantage and IP class. See the SEACMA observation above; datacentre and Tor addresses are served differently. Of 67 papers with a vantage record, 37 state where they measured from.
  • Automation fingerprint. Ad networks run anti-bot checks. [36Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] wrote its own DevTools client “to avoid anti-bot checks implemented by some of the ad networks”, and [35Moti, Zahra; Senol, Asuman; Bostani, Hamid; Zuiderveen Borgesius, Frederik J.; Moonsamy, Veelasha; Mathur, Arunesh; Acar, Gunes (2024): "Targeted and Troublesome: Tracking and Advertising on Children's Websites", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] used Tracker Radar Collector's anti-bot measures. Crawler detection.
  • Blockers. A filter list that detects ads also blocks them, and a browser with tracking protection defeats the retargeting stimulus. [39Roongta, Ritik; Jose, Julia; Habib, Hussam; Greenstadt, Rachel (2025): "Sheep's Clothing, Wolfish Intent: Automated Detection and Evaluation of Problematic 'Allowed' Advertisements", in: Proceedings on Privacy Enhancing Technologies. (DOI)] makes the blocker the treatment; if it is not your treatment, make sure it is off and say so (Browser protection).
  • Clicking. A click costs an advertiser money and can trigger fraud defences. [42Bashir, Muhammad Ahmad; Arshad, Sajjad; Wilson, Christo (2016): "Recommended For You: A First Look at Content Recommendation Networks", in: Proceedings of the ACM Internet Measurement Conference. (DOI)], for one, does not click advertiser URLs and visits the extracted link separately; if you click, say how often and why — Ethics.

Tools, datasets and services that already exist

Liveness checked on 2026-09-23; the check is on External sources. A dead link in a paper is not a dead method.

What From Use it for
AdFisher (dattaamit/info-flow-experiments, renamed from tadatitam/…; last pushed 2021) [6Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] randomised persona experiments with a permutation test
Sunlight and XRay [3Lécuyer, Mathias; Spahn, Riley; Spiliopolous, Yannis; Chaintreau, Augustin; Geambasu, Roxana; Hsu, Daniel J. (2015): "Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)], [49Lécuyer, Mathias; Ducoffe, Guillaume; Lan, Francis; Papancea, Andrei; Petsios, Theofilos; Spahn, Riley; Chaintreau, Augustin; Geambasu, Roxana (2014): "XRay: Enhancing the Web’s Transparency with Differential Correlation", in: Proceedings of the USENIX Security Symposium. (Link)] the statistics of turning ad observations into targeting hypotheses
UW ad scraper (UWCSESecurityLab/adscraper), and the group's ad-data pages used by [38Yeung, Christina; Kohno, Tadayoshi; Roesner, Franziska (2024): "Analyzing the (In)Accessibility of Online Advertisements", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]; data for [41Zeng, Eric; Wei, Miranda; Gregersen, Theo; Kohno, Tadayoshi; Roesner, Franziska (2021): "Polls, Clickbait, and Commemorative \$2 Bills: Problematic Political Advertising on News and Media Websites Around the 2020 U.S. Elections", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [37Yeung, Christina; Iqbal, Umar; O'Neil, Yekaterina Tsipenyuk; Kohno, Tadayoshi; Roesner, Franziska (2023): "Online Advertising in Ukraine and Russia During the 2022 Russian Invasion", in: Proceedings of the ACM Web Conference. (DOI)] crawl-collected creatives with screenshots and landing pages
YouTube ad-settings experiment code and labelled ads (CybersecurityForDemocracy/youtube-ad-settings) [10Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)] a current persona pipeline with teardown
Acceptable Ads crawl and classifier (Racro/AcceptableAds_PETS) [39Roongta, Ritik; Jose, Julia; Habib, Hussam; Greenstadt, Rachel (2025): "Sheep's Clothing, Wolfish Intent: Automated Detection and Evaluation of Problematic 'Allowed' Advertisements", in: Proceedings on Privacy Enhancing Technologies. (DOI)] LLM classification of ad content, with the validation
Alexa Echo ad-targeting code and data [19Iqbal, Umar; Bahrami, Pouneh Nikkhah; Trimananda, Rahmadi; Cui, Hao; Gamero-Garrido, Alexander; Dubois, Daniel J.; Choffnes, David R.; Markopoulou, Athina; Roesner, Franziska; Shafiq, Zubair (2023): "Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] bids as a dependent variable on OpenWPM
App-store ad tools (seemoo-lab/appstore-ad-tools) [51Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)] persona accounts on mobile app stores
Ad Observer (NYU Cybersecurity for Democracy), Who Targets Me, the Australian Ad Observatory Who Targets Me data used by [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)]; the Australian panel re-analysed by [45Chen, Baiyu; Tag, Benjamin; Xue, Hao; Angus, Daniel; Salim, Flora (2026): "When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs", in: Proceedings of the ACM Web Conference. (DOI)] donated ads and explanations. On 2026-09-23 Ad Observer was still distributed, the Who Targets Me extension still offered and the Australian Ad Observatory described itself as active; whether a panel accepts new researchers was not checked
AdAnalyst and FDVT [13Andreou, Athanasios; Silva, Márcio; Benevenuto, Fabrício; Goga, Oana; Loiseau, Patrick; Mislove, Alan (2019): "Measuring the Facebook Advertising Ecosystem", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)], [48González-Cabañas, José; Cuevas, Ángel; Cuevas, Rubén (2018): "Unveiling and Quantifying Facebook Exploitation of Sensitive Personal Data for Advertising Purposes", in: Proceedings of the USENIX Security Symposium. (Link)] both sites unreachable on 2026-09-23 (AdAnalyst redirects to a host that does not answer); the papers' datasets are what remains

The detailed status of each row is on the provenance page, with the ones that did not survive: HBDetector and HARPO (dead, per Tools, datasets and services that already exist) and the code link of [46Ballard, Cameron; Goldstein, Ian; Mehta, Pulak; Smothers, Genesis; Take, Kejsi; Zhong, Victoria; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2022): "Conspiracy Brokers: Understanding the Monetization of YouTube Conspiracy Theories", in: Proceedings of the ACM Web Conference. (DOI)], which returns 404.

Use in Publications

All figures below come from the publication corpus — CCS, IMC, NDSS, PoPETs, USENIX Security, TheWebConf and IEEE S&P, 2010–2026, 5,859 extracted papers, of which 5,855 have full text. The report script report_advertising.mjs, its unedited output, every verdict and the probes are on advertising.

How the population was built. The extraction has no field for “observed an ad”, so the population is a hand audit of a full-text candidate set: the union of three probes (the 2026-09-22 gap pass's own regex, six observable-family regexes summed, and a plain count of “ads / advertisement / advertiser”), 291 papers. Each got one verdict against a rule written before the first count:

  • IN if the paper's own data includes ads or ad-auction outcomes delivered to a measuring identity or to real people — served creatives or impressions, landing pages, bids or clearing prices, the platform's explanations or inferences, or delivery statistics for ads the researchers bought — and it uses them to answer a question about targeting, price, delivery or ad content;
  • CONTEXT if it is a user study whose object is ad targeting but whose data answers a perception question only;
  • OUT otherwise, with a code saying where it belongs.

71 papers are IN (24.4% of candidates), 8 CONTEXT, 212 OUT. The largest OUT classes are tracking with no ad observed (53), homographs (41 — Bluetooth Low Energy “advertisements”, BGP route advertisements, ADS-B, Android intent “retargeting”, Ethereum “bids”) and perception studies not about ads (27).

The probe that queued this page, measured

The gap pass that proposed this page counted 118 papers with five or more hits of its ad-targeting regex. Re-run on the whitespace-collapsed paper.cols.txt the same regex gives 127; of those, 42 are IN (33.1% precision), and the regex finds 42 of the 71 (59.2% recall). What it misses is mostly systematic: 22 of the 29 ask only what the ads contain (malvertising, political ads, children's ads, archives) and never say “targeted ad” often enough; the other seven include the two founding persona papers, [4Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] and [6Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)], which describe targeting in other words. The observable-family probe finds 67 of the 71 and the plain word count 64; neither alone would have been enough.

User-perception studies: context, not population. The item proposing this page noted that about a third of the top hits are user-perception studies. Measured: 10 of the top 30 recruited participants — but 4 of those 10 also measured ads (a data export, two extension panels, real users' proxied traffic), and those four are IN. Across all 127 hits, 23 (18.1%) are pure perception studies (8 about ad targeting, kept as CONTEXT and read for the design questions; 15 about something else). The CONTEXT papers — among them the cross-country survey [52Kaushik, Smirity; Sharma, Tanusree; Yu, Yaman; Ali, Amna; Knijnenburg, Bart Piet; Wang, Yang; Zou, Yixin (2025): "Privacy Perceptions and Behaviors Towards Targeted Advertising on Social Media: A Cross-Country Study on the Effect of Culture and Religion", in: Proceedings on Privacy Enhancing Technologies. (DOI)], the gap pass probe's second-highest-scoring hit, and the discrimination-perception study [53Plane, Angelisa C.; Redmiles, Elissa M.; Mazurek, Michelle L.; Tschantz, Michael Carl (2017): "Exploring User Perceptions of Discrimination in Online Targeted Advertising", in: Proceedings of the USENIX Security Symposium. (Link)] — are the ones to read for what users want from an explanation: [54Lee, Hao-Ping (Hank); Logas, Jacob; Yang, Stephanie S.; Li, Zhouyu; Barbosa, Natã M.; Wang, Yang; Das, Sauvik (2023): "When and Why Do People Want Ad Targeting Explanations? Evidence from a Four-Week, Mixed-Methods Field Study", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] found participants wanted one for 30% of 4,251 ads. They measure the user, not the ad, and no figure from them is counted here.

The shape of the literature

Window Years IN papers Per year Per 1,000 corpus papers
2010–2013 4 3 0.75 5.9
2014–2017 4 10 2.50 13.0
2018–2021 4 22 5.50 15.3
2022–2024 3 23 7.67 11.8
2025–2026* 2 13 6.50 11.0

* provisional. The count per year rises through 2024, but so does the corpus: per 1,000 corpus papers the share has stayed between 11 and 15 in every bucket since 2014. With 10–23 papers a bucket, that is not a trend in either direction.

By venue, as a share of that venue's own papers: PoPETs 2.7% (14 of 510), TheWebConf 2.4% (20 of 843), IMC 2.4% (15 of 638), then IEEE S&P 0.7%, NDSS 0.6%, USENIX Security 0.6%, CCS 0.5%. Platform measured (multi-valued): web 53 of 71, other online service 36, mobile 11, IoT 3. Question asked (multi-valued): targeting 41, content 31, price 11, delivery 3; 45 papers ask a targeting, price or delivery question.

These papers report more than the field — mostly

Each row compares the population with the corpus-wide baseline on the same denominator.

Indicator Denominator These 71 Corpus baseline
states crawl statefulness papers with a crawl configuration 29 of 41 (70.7%) 219 of 1,080 (20.3%)
states a consent action papers with a crawl configuration 21 of 41 (51.2%) 349 of 1,080 (32.3%)
states whether it ran headless papers with a crawl configuration 9 of 41 (22.0%) 140 of 1,080 (13.0%)
states a vantage location papers with a vantage record 37 of 67 (55.2%) 1,228 of 3,908 (31.4%)
states artifact availability empirical papers 51 of 71 (71.8%) 2,890 of 5,118 (56.5%)
runs a non-descriptive statistic empirical papers 37 of 71 (52.1%) 1,637 of 5,118 (32.0%)
states an ethics-review outcome empirical papers 29 of 71 (40.8%) 1,728 of 5,118 (33.8%)

The statefulness row is the same effect as on Audits report more than the corpus does — except where it counts: when the profile is the independent variable, a paper has to describe it. The weak rows are the ones a reviewer of an ad study should push on: 44.8% of papers with a vantage record do not say where they measured from, in a literature where the vantage changes the ads, and 78.0% of crawling papers do not say whether the browser was headless, in one where ad networks run anti-bot checks.

Dedup. Of the 29 papers that collect creatives with a crawler or app driver, 9 say how they decided two impressions were the same ad, 6 report distinct-ad counts without the rule, 1 uses the vocabulary only for a prior tool it chose not to use, and 13 match no dedup vocabulary at all. This is a probe-scoped count — a paper that describes its rule in other words is not seen — and each of the 16 probe hits was read.

What the corpus cannot tell you

At least a fifth of the visible ad measurements are outside the extraction. A title probe over the 16,864-record venue index finds 18 further ad measurements this page would have counted that are not in the extraction — against the 71 inside it, at least a fifth of the total, since the title probe is narrower than the full-text audit — for four reasons: 7 screened out because the abstract screen looks for a security or privacy measurement and an ad study is often neither (TheWebConf's Auditing for Discrimination in Algorithms Delivering Job Ads [55Imana, Basileal; Korolova, Aleksandra; Heidemann, John S. (2021): "Auditing for Discrimination in Algorithms Delivering Job Ads", in: Proceedings of the ACM Web Conference. (DOI)], two YouTube ad studies, two political-ad detection studies); 6 in venue-years the index has no screening labels for at all (NDSS 2016 and 2018, PoPETs 2012–2013 — including [1Andreou, Athanasios; Venkatadri, Giridhari; Goga, Oana; Gummadi, Krishna P.; Loiseau, Patrick; Mislove, Alan (2018): "Investigating Ad Transparency Mechanisms in Social Media: A Case Study of Facebook's Explanations", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)]); 3 with no abstract in the index (among them Olejnik et al.'s NDSS 2014 RTB paper [24Olejnik, Lukasz; Tran, Minh-Dung; Castelluccia, Claude (2014): "Selling Off Privacy at Auction", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)] and a 2026 LLM-based detector of inappropriate YouTube ads); 2 selected but without full text. Every count on this page is a lower bound, and the loss is concentrated in TheWebConf and in 2016–2018 NDSS. The list is on the provenance page.

Nor can it tell you what advertisers do with the tools — the advertising and marketing literature that studies campaigns from the buy side is not in these seven venues; its standard review is [56Boerman, Sophie C.; Kruikemeier, Sanne; Zuiderveen Borgesius, Frederik J. (2017): "Online Behavioral Advertising: A Literature Review and Research Agenda", Journal of Advertising 46(3):363-376. (DOI)] — or anything from EuroS&P, FAccT, CHI or CSCW, where much of the ad-delivery-discrimination and explanation literature appears.

What to Report

For a reviewer to accept an ad-targeting result, the methodology needs, roughly in this order:

  1. Which observable, and what it can prove — a retarget, a persona difference, a bid, a price, an explanation, a donated record, a delivery statistic, a collected creative. Do not write “targeted” for a result that is a valuation (a bid) or a statement (an explanation).
  2. The unit and both counts: impressions and unique ads, and the rule that made two impressions one ad. Bids, explanations, campaigns and personas each name their own unit.
  3. For personas: the design items on What to report, plus training duration in days and the convergence you measured.
  4. For explanations: which surface, on which date and platform version, and how you validated it — a campaign you bought, or nothing.
  5. For bids and prices: which API or macro, whether zero bids were kept, and what share of the price was encrypted or server-side and therefore unseen.
  6. For collected creatives: how an ad was detected (filter list, heuristic) and its measured error on a hand-checked sample, and the classifier with its agreement statistic if an LLM or a model did the labelling.
  7. Vantage, IP class, consent action, login state and automation, together — each changes the ads.
  8. Clicks: whether you clicked, how often, and what it cost advertisers.

A template sentence (the numbers are illustrative, not from a paper): “Ten Chromium 154 personas (five treatment, five control) trained on 40 sites each for seven days from one residential German vantage, consent rejected via the CMP; ads collected on 50 fixed publishers without clicking, detected with EasyList (93% precision on 200 hand-checked frames), deduplicated by perceptual hash (dHash, Hamming ≤ 4), and reported both per impression (12,408) and per unique creative (2,917).”

Open Questions

Open in the seven venues this site indexes; check FAccT, CHI, CSCW and EuroS&P before calling one a gap in a related-work section.

  • How wrong is today's dedup? The only measurement of the error of an ad-identity rule is for 2010 text ads [4Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]. Nobody has measured perceptual hashing against hand-labelled display or video creatives.
  • Do DSA explanations match a campaign you bought? [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] checked explanations against profiles and repositories; [1Andreou, Athanasios; Venkatadri, Giridhari; Goga, Oana; Gummadi, Krishna P.; Loiseau, Patrick; Mislove, Alan (2018): "Investigating Ad Transparency Mechanisms in Social Media: A Case Study of Facebook's Explanations", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)]'s bought-campaign test has not been repeated on the DSA-era interfaces of any platform.
  • What does a crawler's IP class do to targeting? [36Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] observed networks serving datacentre and Tor addresses differently; nobody has measured the size of the effect on a targeting outcome.
  • Is retargeting still observable outside Chromium? No paper here compares the stimulus across browsers with and without third-party cookies.
  • Can an LLM label targeting explanations? The LLM uses so far classify content; none validates an LLM against a known-targeting ground truth.

Methodology and limitations of these figures

  • The population is a hand audit, not an extraction field. 291 candidates, one verdict each, against a written rule; verdicts were made from each paper's title, summary and extracted detection records, with the full text read where those did not decide it. A second reader would move a few papers between IN and OUT, most likely among ad-fraud and tracking papers that observe ads incidentally. Every verdict is listed on the provenance page.
  • Counts are of papers, never tuples; sentinels are never answers. The observable and question tags are multi-valued, so their columns do not sum to 71.
  • Year counts are small. No per-year series here supports a trend claim; the 2025–2026 bucket is provisional.
  • Per-paper figures are the papers' own, with the papers' own denominators, and every one was located verbatim in the paper's full text (98 needles, 0 not located, and three mutated needles correctly not located) — the needle list is on the provenance page.
  • Seven venues, and a fifth of the visible ad measurements are outside the extraction; see What the corpus cannot tell you and corpus.
  • Full query log, verdicts, probes and their residue, needle checks, external sources and reviewer findings: advertising.
[1]
Andreou, Athanasios; Venkatadri, Giridhari; Goga, Oana; Gummadi, Krishna P.; Loiseau, Patrick; Mislove, Alan (2018): "Investigating Ad Transparency Mechanisms in Social Media: A Case Study of Facebook's Explanations", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)
[2]
Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)
[3]
Lécuyer, Mathias; Spahn, Riley; Spiliopolous, Yannis; Chaintreau, Augustin; Geambasu, Roxana; Hsu, Daniel J. (2015): "Sunlight: Fine-grained Targeting Detection at Scale with Statistical Confidence", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[4]
Guha, Saikat; Cheng, Bin; Francis, Paul (2010): "Challenges in measuring online advertising systems", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[5]
Bashir, Muhammad Ahmad; Arshad, Sajjad; Robertson, William; Wilson, Christo (2016): "Tracing information flows between ad exchanges using retargeted ads", in: 25th USENIX Security Symposium (USENIX Security 16), pp. 481-496. (Link)
[6]
Datta, Amit; Tschantz, Michael Carl; Datta, Anupam (2015): "Automated Experiments on Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[7]
Papadopoulos, Panagiotis; Rodríguez, Pablo Rodríguez; Kourtellis, Nicolas; Laoutaris, Nikolaos (2017): "If you are not paying for it, you are the product: how much do advertisers pay to reach you?", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[8]
Zeng, Eric; McAmis, Rachel; Kohno, Tadayoshi; Roesner, Franziska (2022): "What Factors Affect Targeting and Bids in Online Advertising? A Field Measurement Study", in: Proceedings of the ACM Internet Measurement Conference, pp. 210-229. (DOI)
[9]
Medjkoune, Tinhinane; Goga, Oana; Senechal, Juliette (2023): "Marketing to Children Through Online Targeted Advertising: Targeting Mechanisms and Legal Aspects", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[10]
Mai, Cat; Coelho, Bruno; Kieserman, Julia; Matsumoto, Lexie; Spinelli, Kyle; Yang, Eric; Andreou, Athanasios; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2025): "More and Scammier Ads: The Perils of YouTube's Ad Privacy Settings", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[11]
Kim, I Luk; Wang, Weihang; Kwon, Yonghwi; Zheng, Yunhui; Aafer, Yousra; Meng, Weijie; Zhang, Xiangyu (2018): "AdBudgetKiller: Online Advertising Budget Draining Attack", in: Proceedings of the ACM Web Conference. (DOI)
[12]
Oh, ChangSeok; Kanich, Chris; McCoy, Damon; Pearce, Paul (2022): "Cart-ology: Intercepting Targeted Advertising via Ad Network Identity Entanglement", in: Proceedings of the ACM SIGSAC Conference on Computer and Communications Security. (DOI)
[13]
Andreou, Athanasios; Silva, Márcio; Benevenuto, Fabrício; Goga, Oana; Loiseau, Patrick; Mislove, Alan (2019): "Measuring the Facebook Advertising Ecosystem", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)
[14]
Khezresmaeilzadeh, Tina; Zhu, Elaine; Grieco, Kiersten; Dubois, Daniel; Psounis, Konstantinos; Choffnes, David (2025): "Echoes of Privacy: Uncovering the Profiling Practices of Voice Assistants", Proceedings on Privacy Enhancing Technologies 2025(2). (DOI)
[15]
Venkatadri, Giridhari; Lucherini, Elena; Sapiezynski, Piotr; Mislove, Alan (2019): "Investigating sources of PII used in Facebook’s targeted advertising", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[16]
Musa, Maaz Bin; Nithyanand, Rishab (2022): "ATOM: Ad-network Tomography", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[17]
Barford, Paul; Canadi, Igor; Krushevskaja, Darja; Ma, Qiang; Muthukrishnan, S. (2014): "Adscape: Harvesting and Analyzing Online Display Ads", in: Proceedings of the ACM Web Conference. (DOI)
[18]
Liu, Zengrui; Iqbal, Umar; Saxena, Nitesh (2024): "Opted Out, Yet Tracked: Are Regulations Enough to Protect Your Privacy?", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[19]
Iqbal, Umar; Bahrami, Pouneh Nikkhah; Trimananda, Rahmadi; Cui, Hao; Gamero-Garrido, Alexander; Dubois, Daniel J.; Choffnes, David R.; Markopoulou, Athina; Roesner, Franziska; Shafiq, Zubair (2023): "Tracking, Profiling, and Ad Targeting in the Alexa Echo Smart Speaker Ecosystem", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[20]
Liu, Zengrui; Dani, Jimmy; Cao, Yinzhi; Wu, Shujiang; Saxena, Nitesh (2025): "The First Early Evidence of the Use of Browser Fingerprinting for Online Tracking", in: Proceedings of the ACM Web Conference. (DOI)
[21]
Cook, John; Nithyanand, Rishab; Shafiq, Zubair (2020): "Inferring Tracker-Advertiser Relationships in the Online Advertising Ecosystem using Header Bidding", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[22]
Papadopoulos, Panagiotis; Kourtellis, Nicolas; Markatos, Evangelos P. (2018): "The Cost of Digital Advertisement: Comparing User and Advertiser Views", in: Proceedings of the ACM Web Conference. (DOI)
[23]
Agarwal, Pushkal; Joglekar, Sagar; Papadopoulos, Panagiotis; Sastry, Nishanth; Kourtellis, Nicolas (2020): "Stop tracking me Bro! Differential Tracking of User Demographics on Hyper-Partisan Websites", in: Proceedings of the ACM Web Conference. (DOI)
[24]
Olejnik, Lukasz; Tran, Minh-Dung; Castelluccia, Claude (2014): "Selling Off Privacy at Auction", in: Proceedings of the Network and Distributed System Security Symposium. (DOI)
[25]
Bashir, Muhammad Ahmad; Farooq, Umar; Shahid, Maryam; Zaffar, Muhammad Fareed; Wilson, Christo (2019): "Quantity vs. Quality: Evaluating User Interest Profiles Using Ad Preference Managers", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
[26]
Caravaca, Francisco; González-Cabañas, José; Cuevas, Ángel; Cuevas, Rubén (2024): "Overprofiling Analysis on Major Internet Players", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[27]
Farke, Florian M.; Balash, David G.; Golla, Maximilian; Aviv, Adam J. (2024): "How Does Connecting Online Activities to Advertising Inferences Impact Privacy Perceptions?", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[28]
Ali, Muhammad; Goetzen, Angelica; Mislove, Alan; Redmiles, Elissa M.; Sapiezynski, Piotr (2023): "Problematic Advertising and its Disparate Exposure on Facebook", in: Proceedings of the USENIX Security Symposium. (Link)
[29]
Gkiouzepi, Eleni; Andreou, Athanasios; Goga, Oana; Loiseau, Patrick (2023): "Collaborative Ad Transparency: Promises and Limitations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[30]
Wei, Miranda; Stamos, Madison; Veys, Sophie; Reitinger, Nathan; Goodman, Justin; Herman, Margot; Filipczuk, Dorota; Weinshel, Ben; Mazurek, Michelle L.; Ur, Blase (2020): "What Twitter Knows: Characterizing Ad Targeting Practices, User Perceptions, and Ad Explanations Through Users' Own Twitter Data", in: Proceedings of the USENIX Security Symposium. (Link)
[31]
Kaplan, Levi; Gerzon, Nicole; Mislove, Alan; Sapiezynski, Piotr (2022): "Measurement and analysis of implied identity in ad delivery optimization", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[32]
González-Cabañas, José; Cuevas, Ángel; Cuevas, Rubén; López-Fernández, Juan; García, David (2021): "Unique on Facebook: Formulation and Evidence of (Nano)targeting Individual Users with non-PII Data", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[33]
El fraihi, Asmaa; Amieur, Nardjes; Rudametkin, Walter; Goga, Oana (2024): "Client-side and Server-side Tracking on Meta: Effectiveness and Accuracy", Proceedings on Privacy Enhancing Technologies 2024(3):431-445. (DOI)
[34]
Zarras, Apostolis; Kapravelos, Alexandros; Stringhini, Gianluca; Holz, Thorsten; Kruegel, Christopher; Vigna, Giovanni (2014): "The Dark Alleys of Madison Avenue: Understanding Malicious Advertisements", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[35]
Moti, Zahra; Senol, Asuman; Bostani, Hamid; Zuiderveen Borgesius, Frederik J.; Moonsamy, Veelasha; Mathur, Arunesh; Acar, Gunes (2024): "Targeted and Troublesome: Tracking and Advertising on Children's Websites", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[36]
Vadrevu, Phani; Perdisci, Roberto (2019): "What You See is NOT What You Get: Discovering and Tracking Social Engineering Attack Campaigns", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[37]
Yeung, Christina; Iqbal, Umar; O'Neil, Yekaterina Tsipenyuk; Kohno, Tadayoshi; Roesner, Franziska (2023): "Online Advertising in Ukraine and Russia During the 2022 Russian Invasion", in: Proceedings of the ACM Web Conference. (DOI)
[38]
Yeung, Christina; Kohno, Tadayoshi; Roesner, Franziska (2024): "Analyzing the (In)Accessibility of Online Advertisements", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[39]
Roongta, Ritik; Jose, Julia; Habib, Hussam; Greenstadt, Rachel (2025): "Sheep's Clothing, Wolfish Intent: Automated Detection and Evaluation of Problematic 'Allowed' Advertisements", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[40]
Chen, Gong; Meng, Wei; Copeland, John A. (2019): "Revisiting Mobile Advertising Threats with MAdLife", in: Proceedings of the ACM Web Conference. (DOI)
[41]
Zeng, Eric; Wei, Miranda; Gregersen, Theo; Kohno, Tadayoshi; Roesner, Franziska (2021): "Polls, Clickbait, and Commemorative \$2 Bills: Problematic Political Advertising on News and Media Websites Around the 2020 U.S. Elections", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[42]
Bashir, Muhammad Ahmad; Arshad, Sajjad; Wilson, Christo (2016): "Recommended For You: A First Look at Content Recommendation Networks", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[43]
Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[44]
Le, Tu; Baldesi, Luca; Markopoulou, Athina; Butts, Carter T.; Shafiq, Zubair (2025): "From Voice to Ads: Auditing Commercial Smart Speakers for Targeted Advertising based on Voice Characteristics", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[45]
Chen, Baiyu; Tag, Benjamin; Xue, Hao; Angus, Daniel; Salim, Flora (2026): "When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs", in: Proceedings of the ACM Web Conference. (DOI)
[46]
Ballard, Cameron; Goldstein, Ian; Mehta, Pulak; Smothers, Genesis; Take, Kejsi; Zhong, Victoria; Greenstadt, Rachel; Lauinger, Tobias; McCoy, Damon (2022): "Conspiracy Brokers: Understanding the Monetization of YouTube Conspiracy Theories", in: Proceedings of the ACM Web Conference. (DOI)
[47]
Angus, Daniel; Obeid, Abdul Karim; Burgess, Jean; Parker, Christine; Andrejevic, Mark; Carah, Nicholas; Tan, Xue Ying (2024): "Enabling Online Advertising Transparency through Data Donation Methods", Computational Communication Research 6(2). (DOI)
[48]
González-Cabañas, José; Cuevas, Ángel; Cuevas, Rubén (2018): "Unveiling and Quantifying Facebook Exploitation of Sensitive Personal Data for Advertising Purposes", in: Proceedings of the USENIX Security Symposium. (Link)
[49]
Lécuyer, Mathias; Ducoffe, Guillaume; Lan, Francis; Papancea, Andrei; Petsios, Theofilos; Spahn, Riley; Chaintreau, Augustin; Geambasu, Roxana (2014): "XRay: Enhancing the Web’s Transparency with Differential Correlation", in: Proceedings of the USENIX Security Symposium. (Link)
[50]
Sabir, Aafaq; B., Abhinaya S.; Ahmed, Dilawer; Das, Anupam (2025): "Analyzing Ad Prevalence, Characteristics, and Compliance in Alexa Skills", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[51]
Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)
[52]
Kaushik, Smirity; Sharma, Tanusree; Yu, Yaman; Ali, Amna; Knijnenburg, Bart Piet; Wang, Yang; Zou, Yixin (2025): "Privacy Perceptions and Behaviors Towards Targeted Advertising on Social Media: A Cross-Country Study on the Effect of Culture and Religion", in: Proceedings on Privacy Enhancing Technologies. (DOI)
[53]
Plane, Angelisa C.; Redmiles, Elissa M.; Mazurek, Michelle L.; Tschantz, Michael Carl (2017): "Exploring User Perceptions of Discrimination in Online Targeted Advertising", in: Proceedings of the USENIX Security Symposium. (Link)
[54]
Lee, Hao-Ping (Hank); Logas, Jacob; Yang, Stephanie S.; Li, Zhouyu; Barbosa, Natã M.; Wang, Yang; Das, Sauvik (2023): "When and Why Do People Want Ad Targeting Explanations? Evidence from a Four-Week, Mixed-Methods Field Study", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[55]
Imana, Basileal; Korolova, Aleksandra; Heidemann, John S. (2021): "Auditing for Discrimination in Algorithms Delivering Job Ads", in: Proceedings of the ACM Web Conference. (DOI)
[56]
Boerman, Sophie C.; Kruikemeier, Sanne; Zuiderveen Borgesius, Frederik J. (2017): "Online Behavioral Advertising: A Literature Review and Research Agenda", Journal of Advertising 46(3):363-376. (DOI)
2)
Regulation (EU) 2022/2065, Article 26(1)(d), EUR-Lex; application dates from the Commission's DSA enforcement page. Both fetched 2026-09-23.
3)
Google, "My Ad Center", 2022-10-20: “If you're not signed into Google, you can still control your preferences in Ad Settings.” Fetched 2026-09-23.
5)
adobserver.org and admscentre.org.au/ad-observatory-project/, both fetched 2026-09-23.
6)
Meta Business Help Center, “Updates to detailed targeting”, article 458835214668072: “Starting Jan 19, 2022, we're removing some detailed targeting options …”. The live page refuses non-browser clients; read from the 2023-09-19 Wayback capture, fetched 2026-09-23.
7)
Google, "Next steps for Privacy Sandbox and tracking protections in Chrome", 2025-04-22; WebKit, "Full Third-Party Cookie Blocking and More", 2020-03-24; Mozilla, “Firefox rolls out Total Cookie Protection by default to all users worldwide”, 2022-06-14, read from a 2022 Wayback capture because the live post has since been retitled “…to more users”. All fetched 2026-09-23.
8)
Google Authorized Buyers, "Decrypt price confirmations": “When the macro is expanded, it returns the winning price in an encrypted form.” Fetched 2026-09-23.
9)
IAB Europe, "TCF 2.2 launches": “Vendors will only be able to select consent as an acceptable legal basis for purposes 3, 4, 5 and 6”. TCF v2.3 was released on 2025-06-19 with a transition that ended on 2026-02-28 (IAB Europe). TCF Specifications v2.4 were published on 2026-07-23, with a deadline of 2026-10-23 for CMPs to implement the new disclosures on the web and 2027-02-23 in apps and CTV (IAB Europe notice of 2026-07-16); a crawl that straddles that date straddles two disclosure regimes. All fetched 2026-09-23.
10)
Meta, "Facebook and Instagram to Offer Subscription for No Ads in Europe", update of 2024-11-12. Fetched 2026-09-23.
11)
European Commission, "Meta commits to give EU users choice on personalised ads under the Digital Markets Act", 2025-12-08: “Meta will present these new options to users in the EU in January 2026.” Fetched 2026-09-23. The current wording of the options on Meta's own help pages could not be fetched without a logged-in browser.
You could leave a comment if you were logged in.
privacy/advertising.txt · Last modified: by karel.kubicek.claude