User Tools

Site Tools


provenance:design:platforms:ad_archives

This is an old revision of the document!


Provenance: Ad-Transparency Archives

Working notes behind ad_archives: every query with its population, the hand-audit that produced the paper counts, the quote checks, the external sources with fetch dates and the ones that were rejected, and the decisions taken along the way. Corpus-level caveats — the seven-venue scope, the 2025–2026 provisional slice, extraction stability — are on corpus and are not restated here.

This page is a log, not prose. It is for somebody checking a number.

The run

Item Value
Date 2026-09-08
Corpus data/extract/run1/extractions.jsonl, 5,859 extracted papers, 5,855 with paper.cols.txt
Corpus commit the 2026-08-11 extension (8a6b843). Every corpus figure here was derived on that corpus; two figures quoted about the corpus — ~20% free-text stability and 0.9% unlocatable quotes — were measured on the previous 4,322-paper run and are used as orders of magnitude, which What could not be established records.
Venues CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf, IEEE S&P, 2010–2026
Page created design:platforms:ad_archives, new page. First save 48,797 bytes at rev 1788865497; five review-driven revisions took it to ~66 kB.
Report script scripts/adarchives_report.mjs (output below, unedited)
Quote checker scripts/adarchives_quotecheck.py (output below, unedited)
Bibliography 10 entries added to bibliography at rev 1788865485; 0 duplicate keys case-insensitively across all 860 entries; 0 literal @ inside a field; ?purge=true issued on the bibliography before the citing page was saved, and the rendered page shows 15 references against 15 distinct {[key]} markers (100 markers, 200 bibtex_citekey spans — the plugin emits two per marker)
Agents one Opus session did the queries, the paper reading, the external fetches and the drafting; four review sub-agents (three Sonnet, one Fable) — every finding, accepted or rejected, is in Review log

Mistakes made in this run

Listed here rather than buried in the review log, because a provenance page that reads clean is less useful than one that says where it went wrong.

  • A figure on the page contradicted the script that produced it. X's repository was published as 0 papers in five places while the script's own output printed 1. Nothing in the workflow catches that except a reviewer re-reading the output against the page, which is exactly why that pass exists.
  • The page contradicted itself. Its Apple row said 1 paper; a sentence nine lines below said no archive other than Meta's and Google's had been used. Both were visible on one screen.
  • A paper was credited with a method it never used. [1Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)] was placed in the “buy your own ads” row; the paper that bought ads was [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)]. A measured-results row credited the same finding to the wrong paper.
  • A materially wrong access claim. The first draft said only four of thirteen repositories offer bulk access and “the rest are search boxes”. Mozilla's own annex, which the page already cited, records an API for nine of them.
  • An enforcement action against the exact instrument the page recommends was missed. X's repository was fined over on 5 December 2025, and the first draft described it as “current, least gated” with no mention.
  • A verification script silently checked less. Adding quotes re-used a CHECKS key and dropped three existing checks while still reporting “0 NOT FOUND”. A guard now aborts on duplicate keys.
  • “Web-only” was published for Google's Ads Transparency Center in the first save, which [1Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)] contradicts — it is a BigQuery public dataset. Caught by re-reading the paper, not by a reviewer.

No credential or token was printed, logged or exposed during the run.

No ~~DISCUSSION~~ block on this page, and this is the decision that sets the default. Comments belong on the content page, where a reader who disagrees with a claim will be. A provenance page is an audit trail; a thread on it would be a second, disconnected discussion of the same claims. Subsequent provenance pages should follow this unless there is a reason not to.

Judgement calls

Decision Alternative a reasonable person would have taken Why this way
A new page rather than a section on platforms. Add an “ad archives” section to the parent, which already has five paragraphs on them. The parent is organised by route into a platform; the archive needs to be organised by archive, because the thirteen repositories differ on scope, retention and bulk access in ways that decide feasibility. That comparison is a table the parent cannot carry, and the page id was already fixed on roadmap and gated by scripts/sitemap.mjs.
Kept the id design:platforms:ad_archives. privacy:advertising or design:ad_archives. The id was decided on roadmap on 2026-09-07 and sitemap.mjs gates on it. Changing it would have required editing the roadmap row in the same sitting; there was no reason to.
Reported 11 papers, not 69. Publish the parent's 69 for consistency between sibling pages. 69 is a probe artefact: the probe cannot separate Meta's Ad Library from an Android advertising SDK, and 44 of the 69 are the SDK sense. Publishing a number known to be wrong for consistency's sake is the worst option available. The parent page has been corrected in the same sitting.
Counted “the archive is the object of study” as use of the archive. Count only papers that treated the archive as a data source. [3Edelson, Laura; Lauinger, Tobias; McCoy, Damon (2020): "A Security Analysis of the Facebook Ad Library", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)], [4Le Pochat, Victor; Edelson, Laura; Van Goethem, Tom; Joosen, Wouter; McCoy, Damon; Lauinger, Tobias (2022): "An Audit of Facebook's Political Ad Policy Enforcement", in: Proceedings of the USENIX Security Symposium. (Link)] and [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] all query the archive at scale and measure it; excluding them would drop the four most method-bearing papers on the page. [5Gkiouzepi, Eleni; Andreou, Athanasios; Goga, Oana; Loiseau, Patrick (2023): "Collaborative Ad Transparency: Promises and Limitations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] is the marginal call: it critiques ad libraries and proposes an alternative rather than querying one, and it is counted. Excluding it gives 10 rather than 11.
Called “archive as ground truth” superseded. Call it current — most of the corpus does it. One 2026 paper states outright that it tried and could not ([1Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)]), and four papers spanning 2020–2026 measure the archive's error rate rather than assuming it away. The direction of travel is unambiguous even at n=11.
Did not put the Mozilla/Check First stress test in the page's main table. Use its per-platform assessments as the coverage comparison. It is dated “as of March 18, 2024” and this page is dated 2026-09-08. Its inventory of thirteen repositories was reused (and each URL re-checked); its per-platform judgements are cited as of their own date and not presented as current.
Cut the per-year series at 2026 with a provisional label rather than at 2024. Cut at 2024, as corpus recommends for trends. The counts are 1–3 papers a year: this is a list of eleven papers, not a trend, and dropping 2025–2026 would drop the only two papers that measure a non-Meta archive. The provisional label is stated wherever the years appear.
Published the google-ads-transparency-center dataset on a paper's authority. Leave it out, since the vendor's own docs do not mention it. [1Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)] queried it, names two of its fields and gives a retrieval date, and the alternative was to publish “web-only”, which is what the first draft said and is wrong. The uncertainty is stated in the footnote on the page, not only here.
Linked interrater_agreement rather than the queued statistics:annotation. Link the queued id, which is legal because it is declared on roadmap. No content page on this wiki currently links a queued page, and a red link on a brand-new page reads as a defect to a reader who has not read the roadmap.

Queries, with populations and denominators

Every figure on the page traces to a row here. The script is scripts/adarchives_report.mjs; its unedited output follows.

Denominators used

Population N Definition
extracted papers 5,859 rows in extract/run1/extractions.jsonl
papers with extractable full text 5,855 fulltext/<year>/<venue>/<slug>/paper.cols.txt exists. This is the denominator for every probe on the page.
papers that drew a study population 5,712 population array non-empty
advertising-measurement candidate set 180 at least one detection.phenomenon matching an advertising-shaped regex. A floor: free text, ~20% run-to-run stability, and a paper that measured ads without an ad-shaped word in that field is invisible.
the audited set 11 hand-audited from a 29-paper candidate set; see below

Probe 1: reproducing and decomposing the parent page's figure

/Ad Library|Ads? Archive API/i  over 5,855 full texts  ->  69 papers
  of those, proper-noun "Ad Library"/"Ad Libraries" present (case-sensitive)  ->  25
  of those, only lowercase "ad library"/"ad libraries"                        ->  44   (Android advertising SDKs)
  full texts containing "read librar" (a substring the probe also matches)    ->   3

The homonym is the whole story. In mobile-security writing an ad library is an advertising SDK compiled into an Android app; towards-https-everywhere-on-android-we-are-not-there-yet (USENIX 2020) has a table column headed “Ad Library” listing AdMob, Facebook Audience Network and Unity. A capital L is not a sufficient discriminator either: six papers capitalise it only inside a bibliography entry for Book, Pridgen and Wallach, Longitudinal Analysis of Android Ad Library Permissions (MoST 2013).

Probe 1b: screening the 44 lowercase-only papers

The 44 are outside the candidate set by construction, because the candidate term is case-sensitive. They were therefore never hand-audited one by one, and an unaudited stratum inside a published correction is exactly the kind of gap that should be stated rather than glossed. Section 1b of the script screens them: any lowercase mention within 80 characters of a transparency-archive word is printed for a human to read. Nine fired, and all nine were read.

Paper Verdict
NDSS 2017 obfuscation-resilient-privacy-leak-detection… SDK — emulator detection by ad libraries
WWW 2019 a-multi-modal-neural-embeddings-approach… SDK — “ad library difference” between app pairs
CCS 2021 this-sneaky-piggy-went-to-the-android-ad-market… SDK — “Google's interstitial ad library”
NDSS 2021 the-abuser-inside-apps… SDK — providers “harness their ad libraries” to commit fraud
USENIX 2021 adcube-webvr-ad-fraud… SDK — “a third-party JS ad library in Secure ECMAScript”
USENIX 2023 problematic-advertising-and-its-disparate-exposure… bibliography entry. USENIX's small-caps reference style lowercases titles, so [6Ali, Muhammad; Goetzen, Angelica; Mislove, Alan; Redmiles, Elissa M.; Sapiezynski, Piotr (2023): "Problematic Advertising and its Disparate Exposure on Facebook", in: Proceedings of the USENIX Security Symposium. (Link)]'s citation of [3Edelson, Laura; Lauinger, Tobias; McCoy, Damon (2020): "A Security Analysis of the Facebook Ad Library", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] lands in the lowercase stratum. The paper uses Ad Observer, not the archive.
WWW 2024 in-security-of-file-uploads-in-node-js not an ad library at all — “popular file uplo[ad librar]ies (e.g., express-fileupload)”. A pure substring collision in a paper with no advertising content.
NDSS 2025 careful-about-what-app-promotion-ads-recommend… SDK — “Google AdMob, which is the most widely used ad library”
NDSS 2026 cross-boundary-mobile-tracking… SDK — “Google Ads' dominance, as a standalone ad library”

The remaining 35 have no transparency-archive word within 80 characters of the mention. This is a screen, not an audit, and it is the only recall evidence for that stratum. It found no ad-archive user among the 44. The script throws if a suspect has no hand verdict, so widening the suspect regex forces the reading rather than silently enlarging a clean-looking result.

Two by-products worth recording. First, the “upload libraries” hit means the parent page's 69 included at least one paper about Node.js file uploads, which is a stronger illustration of the probe's problem than the SDK collision. Second, /read librar/i matches 3 full texts corpus-wide, so that collision was available to fire and, on these 69, did not.

Probe 2: the candidate set

Seven terms, unioned. Case-sensitivity on the first is deliberate.

  26  proper-noun Ad Library / Ad Libraries (case-SENSITIVE)
   8  Ad Library API / Ads Archive API
  10  ad archive / ad repository / advertising archive / archive of ads
   4  Google Ads Transparency Center / google_political_ads / Google political ad*
   0  TikTok Commercial Content Library / API
   3  political ad archive / library / repository
   4  Who Targets Me
 ---
  29  UNION (the candidate set)

Per year: 2013:2 2014:1 2015:2 2017:2 2018:3 2019:1 2020:5 2021:1 2022:4 2023:4 2024:1 2025:1 2026:2.

A widened probe is a recall claim, so the width was chosen before the audit and the script asserts that the audit covers the union exactly — see The invariant.

Probe 3: the structured-field cross-check (recall)

Independently of full text, an archive-shaped regex was run over the structured extraction fields tools[].name, otherToolsMentioned, population[].sourceList, detection[].phenomenon/technique/metric and classification[].resourceName/method/groundTruth. It returned 16 papers. Fourteen were already in the full-text candidate set. Two were not, and both were kept as the donation-corpus complement rather than as archive users:

  • problematic-advertising-and-its-disparate-exposure-on-facebook (USENIX 2023) — tools.name = “Ad Observer”. Uses the donation extension, not the archive. On the page as [6Ali, Muhammad; Goetzen, Angelica; Mislove, Alan; Redmiles, Elissa M.; Sapiezynski, Piotr (2023): "Problematic Advertising and its Disparate Exposure on Facebook", in: Proceedings of the USENIX Security Symposium. (Link)].
  • when-ads-become-profiles… (TheWebConf 2026) — population.sourceList = “Australian Ad Observatory”. Donated corpus. On the page as [7Chen, Baiyu; Tag, Benjamin; Xue, Hao; Angus, Daniel; Salim, Flora (2026): "When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs", in: Proceedings of the ACM Web Conference. (DOI)].

That cross-check is the only recall evidence for the audited set, and it found no missed archive user. It is not proof of completeness: a paper that queried an archive without naming it in either full text or a structured field would be invisible to both.

Probe 3b: which archive each audited paper actually used

Assigned by hand from each paper's methods section, in an ARCHIVE_USED map in the script, with an invariant that throws if an audited paper has no assignment. A regex cannot produce this split, and the first draft of the page got it wrong by trying:

Archive Audited papers
Meta 8
Google 3
X 1
Apple 1
none — critiques ad libraries and queries Facebook's Ads Manager instead 1 ([5Gkiouzepi, Eleni; Andreou, Athanasios; Goga, Oana; Loiseau, Patrick (2023): "Collaborative Ad Transparency: Promises and Limitations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)])

A paper can use several, so these do not sum to 11. Two reasons the hand assignment was necessary:

  • [8Zeng, Eric; Wei, Miranda; Gregersen, Theo; Kohno, Tadayoshi; Roesner, Franziska (2021): "Polls, Clickbait, and Commemorative \$2 Bills: Problematic Political Advertising on News and Media Websites Around the 2020 U.S. Elections", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] writes “Google's (or others') political ad transparency reports”. The parenthetical breaks any \bGoogle('s)? political ad\w*\b pattern, so the name probe in Probe 4 misses it while the page's own per-paper table (correctly) assigns it to Google.
  • [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] studies four platforms' repositories in one paper — Facebook, Instagram, YouTube and X — so it is counted under Meta, Google and X.

This is where a reviewer caught the page's worst error. The first draft said “nobody in these seven venues has published from any archive other than Meta's and Google's” and gave X and TikTok as 0 papers each — while the script's own Probe 4 output printed 1 for X, and the page's own Apple row printed 1. Reading the papers settled it: [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] has a section headed 10.4 X's Ad Repository and reports “In total, we identified advertisers with repository entries for 2,190 ads. Of these, 1,458 ads were successfully matched using Tweet IDs”; [1Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)] reports “Apple provides a web interface and an API for its Ad Repository” and, notably, “We did not detect missing entries by comparing randomly selected samples of our dataset to Apple's Ad Repository”. The “0” was in five places on the page and is a textbook instance of a published figure contradicting the script that produced it.

Probe 4: which archive, by name

The zeros are load-bearing, so each regex is printed in full in the script.

Archive named Papers Years
Meta / Facebook Ad Library (proper noun, or API/Report) 13 2018:1 2020:2 2021:1 2022:3 2023:3 2024:1 2025:1 2026:1
Google Ads Transparency Center / political-ads BigQuery 3 2020:1 2023:1 2026:1
TikTok Commercial Content Library / API 0
X / Twitter ads repository or ad transparency centre 1 2026:1
LinkedIn Ad Library 0
Snapchat Ads Gallery / political ads library 0
Apple App Store Ad Repository 1 2026:1
Microsoft / Bing Ad Library 0
Pinterest / Booking.com / AliExpress / Zalando ads repository 0
the EU “European repository for online political advertisements” 0
Digital Services Act (any mention) 29 2023:2 2024:7 2025:7 2026:13
DSA Article 39 within 400 characters of “Digital Services Act” or “DSA” 4 2023:1 2024:1 2026:2
the EU political-advertising regulation, any phrasing 0 in substance 1 hit, hand-checked: two bibliography titles in [4Le Pochat, Victor; Edelson, Laura; Van Goethem, Tom; Joosen, Wouter; McCoy, Damon; Lauinger, Tobias (2022): "An Audit of Facebook's Political Ad Policy Enforcement", in: Proceedings of the USENIX Security Symposium. (Link)]

The Meta row is 13 while the audited set has 8 Meta users, because the row includes bibliography-only mentions. The 13 is not published on the content page for that reason; the audited 8 is.

Probe 5: the donation complement

Named Papers Years
Who Targets Me 4 2020:1 2022:1 2023:1 2026:1
Ad Observer / an Ad Observatory 5 2022:1 2023:3 2026:1
Australian Ad Observatory 1 2026:1
data donation, any phrasing 25 2019:1 2023:2 2024:5 2025:11 2026:6

The hand-audit

All 29 candidates were read in their own paper.cols.txt at every match site. Verdicts live in the AUDIT map inside scripts/adarchives_report.mjs so that they are versioned with the counts they produce.

Verdict Papers Meaning
instrument 11 the archive is a data source, or the archive itself is the object measured
reference 8 the archive appears only in a bibliography entry or a related-work sentence
sdk 8 “ad library” means an Android advertising SDK
other-sense 2 an ad archive in a different sense entirely

Precision of the candidate probe: 11/29 = 37.9%. Precision of the parent page's probe for the same question: 10/69 = 14.5%, and it also missed [1Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)], which uses Google's archive and never writes “Ad Library”.

Rejected candidates, in full

This is the residue. It is printed here rather than left in a script output because a residue nobody reads is not a residue.

Verdict Paper Why rejected
other-sense WWW 2013 the-cost-of-annoying-ads “Adverlicious, an online display advertising archive” supplied ad stimuli for a crowdworker experiment. Not a platform transparency archive.
reference CCS 2013 the-impact-of-vendor-customizations-on-android-security bibliography: Book et al., Longitudinal Analysis of Android Ad Library Permissions
reference WWW 2014 reconciling-mobile-app-privacy-and-usability-on-smartphones… same bibliography entry
sdk NDSS 2015 introducing-privacy-threats-from-ad-libraries-to-android-users… Android advertising SDKs; the paper's own detection.phenomenon is “Ad-library prevalence”
other-sense PETS 2015 automated-experiments-on-ad-privacy-settings “the ad repository lacking ads relevant to the interest” — the ad ecosystem's inventory. The paper is a Google Ad Settings audit.
reference CCS 2017 identifying-open-source-license-violation-and-1-day-security-risk… same bibliography entry
reference CCS 2017 keep-me-updated-an-empirical-study-of-third-party-library-updatability… same bibliography entry
sdk IMC 2018 beyond-google-play-a-large-scale-comparative-study-of-chinese… Android advertising SDKs
sdk PETS 2018 nomoads-effective-and-efficient-cross-app-mobile-ad-blocking Android advertising SDKs
reference PETS 2018 panoptispy-characterizing-audio-and-video-exfiltration… same bibliography entry
sdk CCS 2019 deepintent-deep-icon-behavior-learning… Android advertising SDKs
sdk PETS 2020 nomoats-towards-automatic-detection-of-mobile-tracking Android advertising SDKs
sdk USENIX 2020 towards-https-everywhere-on-android-we-are-not-there-yet a table column headed “Ad Library” listing AdMob, Facebook Audience Network, Unity
sdk WWW 2020 maddroid-characterizing-and-detecting-devious-ad-contents… a section headed “Mobile Ad Library Detection and Analysis”
sdk PETS 2022 analyzing-the-monetization-ecosystem-of-stalkerware a figure axis reading “# of Ad Libraries”
reference PETS 2022 charting-app-developers-journey-through-privacy-regulation-features… bibliography entry for [3Edelson, Laura; Lauinger, Tobias; McCoy, Damon (2020): "A Security Analysis of the Facebook Ad Library", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] only
reference WWW 2022 conspiracy-brokers-understanding-the-monetization-of-youtube… bibliography entry for [3Edelson, Laura; Lauinger, Tobias; McCoy, Damon (2020): "A Security Analysis of the Facebook Ad Library", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] only
reference CCS 2023 marketing-to-children-through-online-targeted-advertising… bibliography entry for [3Edelson, Laura; Lauinger, Tobias; McCoy, Damon (2020): "A Security Analysis of the Facebook Ad Library", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]; the study measures ad explanations, not the archive

The invariant

The script throws if AUDIT and the candidate union differ in either direction:

if (missing.length || extra.length) {
  throw new Error(`AUDIT does not match the candidate set.
    unaudited candidates: ...
    audit rows with no candidate: ...`);
}

This exists because widening a probe is the standard way an unaudited paper enters a published count. Tested by deleting one AUDIT row and one candidate term; both directions throw.

Rejected probes

Recorded so the next run does not re-add them.

Probe Hits Why rejected
/TTPA/i 71 matches the substring inside “httpa…” — HTTPA, “http api”. The spelled-out name and 2024/900 are probed instead, and return zero.
/Google Transparency Report/i 28 that report is about government requests and HTTPS adoption, not ads. It put censorship and TLS papers into an ad-archive candidate set in the first draft.
/\bFORT\b|Facebook Open Research/ 4 intended to find Meta's FORT ad-targeting dataset. FORT collides with unrelated all-caps tokens and OCR artefacts; never hand-audited, never used.
/\bArticle 39\b|\bArt\.? 39\b/ alone 7 collides with GDPR Art. 39 and ACM reference numbering — it returns image-disguising-for-privacy-preserving-deep-learning (CCS 2018) and banned-books-analysis-of-censorship-on-amazon-com (PETS 2026), neither about ad repositories. The DSA-context variant returns 4 and is the one published.
/Ad Library|Ads? Archive API/i 69 the parent page's probe. 14.5% precision for this question; kept in the script only to reproduce and decompose the published figure.

Quote and figure checks

scripts/adarchives_quotecheck.py checks 83 strings against the cited paper's own text: every phrase quoted on the content page, and every paper-sourced figure a reviewer would plausibly pull on. 83 found, 0 not found. The unedited output is below.

What it does not cover, stated plainly. It is not “every numeral on the page”. A generic reviewer listed roughly a dozen figures the first version of the checker had missed — [4Le Pochat, Victor; Edelson, Laura; Van Goethem, Tom; Joosen, Wouter; McCoy, Damon; Lauinger, Tobias (2022): "An Audit of Facebook's Political Ad Policy Enforcement", in: Proceedings of the USENIX Security Symposium. (Link)]'s 45% and 116,963, [7Chen, Baiyu; Tag, Benjamin; Xue, Hao; Angus, Daniel; Salim, Flora (2026): "When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs", in: Proceedings of the ACM Web Conference. (DOI)]'s 76.38% and 41.60%, [9Coelho, Bruno; Lauinger, Tobias; Edelson, Laura; Goldstein, Ian; McCoy, Damon (2023): "Propaganda Política Pagada: Exploring U.S. Political Facebook Ads en Español", in: Proceedings of the ACM Web Conference. (DOI)]'s 1.90% and 2.25%, [10Capozzi, Arthur; Morales, Gianmarco De Francisci; Mejova, Yelena; Monti, Corrado; Panisson, André (2023): "The Thin Ideology of Populist Advertising on Facebook during the 2019 EU Elections", in: Proceedings of the ACM Web Conference. (DOI)]'s +50%, [5Gkiouzepi, Eleni; Andreou, Athanasios; Goga, Oana; Loiseau, Patrick (2023): "Collaborative Ad Transparency: Promises and Limitations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]'s 420 and 51,681 and 86-of-1,021, [11Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]'s 82 still-active pages, [1Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)]'s 249 accounts and [8Zeng, Eric; Wei, Miranda; Gregersen, Theo; Kohno, Tadayoshi; Roesner, Franziska (2021): "Polls, Clickbait, and Commemorative \$2 Bills: Problematic Political Advertising on News and Media Websites Around the 2020 U.S. Elections", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]'s three site lists — probed them by hand and found them all correct. They have since been added, which is what took the count from 67 to 83. Figures still not in the checker are read directly from the extraction's prevalence and population fields, which carry their own evidence quote; those quotes are what the checker verifies for the headline figures.

The checker tries three renderings in order and two match modes per rendering, and it needs all of them:

Match mode Strings
paper.cols.txt verbatim 68
paper.cols.txt folded (lowercase alphanumerics, fi/flf) 2
pypdf extracted live from paper.pdf, verbatim 11
pypdf folded 2

15 of 83 strings needed something other than a verbatim paper.cols.txt match, and two of those exist in no corpus rendering at all — [11Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]'s model name and its “82 were still active” both appear only in a live pypdf extraction. Three findings from that:

  • paper.cols.txt is still spliced for [3Edelson, Laura; Lauinger, Tobias; McCoy, Damon (2020): "A Security Analysis of the Facebook Ad Library", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]. Its 15.8%-of-spend sentence renders as “We find that 15.8 % of total ad spend in the Ad Library cannot to correctly disclose…” — a column splice mid-sentence. Every figure quoted from that paper on the content page was therefore re-verified against a live pypdf extraction, where all of them appear verbatim: 9.7%, 68.3%, $37 M, 6% of spend, 15.8%, $98.2 M, 23%, 16,160 ads, the 16 clusters and $3,867,613. The extraction's evidence quotes were right and the .cols rendering was wrong — the opposite of the usual failure.
  • Two papers drop the fi ligature entirely, so [10Capozzi, Arthur; Morales, Gianmarco De Francisci; Mejova, Yelena; Monti, Corrado; Panisson, André (2023): "The Thin Ideology of Populist Advertising on Facebook during the 2019 EU Elections", in: Proceedings of the ACM Web Conference. (DOI)] literally reads “we fnd 44 949 ads” and [9Coelho, Bruno; Lauinger, Tobias; Edelson, Laura; Goldstein, Ian; McCoy, Damon (2023): "Propaganda Política Pagada: Exploring U.S. Political Facebook Ads en Español", in: Proceedings of the ACM Web Conference. (DOI)] “our fnal data set”. A verbatim check on the correct sentence FAILS. Hence the fold.
  • [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] writes its own sample size with a mathematical-italic N (U+1D441), not ASCII N. NFKC normalisation was added to the checker for this; without it, a correct quote of “N = 48,511” is reported as not found.

A fourth finding is about the checker itself. Sixteen strings were added after the review passes, and one of the additions re-used a CHECKS key that was already present. A repeated Python dict key silently discards the earlier value, so three previously-passing quotes vanished from the run and the total went from 51 to 64 instead of 67 — with a clean “0 NOT FOUND” both times. A _guard_no_duplicate_keys function now counts the keys in the file's own source and aborts; it was mutation-tested by duplicating a key, and it fires. This is the same failure mode as an in-place generator eating a section: the check passed, and it passed on less.

External sources

Every load-bearing external claim on the page, with what was fetched and how. Fetches on 2026-09-08 unless noted. curl means curl -sL with a desktop Chrome User-Agent; “headless” means Playwright's own Chromium build.

Claim on the page Source How verified
Meta Ad Library scope: political ads worldwide for a rolling 7 years, all UK/EU ads for 1 year; identity and location verification at facebook.com/ID with “It can take a few days to confirm the information you submit”; and the per-category field lists the content page's access section relies on https://www.facebook.com/ads/library/api/ headless. curl returns HTTP 403. The page geo-localises — the first fetch came back in German — so it was re-fetched with locale=en_US and an en-US Accept-Language. Verbatim: “Ads about social issues, election or politics that were delivered anywhere in the world during the past 7 years” and “Ads of any type that were delivered to the United Kingdom (UK) or European Union during the past year”.
Google political archive: verified election advertisers only, 7 years, BigQuery, non-unique impressions, mutable history https://adstransparency.google.com/political/faq headless, all 22 FAQ items expanded by clicking. Verbatim: “Election ads … that were paid for by verified election advertisers”; “remain available for seven years after the ad's last impression”; “The ads are published as a public data set on Google Cloud BigQuery”; “if it was shown to the same user 10 times, that counts as 10 and not 1”; “historical data may slightly fluctuate”; “published after a delay of 90 days”.
Ads Transparency Center scope (Search, Display, Gmail, YouTube) https://support.google.com/adspolicy/answer/13733850 headless. Verbatim: “a searchable repository of advertisers and the ads they've served on Google platforms like Search, Display, Gmail, and YouTube”.
The political BigQuery dataset id, table and start year https://cloud.google.com/blog/topics/developers-practitioners/how-get-started-political-ads-transparency-report-dataset WebFetch. Google Cloud blog dated 25 January 2023. Gives bigquery-public-data.google_political_ads, table creative_stats, and “which begins in 2018 and retains ads for 7 years”. Secondary to the product but published by the vendor; the 7-year retention agrees independently with the product FAQ.
That the whole Ads Transparency Center — not only the political subset — is a BigQuery public dataset, and that its web and BigQuery surfaces disagree [1Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)] §7.1, plus https://console.cloud.google.com/marketplace/details/bigquery-public-data/google-ads-transparency-center Partly unverified, and the page says so. The console page returned HTTP 200 with a 3 MB JavaScript shell containing no dataset metadata (curl and headless both), and Playwright timed out on the product variant, so the listing could not be rendered. The dataset is attested by the paper, which queried it and names its surface_serving_stats and creative_page_url fields and cites the same URL retrieved 28 February 2025. Google's own launch post for the centre (https://blog.google/technology/ads/announcing-the-launch-of-the-new-ads-transparency-center/, 29 March 2023, WebFetch) describes it as “a searchable hub” and says nothing about BigQuery, so the vendor's own documentation does not corroborate it. This was caught late: the first draft of the page said the centre was “web-only”, which the paper contradicts.
The Commission designation list now carries 28 active designations, “Information updated on 31 August 2026” https://digital-strategy.ec.europa.eu/en/policies/list-designated-vlops-and-vloses WebFetch, 2026-09-08, after a reviewer flagged that the first draft's “designated in April 2023” framing was stale. Confirms Shein 26.04.2024, Temu 31.05.2024, WhatsApp 26.01.2026, Reddit / Roblox / ChatGPT (VLOSE) all 31.08.2026, and Stripchat terminated 27.05.2025.
X fined €120 million on 5 December 2025, the ad repository one of three grounds https://digital-strategy.ec.europa.eu/en/news/commission-fines-x-eu120-million-under-digital-services-act WebFetch. Page carries “Last update 16 January 2026”. Verbatim: “X's advertisement repository fails to meet the transparency and accessibility requirements of the DSA”; “X incorporates design features and access barriers, such as excessive delays in processing, which undermine the purpose of ad repositories”; “lacks critical information, such as the content and topic of the advertisement, as well as the legal entity paying for it”. platforms already cited this decision for the researcher access ground; the ad-repository ground is new to this page.
TikTok's ad repository preliminarily in breach 15 May 2025; binding commitments accepted 5 December 2025 https://digital-strategy.ec.europa.eu/en/news/commission-preliminarily-finds-tiktoks-ad-repository-breach-digital-services-act and https://digital-strategy.ec.europa.eu/en/news/commission-accepts-tiktoks-commitments-advertising-transparency-under-digital-services-act WebSearch restricted to ec.europa.eu to find the pages, then WebFetch on the commitments page (which carries “Last update 1 April 2026”) for the four verbatim commitments. The May 2025 preliminary-findings wording was read from the search result's extract of the Commission page rather than from a direct fetch of it, so it is a weaker verification than the others in this table.
Snapchat's Ad Gallery unavailable 99.7% of the time, fixed 29 April 2026 https://aiforensics.org/work/snapchat-dsa-ad-transparency WebFetch. Dated 13-05-2026. AI Forensics is a European non-profit that audits platforms; this is a research organisation's report, not a regulator's decision, and the page labels it as such. It also falsifies the first draft's claim that Mozilla's 2024 test was “the only source that has tried them side by side” for the per-platform case.
Amazon's Art. 39 obligation upheld: interim suspension set aside 27 March 2024 https://curia.europa.eu/jcms/upload/docs/application/pdf/2024-03/cp240060en.pdf curl HTTP 200, extracted with pypdf. Court of Justice press release No 60/24. Verbatim: “Amazon's request to suspend its obligation to make an advertisement repository publicly available is rejected”. The underlying order was Case T-367/23 R of 27 September 2023, cited in the Mozilla report as its reason for not testing Amazon. The outcome of Amazon's underlying action for annulment was not established.
Amazon's Ad Library API is free, EU-only, 1,000 results per request https://advertising.amazon.com/API/docs/en-us/guides/ad-library/overview HTTP 200 to curl but JavaScript-rendered, so read with a headless browser. Verbatim: “The Ad Library API is available at no cost”, “accessible to anyone who visits Amazon EU stores”, “up to 1,000 results for each request”, and the nine EU marketplaces. https://advertising-api-eu.amazon.com/ returns HTTP 400 to an unauthenticated GET.
Mozilla's annex records an ads-repository API for nine of the thirteen the same 55-page PDF, page 53 pypdf. The annex table's third column is Ads Repository API. This falsified the first draft's claim that only four of the thirteen offer bulk access and that “the rest are search boxes”.
TikTok Commercial Content Library is EU-only https://developers.tiktok.com/products/commercial-content-api WebFetch. Verbatim: “in this phase we are ONLY including data from EU countries, while a researcher/professional who is requesting it can be located in any country”. Matches what platforms recorded on 2026-08-27.
X Ads Repository offers a full-dataset CSV with no application https://ads.x.com/ads-repository headless. Renders “Download the full commercial communications dataset: Commercial Communications CSV” plus a country and date filter. curl returns 200 with a 2,368-byte JS shell. A stale in-page banner about an Ad Manager migration “on March 26th” was ignored.
DSA Art. 39: repository, API, one year after last presentation, mandated fields, and Art. 39(3) redaction of removed ads https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32022R2065 headless. EUR-Lex returns HTTP 202 with a zero-byte body to curl — the classic “empty file that is really a bot challenge”. Art. 39(3) quoted verbatim.
Reg. (EU) 2024/900 applies from 10 October 2025; mandates a European repository; 7-year notice retention https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX%3A32024R0900 headless. Verbatim: “It shall apply from 10 October 2025” and “the European repository for online political advertisements”.
Implementing Reg. (EU) 2026/818 of 9 April 2026 specifies the repository's data structure and API https://eur-lex.europa.eu/eli/reg_impl/2026/818/oj/eng and CELEX 32026R0818 headless. Record reads “OJ L, 2026/818, 10.4.2026”, “In force”; Art. 2 gives entry into force on the twentieth day after publication. Also names Implementing Reg. (EU) 2025/1410 of 9 July 2025 on transparency-notice formats.
Meta ended EU political ads on 6 October 2025 https://about.fb.com/news/2025/07/ending-political-electoral-and-social-issue-advertising-in-the-eu/ and https://developers.facebook.com/blog/post/2025/10/06/prohibiting-ads-about-social-issues-elections-or-politics-from-running-in-the-eu-due-to-regulation/ WebFetch. Post dated 25 July 2025, updated 6 October 2025. Verbatim: “Beginning on October 6, 2025 at 6:00 PM Central European Time (CET) … social issue, electoral and political ads will no longer be able to be delivered in the EU and associated territories.
Google stopped EU political ads before the TTPA https://blog.google/company-news/inside-google/around-the-globe/google-europe/political-advertising-in-eu/ and https://support.google.com/adspolicy/answer/16409999 WebFetch and headless. Blog post dated 14 November 2024. Verbatim: “Google will stop serving political advertising in the EU before the TTPA enters into force in October 2025”. The policy change log confirms “In September 2025, Google will update the Political content policy to include the following regional restriction for the European Union (EU)” and defines the restricted category by reference to Reg. 2024/900.
Mozilla / Check First stress test: 13 services, dated 16 April 2024, “even the best approaches don't meet our baseline” https://www.mozillafoundation.org/en/research/library/full-disclosure-stress-testing-tech-platforms-ad-repositories/ and the 55-page PDF headless for the page (WebFetch got HTTP 403), curl for the PDF, then pypdf. Testing “as of March 18, 2024”. The repository inventory and its URLs come from this report's annex; each URL was independently re-checked (below).
Who Targets Me and Ad Observer still operating https://whotargets.me/ and https://adobserver.org/ curl HTTP 200 on both; WebFetch for content. Ad Observer states it “copies the ads you see on Facebook and YouTube”.
Australian Ad Observatory phase one closed March 2024, phase two 2024–2027 https://www.admscentre.org.au/adobservatory/ WebFetch. Verbatim: “The first phase of the Australian Ad Observatory project was completed in March 2024”; “328,107 unique ads”; “1,909 citizens”. adobservatory.org does not resolve.

Repository URLs re-checked one by one

The inventory came from the Mozilla annex; the statuses are ours, from curl -sL with %{http_code} printed. 403 is a bot block, not absence.

https://www.aliexpress.com/p/ad-search-page/index.html                  200
https://advertising.amazon.com/API/docs/en-us/guides/ad-library/overview 200
https://adstransparency.google.com/                                     200
https://adrepository.apple.com/                                         200
https://adlibrary.ads.microsoft.com/                                    200
https://www.booking.com/ad-repository.html                              200
https://www.linkedin.com/ad-library/home                                403   (bot block)
https://www.facebook.com/ads/library/                                    403   (bot block; verified separately with a headless browser)
https://ads.pinterest.com/ads-repository                                200
https://adsgallery.snap.com/                                            200
https://library.tiktok.com/ads                                          200
https://en.zalando.de/ads-repository/                                    000   (connection failure; headless shows "Site can't be reached right now")
https://ads.x.com/ads-repository                                        200

Sources rejected

The rejections matter more than the acceptances: they are what stops the next run re-adding an SEO listicle.

Rejected Why
adlibrary.com, admapix.com, adsuploader.com, swipekit.app, gemfind.com, parse.bot The first web search for Meta Ad Library scope and TikTok Commercial Content Library access returned almost nothing else. All are affiliate or SEO content marketing dated “2026 Guide”. Several stated Meta retention and TikTok access rules that are close to correct, which makes them more dangerous, not less. Every such claim was instead taken from the provider's own documentation.
techtimes.com, techcrunch.com on TikTok's 2023 ads library launch Secondary reporting where a primary page exists.
newsroom.tiktok.com/expanding-tiktoks-research-api-and-commercial-content-library The URL resolves to the newsroom index, which on 2026-09-08 served an unrelated post dated 4 September 2026. A URL that silently becomes a different article cannot be cited.
Any claim about Reddit, or about ICWSM/CHI/FAccT publication volumes Out of scope, and not verified. The page says eleven papers is a count for seven venues and does not guess at the wider field's size.
Search-engine summaries of the Commission's political-ads implementing act The summary named Implementing Reg. (EU) 2026/818 correctly, but the date, OJ citation and entry-into-force were taken from EUR-Lex directly.
Search-engine summaries of the Google BigQuery ads datasets Every non-Google result for google_ads_transparency_center was affiliate content. The dataset id and table name come from Google's own Cloud blog; the all-ads dataset from a peer-reviewed paper that queried it.

What could not be established

  • Retention windows for most of the thirteen repositories. Only Meta's and Google's state one in their own documentation. The page leaves those cells “not stated” rather than inferring one year from DSA Art. 39, because Art. 39 sets a floor and platforms may exceed it.
  • Zalando's repository. Connection failure to curl and an error page to a headless browser. The page makes no claim about it; it is grouped in a row with the retail platforms and flagged in the footnote.
  • Whether the European repository for online political advertisements is live and queryable. Implementing Reg. (EU) 2026/818 specifies its structure; whether a portal exists that a researcher can pull from was not established, and the page says only that it is specified.
  • Whether X's CSV is complete or usable. The download was not exercised. The page claims only that the page offers it without an application.
  • The google-ads-transparency-center BigQuery dataset could not be independently confirmed. The Cloud Console marketplace listing is an authenticated JavaScript application. The page relies on [1Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)], which queried the dataset and names two of its fields, and says openly that it does so.
  • API availability for LinkedIn, Bing, Snap, Pinterest, Booking, AliExpress, Zalando and Amazon. Only the Mozilla report's March 2024 assessment covers these, and it is cited as of that date. Amazon publishes ad-library API docs (HTTP 200), which was not investigated further.
  • Whether any paper used an archive without naming it. The full-text and structured-field probes agree, which is weak evidence of completeness, not proof.
  • The 20%-free-text-stability and 0.9%-unlocatable-quote figures quoted from corpus were measured on the previous, 4,322-paper run and have not been re-measured. They are treated here as orders of magnitude.
  • Whether DSA Article 39 itself is under live challenge. A reviewer reported, from EUR-Lex's “Affected by case” list for CELEX 32022R2065, three pending General Court actions specifically challenging Art. 39's applicability — T-138/24 and T-139/24 (Aylo Freesites) and T-486/24 (NKL Associates) — with interim relief refused and an appeal dismissed as C-511/24 P(R). This was not independently verified: EUR-Lex serves HTTP 202 with an empty body to curl, the ALL-view page did not render the case list in a headless browser, and InfoCuria's result list is JavaScript-loaded and did not populate. The claim is therefore not on the content page, and closing it is a TODO. Note the contrast with Amazon's litigation, which is on the page because the Court's own press release is a fetchable PDF.
  • Temu's ad repository contents. https://www.temu.com/ads-repository returns HTTP 200 to curl (2,903 bytes) but a headless browser is redirected to a bot-verification challenge. The page therefore says the path exists and makes no claim about what is in it. Repository URLs for Shein and Roblox were guessed and returned 404; no URL was established for either.
  • Whether X's repository has been remediated since the December 2025 decision. A reviewer reported, from secondary sources, an appeal filed 16 February 2026 and an action plan submitted 16 July 2026 judged “partially adequate”. None of that was verified against a primary source and none of it is on the page.

TODOs

  • Exercise X's Commercial Communications CSV and record its size, fields and completeness. If it is what the page says it is, it is the cheapest large ad dataset on this wiki and the page currently under-sells it because nobody has tried.
  • Check whether the European repository portal exists and is queryable; re-fetch the Commission's political-advertising page in a few months.
  • Re-verify the thirteen repository statuses at the next refresh. Six months is long enough for two of them to move.
  • Measure DSA-repository completeness for commercial ads. Every error figure on the page is about a political classifier. [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)]'s 48% is about matching, not about ingestion.
  • Reconcile the archive count with platforms at every future refresh. The parent's 69 was corrected in this sitting; the correction will decay if either page is refreshed alone.

Changes made to other pages in this sitting

Page Change
bibliography 10 entries added before </bibtex>; rev 1788543026 → 1788865485. Duplicate-key scan case-insensitive over all 860 entries: none. DOI and title scan against the existing file for all 10: none. Cache purged with ?purge=true before publishing the citing page.
platforms the “69 papers name an Ad Library or ad archive” figure corrected in four places to the audited 11 (the Ways In route table, the Which Methods Are Current row, the researcher-routes paragraph and the per-company-pages section), with the homonym named in a <WRAP tip>; the new page linked from The Ways In, Which Methods Are Current, Should There Be Facebook, Twitter, TikTok and Amazon Pages?, Open Questions and Related Pages — that page had deliberately named it without linking so as not to create a fresh red link. Only the ad-archive figure was re-derived; the rest of that page's numbers were not re-run, and the edit says so on the page.
roadmap the design:platforms:ad_archives row removed from Queued — required, not cosmetic: scripts/sitemap.mjs reports a written-but-still-queued row as a failure and exits non-zero. A disposition row with the corrected count was added to Assessed.
roadmap a new section 3b recording the dequeue and the count correction, and the general lesson that a candidate-set count can over-count by a factor of six when the probe term is a homonym.
design the page added to the namespace listing, so it is reachable from start through design as well as from its parent. start itself lists namespaces rather than pages and needed no edit.

Rendered-page checks

Run with ?purge=true after each save, against the rendered DOM rather than the source.

Check Content page This page
bibtex_citekey spans 202 (= 2 per marker × 101 markers; the plugin emits two) 121
distinct {[key]} markers in source 15 13 (plus one deliberately '' -escaped literal, which correctly did **not** become a citation) | | entries in the rendered ''bibtex_references'' list | **15** | **13** | | tables | 10 | 18 | | ''<WRAP>'' boxes rendered | 5 | 2 (the other two ''<WRAP>'' strings on this page are ''''-escaped literals inside table cells, and both rendered as text without truncating the table — worth recording, because inline monospace alone does **not** escape a plugin tag) | | ''<pre>'' blocks | 0 | 8 (4 ''<file>'' + 4 ''<code>'') | | footnote open/close/rendered | 25 / 25 / 25 | — | | red links (''wikilink2'') | **0** | **0** | | dangling in-page anchors | 0 | 0 | Two things that check out only if you look at the right thing: * **A pipe inside a table cell survives if it is inside a '''' span.** Five rows across the two pages put a regex containing ''|'' in a cell as ''''...''''; each rendered as one cell with the pipe intact. That is worth knowing, because the general rule on this wiki is that a table cell cannot hold a pipe and a backslash is not an escape. * **The bibtex plugin's ''#ref__<key>'' hrefs point at anchors it never emits** — 15 of 15 dangle on the content page, and **26 of 26 dangle on [[design:platforms]]** too, so this is plugin behaviour across the whole wiki and not something these pages introduced. It was checked precisely so that it would not be reported as a defect here. ===== Review log ===== Four reviewers, each handed the page text, the report script, its unedited output, the quote-check output and these notes, and each told explicitly that the author's context might not be exhaustive. The three focused passes ran in parallel first; the generic pass ran after their findings were applied. ==== Pass 1 — figures against the script (''model: sonnet'') ==== Handed the page text, both scripts, both outputs and these notes. Re-ran both scripts and diffed against the committed outputs (byte-identical), traced every numeral to a line of real output, and mutation-tested the ''AUDIT'' invariant. ^ # ^ Finding ^ Verdict ^ Action ^ | 1 | **"X Ads Repository: 0 papers" contradicts the script**, which prints ''1'' for X in Probe 4. {[benzaamia2026_year]} has a section headed //10.4 X's Ad Repository// and matched "//1,458//" of "//2,190//" repository entries by Tweet ID. The 0 was repeated in five places, including the headline "the zeros are the finding" claim. | **Accepted — the worst error in the run.** | Fixed in all five places; the X row now reads 1, and the "nobody but Meta and Google" claim is replaced by the script-produced Meta 8 / Google 3 / X 1 / Apple 1 split. | | 2 | **The page contradicted itself**: its Apple row said 1 paper while a sentence nine lines later said "none used any of the other eleven DSA repositories". {[breuer2026_ad]} genuinely used Apple's repository and, uniquely, found no missing entries in a sampled check. | **Accepted.** | Fixed, and Apple's positive result was promoted into a ''<WRAP tip>'' — it is the only counter-example on a page that otherwise reads as uniformly negative. | | 3 | "**six of them cite {[edelson2020_security]} and nothing else**" — the audit shows **three**; the other five cite //Longitudinal Analysis of Android Ad Library Permissions//. | **Accepted.** | Corrected, and the correction makes the homonym point sharper rather than weaker. | | 4 | The new access section claimed "**the two with no friction at all have none** [of the published research]" — but Google's frictionless BigQuery route has 3 papers and X's has 1. | **Accepted.** | The friction-versus-usage claim is reversed: access friction does //not// explain the distribution. | | 5 | {[bryson2025_characterizing]} taxonomises **22** systems but its usability arm covers **eight**, "//25 per ATS//" for 200 participants; the page attached the 22 to the usability finding. | **Accepted.** | Reworded, and the eight-system phrase added to the quote checker. | | 6 | The Meta/Google split was editorial, not script-produced, and a regex over full text cannot reproduce it: {[zeng2021_polls]}'s "//Google's (or others') political ad transparency reports//" defeats the name pattern, and {[benzaamia2026_year]} uses four platforms at once. | **Accepted.** | An ''ARCHIVE_USED'' hand-assignment map with a throwing invariant was added to the script, so the split on the page now traces to script output. | | 7 | The "growing: 2 in 2020, 4 in 2023–2024, 3 in 2025–2026" phrasing skipped the two single-paper years and presented a bumpy series as a growth curve. | **Accepted.** | Replaced with the full per-year series and an explicit "no trend worth naming". | | 8 | Three quotes from {[breuer2026_ad]} added to the page after the checker was written were not covered by it, so "51 found, 0 not found" no longer covered every quote. | **Accepted.** | Sixteen strings added, taking the checker to 67. Adding them exposed a **duplicate ''CHECKS'' key** that silently dropped three existing checks — a guard now aborts on that, and was mutation-tested. | | 9 | The {[venkatadri2019_auditing]} row is outside both scripts' coverage; the reviewer hand-verified it. | **Accepted as a coverage gap.** | Its three figures added to the checker. | | — | The ''AUDIT'' invariant fires in both directions when mutated; counts are of papers not tuples; the case-sensitive split reproduces exactly (25 + 44 = 69); no arithmetic errors. | Confirmed. | No action. | ==== Pass 2 — citations and quotes (''model: sonnet'') ==== ^ # ^ Finding ^ Verdict ^ Action ^ | 1 | All 15 citekeys resolve, no duplicate or case-colliding keys across all 860 bibliography entries, no literal ''@'' in a field. All 10 new entries verified against each paper's own PDF first page for author list **and order**, plus independently against Crossref; all 7 DOIs resolve; both USENIX landing pages return 200 with matching author order. | Confirmed. | No action. | | 2 | **{[breuer2026_ad]} is bucketed under "your own ad campaign" and it never buys an ad.** Its method is persona accounts plus independent observation, then matching against the archive. {[gkiouzepi2023_collaborative]} and {[venkatadri2019_auditing]} genuinely fit that row. | **Accepted.** Checked directly: ''/as an advertiser/%% returns 0 hits in [1Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)] and 1 in [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] — “Acting as an advertiser, we created a YouTube campaign targeting users in France”. The own-campaign paper was [2Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)] all along. Swapped in two tables, and a measured-results row that credited the location-only-campaign finding to the wrong paper was split in two.
3 Amazon's absence from Mozilla's tested set has a reason the page omits: “We did not examine Amazon Store because Amazon had been granted an exemption from making its ad repository publicly available by the European Court of Justice” (Case T-367/23 R). Accepted, and it opened a better line than the reviewer proposed. Following it to the Court's own press release showed the interim order was set aside on 27 March 2024, so Amazon does publish a repository — a free API with no papers using it. Amazon promoted to its own table row; the litigation history added to the Art. 39 section; the Mozilla footnote now states the reason.
4 Mozilla's “our accuracy testing…” quote is silently lower-cased mid-sentence without brackets. Accepted as trivial. Left as is; it is standard practice and changes no meaning.
5 Two Google/X quotes could not be verified by the reviewer because both pages are JavaScript-only. Not a finding against the page. Recorded here as a tooling limit. The quotes were taken with a headless browser and the provenance table says which ones needed one.

Pass 3 — external currency (''model: sonnet'')

The most productive pass. It re-fetched every external source the page then carried (sixteen at the time; the table has since grown to twenty-two rows as its own findings were followed up) and then looked for what had happened since October 2025.

# Finding Verdict Action
1 X was fined €120 million on 5 December 2025 and its ad repository is one of the three grounds — the page described the same repository as “current, least gated” with no mention of it. Accepted; the largest omission found. Verified against the Commission's own page. A whole new subsection, The enforcement record is a published error budget, and the Which Methods Are Current row rewritten.
2 TikTok's repository was preliminarily found in breach on 15 May 2025, and binding commitments were accepted on 5 December 2025. Accepted. Added to the same subsection, with the consequence spelled out: a study spanning December 2025 has an instrument change in the middle, not a platform change.
3 AI Forensics audited Snapchat's Ad Gallery (13 May 2026) and found it “unavailable 99.7% of the time”, fixed on 29 April 2026. This also falsifies the page's claim that Mozilla 2024 was the only side-by-side look. Accepted. Added, labelled as a research organisation's report rather than a regulator's decision.
4 The designation list has grown to 28 active designations (“Information updated on 31 August 2026”), adding Shein, Temu, WhatsApp, Reddit, Roblox and ChatGPT-as-VLOSE; Temu already serves an /ads-repository path. So “thirteen” is a floor. Accepted. The framing now says floor and names the additions with dates; Temu's contents are explicitly not verified (bot challenge).
5 Three pending General Court actions challenge Art. 39 itself (T-138/24, T-139/24 Aylo; T-486/24 NKL), per EUR-Lex's affected-by-case list. Not published. Could not be verified from a primary source: EUR-Lex returns 202-with-empty-body to curl, the case list did not render headless, and InfoCuria's results are JavaScript-loaded. Recorded under What could not be established with the exact reason, plus a TODO.
6 Zalando's failure mode changed between the author's fetch and the reviewer's: connection failure versus a branded HTTP 403. Accepted. The footnote now reports both observations and says the mechanism is unstable. The “no claim” abstention was already right.
7 TikTok's own docs also say the library will be “made available globally and will include ads data from EU to begin with” — an expansion plan the page does not mention. Not added. A stated intention with no date is not a fact a 2026 study can plan around, and the page already dates the EU-only scope. Recorded here as a rejection.
8 socialmediatransparency.org maintains an “Article 39 Scoreboard”. Rejected as a source. Third-party, authorship not verified, and the reviewer itself marked it COULD-NOT-VERIFY. Not cited.
All sixteen source URLs, both regulations, Impl. Reg. 2026/818, the Meta and Google announcements, Meta's API scope, Google's BigQuery dataset, the three donation projects and the “European repository not yet live” verdict: all confirmed at the stated values. Confirmed. No action.

Pass 4 — generic (''model: fable'')

Run after the three focused passes were applied, with no checklist, over the content page and this one.

# Finding Verdict Action
1 The 7-year retention is a rolling window, so the page's “the historical ads stay in the archives” is already false for its own foundational papers. Meta's own text is “delivered anywhere in the world during the past 7 years”; [10Capozzi, Arthur; Morales, Gianmarco De Francisci; Mejova, Yelena; Monti, Corrado; Panisson, André (2023): "The Thin Ideology of Populist Advertising on Facebook during the 2019 EU Elections", in: Proceedings of the ACM Web Conference. (DOI)]'s April–May 2019 ads and [3Edelson, Laura; Lauinger, Tobias; McCoy, Damon (2020): "A Security Analysis of the Facebook Ad Library", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)]'s May 2018 – June 2019 corpus have aged out. Accepted, and it produced the single best paragraph added in the whole run. Rewritten as “closing corpora, not frozen ones”, with the reproducibility lesson the reviewer identified: the foundational papers of this literature can no longer be reproduced against the live archive, so deposit the pull and not the query.
2 The two-sided-error table refutes the sentence introducing it: three of five rows have “—” in the false-positive column, and [4Le Pochat, Victor; Edelson, Laura; Van Goethem, Tom; Joosen, Wouter; McCoy, Damon; Lauinger, Tobias (2022): "An Audit of Facebook's Political Ad Policy Enforcement", in: Proceedings of the USENIX Security Symposium. (Link)]'s 0.85% US miss rate is not “large”. Accepted. Reworded to “material and strongly region-dependent — but only two of the five measured both directions”, and the asymmetry turned into the point.
3 The count violates the page's own inclusion rule. [5Gkiouzepi, Eleni; Andreou, Athanasios; Goga, Oana; Loiseau, Patrick (2023): "Collaborative Ad Transparency: Promises and Limitations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] queries no archive, the provenance page admits it is the marginal call, and the content page never said so while carrying 11 onto the parent. Accepted. The stated rule now has a third clause, and the page says outright that excluding it gives 10, so a reader who prefers the narrow rule can.
4 “Eight more mention one only in their bibliography” — five of the eight cite an SDK paper and never mention an archive at all. Accepted. Corrected to “three more”.
5 [11Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]'s 92.3% is a labelling miss, not an archive miss — under the DSA all-EU scope those ads are in the archive — and 92.3% is a bolded computed complement. Accepted; this was the subtlest finding of the run. The row now says “in the archive but not labelled political”, and the derived number is unbolded.
6 48% is stated three ways: “missing from the repository” in the table and reading list, “really 48% we could not match” in the body. Accepted. Carried as ≤48%, an upper bound everywhere it appears.
7 The friction-versus-usage argument conflates route with archive and ignores that Meta's archive is five years older than X's CSV. Accepted. Reworded to name age and the political-ads agenda as confounders, and to say the distribution is not a ranking of the archives.
8 “The four this page can confirm hands-on” — nothing was exercised. Accepted. “Confirmed from the provider's own current documentation”, plus an explicit “none of them was exercised for this page”.
9 “the cheapest bulk archive access described anywhere on this wiki” is an unchecked wiki-wide superlative that also contradicts the X row. Accepted. Superlative dropped.
10 “when they tried them all in 2024” is contradicted by its own footnote (twelve tested, Amazon not examined). Accepted. “Listed in 2024, twelve of which they tested.”
11 Parent and child publish different counts for the same DSA probe — 25 versus 29. Accepted. Checked: case-sensitive gives 25, insensitive 29, and the four extras are lowercase bibliography titles. The child now publishes 25 with a footnote naming the difference, so the two pages agree.
12 “Google's political section is worldwide” is unsupported. Accepted — it was wrong. Google's policy: “Google has different requirements for political and election advertising based on region”, with named per-region sections; Google Cloud's own figure is “over 35 countries”. Corrected, with both sources footnoted and the earlier error named in the footnote.
13 “still operating” rests on an HTTP 200 on a homepage. Accepted. “Still had live sites … which is not the same as still collecting, and neither was tested.”
14 “IEEE S&P 2026 and TheWebConf 2026 abstracts are not in OpenAlex” overstates — the corpus holds 28 and 67 of those papers respectively, and the page cites one. Accepted. Replaced with corpus's own numbers: 563 papers have a DOI and no abstract, dominated by WWW 2026 (324) and IEEE S&P 2026 (194).
15 The lead callout is wiki bookkeeping, and the same homonym story is told twice. Accepted. The lead is now four design takeaways; the count correction moved to a second, smaller box pointing at the section that already carries the audit. The zero-user list was cut from five appearances to three.
16 The access section drifts into API mechanics; and its facebook.com/ID claims were not in the provenance source table. Partly accepted. The Meta script repository was cut. The identity-and-location verification was kept — it is not API trivia, it is the fact that this access is not anonymous — and the row for that page in the source table was made explicit about which quotes come from it.
17 “Treating an archive as ground truth: superseded” is a statement about the literature dressed as a verdict, and “most of the corpus does it” is not right either. Accepted. The verdict is now “don't”, with the reason, and an explicit note that using an archive as a corpus is a different and current thing.
18 The provenance page overstates its own coverage: “every paper-sourced figure” was not true, and the reviewer listed about a dozen it missed (and hand-checked them — all correct). Accepted. Sixteen of them added to the checker (67 → 83), and the coverage claim rewritten to say what it does and does not cover. The reviewer also found that two [11Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] strings exist in no corpus rendering, only in a live pypdf extraction.
19 “every figure here is derived on that corpus, none carried over” versus two figures admitted elsewhere as carried over. Accepted. The run table now distinguishes corpus figures from figures about the corpus.
20 The published script's own comment contradicts its output — “eight are the Android SDK sense” against 7 + 1 + 1. Accepted. Published code is a citation surface. Comment corrected, and it now also says the other 35 were screened rather than read.
21 Stale counts on the provenance page: “all sixteen external sources” (22 rows now), “corrected in three places” (four). Accepted. Both corrected.
The reviewer also listed what is working and should not be touched: the two-sided-error table with a denominator per row, the denominator sentence box, the Apple positive control, the enforcement-record section, the four-instruments table, the “screened rather than audited” language, What to Read First, What to Report, and the parent edit's “only this figure was re-derived” sentence. Recorded so a later pass does not “fix” them.

What the review layer was worth

Between them the three focused passes found one figure that flatly contradicted the script it came from (X: 0 versus 1), one internal self-contradiction (Apple), one paper credited with a method it never used ([1Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)] as an advertiser), one materially wrong access claim (bulk access limited to four of thirteen), and one omission large enough to change the page's advice (the €120 million decision naming this very instrument). Two of those were only findable by reading the papers, and two only by fetching. None would have been caught by re-reading the draft.

The rejections are worth as much: TikTok's undated global-expansion intention and a third-party Art. 39 scoreboard were both offered and both declined, and a reported Art. 39 litigation thread was left off the page because it could not be reached through a primary source.

The report script

adarchives_report.mjs
#!/usr/bin/env node
// Report script for design:platforms:ad_archives.
//
//   node scripts/adarchives_report.mjs
//
// Every figure on the page comes from here, with its denominator printed next to
// it. Three things this script enforces rather than trusting the caller with:
//
//   1. "Ad Library" is a HOMONYM. In mobile-security writing an "ad library" is
//      an advertising SDK linked into an Android app; in platform-transparency
//      writing "the Ad Library" is Meta's public ad archive. The probe the
//      design:platforms page used (/Ad Library|Ads? Archive API/i) cannot tell
//      them apart, and additionally matches the substring inside "re[ad librar]y".
//      Section 1 decomposes its 69 hits so the collision is visible.
//   2. The candidate set is hand-audited paper by paper and the verdicts live in
//      this file (AUDIT below), each with a reason. Precision is printed, and
//      every rejected candidate is printed in full as the residue.
//   3. An assertion fires if AUDIT does not cover the candidate set exactly, so
//      a widened probe cannot silently acquire unaudited papers.
//
// Counts are of PAPERS, never tuples. Populations are named before numerators.
// All regexes are registered up front and evaluated in ONE pass over the full
// texts: holding 5,855 papers in memory at once costs gigabytes and minutes.
import fs from 'node:fs';
import path from 'node:path';
import { dataRoot, loadExtractions } from './lib.mjs';
 
const rows = loadExtractions();
const tp = (r) => path.join(dataRoot(), `fulltext/${r.year}/${r.venue}/${r.slug}/paper.cols.txt`);
const withText = rows.filter((r) => fs.existsSync(tp(r)));
 
const REG = new Map(); // source string -> {re, hits:Set<slug>}
const reg = (re) => { const k = re.source + '/' + re.flags; if (!REG.has(k)) REG.set(k, { re, hits: new Set() }); return re; };
const hits = (re) => REG.get(re.source + '/' + re.flags).hits;
const papers = (re) => withText.filter((r) => hits(re).has(r.slug));
 
// ---------------------------------------------------------------- the probes
const OLD_PROBE = reg(/Ad Library|Ads? Archive API/i);
const PROPER = reg(/\bAds? Librar(y|ies)\b/); // case-sensitive: proper noun
// Substring collisions the old probe also matches: "re[ad librar]y" and
// "uplo[ad librar]ies". The second one really did fire, on WWW 2024.
const READ_COLLISION = reg(/read librar/i);
const UPLOAD_COLLISION = reg(/upload librar/i);
 
const CANDIDATE = {
  'proper-noun Ad Library / Ad Libraries (case-sensitive)': reg(/\bAds? Librar(y|ies)\b/),
  'Ad Library API / Ads Archive API': reg(/\bAds? Library API\b|\bAds? Archive API\b/i),
  'ad archive / ad repository / advertising archive': reg(/\bads? archive\b|\bad repositor(y|ies)\b|\badvertis\w* (archive|repositor\w+)\b|\barchive of ads\b/i),
  'Google Ads Transparency Center / google_political_ads': reg(/\bAds? Transparency Cent(er|re)\b|\bgoogle_political_ads\b|\bGoogle('s)? Political Ad\w*\b|\bPolitical Advertising on Google\b/i),
  'TikTok Commercial Content Library / API': reg(/\bCommercial Content (Library|API)\b/i),
  'political ad archive / library / repository': reg(/\bpolitical ad(vertising|s)? (archive|library|repositor\w+)\b/i),
  'Who Targets Me': reg(/\bWho Targets Me\b/i),
};
 
const ARCHIVES = {
  'Meta / Facebook Ad Library (proper noun or API/Report)': reg(/\b(Facebook|Meta|FB)['’s]*\s+Ads? Librar(y|ies)\b|\bthe Ads? Librar(y|ies)\b|\bAds? Library (API|Report)\b/),
  'Google Ads Transparency Center / political-ads BigQuery': reg(/\bAds? Transparency Cent(er|re)\b|\bgoogle_political_ads\b|\bGoogle('s)? political ad(vertising)? (archive|library|repositor\w+|transparency)\b/i),
  'TikTok Commercial Content Library / API': reg(/\bCommercial Content (Library|API)\b/i),
  'X / Twitter ads repository or ad transparency centre': reg(/\b(X|Twitter)['’s]*\s+(ads? )?(repositor\w+|transparency cent(er|re))\b|\bTwitter Ads? Transparency\b/i),
  'LinkedIn Ad Library': reg(/\bLinkedIn['’s]*\s*Ads? Librar\w+\b|\bLinkedIn ads? (archive|repositor\w+)\b/i),
  'Snapchat Ads Gallery / political ads library': reg(/\bSnap(chat)?['’s]*\s*(Political )?Ads? (Librar\w+|Gallery|archive|repositor\w+)\b/i),
  'Apple App Store Ad Repository': reg(/\bApple['’s]*\s*(App Store )?Ads? (Repositor\w+|Librar\w+)\b|\badrepository\.apple\.com\b/i),
  'Microsoft / Bing Ad Library': reg(/\bBing Ads? Librar\w+\b|\badlibrary\.ads\.microsoft\.com\b/i),
  'Amazon Store Ad Library API': reg(/\bAmazon['’s]*\s*(Store )?Ads? Librar\w+\b|\badvertising-api-eu\.amazon\.com\b/i),
  'Pinterest / Booking.com / AliExpress / Zalando ads repository': reg(/\b(Pinterest|Booking\.com|AliExpress|Zalando)['’s]*\s*ads? repositor\w+\b/i),
  'the EU "European repository for online political advertisements"': reg(/European repositor\w+ for online political advertis\w+/i),
  // the bare "Article 39" probe collides: CCS 2018 image-disguising and PETS
  // 2026 banned-books are not about ad repositories. The DSA-context probe
  // below is the one to read.
  'DSA Article 39 (any mention of the article number, COLLIDES)': reg(/\bArticle 39\b|\bArt\.? 39\b/i),
  'Regulation (EU) 2024/900 (political-advertising regulation, TTPA)': reg(/2024\/900|Transparency and Targeting of Political Advertising/i),
  'Digital Services Act (any mention)': reg(/Digital Services Act/i),
  'DSA Article 39 within 400 chars of "Digital Services Act" or "DSA"': reg(/Digital Services Act[\s\S]{0,400}Article 39|Article 39[\s\S]{0,400}Digital Services Act|\bDSA\b[\s\S]{0,200}Article 39|Article 39[\s\S]{0,200}\bDSA\b/i),
  // hand-audited 2026-09-08: the single hit is two bibliography titles in
  // lepochat2022 ("Tradeoffs in Automated Political Advertising Regulation",
  // "...the regulation of political advertising"), not a discussion of the EU
  // Regulation. Read this row as ZERO papers.
  'the EU political-advertising regulation, any phrasing (1 hit = bibliography titles only)': reg(/political advertising regulation|regulation on [\w\s]{0,25}political advertising|Political Ads Regulation|Transparency and Targeting of Political Advertising|2024\s*\/\s*900/i),
};
 
const DONATION = {
  'Who Targets Me': reg(/\bWho Targets Me\b/i),
  'Ad Observer / NYU Ad Observatory': reg(/\bAd Observer\b|\bAd Observator\w+\b/i),
  'Australian Ad Observatory': reg(/\bAustralian Ad Observator\w+\b/i),
  'data donation (any phrasing)': reg(/data donation|donated (their )?data|donation (extension|panel)/i),
};
 
// The 44 papers that contain only the lowercase "ad library" are OUTSIDE the
// candidate set by construction (the candidate term is case-sensitive), so they
// were never hand-audited one by one. Section 1b screens them instead: any
// lowercase mention that sits within 80 characters of a transparency-archive
// word is printed for a human to read. Nine fired and all nine were read: seven
// are the Android SDK sense, one is a lowercase bibliography entry (USENIX's
// small-caps reference style lowercases titles) and one is a pure substring
// collision in a paper about Node.js file uploads ("uplo[ad librar]ies"). No
// archive user hides in the 44 -- but the other 35 were screened, not read.
const LOWERCASE_SUSPECT = reg(/(facebook|meta|google|political|transparency|archive|repositor)[^.]{0,80}ad librar|ad librar[^.]{0,80}(facebook|meta|google|political|transparency|archive|repositor)/i);
 
const REJECTED = {
  '/TTPA/i': [reg(/TTPA/i), 'matches the substring inside "httpa..." (HTTPA, "http api"). Useless as written; the spelled-out name is probed above instead.'],
  '/\\bFORT\\b|Facebook Open Research/': [reg(/\bFORT\b|Facebook Open Research/), 'FORT collides with unrelated all-caps tokens and OCR artefacts. Never hand-audited, never used.'],
  '/Google Transparency Report/i': [reg(/Google Transparency Report/i), 'that report covers government requests and HTTPS adoption, not ads. It put censorship and TLS papers into an ad-archive candidate set.'],
};
 
// ------------------------------------------------------------------ one pass
for (const r of withText) {
  const t = fs.readFileSync(tp(r), 'utf8').replace(/\s+/g, ' ');
  for (const { re, hits: h } of REG.values()) if (re.test(t)) h.add(r.slug);
}
 
const perYear = (list) => Object.entries(list.reduce((a, r) => ((a[r.year] = (a[r.year] || 0) + 1), a), {}))
  .sort().map(([y, n]) => `${y}:${n}`).join(' ');
const h1 = (s) => console.log(`\n${'='.repeat(78)}\n${s}\n${'='.repeat(78)}`);
 
h1('0. DENOMINATORS');
console.log(`extracted papers in data/extract/run1        : ${rows.length}`);
console.log(`  of which have paper.cols.txt (probe base)  : ${withText.length}`);
console.log(`papers that drew a study population          : ${rows.filter((r) => r.population.length > 0).length}`);
// The advertising-measurement floor. detection.phenomenon is free text and only
// ~20% stable run-to-run, so this is a candidate set and a FLOOR, never a
// population: a paper that measured ads without using an ad-shaped word here is
// invisible to it.
const AD_PHENOM = /\badvert|\bads?\b|ad[- ](network|delivery|targeting|tech|exchange|impression|creative|auction)|targeted ad|programmatic|header bidding|\brtb\b/i;
const adPapers = rows.filter((r) => (r.detection || []).some((d) => d.phenomenon && AD_PHENOM.test(d.phenomenon)));
const adSlugs = new Set(adPapers.map((r) => r.slug));
console.log(`papers with an advertising-shaped detection.phenomenon (candidate set, FLOOR): ${adPapers.length}`);
console.log(`  per year: ${perYear(adPapers)}`);
 
h1('1. THE PROBE design:platforms USED, DECOMPOSED');
const oldHits = papers(OLD_PROBE);
console.log(`/Ad Library|Ads? Archive API/i over ${withText.length} full texts : ${oldHits.length} papers`);
console.log('  (this is the published "69 papers name an Ad Library or ad archive" figure)');
console.log(`  proper-noun "Ad Library" present (capital L)   : ${oldHits.filter((r) => hits(PROPER).has(r.slug)).length}`);
console.log(`  only lowercase "ad library/libraries"          : ${oldHits.filter((r) => !hits(PROPER).has(r.slug)).length}  <- screened in section 1b`);
console.log(`  full texts containing "read librar" (a substring the probe also matches)   : ${papers(READ_COLLISION).length}`);
console.log(`  full texts containing "upload librar" (ditto -- this one fires in the 69)  : ${papers(UPLOAD_COLLISION).length}`);
console.log('  A capital L is not sufficient on its own: bibliography entries for a paper');
console.log('  titled "Longitudinal Analysis of Android Ad Library Permissions" capitalise too.');
 
h1('1b. SCREENING THE 44 LOWERCASE-ONLY PAPERS');
const lowOnly = oldHits.filter((r) => !hits(PROPER).has(r.slug));
const suspects = lowOnly.filter((r) => hits(LOWERCASE_SUSPECT).has(r.slug));
console.log(`lowercase-only papers                                  : ${lowOnly.length}`);
console.log(`  of those, a lowercase mention near an archive word    : ${suspects.length}  (each read by hand)`);
const LOWERCASE_VERDICTS = {
  'obfuscation-resilient-privacy-leak-detection-for-mobile-apps-through-differentia': 'sdk — emulator detection by ad libraries',
  'a-multi-modal-neural-embeddings-approach-for-detecting-mobile-counterfeit-apps': 'sdk — "ad library difference" between app pairs',
  'this-sneaky-piggy-went-to-the-android-ad-market-misusing-mobile-sensors-for-stea': 'sdk — "Google\'s interstitial ad library"',
  'the-abuser-inside-apps-finding-the-culprit-committing-mobile-ad-fraud': 'sdk — providers "harness their ad libraries" to commit fraud',
  'adcube-webvr-ad-fraud-and-practical-confinement-of-third-party-ads': 'sdk — "a third-party JS ad library in Secure ECMAScript"',
  'problematic-advertising-and-its-disparate-exposure-on-facebook': 'REFERENCE — "A security analysis of the facebook ad library" in a USENIX small-caps bibliography entry. Not an archive user; it uses the Ad Observer donation extension (see section 6).',
  'in-security-of-file-uploads-in-node-js': 'NOT AN AD LIBRARY — "popular file uplo[ad librar]ies (e.g., express-fileupload)". A pure substring collision, and proof that the old probe counts papers with no advertising content at all.',
  'careful-about-what-app-promotion-ads-recommend-detecting-and-explaining-malware-promotion-via-app-promotion-graph': 'sdk — "Google AdMob, which is the most widely used ad library"',
  'cross-boundary-mobile-tracking-exploring-java-to-javascript-information-diffusion-in-webviews': 'sdk — "Google Ads\' dominance, as a standalone ad library"',
};
const unverdicted = suspects.filter((r) => !LOWERCASE_VERDICTS[r.slug]);
if (unverdicted.length) throw new Error(`lowercase suspects with no hand verdict: ${unverdicted.map((r) => r.slug).join(', ')}`);
for (const r of suspects) console.log(`    ${r.venue} ${r.year} ${r.slug.slice(0, 46)}\n        ${LOWERCASE_VERDICTS[r.slug]}`);
const lowRefs = suspects.filter((r) => LOWERCASE_VERDICTS[r.slug].startsWith('REFERENCE'));
const lowNotAd = suspects.filter((r) => LOWERCASE_VERDICTS[r.slug].startsWith('NOT AN AD'));
console.log(`\n  -> of the ${lowOnly.length} lowercase-only papers, ${suspects.length} were read by hand:`);
console.log(`       ${suspects.length - lowRefs.length - lowNotAd.length} Android advertising SDK, ${lowRefs.length} bibliography entry, ${lowNotAd.length} substring collision with no advertising content.`);
console.log(`     the other ${lowOnly.length - suspects.length} have no transparency-archive word within 80 characters of the mention.`);
console.log('  -> That is a SCREEN, not an audit. It found no ad-archive user among the 44, and it is the only');
console.log('     recall evidence for the lowercase-only stratum.');
 
h1('2. CANDIDATE PROBE (per term, then unioned)');
for (const [k, re] of Object.entries(CANDIDATE)) console.log(`${String(papers(re).length).padStart(4)}  ${k}`);
const candidates = withText.filter((r) => Object.values(CANDIDATE).some((re) => hits(re).has(r.slug)));
console.log(`\nUNION (candidate set)                        : ${candidates.length} papers`);
console.log(`  per year: ${perYear(candidates)}`);
 
h1('3. HAND-AUDIT OF EVERY CANDIDATE');
// verdict: 'instrument'  = the archive is a data source, or the archive itself
//                          is the object measured
//          'sdk'         = "ad library" means an Android advertising SDK
//          'reference'   = the archive appears only in the bibliography or a
//                          related-work sentence; nothing was collected from it
//          'other-sense' = an ad archive in some other sense (a stimulus corpus,
//                          or the ad ecosystem's own inventory)
const AUDIT = {
  'the-cost-of-annoying-ads': ['other-sense', 'Adverlicious, "an online display advertising archive", supplied ad stimuli for a crowdworker experiment; not a platform transparency archive'],
  'the-impact-of-vendor-customizations-on-android-security': ['reference', 'bibliography entry: Book et al., "Longitudinal Analysis of Android Ad Library Permissions"'],
  'reconciling-mobile-app-privacy-and-usability-on-smartphones-could-user-privacy-p': ['reference', 'same bibliography entry'],
  'introducing-privacy-threats-from-ad-libraries-to-android-users-through-privacy-g': ['sdk', 'the paper is about Android advertising SDKs; its own detection.phenomenon is "Ad-library prevalence"'],
  'automated-experiments-on-ad-privacy-settings': ['other-sense', '"the ad repository lacking ads relevant to the interest" - the ad ecosystem\'s inventory, not an archive. The paper is a Google Ad Settings audit'],
  'identifying-open-source-license-violation-and-1-day-security-risk-at-large-scale': ['reference', 'same bibliography entry'],
  'keep-me-updated-an-empirical-study-of-third-party-library-updatability-on-androi': ['reference', 'same bibliography entry'],
  'beyond-google-play-a-large-scale-comparative-study-of-chinese-android-app-market': ['sdk', 'Android advertising SDKs'],
  'nomoads-effective-and-efficient-cross-app-mobile-ad-blocking': ['sdk', 'Android advertising SDKs'],
  'panoptispy-characterizing-audio-and-video-exfiltration-from-android-applications': ['reference', 'same bibliography entry'],
  'deepintent-deep-icon-behavior-learning-for-detecting-intention-behavior-discrepa': ['sdk', 'Android advertising SDKs'],
  'nomoats-towards-automatic-detection-of-mobile-tracking': ['sdk', 'Android advertising SDKs'],
  'towards-https-everywhere-on-android-we-are-not-there-yet': ['sdk', 'a table column headed "Ad Library" listing AdMob, Facebook Audience Network, Unity - SDKs'],
  'facebook-ads-monitor-an-independent-auditing-system-for-political-ads-on-faceboo': ['instrument', 'built an independent collector and compared it against the Facebook Ad Library for Brazil'],
  'maddroid-characterizing-and-detecting-devious-ad-contents-for-android-apps': ['sdk', 'a section headed "Mobile Ad Library Detection and Analysis"'],
  'polls-clickbait-and-commemorative-2-bills-problematic-political-advertising-on-n': ['instrument', 'crawled 1,000 political ads from the Google political ad archive to balance a classifier, and cites Facebook ad-archive incompleteness'],
  'a-security-analysis-of-the-facebook-ad-library': ['instrument', 'the archive itself is the object: Ad Library API + Ad Library Report, 3,685,558 ads'],
  'analyzing-the-monetization-ecosystem-of-stalkerware': ['sdk', 'a figure axis reading "# of Ad Libraries" - Android SDKs per app'],
  'charting-app-developers-journey-through-privacy-regulation-features-in-ad-networ': ['reference', 'bibliography entry for Edelson et al. 2020 only'],
  'an-audit-of-facebooks-political-ad-policy-enforcement': ['instrument', 'Ad Library web portal + API + snapshot tool; 33.8 M ads observed'],
  'conspiracy-brokers-understanding-the-monetization-of-youtube-conspiracy-theories': ['reference', 'bibliography entry for Edelson et al. 2020 only'],
  'marketing-to-children-through-online-targeted-advertising-targeting-mechanisms-a': ['reference', 'bibliography entry for Edelson et al. 2020; the study measures ad explanations, not the archive'],
  'propaganda-politica-pagada-exploring-u-s-political-facebook-ads-en-espanol': ['instrument', 'Facebook Ad Library API, 5.2 M ads narrowed to 4.7 M'],
  'the-thin-ideology-of-populist-advertising-on-facebook-during-the-2019-eu-electio': ['instrument', 'Meta Ad Library API, 44,949 ads from 57 party pages'],
  'beyond-the-guidelines-assessing-metas-political-ad-moderation-in-the-eu': ['instrument', 'Meta Ad Library API under the DSA all-EU-ads expansion'],
  'scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns': ['instrument', 'secondary use: analysed fraudulent-shop ads drawn from the Facebook Ad Library'],
  'a-year-under-the-dsa-ad-transparencys-uneven-landscape': ['instrument', 'the DSA ad repositories of four platforms are the object; the collection instrument is donation (Who Targets Me)'],
  'ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis': ['instrument', 'Google Ads Transparency Center via the BigQuery public dataset, used as intended ground truth and found deficient'],
  'collaborative-ad-transparency-promises-and-limitations': ['instrument', 'ad libraries are the baseline being critiqued; proposes collaborative inference as the alternative'],
};
// Which archive each audited paper actually used, assigned by hand from the
// paper's own methods section. A regex over full text CANNOT produce this
// split: zeng2021 writes "Google's (or others') political ad transparency
// reports", whose parenthetical defeats a \bGoogle('s)? political ad\b
// pattern, and benzaamia2026 studies four platforms' repositories at once.
// Keyed by slug so the invariant below forces every audited paper to have one.
const ARCHIVE_USED = {
  'a-security-analysis-of-the-facebook-ad-library': ['Meta'],
  'facebook-ads-monitor-an-independent-auditing-system-for-political-ads-on-faceboo': ['Meta'],
  'polls-clickbait-and-commemorative-2-bills-problematic-political-advertising-on-n': ['Google'],
  'an-audit-of-facebooks-political-ad-policy-enforcement': ['Meta'],
  'collaborative-ad-transparency-promises-and-limitations': ['(none: critiques ad libraries, queries Facebook Ads Manager instead)'],
  'propaganda-politica-pagada-exploring-u-s-political-facebook-ads-en-espanol': ['Meta'],
  'the-thin-ideology-of-populist-advertising-on-facebook-during-the-2019-eu-electio': ['Meta'],
  'beyond-the-guidelines-assessing-metas-political-ad-moderation-in-the-eu': ['Meta'],
  'scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns': ['Meta'],
  'a-year-under-the-dsa-ad-transparencys-uneven-landscape': ['Meta', 'Google', 'X'],
  'ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis': ['Google', 'Apple'],
};
 
const cslugs = new Set(candidates.map((r) => r.slug));
const missing = [...cslugs].filter((s) => !AUDIT[s]);
const extra = Object.keys(AUDIT).filter((s) => !cslugs.has(s));
if (missing.length || extra.length) {
  throw new Error(`AUDIT does not match the candidate set.\n  unaudited candidates: ${JSON.stringify(missing, null, 1)}\n  audit rows with no candidate: ${JSON.stringify(extra, null, 1)}`);
}
console.log(`invariant OK: all ${candidates.length} candidates carry a hand verdict\n`);
const byVerdict = {};
for (const r of candidates) (byVerdict[AUDIT[r.slug][0]] ||= []).push(r);
for (const [v, list] of Object.entries(byVerdict).sort((a, b) => b[1].length - a[1].length)) console.log(`${String(list.length).padStart(3)}  ${v}`);
const instrument = (byVerdict['instrument'] || []).sort((a, b) => a.year - b.year || a.venue.localeCompare(b.venue));
console.log(`\nPRECISION of the candidate probe for "archive as instrument or object": ${instrument.length}/${candidates.length} = ${(100 * instrument.length / candidates.length).toFixed(1)}%`);
const oldGenuine = instrument.filter((r) => hits(OLD_PROBE).has(r.slug)).length;
console.log(`PRECISION of the old /Ad Library|Ads? Archive API/i probe for the same question: ${oldGenuine}/${oldHits.length} = ${(100 * oldGenuine / oldHits.length).toFixed(1)}%`);
console.log(`  audited-genuine papers the old probe MISSED: ${instrument.filter((r) => !hits(OLD_PROBE).has(r.slug)).map((r) => `${r.venue} ${r.year} ${r.slug.slice(0, 40)}`).join('; ') || 'none'}`);
console.log('\n--- REJECTED CANDIDATES, IN FULL (the residue) ---');
for (const r of candidates.filter((x) => AUDIT[x.slug][0] !== 'instrument').sort((a, b) => a.year - b.year)) {
  console.log(`  ${AUDIT[r.slug][0].padEnd(12)} ${r.venue} ${r.year} ${r.slug}\n        ${AUDIT[r.slug][1]}`);
}
 
h1('4. THE AUDITED SET: an ad-transparency archive as instrument or as object');
console.log(`papers: ${instrument.length} of ${withText.length} with full text = ${(100 * instrument.length / withText.length).toFixed(2)}%`);
const both = instrument.filter((r) => adSlugs.has(r.slug)).length;
console.log(`  also inside the ${adPapers.length}-paper advertising candidate set: ${both} = ${(100 * both / adPapers.length).toFixed(1)}% of it`);
console.log(`per year: ${perYear(instrument)}`);
console.log(`per venue: ${Object.entries(instrument.reduce((a, r) => ((a[r.venue] = (a[r.venue] || 0) + 1), a), {})).sort((a, b) => b[1] - a[1]).map(([v, n]) => `${v}:${n}`).join(' ')}`);
for (const r of instrument) console.log(`  ${r.venue.padEnd(8)} ${r.year}  ${r.slug}`);
 
h1('4b. WHICH ARCHIVE EACH AUDITED PAPER ACTUALLY USED (hand-assigned)');
const noArchive = instrument.filter((r) => !ARCHIVE_USED[r.slug]);
if (noArchive.length) throw new Error(`audited papers with no ARCHIVE_USED assignment: ${noArchive.map((r) => r.slug).join(', ')}`);
const byArchive = {};
for (const r of instrument) for (const a of ARCHIVE_USED[r.slug]) (byArchive[a] ||= []).push(r);
for (const [a, list] of Object.entries(byArchive).sort((x, y) => y[1].length - x[1].length)) {
  console.log(`${String(list.length).padStart(3)}  ${a}`);
  for (const r of list) console.log(`       ${r.venue} ${r.year} ${r.slug.slice(0, 56)}`);
}
console.log(`\n  a paper can use more than one, so these do not sum to ${instrument.length}:`);
console.log(`  ${Object.entries(byArchive).sort((x, y) => y[1].length - x[1].length).map(([a, l]) => `${a} ${l.length}`).join(', ')}`);
console.log('  archives with ZERO audited users: TikTok Commercial Content Library, LinkedIn, Microsoft/Bing,');
console.log('  Snapchat, Amazon, Pinterest, Booking.com, AliExpress, Zalando, and the EU political-ad repository.');
 
h1('5. WHICH ARCHIVE, BY NAME (the zeros are the finding)');
for (const [k, re] of Object.entries(ARCHIVES)) {
  const h = papers(re);
  console.log(`${String(h.length).padStart(4)}  ${k}  [${perYear(h) || 'none'}]`);
  if (h.length && h.length <= 8) for (const r of h) console.log(`        ${r.venue} ${r.year} ${r.slug.slice(0, 56)}`);
}
 
h1('6. THE COMPLEMENT: DONATED AND CROWDSOURCED AD CORPORA');
for (const [k, re] of Object.entries(DONATION)) {
  const h = papers(re);
  console.log(`${String(h.length).padStart(4)}  ${k}  [${perYear(h) || 'none'}]`);
  if (h.length <= 8) for (const r of h) console.log(`        ${r.venue} ${r.year} ${r.slug.slice(0, 56)}`);
}
 
h1('7. MEASURED RESULTS FROM THE AUDITED SET, WITH THE PAPERS OWN DENOMINATORS');
for (const r of instrument) {
  console.log(`\n--- ${r.venue} ${r.year} ${r.slug}`);
  for (const p of r.population || []) console.log(`    POPULATION n=${p.n} unit=${p.unit} source=${JSON.stringify(p.sourceList)} sampling=${p.samplingMethod}`);
  for (const d of r.detection || []) {
    if (d.prevalence === null || d.prevalence === undefined) continue;
    console.log(`    ${d.phenomenon}\n        metric    : ${d.metric}\n        prevalence: ${d.prevalence}\n        quote (§${d.evidence?.section}): "${d.evidence?.quote}"`);
  }
}
 
h1('8. REJECTED PROBES (recorded so they are not re-added)');
for (const [k, [re, why]] of Object.entries(REJECTED)) console.log(`${String(papers(re).length).padStart(5)} papers  ${k}\n            REJECTED: ${why}`);

The report script's output

adarchives_report-output.txt
==============================================================================
0. DENOMINATORS
==============================================================================
extracted papers in data/extract/run1        : 5859
  of which have paper.cols.txt (probe base)  : 5855
papers that drew a study population          : 5712
papers with an advertising-shaped detection.phenomenon (candidate set, FLOOR): 180
  per year: 2010:2 2011:5 2012:4 2013:3 2014:8 2015:10 2016:7 2017:7 2018:8 2019:16 2020:17 2021:11 2022:20 2023:19 2024:12 2025:22 2026:9

==============================================================================
1. THE PROBE design:platforms USED, DECOMPOSED
==============================================================================
/Ad Library|Ads? Archive API/i over 5855 full texts : 69 papers
  (this is the published "69 papers name an Ad Library or ad archive" figure)
  proper-noun "Ad Library" present (capital L)   : 25
  only lowercase "ad library/libraries"          : 44  <- screened in section 1b
  full texts containing "read librar" (a substring the probe also matches)   : 3
  full texts containing "upload librar" (ditto -- this one fires in the 69)  : 1
  A capital L is not sufficient on its own: bibliography entries for a paper
  titled "Longitudinal Analysis of Android Ad Library Permissions" capitalise too.

==============================================================================
1b. SCREENING THE 44 LOWERCASE-ONLY PAPERS
==============================================================================
lowercase-only papers                                  : 44
  of those, a lowercase mention near an archive word    : 9  (each read by hand)
    NDSS 2017 obfuscation-resilient-privacy-leak-detection-f
        sdk — emulator detection by ad libraries
    WWW 2019 a-multi-modal-neural-embeddings-approach-for-d
        sdk — "ad library difference" between app pairs
    CCS 2021 this-sneaky-piggy-went-to-the-android-ad-marke
        sdk — "Google's interstitial ad library"
    NDSS 2021 the-abuser-inside-apps-finding-the-culprit-com
        sdk — providers "harness their ad libraries" to commit fraud
    USENIX 2021 adcube-webvr-ad-fraud-and-practical-confinemen
        sdk — "a third-party JS ad library in Secure ECMAScript"
    USENIX 2023 problematic-advertising-and-its-disparate-expo
        REFERENCE — "A security analysis of the facebook ad library" in a USENIX small-caps bibliography entry. Not an archive user; it uses the Ad Observer donation extension (see section 6).
    WWW 2024 in-security-of-file-uploads-in-node-js
        NOT AN AD LIBRARY — "popular file uplo[ad librar]ies (e.g., express-fileupload)". A pure substring collision, and proof that the old probe counts papers with no advertising content at all.
    NDSS 2025 careful-about-what-app-promotion-ads-recommend
        sdk — "Google AdMob, which is the most widely used ad library"
    NDSS 2026 cross-boundary-mobile-tracking-exploring-java-
        sdk — "Google Ads' dominance, as a standalone ad library"

  -> of the 44 lowercase-only papers, 9 were read by hand:
       7 Android advertising SDK, 1 bibliography entry, 1 substring collision with no advertising content.
     the other 35 have no transparency-archive word within 80 characters of the mention.
  -> That is a SCREEN, not an audit. It found no ad-archive user among the 44, and it is the only
     recall evidence for the lowercase-only stratum.

==============================================================================
2. CANDIDATE PROBE (per term, then unioned)
==============================================================================
  26  proper-noun Ad Library / Ad Libraries (case-sensitive)
   8  Ad Library API / Ads Archive API
  10  ad archive / ad repository / advertising archive
   4  Google Ads Transparency Center / google_political_ads
   0  TikTok Commercial Content Library / API
   3  political ad archive / library / repository
   4  Who Targets Me

UNION (candidate set)                        : 29 papers
  per year: 2013:2 2014:1 2015:2 2017:2 2018:3 2019:1 2020:5 2021:1 2022:4 2023:4 2024:1 2025:1 2026:2

==============================================================================
3. HAND-AUDIT OF EVERY CANDIDATE
==============================================================================
invariant OK: all 29 candidates carry a hand verdict

 11  instrument
  8  reference
  8  sdk
  2  other-sense

PRECISION of the candidate probe for "archive as instrument or object": 11/29 = 37.9%
PRECISION of the old /Ad Library|Ads? Archive API/i probe for the same question: 10/69 = 14.5%
  audited-genuine papers the old probe MISSED: PETS 2026 ad-personalization-and-transparency-in-m

--- REJECTED CANDIDATES, IN FULL (the residue) ---
  other-sense  WWW 2013 the-cost-of-annoying-ads
        Adverlicious, "an online display advertising archive", supplied ad stimuli for a crowdworker experiment; not a platform transparency archive
  reference    CCS 2013 the-impact-of-vendor-customizations-on-android-security
        bibliography entry: Book et al., "Longitudinal Analysis of Android Ad Library Permissions"
  reference    WWW 2014 reconciling-mobile-app-privacy-and-usability-on-smartphones-could-user-privacy-p
        same bibliography entry
  sdk          NDSS 2015 introducing-privacy-threats-from-ad-libraries-to-android-users-through-privacy-g
        the paper is about Android advertising SDKs; its own detection.phenomenon is "Ad-library prevalence"
  other-sense  PETS 2015 automated-experiments-on-ad-privacy-settings
        "the ad repository lacking ads relevant to the interest" - the ad ecosystem's inventory, not an archive. The paper is a Google Ad Settings audit
  reference    CCS 2017 identifying-open-source-license-violation-and-1-day-security-risk-at-large-scale
        same bibliography entry
  reference    CCS 2017 keep-me-updated-an-empirical-study-of-third-party-library-updatability-on-androi
        same bibliography entry
  sdk          IMC 2018 beyond-google-play-a-large-scale-comparative-study-of-chinese-android-app-market
        Android advertising SDKs
  sdk          PETS 2018 nomoads-effective-and-efficient-cross-app-mobile-ad-blocking
        Android advertising SDKs
  reference    PETS 2018 panoptispy-characterizing-audio-and-video-exfiltration-from-android-applications
        same bibliography entry
  sdk          CCS 2019 deepintent-deep-icon-behavior-learning-for-detecting-intention-behavior-discrepa
        Android advertising SDKs
  sdk          PETS 2020 nomoats-towards-automatic-detection-of-mobile-tracking
        Android advertising SDKs
  sdk          USENIX 2020 towards-https-everywhere-on-android-we-are-not-there-yet
        a table column headed "Ad Library" listing AdMob, Facebook Audience Network, Unity - SDKs
  sdk          WWW 2020 maddroid-characterizing-and-detecting-devious-ad-contents-for-android-apps
        a section headed "Mobile Ad Library Detection and Analysis"
  sdk          PETS 2022 analyzing-the-monetization-ecosystem-of-stalkerware
        a figure axis reading "# of Ad Libraries" - Android SDKs per app
  reference    PETS 2022 charting-app-developers-journey-through-privacy-regulation-features-in-ad-networ
        bibliography entry for Edelson et al. 2020 only
  reference    WWW 2022 conspiracy-brokers-understanding-the-monetization-of-youtube-conspiracy-theories
        bibliography entry for Edelson et al. 2020 only
  reference    CCS 2023 marketing-to-children-through-online-targeted-advertising-targeting-mechanisms-a
        bibliography entry for Edelson et al. 2020; the study measures ad explanations, not the archive

==============================================================================
4. THE AUDITED SET: an ad-transparency archive as instrument or as object
==============================================================================
papers: 11 of 5855 with full text = 0.19%
  also inside the 180-paper advertising candidate set: 10 = 5.6% of it
per year: 2020:2 2021:1 2022:1 2023:3 2024:1 2025:1 2026:2
per venue: WWW:3 IEEE-SP:2 IMC:2 PETS:2 USENIX:1 NDSS:1
  IEEE-SP  2020  a-security-analysis-of-the-facebook-ad-library
  WWW      2020  facebook-ads-monitor-an-independent-auditing-system-for-political-ads-on-faceboo
  IMC      2021  polls-clickbait-and-commemorative-2-bills-problematic-political-advertising-on-n
  USENIX   2022  an-audit-of-facebooks-political-ad-policy-enforcement
  IEEE-SP  2023  collaborative-ad-transparency-promises-and-limitations
  WWW      2023  propaganda-politica-pagada-exploring-u-s-political-facebook-ads-en-espanol
  WWW      2023  the-thin-ideology-of-populist-advertising-on-facebook-during-the-2019-eu-electio
  IMC      2024  beyond-the-guidelines-assessing-metas-political-ad-moderation-in-the-eu
  NDSS     2025  scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns
  PETS     2026  a-year-under-the-dsa-ad-transparencys-uneven-landscape
  PETS     2026  ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis

==============================================================================
4b. WHICH ARCHIVE EACH AUDITED PAPER ACTUALLY USED (hand-assigned)
==============================================================================
  8  Meta
       IEEE-SP 2020 a-security-analysis-of-the-facebook-ad-library
       WWW 2020 facebook-ads-monitor-an-independent-auditing-system-for-
       USENIX 2022 an-audit-of-facebooks-political-ad-policy-enforcement
       WWW 2023 propaganda-politica-pagada-exploring-u-s-political-faceb
       WWW 2023 the-thin-ideology-of-populist-advertising-on-facebook-du
       IMC 2024 beyond-the-guidelines-assessing-metas-political-ad-moder
       NDSS 2025 scammagnifier-piercing-the-veil-of-fraudulent-shopping-w
       PETS 2026 a-year-under-the-dsa-ad-transparencys-uneven-landscape
  3  Google
       IMC 2021 polls-clickbait-and-commemorative-2-bills-problematic-po
       PETS 2026 a-year-under-the-dsa-ad-transparencys-uneven-landscape
       PETS 2026 ad-personalization-and-transparency-in-mobile-ecosystems
  1  (none: critiques ad libraries, queries Facebook Ads Manager instead)
       IEEE-SP 2023 collaborative-ad-transparency-promises-and-limitations
  1  X
       PETS 2026 a-year-under-the-dsa-ad-transparencys-uneven-landscape
  1  Apple
       PETS 2026 ad-personalization-and-transparency-in-mobile-ecosystems

  a paper can use more than one, so these do not sum to 11:
  Meta 8, Google 3, (none: critiques ad libraries, queries Facebook Ads Manager instead) 1, X 1, Apple 1
  archives with ZERO audited users: TikTok Commercial Content Library, LinkedIn, Microsoft/Bing,
  Snapchat, Amazon, Pinterest, Booking.com, AliExpress, Zalando, and the EU political-ad repository.

==============================================================================
5. WHICH ARCHIVE, BY NAME (the zeros are the finding)
==============================================================================
  13  Meta / Facebook Ad Library (proper noun or API/Report)  [2018:1 2020:2 2021:1 2022:3 2023:3 2024:1 2025:1 2026:1]
   3  Google Ads Transparency Center / political-ads BigQuery  [2020:1 2023:1 2026:1]
        IEEE-SP 2020 a-security-analysis-of-the-facebook-ad-library
        PETS 2026 ad-personalization-and-transparency-in-mobile-ecosystems
        IEEE-SP 2023 collaborative-ad-transparency-promises-and-limitations
   0  TikTok Commercial Content Library / API  [none]
   1  X / Twitter ads repository or ad transparency centre  [2026:1]
        PETS 2026 a-year-under-the-dsa-ad-transparencys-uneven-landscape
   0  LinkedIn Ad Library  [none]
   0  Snapchat Ads Gallery / political ads library  [none]
   1  Apple App Store Ad Repository  [2026:1]
        PETS 2026 ad-personalization-and-transparency-in-mobile-ecosystems
   0  Microsoft / Bing Ad Library  [none]
   0  Amazon Store Ad Library API  [none]
   0  Pinterest / Booking.com / AliExpress / Zalando ads repository  [none]
   0  the EU "European repository for online political advertisements"  [none]
   7  DSA Article 39 (any mention of the article number, COLLIDES)  [2018:1 2022:1 2023:1 2024:1 2026:3]
        CCS 2018 image-disguising-for-privacy-preserving-deep-learning
        PETS 2022 revisiting-identification-issues-in-gdpr-right-of-access
        CCS 2023 marketing-to-children-through-online-targeted-advertisin
        IMC 2024 beyond-the-guidelines-assessing-metas-political-ad-moder
        PETS 2026 banned-books-analysis-of-censorship-on-amazon-com
        PETS 2026 a-year-under-the-dsa-ad-transparencys-uneven-landscape
        PETS 2026 ad-personalization-and-transparency-in-mobile-ecosystems
   0  Regulation (EU) 2024/900 (political-advertising regulation, TTPA)  [none]
  29  Digital Services Act (any mention)  [2023:2 2024:7 2025:7 2026:13]
   4  DSA Article 39 within 400 chars of "Digital Services Act" or "DSA"  [2023:1 2024:1 2026:2]
        CCS 2023 marketing-to-children-through-online-targeted-advertisin
        IMC 2024 beyond-the-guidelines-assessing-metas-political-ad-moder
        PETS 2026 a-year-under-the-dsa-ad-transparencys-uneven-landscape
        PETS 2026 ad-personalization-and-transparency-in-mobile-ecosystems
   1  the EU political-advertising regulation, any phrasing (1 hit = bibliography titles only)  [2022:1]
        USENIX 2022 an-audit-of-facebooks-political-ad-policy-enforcement

==============================================================================
6. THE COMPLEMENT: DONATED AND CROWDSOURCED AD CORPORA
==============================================================================
   4  Who Targets Me  [2020:1 2022:1 2023:1 2026:1]
        WWW 2020 facebook-ads-monitor-an-independent-auditing-system-for-
        USENIX 2022 an-audit-of-facebooks-political-ad-policy-enforcement
        PETS 2026 a-year-under-the-dsa-ad-transparencys-uneven-landscape
        IEEE-SP 2023 collaborative-ad-transparency-promises-and-limitations
   5  Ad Observer / NYU Ad Observatory  [2022:1 2023:3 2026:1]
        USENIX 2022 an-audit-of-facebooks-political-ad-policy-enforcement
        WWW 2023 propaganda-politica-pagada-exploring-u-s-political-faceb
        USENIX 2023 problematic-advertising-and-its-disparate-exposure-on-fa
        WWW 2026 when-ads-become-profiles-uncovering-the-invisible-risk-o
        IEEE-SP 2023 collaborative-ad-transparency-promises-and-limitations
   1  Australian Ad Observatory  [2026:1]
        WWW 2026 when-ads-become-profiles-uncovering-the-invisible-risk-o
  25  data donation (any phrasing)  [2019:1 2023:2 2024:5 2025:11 2026:6]

==============================================================================
7. MEASURED RESULTS FROM THE AUDITED SET, WITH THE PAPERS OWN DENOMINATORS
==============================================================================

--- IEEE-SP 2020 a-security-analysis-of-the-facebook-ad-library
    POPULATION n=126013 unit=web-pages source="Facebook's Ad Library Report" sampling=exhaustive
    POPULATION n=3685558 unit=other source="Facebook's Ad Library API" sampling=exhaustive
    Political-ad archive coverage
        metric    : unretrievable advertisements
        prevalence: 16,160 ads could not be retrieved; 7,817 additional ads were absent from the latest report
        quote (§methodology): "Overall, we could not retrieve 16,160 ads on 7,515 pages using the API, despite repeated attempts."
    Missing disclosure strings
        metric    : share of ads, pages, and spend
        prevalence: 9.7% of ads; 68.3% of pages; at least $37 million in spend
        quote (§results): "9.7 % of all ads in the Ad Library do not include a disclosure string. Advertisers spent at least $ 37 million on such ads."
    Disclosure-string fragmentation
        metric    : misattributed ad spend
        prevalence: 15.8% of total ad spend, or $98.2 million, was missing or potentially misattributed
        quote (§results): "We find that 15.8 % of total ad spend in the Ad Library cannot be attributed due to missing disclosure strings, or would be misattributed due to disclosure string fragmentation."
    Nonconforming disclosure strings
        metric    : share of evaluated disclosure strings
        prevalence: 23% appeared not to conform to Facebook's stated policy
        quote (§results): "In total, 23 % of the disclosure strings that we evaluated appeared to not conform to Facebook's stated policy."
    Undeclared coordinated advertising
        metric    : number of advertiser clusters
        prevalence: 172 clusters met the threshold; 16 were likely inauthentic communities
        quote (§results): "Overall, 172 clusters of advertisers met our threshold for undeclared coordinated behavior. We performed a manual review of these clusters."
    Coordinated advertiser activity
        metric    : cluster lifespan and spend
        prevalence: 16 likely inauthentic-community clusters spent $3,867,613 on 19,526 ads
        quote (§introduction): "Specifically, we found 16 clusters of likely inauthentic communities that spent $ 3,867,613 on a total of 19,526 ads."

--- WWW 2020 facebook-ads-monitor-an-independent-auditing-system-for-political-ads-on-faceboo
    POPULATION n=239000 unit=other source="volunteers using Facebook" sampling=convenience
    POPULATION n=100778 unit=other source="Facebook Ad Library" sampling=exhaustive
    POPULATION n=10000 unit=other source="FbAdLibrary dataset" sampling=random
    POPULATION n=10000 unit=other source="AdCollector dataset" sampling=random
    POPULATION n=38110 unit=other source="AdCollector dataset" sampling=purposive
    political-ad detection
        metric    : accuracy, AUC, Macro-F1, true-positive rate
        prevalence: 835 of 38,110 ads, approximately 2%, classified as political
        quote (§results): "Using a threshold of 0.97, our CNN model classifies 835 ads as political out of the 38,110 ads we tested - 2% of the ads are political."
    undisclosed political advertisements
        metric    : share of detected ads absent from Ad Library
        prevalence: Only 34 of 835 detected ads had a corresponding Ad Library ad
        quote (§results): "Only 34 of the 835 ads have a corresponding ad in FbAdLibrary."
    political-ad legal disclaimer
        metric    : count of ads and advertisers
        prevalence: 90 ads from 53 advertisers contained the required keywords and tax IDs
        quote (§results): "We identified 90 such ads coming from 53 advertisers in Political Ads."

--- IMC 2021 polls-clickbait-and-commemorative-2-bills-problematic-political-advertising-on-n
    POPULATION n=6144 unit=websites source="Tranco" sampling=stratified
    POPULATION n=1344 unit=websites source="custom seed list" sampling=purposive
    POPULATION n=745 unit=websites source="custom seed list" sampling=stratified
    political advertisements
        metric    : share of all collected ads
        prevalence: 67,501 ads (8,836 unique), or 3.9% of the overall dataset
        quote (§results): "Our political ad classifier and qualitative coding, detected 67,501 ads (8,836 unique) with political content, or 3.9% of the overall dataset."
    misleading political polls
        metric    : count of poll, petition, or survey ads
        prevalence: 7,602 ads
        quote (§results): "Poll, Petition, or Survey 7,602"
    political clickbait and sponsored articles
        metric    : count of sponsored articles or direct article links
        prevalence: 25,103 ads (85.4% of political news and media ads)
        quote (§results): "most political news and media ads were sponsored content or links to articles (25,103 ads, 85.4%)."
    system-popup impersonation ads
        metric    : count of ads
        prevalence: 162 ads
        quote (§appendix): "We found 162 ads of this style in our dataset."

--- USENIX 2022 an-audit-of-facebooks-political-ad-policy-enforcement
    POPULATION n=265824 unit=other source="Ad Library Report" sampling=seed-and-crawl
    POPULATION n=33828769 unit=other source="Facebook Ad Library web portal" sampling=exhaustive
    POPULATION n=215030 unit=other source="Facebook Ad Library Report" sampling=exhaustive
    Facebook political-ad detection prevalence
        metric    : share of observed political ads detected post-hoc
        prevalence: 1.7% of 4.2 million observed political ads
        quote (§results): "Within our measurement data, 72,678 ads were marked at some point as 'detected,' i.e., political but not properly declared, within 14 days"
    False-positive political-ad detections
        metric    : false-positive rate among detected ads
        prevalence: 55% of detected U.S. ads
        quote (§results): "Across our sample, a majority of detected ads (55%) should not have been enforced upon"
    Missed political ads
        metric    : false-negative rate
        prevalence: 116,963 ads missed worldwide; 4.5% false-negative rate
        quote (§results): "Across all political advertisers worldwide, we find a false negative rate of 4.5%"
    Country variation in missed ads
        metric    : country-level false-negative rate
        prevalence: 0.85% in the United States versus up to 45% elsewhere
        quote (§results): "Facebook misses the fewest ads in the United States (0.85% false negatives), whereas enforcement can be considerably worse in other countries"
    Detection delay
        metric    : time until detection
        prevalence: 40% detected within less than one day; median under two days
        quote (§results): "40% of ads were detected within less than 1 day, with the median activity period being less than 2 days."

--- IEEE-SP 2023 collaborative-ad-transparency-promises-and-limitations
    POPULATION n=420 unit=human-participants source="custom recruitment through mainstream media" sampling=convenience
    POPULATION n=51681 unit=other source="Facebook Ads Manager" sampling=exhaustive
    POPULATION n=1000000 unit=other source="Facebook population statistics derived from Dp" sampling=pre-existing-dataset
    POPULATION n=1000 unit=other source="random targeting-formula selection from 51,681 combinations" sampling=random
    advertiser targeting formula
        metric    : inference accuracy
        prevalence: 17 of 45 experiments inferred the formula correctly; 65% when at least ten users received the ad.
        quote (§abstract): "Our results show that we infer the targeting formula correctly in 17 cases. However, if we only look at experiments where our ad has been received by ten or more monitored users, the accuracy increases to 65%."
    simulation-to-platform inference consistency
        metric    : proportion of experiments with consistent accuracy
        prevalence: 82.2% of all ad experiments
        quote (§introduction): "The accuracy is consistent in 82.2% of all ad experiments."
    off-target ad delivery
        metric    : cases among monitored ad receptions
        prevalence: 86 of 1,021 receptions across 45 experiments
        quote (§results): "We observed several cases where some monitored users received our ad even if their ui did not contain all the targeting attributes we specified. This happened for 86 cases (out of a total of 1021 cases)."

--- WWW 2023 propaganda-politica-pagada-exploring-u-s-political-facebook-ads-en-espanol
    POPULATION n=5200000 unit=other source="Facebook Ad Library API" sampling=exhaustive
    POPULATION n=4700000 unit=other source="Facebook Ad Library API" sampling=exhaustive
    ad-creative language
        metric    : share of ad creatives
        prevalence: 97.95% English, 1.58% Spanish, and 0.07% French
        quote (§methodology): "The vast majority of ad creatives are in English (97.95 %), followed by Spanish (1.58 %) and French (0.07 %)."
    Spanish advertising spending
        metric    : share of total spending
        prevalence: Spanish ads accounted for 1.90% of spending
        quote (§results): "Spanish ads make up a higher 1.90 % of spending, and a disproportionately higher 2.25 % of impressions"
    advertiser composition
        metric    : share of spending among top advertisers
        prevalence: Government agencies represented 29.2% of Spanish top-100 spending
        quote (§results): "Government agencies in the Spanish top 100 were by far the largest type of ads (29.2 % of Spanish spending)"
    multilingual ad topics
        metric    : share of spending per topic
        prevalence: Elections/Voting/Court System represented 25.7% of Spanish spending and 23.2% of English spending
        quote (§results): "advertisers spent most on the "Elections/Voting/Court System" topic (25.7 % of all Spanish spending, and 23.2 % of English)."
    candidate topic coverage
        metric    : topics exceeding 1% of budget
        prevalence: Biden covered six English topics but three Spanish topics; Trump covered eight English topics but three Spanish topics
        quote (§results): "in English, Biden talked about 6 of our 14 topics ... while only speaking about the latter three in Spanish."

--- WWW 2023 the-thin-ideology-of-populist-advertising-on-facebook-during-the-2019-eu-electio
    POPULATION n=46 unit=other source="PopuList" sampling=purposive
    POPULATION n=44949 unit=other source="Meta Ad Library" sampling=exhaustive
    POPULATION n=57 unit=other source="Facebook pages manually identified for parties and leaders" sampling=purposive
    political advertising volume and reach
        metric    : number of ads, estimated impressions, and expenditure share
        prevalence: 44,949 ads from 39 parties; PopuList parties generated 40% of impressions
        quote (§dataset): "Overall, we find 44 949 ads from the selected five countries."
    audience demographic differences
        metric    : male-to-female odds ratio and Wasserstein distance
        prevalence: Populist parties reached 50% more male audiences and older audiences with Wasserstein distance 7.92
        quote (§results): "Overall, populist, far-right, and eurosceptic parties tend to reach more male audiences (+50%, +79%, +70% respectively)."
    issue emphasis in political ads
        metric    : issue and sub-issue coefficients
        prevalence: Populist parties emphasized monetary policy, bureaucracy and reforms, and security
        quote (§results): "In general, these parties focus more on the Euro, bureaucracy, illegal immigration, law & order, and institutions such as police and the military"
    cross-country demographic similarity
        metric    : F1 score and cosine similarity
        prevalence: Country differences were more prominent than demographic similarities
        quote (§appendix): "Most of the F1 scores are quite low. This finding suggests that for the studied tags, differences across European countries are more prominent than any similarities in their demographics."

--- IMC 2024 beyond-the-guidelines-assessing-metas-political-ad-moderation-in-the-eu
    POPULATION n=null unit=other source="Meta Ad Library API" sampling=exhaustive
    POPULATION n=29500000 unit=other source="Meta Ad Library API" sampling=exhaustive
    undeclared political advertisements
        metric    : estimated daily volume and moderation recall
        prevalence: 1,217 undeclared political ads launched daily in the 16 EU countries
        quote (§results): "The model trained on Meta's guidelines provides a conservative estimate of 1,217 undeclared political ads launched daily; see the breakdown by country in the Appendix."
    Meta political-ad moderation
        metric    : moderation recall
        prevalence: Only 7.7% of undeclared political ads were moderated
        quote (§results): "Across the 16 countries, only 7.7% of undeclared political ads (excluding news-related pages) were moderated as political by Meta"
    false-positive political moderation
        metric    : weighted proportion of moderated ads outside guidelines
        prevalence: 60.4% of ads moderated by Meta did not align with its political-ad criteria
        quote (§results): "Similarly, 60.4% of ads moderated by Meta did not align with Meta's criteria for political advertising"
    repeated moderation violations
        metric    : number of repeatedly violating pages and subsequent activity
        prevalence: 92 pages were moderated at least 28 times; 82 remained active
        quote (§results): "we identified 92 advertising pages that had been moderated at least 28 times... Of the 92 pages... 82 were still active as of March 29th, 2024"
    moderation circumvention
        metric    : reported strategy occurrences
        prevalence: “Monsieur 25 %” appeared in more than 250 ads
        quote (§results): "Monsieur 25 % was used in more than 250 ads to refer to the French president Emmanuel Macron"

--- NDSS 2025 scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns
    POPULATION n=1155237 unit=domains source="daily feed of newly registered domains" sampling=not-stated
    fraudulent shopping websites
        metric    : number of identified websites
        prevalence: 46,746 of 1,155,237 collected domains
        quote (§dataset): "we collected 1,155,237 domains, with 46,746 identified as potential fraudulent shopping websites using the ML-based classifier."
    merchant-ID reuse across scam domains
        metric    : unique merchant IDs
        prevalence: 5,278 merchant IDs from 41,863 completed checkouts
        quote (§dataset): "Ultimately, out of 41,863 completed checkouts, S CAM M AGNIFIER was able to extract merchant IDs for three different payment processors for 5,278 total."
    historical merchant-domain linkage
        metric    : linked domains
        prevalence: 14,394 domains linked to identified fraudulent merchants
        quote (§introduction): "Intriguingly, 14,394 domains are connected to these merchants."
    checkout redirection
        metric    : fraudulent websites with intermediary redirects
        prevalence: 263 fraudulent shopping websites
        quote (§results): "The results showed that 263 fraudulent shopping websites redirected to an intermediary domain, which we refer to as B."
    advertising referral traffic
        metric    : share of users
        prevalence: 28.78% Facebook, 21.10% Google, and 9.38% Bing
        quote (§results): "28.78% of users reached fraudulent shopping websites through advertisements on Facebook (including Instagram), 21.10% from Google (Ads or search results), and 9.38% from Bing."
    rapid scam monetization
        metric    : share monetized within one year
        prevalence: 97.73% monetized in less than a year
        quote (§results): "Specifically 20.17% of fraudulent shopping websites have transactions within 10 days after their creation date and 97.73% are monetized in less than a year."
    browser-extension scam detection
        metric    : detection rate
        prevalence: 76.74% (66 of 86 expert-labeled fraudulent websites)
        quote (§evaluation): "The integrated approach achieved a much higher detection rate of 76.74% (66 detected fraudulent websites) compared to Beyond Phish's standalone performance of 59.30%"
    shared website content
        metric    : number of clusters
        prevalence: ten categories
        quote (§discussion): "We performed a simple clustering method on fraudulent shopping websites screenshots to cluster them into ten categories."

--- PETS 2026 a-year-under-the-dsa-ad-transparencys-uneven-landscape
    POPULATION n=48511 unit=other source="Who Targets Me (WTM)" sampling=pre-existing-dataset
    POPULATION n=14000 unit=other source="Who Targets Me (WTM)" sampling=random
    POPULATION n=38011 unit=other source="Who Targets Me (WTM)" sampling=exhaustive
    WAIST transparency
        metric    : share of explanations by targeting category and transparency dimension
        prevalence: 48,511 WAIST notices across Facebook, Instagram, YouTube, and X
        quote (§introduction): "Our analysis covers N = 48, 511 ad explanations collected across four platforms: Facebook, Instagram, YouTube, and X"
    YouTube explanation granularity
        metric    : share of explanations classified as broad
        prevalence: 98.9% of YouTube explanation texts cited only the main targeting form
        quote (§results): "98.9% of all explanation texts cite only the main targeting form (e.g., "Your activity"), and do not detail the exact attributes used."
    YouTube targeting omission
        metric    : number of shared ads across profiles
        prevalence: 11 unique ads appeared across all profiles; search terms were omitted for three profile configurations
        quote (§methodology): "Across 60 search queries, we identified 11 unique ads that were shown across all profiles"
    Repository matching discrepancies
        metric    : missing-category and discrepant-attribute rates
        prevalence: Google and X omitted substantial portions of user-visible targeting attributes
        quote (§methodology): "Each user-facing ad explanation is linked to its corresponding repository record when a unique match can be determined using shared identifiers"
    Google repository attribution
        metric    : number of additional repository targeting categories
        prevalence: Three categories listed, including demographic and contextual signals not selected by the advertiser
        quote (§results): "When the campaign appeared in the repository, it listed three targeting parameters: "Geographical location", "Demographical information", and "Contextual signals""
    Facebook transparency evolution
        metric    : transparency dimensions across regulatory and product phases
        prevalence: 38,011 Facebook WAIST notices analyzed
        quote (§results): "We analyze 38,011 Facebook WAIST objects collected between September 2017 and March 2025"
    Public repository completeness
        metric    : successful match rate
        prevalence: 181/351 Facebook matches and 399/438 Instagram matches
        quote (§results): "This yielded 181 out of 351 matches on Facebook (48% missing from the repository) and 399 out of 438 on Instagram (9% missing)."

--- PETS 2026 ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis
    POPULATION n=249 unit=other source="custom account creation" sampling=purposive
    POPULATION n=30 unit=mobile-apps source="top free app charts and curated app lists" sampling=purposive
    ad personalization
        metric    : category-distribution differences and item frequencies
        prevalence: 257,820 ads and 378,057 recommendations in main experiments
        quote (§evaluation): "In total, we extracted 257,820 ads and 378,057 recommendations as part of our main experiments."
    recommendation personalization
        metric    : Jaccard similarity and permutation-test significance
        prevalence: Interest-based treatment-control recommendation differences were consistently significant on Google Play
        quote (§results): "The tests for the interest-based persona are most pronounced. Here, the differences between control and treatment are always significant."
    sensitive-topic targeting
        metric    : presence of sensitive-category items
        prevalence: Sensitive personas received related recommendations but no related ads
        quote (§discussion): "Users can mark a set of topics as sensitive and do not receive any related ads. Nevertheless, our analysis proves this does not apply to recommendations"
    advertisement-transparency inconsistencies
        metric    : repository/UI consistency
        prevalence: Google Play Store ads were absent from Google's Ad Transparency dataset
        quote (§results): "We hence conclude that the ads shown in Google's Play Store are entirely missing from the Ad Transparency dataset"

==============================================================================
8. REJECTED PROBES (recorded so they are not re-added)
==============================================================================
   71 papers  /TTPA/i
            REJECTED: matches the substring inside "httpa..." (HTTPA, "http api"). Useless as written; the spelled-out name is probed above instead.
    4 papers  /\bFORT\b|Facebook Open Research/
            REJECTED: FORT collides with unrelated all-caps tokens and OCR artefacts. Never hand-audited, never used.
   28 papers  /Google Transparency Report/i
            REJECTED: that report covers government requests and HTTPS adoption, not ads. It put censorship and TLS papers into an ad-archive candidate set.

The quote checker

adarchives_quotecheck.py
#!/usr/bin/env python3
"""Quote and figure spot-check for design:platforms:ad_archives.
 
    python3 scripts/adarchives_quotecheck.py
 
Every phrase quoted on the page, and every paper-sourced figure that a reader
could check, is looked for in the cited paper's own text. Coverage is the
figures a reviewer would actually pull on: it does not extend to every numeral
in every table, and the provenance page says which are covered.  Three renderings are tried in order, because no
single one is reliable for two-column ACM/IEEE PDFs:
 
  1. ``paper.cols.txt``  -- the de-columned rendering the corpus ships
  2. ``paper.norm.txt``  -- the raw single-stream rendering
  3. ``pypdf``           -- extracted here, live, from ``paper.pdf``
 
and each rendering is tried twice: verbatim after NFKC/whitespace/quote/dash
normalisation, then folded to lowercase alphanumerics with ``fi``/``fl``
collapsed to ``f``.  The fold is not cosmetic: several of these PDFs drop the
``fi`` ligature outright, so the paper literally reads "we fnd 44 949 ads", and
a verbatim check on the published sentence FAILS on a correct quote.
 
Exits non-zero on any NOT-FOUND.
"""
import re
import sys
import pathlib
import unicodedata
 
ROOTS = [pathlib.Path("/workspace/publications_dataset/data"),
         pathlib.Path("/workspace/publications_dataset")]
ROOT = next(r for r in ROOTS if (r / "extract/run1/extractions.jsonl").exists())
FT = ROOT / "fulltext"
 
 
def norm(s: str) -> str:
    # NFKC folds the mathematical-italic letters that LaTeX emits for inline
    # maths: benzaamia2026 writes its own sample size as U+1D441 ("\U0001d441 =
    # 48,511"), not ASCII "N", so a verbatim check on "N = 48,511" fails.
    s = unicodedata.normalize("NFKC", s)
    s = s.replace("fi", "fi").replace("fl", "fl")
    s = s.replace("‘", "'").replace("’", "'").replace("ʼ", "'")
    s = s.replace("“", '"').replace("”", '"')
    s = re.sub(r"[‐-―−]", "-", s)
    return re.sub(r"\s+", " ", s).strip()
 
 
def fold(s: str) -> str:
    s = norm(s).lower().replace("fi", "f").replace("fl", "f")
    return re.sub(r"[^a-z0-9]", "", s)
 
 
_cache: dict[str, dict[str, str]] = {}
 
 
def renderings(paper: str) -> dict[str, str]:
    if paper in _cache:
        return _cache[paper]
    venue, year, slug = paper.split("/")
    d = FT / year / venue / slug
    out = {}
    for name, f in (("cols", "paper.cols.txt"), ("norm", "paper.norm.txt")):
        if (d / f).exists():
            out[name] = norm((d / f).read_text(encoding="utf8", errors="replace"))
    if (d / "paper.pdf").exists():
        from pypdf import PdfReader
        out["pypdf"] = norm("\n".join(p.extract_text() or "" for p in PdfReader(d / "paper.pdf").pages))
    _cache[paper] = out
    return out
 
 
CHECKS = {
    # ---- the archive's own error characteristics
    "IEEE-SP/2020/a-security-analysis-of-the-facebook-ad-library": [
        "Out of all 126,013 pages with ads in the Ad Library",
        "86,150 (68.3 %) ran at least one undisclosed ad that was subsequently detected and added to the Ad Library",
        "9.7 % of all ads in the Ad Library do not include a disclosure string",
        "Advertisers spent at least $ 37 million on such ads, which is 6 % of the total spend during the study period",
        "We find that 15.8 % of total ad spend in the Ad Library cannot be attributed due to missing disclosure strings, or would be misattributed due to disclosure string fragmentation",
        "for a total misattributed spend of $ 98.2 M",
        "In total, 23 % of the disclosure strings that we evaluated appeared to not conform to Facebook's stated policy",
        "Overall, we could not retrieve 16,160 ads on 7,515 pages using the API, despite repeated attempts",
        "we found 16 clusters of likely inauthentic communities that spent $ 3,867,613 on a total of 19,526 ads",
        "Our dataset contains 3,685,558 ads created during the study period",
    ],
    # ---- the archive misses ads that ran
    "WWW/2020/facebook-ads-monitor-an-independent-auditing-system-for-political-ads-on-faceboo": [
        "Only 34 of the 835 ads have a corresponding ad in FbAdLibrary.",
        "our CNN model classifies 835 ads as political out of the 38,110 ads we tested",
        "not all political ads we detected were present in the Facebook Ad Library for political ads",
    ],
    # ---- two-sided error, measured on both sides at once
    "USENIX/2022/an-audit-of-facebooks-political-ad-policy-enforcement": [
        "Our comprehensive and representative data set contains 4.2 million political and 29.6 million non-political ads from all 215,030 pages",
        "In total, we observed 33.8 million unique ads during our measurement",
        "Across our sample, a majority of detected ads (55%) should not have been enforced upon",
        "Across all political advertisers worldwide, we find a false negative rate of 4.5%",
        "Facebook misses the fewest ads in the United States (0.85% false negatives), whereas enforcement can be considerably worse in other countries",
        "40% of ads were detected within less than 1 day, with the median activity period being less than 2 days",
        # added after a reviewer listed page figures the checker did not cover
        "116,963",
        "45%",
        "retrieves all currently active ads running on Facebook's core advertising platforms",
    ],
    "IMC/2024/beyond-the-guidelines-assessing-metas-political-ad-moderation-in-the-eu": [
        "In August 2023, complying with the European Union's Digital Services Act (DSA), Meta expanded its Ad Library to archive all advertisements targeting individuals within the EU",
        "Across the 16 countries, only 7.7% of undeclared political ads (excluding news-related pages) were moderated as political by Meta",
        "Similarly, 60.4% of ads moderated by Meta did not align with Meta's criteria for political advertising",
        "The model trained on Meta's guidelines provides a conservative estimate of 1,217 undeclared political ads launched daily",
        "The 29.5 million ads run in January and February 2024 will be considered for inference purposes",
        "we identified 92 advertising pages that had been moderated at least 28 times",
        "82 were still active",
        # the page calls this LLM classification; check the model is named
        "gpt-4-turbo",
    ],
    # ---- coverage measured against a donated ground truth
    "PETS/2026/a-year-under-the-dsa-ad-transparencys-uneven-landscape": [
        "This yielded 181 out of 351 matches on Facebook (48% missing from the repository) and 399 out of 438 on Instagram (9% missing).",
        "We rely on a dataset of N = 48, 511 ads collected through Who Targets Me",
        "98.9% of all explanation texts cite only the main targeting form",
        "An additional dataset was constructed with all available Facebook data, comprising 38,011 WAIST notices from September 2017",
        # the X Ads Repository use, and the own-campaign experiment: both were
        # mis-assigned in the first draft of the page
        "In this section, we analyze X's public ad repository",
        "In total, we identified advertisers with repository entries for 2,190 ads. Of these, 1,458 ads were successfully matched using Tweet IDs",
        "Acting as an advertiser, we created a YouTube campaign targeting users in France",
        "it listed three targeting parameters",
    ],
    "PETS/2026/ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis": [
        "We hence conclude that the ads shown in Google's Play Store are entirely missing from the Ad Transparency dataset",
        "Google's Ads Transparency Center can be accessed using either the web interface",
        "We initially intended to refer to these datasets as ground truth to evaluate our experiments but found them affected by several issues",
        # the two-surfaces finding, added after a reviewer noticed it was quoted
        # on the page but not covered here
        "The data shown on the web interface seems incomplete compared to the records in the BigQuery dataset",
        "multiple instances where the web interface does not show any text ads for an advertiser in the Play Store. The dataset, however, lists multiple matching entries",
        "resolves to a page-not-found error",
        "surface_serving_stats",
        # the Apple positive control
        "Apple provides a web interface",
        "We did not detect missing entries by comparing randomly selected samples of our dataset to Apple's Ad Repository",
        "This information is explicitly required as part of the ad repositories by Article 39(2), point (g) DSA",
        "In total, we extracted 257,820 ads and 378,057 recommendations as part of our main experiments",
        "We created 249 Apple or Google Accounts using 213 eSIMs",
    ],
    # ---- studies that used an archive as a corpus
    "WWW/2023/the-thin-ideology-of-populist-advertising-on-facebook-during-the-2019-eu-electio": [
        "Overall, we find 44 949 ads from the selected five countries.",
        "We use the Meta Ad Library API",
        "populist, far-right, and eurosceptic parties tend to reach more male audiences",
        "57 Facebook pages",
    ],
    "WWW/2023/propaganda-politica-pagada-exploring-u-s-political-facebook-ads-en-espanol": [
        "This leaves us with our final data set of 4.7 M ads with 1.5 B",
        "We retrieve the set of political ads that ran on Facebook in 2020 from the Facebook Ad Library",
        "The vast majority of ad creatives are in English (97.95 %), followed by Spanish (1.58 %)",
        "Spanish ads make up a higher 1.90 % of spending, and a disproportionately higher 2.25 % of impressions",
    ],
    "IMC/2021/polls-clickbait-and-commemorative-2-bills-problematic-political-advertising-on-n": [
        "Prior work has found that Facebook's ad archives are incomplete and use a limited definition of \"political\"",
        "crawling 1,000 political ads from the Google",
        "detected 67,501 ads (8,836 unique) with political content, or 3.9% of the overall dataset",
        "We identified 6,144 mainstream news websites in the Tranco Top 1 million",
        "1,344 websites which we refer to as \"misinformation websites\"",
        "we truncated the list to 745 sites",
    ],
    "NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns": [
        "Analyzing fraudulent shopping websites ads from the Facebook Ad Library",
    ],
    # ---- the alternative to an archive
    "IEEE-SP/2023/collaborative-ad-transparency-promises-and-limitations": [
        "if we only look at experiments where our ad has been received by ten or more monitored users, the accuracy increases to 65%",
        "Similar limitations exist in Political Ad Libraries (public repositories aggregating political ads run on the platform)",
        "we infer the targeting formula correctly in 17 cases",
        "420 active users",
        "51, 681",
        "86 cases (out of a total of 1021 cases)",
    ],
    "WWW/2026/when-ads-become-profiles-uncovering-the-invisible-risk-of-web-advertising-at-sca": [
        "The full dataset made available to us comprises over 700,000 ad observations collected between 2021 and 2023 including demographic information from over 2,000 Australian Facebook users.",
        "891 unique users, 63,864 total ad sessions and 435,314 total ad impressions",
        "76.38% accuracy",
        "41.60%",
    ],
    # ---- the per-user explanation, not the archive
    "IEEE-SP/2025/characterizing-the-usability-and-usefulness-of-u-s-ad-transparency-systems": [
        "more than two-thirds of all questions and more than half of desired actions went unanswered or unaddressed",
        "74 (37%) participants reported that navigation was difficult",
        "We recruited a total of 200 participants, 25 per ATS",
        # the taxonomy covers 22 systems, the usability arm only eight -- a
        # reviewer caught the page attaching the 22 to the usability finding
        "explore the ATS of one of eight representative",
    ],
    "USENIX/2023/problematic-advertising-and-its-disparate-exposure-on-facebook": [
        "Ad Observer",
    ],
    # cited on the page as the contrast case (the ad platform itself, not the
    # archive), so outside the report script's audited set and checked here.
    "WWW/2019/auditing-offline-data-brokers-via-facebooks-advertising-platform": [
        "above 90% in the U.S.",
        "81.3%",
        "74.4%",
    ],
}
 
 
def _guard_no_duplicate_keys() -> None:
    """A repeated dict literal key silently discards the earlier value, so an
    edit that re-adds a paper drops its existing checks with no error. Count the
    keys in this file's own source instead of trusting the dict."""
    src = pathlib.Path(__file__).read_text(encoding="utf8")
    keys = re.findall(r'^    "([A-Za-z0-9/_.-]+)":', src, re.M)
    dupes = {k for k in keys if keys.count(k) > 1}
    if dupes:
        sys.exit(f"ABORT: duplicate CHECKS keys, earlier quotes would be silently dropped: {sorted(dupes)}")
    if len(keys) != len(CHECKS):
        sys.exit(f"ABORT: {len(keys)} keys in source but {len(CHECKS)} in the dict.")
 
 
def main() -> int:
    _guard_no_duplicate_keys()
    fails = 0
    total = 0
    for paper, quotes in CHECKS.items():
        rend = renderings(paper)
        print(f"\n### {paper}")
        for q in quotes:
            total += 1
            how = None
            for name in ("cols", "norm", "pypdf"):
                if name not in rend:
                    continue
                if norm(q) in rend[name]:
                    how = f"{name}-exact"
                    break
                if fold(q) in fold(rend[name]):
                    how = f"{name}-folded"
                    break
            if how is None:
                how = "NOT FOUND"
                fails += 1
            print(f"  [{how:12}] {q[:112]}")
    print(f"\n{total} strings checked, {fails} NOT FOUND.")
    return 1 if fails else 0
 
 
if __name__ == "__main__":
    sys.exit(main())

The quote checker's output

adarchives_quotecheck-output.txt
### IEEE-SP/2020/a-security-analysis-of-the-facebook-ad-library
  [cols-exact  ] Out of all 126,013 pages with ads in the Ad Library
  [pypdf-exact ] 86,150 (68.3 %) ran at least one undisclosed ad that was subsequently detected and added to the Ad Library
  [pypdf-exact ] 9.7 % of all ads in the Ad Library do not include a disclosure string
  [pypdf-exact ] Advertisers spent at least $ 37 million on such ads, which is 6 % of the total spend during the study period
  [pypdf-exact ] We find that 15.8 % of total ad spend in the Ad Library cannot be attributed due to missing disclosure strings, 
  [pypdf-exact ] for a total misattributed spend of $ 98.2 M
  [cols-exact  ] In total, 23 % of the disclosure strings that we evaluated appeared to not conform to Facebook's stated policy
  [cols-exact  ] Overall, we could not retrieve 16,160 ads on 7,515 pages using the API, despite repeated attempts
  [pypdf-exact ] we found 16 clusters of likely inauthentic communities that spent $ 3,867,613 on a total of 19,526 ads
  [cols-exact  ] Our dataset contains 3,685,558 ads created during the study period

### WWW/2020/facebook-ads-monitor-an-independent-auditing-system-for-political-ads-on-faceboo
  [pypdf-exact ] Only 34 of the 835 ads have a corresponding ad in FbAdLibrary.
  [pypdf-exact ] our CNN model classifies 835 ads as political out of the 38,110 ads we tested
  [cols-exact  ] not all political ads we detected were present in the Facebook Ad Library for political ads

### USENIX/2022/an-audit-of-facebooks-political-ad-policy-enforcement
  [cols-exact  ] Our comprehensive and representative data set contains 4.2 million political and 29.6 million non-political ads 
  [cols-exact  ] In total, we observed 33.8 million unique ads during our measurement
  [cols-exact  ] Across our sample, a majority of detected ads (55%) should not have been enforced upon
  [cols-exact  ] Across all political advertisers worldwide, we find a false negative rate of 4.5%
  [cols-exact  ] Facebook misses the fewest ads in the United States (0.85% false negatives), whereas enforcement can be consider
  [cols-exact  ] 40% of ads were detected within less than 1 day, with the median activity period being less than 2 days
  [cols-exact  ] 116,963
  [cols-exact  ] 45%
  [cols-exact  ] retrieves all currently active ads running on Facebook's core advertising platforms

### IMC/2024/beyond-the-guidelines-assessing-metas-political-ad-moderation-in-the-eu
  [cols-exact  ] In August 2023, complying with the European Union's Digital Services Act (DSA), Meta expanded its Ad Library to 
  [cols-exact  ] Across the 16 countries, only 7.7% of undeclared political ads (excluding news-related pages) were moderated as 
  [cols-exact  ] Similarly, 60.4% of ads moderated by Meta did not align with Meta's criteria for political advertising
  [cols-exact  ] The model trained on Meta's guidelines provides a conservative estimate of 1,217 undeclared political ads launch
  [cols-exact  ] The 29.5 million ads run in January and February 2024 will be considered for inference purposes
  [cols-exact  ] we identified 92 advertising pages that had been moderated at least 28 times
  [cols-exact  ] 82 were still active
  [cols-exact  ] gpt-4-turbo

### PETS/2026/a-year-under-the-dsa-ad-transparencys-uneven-landscape
  [cols-exact  ] This yielded 181 out of 351 matches on Facebook (48% missing from the repository) and 399 out of 438 on Instagra
  [cols-exact  ] We rely on a dataset of N = 48, 511 ads collected through Who Targets Me
  [cols-exact  ] 98.9% of all explanation texts cite only the main targeting form
  [pypdf-folded] An additional dataset was constructed with all available Facebook data, comprising 38,011 WAIST notices from Sep
  [cols-exact  ] In this section, we analyze X's public ad repository
  [cols-exact  ] In total, we identified advertisers with repository entries for 2,190 ads. Of these, 1,458 ads were successfully
  [cols-exact  ] Acting as an advertiser, we created a YouTube campaign targeting users in France
  [cols-exact  ] it listed three targeting parameters

### PETS/2026/ad-personalization-and-transparency-in-mobile-ecosystems-a-comparative-analysis
  [cols-exact  ] We hence conclude that the ads shown in Google's Play Store are entirely missing from the Ad Transparency datase
  [cols-exact  ] Google's Ads Transparency Center can be accessed using either the web interface
  [cols-exact  ] We initially intended to refer to these datasets as ground truth to evaluate our experiments but found them affe
  [cols-exact  ] The data shown on the web interface seems incomplete compared to the records in the BigQuery dataset
  [cols-exact  ] multiple instances where the web interface does not show any text ads for an advertiser in the Play Store. The d
  [cols-exact  ] resolves to a page-not-found error
  [cols-exact  ] surface_serving_stats
  [cols-exact  ] Apple provides a web interface
  [cols-exact  ] We did not detect missing entries by comparing randomly selected samples of our dataset to Apple's Ad Repository
  [cols-exact  ] This information is explicitly required as part of the ad repositories by Article 39(2), point (g) DSA
  [cols-exact  ] In total, we extracted 257,820 ads and 378,057 recommendations as part of our main experiments
  [cols-exact  ] We created 249 Apple or Google Accounts using 213 eSIMs

### WWW/2023/the-thin-ideology-of-populist-advertising-on-facebook-during-the-2019-eu-electio
  [cols-folded ] Overall, we find 44 949 ads from the selected five countries.
  [cols-exact  ] We use the Meta Ad Library API
  [cols-exact  ] populist, far-right, and eurosceptic parties tend to reach more male audiences
  [cols-exact  ] 57 Facebook pages

### WWW/2023/propaganda-politica-pagada-exploring-u-s-political-facebook-ads-en-espanol
  [cols-folded ] This leaves us with our final data set of 4.7 M ads with 1.5 B
  [cols-exact  ] We retrieve the set of political ads that ran on Facebook in 2020 from the Facebook Ad Library
  [cols-exact  ] The vast majority of ad creatives are in English (97.95 %), followed by Spanish (1.58 %)
  [pypdf-exact ] Spanish ads make up a higher 1.90 % of spending, and a disproportionately higher 2.25 % of impressions

### IMC/2021/polls-clickbait-and-commemorative-2-bills-problematic-political-advertising-on-n
  [cols-exact  ] Prior work has found that Facebook's ad archives are incomplete and use a limited definition of "political"
  [cols-exact  ] crawling 1,000 political ads from the Google
  [cols-exact  ] detected 67,501 ads (8,836 unique) with political content, or 3.9% of the overall dataset
  [cols-exact  ] We identified 6,144 mainstream news websites in the Tranco Top 1 million
  [cols-exact  ] 1,344 websites which we refer to as "misinformation websites"
  [cols-exact  ] we truncated the list to 745 sites

### NDSS/2025/scammagnifier-piercing-the-veil-of-fraudulent-shopping-website-campaigns
  [pypdf-exact ] Analyzing fraudulent shopping websites ads from the Facebook Ad Library

### IEEE-SP/2023/collaborative-ad-transparency-promises-and-limitations
  [cols-exact  ] if we only look at experiments where our ad has been received by ten or more monitored users, the accuracy incre
  [pypdf-exact ] Similar limitations exist in Political Ad Libraries (public repositories aggregating political ads run on the pl
  [cols-exact  ] we infer the targeting formula correctly in 17 cases
  [cols-exact  ] 420 active users
  [cols-exact  ] 51, 681
  [pypdf-folded] 86 cases (out of a total of 1021 cases)

### WWW/2026/when-ads-become-profiles-uncovering-the-invisible-risk-of-web-advertising-at-sca
  [cols-exact  ] The full dataset made available to us comprises over 700,000 ad observations collected between 2021 and 2023 inc
  [cols-exact  ] 891 unique users, 63,864 total ad sessions and 435,314 total ad impressions
  [cols-exact  ] 76.38% accuracy
  [cols-exact  ] 41.60%

### IEEE-SP/2025/characterizing-the-usability-and-usefulness-of-u-s-ad-transparency-systems
  [cols-exact  ] more than two-thirds of all questions and more than half of desired actions went unanswered or unaddressed
  [cols-exact  ] 74 (37%) participants reported that navigation was difficult
  [cols-exact  ] We recruited a total of 200 participants, 25 per ATS
  [cols-exact  ] explore the ATS of one of eight representative

### USENIX/2023/problematic-advertising-and-its-disparate-exposure-on-facebook
  [cols-exact  ] Ad Observer

### WWW/2019/auditing-offline-data-brokers-via-facebooks-advertising-platform
  [cols-exact  ] above 90% in the U.S.
  [cols-exact  ] 81.3%
  [cols-exact  ] 74.4%

83 strings checked, 0 NOT FOUND.

References

[1]
Breuer, David; Becker, Lucas; Hollick, Matthias (2026): "Ad Personalization and Transparency in Mobile Ecosystems: A Comparative Analysis of Google's and Apple's EU App Stores", Proceedings on Privacy Enhancing Technologies 2026(1):604-630. (DOI)
[2]
Benzaamia, Abir; El Fraihi, Asmaa; Abdelaziz, Ines; Goga, Oana (2026): "A Year Under the DSA: Ad Transparency's Uneven Landscape", Proceedings on Privacy Enhancing Technologies 2026(2):517-532. (DOI)
[3]
Edelson, Laura; Lauinger, Tobias; McCoy, Damon (2020): "A Security Analysis of the Facebook Ad Library", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[4]
Le Pochat, Victor; Edelson, Laura; Van Goethem, Tom; Joosen, Wouter; McCoy, Damon; Lauinger, Tobias (2022): "An Audit of Facebook's Political Ad Policy Enforcement", in: Proceedings of the USENIX Security Symposium. (Link)
[5]
Gkiouzepi, Eleni; Andreou, Athanasios; Goga, Oana; Loiseau, Patrick (2023): "Collaborative Ad Transparency: Promises and Limitations", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
[6]
Ali, Muhammad; Goetzen, Angelica; Mislove, Alan; Redmiles, Elissa M.; Sapiezynski, Piotr (2023): "Problematic Advertising and its Disparate Exposure on Facebook", in: Proceedings of the USENIX Security Symposium. (Link)
[7]
Chen, Baiyu; Tag, Benjamin; Xue, Hao; Angus, Daniel; Salim, Flora (2026): "When Ads Become Profiles: Uncovering the Invisible Risk of Web Advertising at Scale with LLMs", in: Proceedings of the ACM Web Conference. (DOI)
[8]
Zeng, Eric; Wei, Miranda; Gregersen, Theo; Kohno, Tadayoshi; Roesner, Franziska (2021): "Polls, Clickbait, and Commemorative \$2 Bills: Problematic Political Advertising on News and Media Websites Around the 2020 U.S. Elections", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
[9]
Coelho, Bruno; Lauinger, Tobias; Edelson, Laura; Goldstein, Ian; McCoy, Damon (2023): "Propaganda Política Pagada: Exploring U.S. Political Facebook Ads en Español", in: Proceedings of the ACM Web Conference. (DOI)
[10]
Capozzi, Arthur; Morales, Gianmarco De Francisci; Mejova, Yelena; Monti, Corrado; Panisson, André (2023): "The Thin Ideology of Populist Advertising on Facebook during the 2019 EU Elections", in: Proceedings of the ACM Web Conference. (DOI)
[11]
Bouchaud, Paul; Liénard, Jean F. (2024): "Beyond the Guidelines: Assessing Meta's Political Ad Moderation in the EU", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
provenance/design/platforms/ad_archives.1788868190.txt.gz · Last modified: by karel.kubicek.claude