User Tools

Site Tools


design:platforms:ad_archives

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
design:platforms:ad_archives [2026/09/08 11:33] – Self-review: four/five error-measuring papers, merge duplicated designation footnote, correct the cross-platform and commercial-completeness open questions in light of Apple, and the quote-check count. Authored by Claude karel.kubicek.claudedesign:platforms:ad_archives [2026/09/08 11:43] (current) – Generic-review fixes: the 7-year window is rolling, so Edelson's and Capozzi's corpora have already aged out of the live archive; Google's political section is region-limited not worldwide; 48pc restated as an upper bound; Bouchaud's 92.3pc is a labelling karel.kubicek.claude
Line 8: Line 8:
  
 <WRAP important> <WRAP important>
-**The corpus figure this page's parent published was wrong, and the correction is the point.** [[Design:Platforms]] reports "**69** papers name an Ad Library or ad archive", from a full-text probe for %%/Ad Library|Ads? Archive API/i%%. That probe is defeated by a homonym. In mobile-security writing an "ad library" is an **advertising SDK linked into an Android app**; in platform-transparency writing "the Ad Library" is **Meta's public ad archive**. Of the 69, **44** contain only the lowercase sense, and the probe additionally matches the substring inside "re[ad librar]y" and "uplo[ad librar]ies" — one of the 69 is a paper about Node.js file uploads.+**Four things to take away before you design anything.**
  
-Hand-auditing every candidate paper against its own full text gives the real number: of the **5,855** papers in the [[literature:corpus|publication corpus]] with extractable full text (seven venues, 2010–2026), **11** use an ad-transparency archive as a data source or measure the archive itself. Eight more mention one only in their bibliography. The old probe's precision for this question was **14.5%**. +  **The archive's inclusion rule is a classifier the platform runs and does not publish**, so a prevalence you compute on it is a prevalence among ads that classifier admitted. Both directions of its error land in your denominator, and the field has measured both: **55%** false positives and **4.5%** false negatives on Facebook's political enforcement {[lepochat2022_audit]}, **7.7%** moderation recall and **60.4%** over-inclusion in the EU {[bouchaud2024_beyond]}
- +  - **The archive cannot tell you what it is missing**, so an archive study needs a second, independent view of the same ads — a donation panel, your own campaignor your own collectorSee [[#Getting the Other Side of the Gap]]. 
-Eleven papers is a small literature, and this page says so rather than inflating it. But it is **coherent** eleven — five of them measure the archive'own accuracyand they agreeThe full per-year series is 2020:2, 2021:1, 2022:1, 2023:3, 2024:1, 2025:12026:2 — one to three papers a year with no trend worth naming, and 2025–2026 are the provisional venue-years.+  **If your subject is EU political advertisingthe instrument closed on 6 October 2025**and the window it holds is rolling rather than frozen. See [[#The Regulatory Floor, and the October 2025 Cliff]]. 
 +  - **A regulator has already itemised what two of these repositories are missing.** Read the enforcement record before you write your limitations section: [[#The enforcement record is a published error budget]].
 </WRAP> </WRAP>
  
 <WRAP tip> <WRAP tip>
-**If you are planning an EU political-ad study in 2026read [[#The Regulatory Floor, and the October 2025 Cliff]] first.** Both Meta and Google stopped carrying political advertising in the EU in October 2025 in response to Regulation (EU) 2024/900. The historical ads stay in the archivesnew ones do not arriveEvery EU political-ad paper in this corpus studies window that has closed.+**This page also corrects its parent's paper countfrom 69 to 11.** [[Design:Platforms]] reported "**69** papers name an Ad Library or ad archive" from a probe for %%/Ad Library|Ads? Archive API/i%%, and that probe cannot separate Meta's **Ad Library** from an Android **advertising SDK** — the ordinary meaning of "ad library" in mobile-security writing. Hand-auditing every candidate gives **11** papers of **5,855** with extractable full text that use an ad-transparency archive or measure one. The audit, the residue and the arithmetic are in [[#The inclusion rule, and the homonym that decides the count]]the parent has been corrected. 
 + 
 +Eleven is a small literature and this page does not inflate itBut five of the eleven measure the archives' own accuracy and none finds them accurate, and the per-year series — 2020:2, 2021:1, 2022:1, 2023:3, 2024:1, 2025:1, 2026:2 — is one to three papers year with no trend worth naming.
 </WRAP> </WRAP>
  
Line 26: Line 29:
   * {[lepochat2022_audit]} (USENIX Security 2022) — **both error directions, measured simultaneously**, on 33.8 million ads retrieved from the Ad Library. A **55%** false-positive rate among the ads Facebook itself flagged as political in the US, a **4.5%** false-negative rate worldwide, and — the finding that should change your study design — false negatives of **0.85%** in the United States against "//considerably worse in other countries//". A single global recall figure hides a factor of fifty.   * {[lepochat2022_audit]} (USENIX Security 2022) — **both error directions, measured simultaneously**, on 33.8 million ads retrieved from the Ad Library. A **55%** false-positive rate among the ads Facebook itself flagged as political in the US, a **4.5%** false-negative rate worldwide, and — the finding that should change your study design — false negatives of **0.85%** in the United States against "//considerably worse in other countries//". A single global recall figure hides a factor of fifty.
   * {[bouchaud2024_beyond]} (IMC 2024) — the same two-sided analysis on the **DSA-expanded** archive, which since August 2023 holds "//all advertisements targeting individuals within the EU//". Across 16 EU countries, "//only 7.7% of undeclared political ads ... were moderated as political by Meta//" while "//60.4% of ads moderated by Meta did not align with Meta's criteria for political advertising//". Read it for how to calibrate a classifier against the platform's own published guidelines rather than against your intuition.   * {[bouchaud2024_beyond]} (IMC 2024) — the same two-sided analysis on the **DSA-expanded** archive, which since August 2023 holds "//all advertisements targeting individuals within the EU//". Across 16 EU countries, "//only 7.7% of undeclared political ads ... were moderated as political by Meta//" while "//60.4% of ads moderated by Meta did not align with Meta's criteria for political advertising//". Read it for how to calibrate a classifier against the platform's own published guidelines rather than against your intuition.
-  * {[benzaamia2026_year]} (PoPETs 2026) — **the archives measured against a donated ground truth**, across four platforms at once. Matching 48,511 real ads collected through the //Who Targets Me// extension back to the platforms' DSA repositories yields "//181 out of 351 matches on Facebook (48% missing from the repository) and 399 out of 438 on Instagram (9% missing)//". Two repositories under the same legal obligation, five times apart on completeness.+  * {[benzaamia2026_year]} (PoPETs 2026) — **the archives measured against a donated ground truth**, across four platforms at once. Matching real ads collected through the //Who Targets Me// extension back to the platforms' DSA repositories yields "//181 out of 351 matches on Facebook (48% missing from the repository) and 399 out of 438 on Instagram (9% missing)//" — read those as upper bounds on missingness, since an unmatched ad is not a proven absence. Two repositories under the same legal obligation, five times apart. Also the only paper here that bought its own ads to test what the repository says about them.
   * {[gkiouzepi2023_collaborative]} (IEEE S&P 2023) — **the alternative**, and its limits. Argues that ad libraries cannot deliver targeting transparency and proposes inferring the targeting formula from a panel of monitored users instead. Honest about the cost: correct inference in 17 of 45 controlled campaigns, rising to "//65%//" only "//where our ad has been received by ten or more monitored users//".   * {[gkiouzepi2023_collaborative]} (IEEE S&P 2023) — **the alternative**, and its limits. Argues that ad libraries cannot deliver targeting transparency and proposes inferring the targeting formula from a panel of monitored users instead. Honest about the cost: correct inference in 17 of 45 controlled campaigns, rising to "//65%//" only "//where our ad has been received by ten or more monitored users//".
  
Line 45: Line 48:
 ===== The Archives, One by One ===== ===== The Archives, One by One =====
  
-[[#The Regulatory Floor, and the October 2025 Cliff|DSA Article 39]] requires an ad repository of every platform the European Commission designates as a very large online platform or search engine, and the inventory below is the **thirteen** repositories Mozilla and Check First listed when they tried them all in 2024. Read that as //the set somebody has actually exercised//, and treat it as a **floor**. The Commission's designation list has grown since: as of its own "Information updated on 31 August 2026" it carries **28 active designations**, adding Shein (26 April 2024), Temu (31 May 2024), WhatsApp (26 January 2026) and Reddit, Roblox and ChatGPT-as-a-search-engine (all 31 August 2026), while Stripchat's designation was terminated on 27 May 2025.((https://digital-strategy.ec.europa.eu/en/policies/list-designated-vlops-and-vloses — fetched 2026-09-08, "Information updated on 31 August 2026". Newly designated platforms get a compliance deadline four months after designation, so the August 2026 cohort is not yet obliged. https://www.temu.com/ads-repository already returns HTTP 200; a headless browser was redirected to a bot-verification challenge, so its **contents** were not verified here. Guessed paths for Shein and Roblox returned 404, and no repository URL was established for them. Mozilla and Check First tested **twelve services from eleven companies** — "//AliExpress, Apple App Store, Bing, Booking.com, Alphabet (Google Search and YouTube), LinkedIn, Meta (Facebook and Instagram), Pinterest, Snapchat, TikTok, X, and Zalando//" — while their annex of repository links names **thirteen**, adding Amazon. Amazon was excluded from the testing for a specific reason the report states: "//We did not examine Amazon Store because Amazon had been granted an exemption from making its ad repository publicly available by the European Court of Justice//" — an interim order that was set aside six months later; see [[#DSA Article 39 is the floor]]. The table below follows the annex.))+[[#The Regulatory Floor, and the October 2025 Cliff|DSA Article 39]] requires an ad repository of every platform the European Commission designates as a very large online platform or search engine, and the inventory below is the **thirteen** repositories Mozilla and Check First listed in 2024, twelve of which they tested. Read that as //the set somebody has actually exercised//, and treat it as a **floor**. The Commission's designation list has grown since: as of its own "Information updated on 31 August 2026" it carries **28 active designations**, adding Shein (26 April 2024), Temu (31 May 2024), WhatsApp (26 January 2026) and Reddit, Roblox and ChatGPT-as-a-search-engine (all 31 August 2026), while Stripchat's designation was terminated on 27 May 2025.((https://digital-strategy.ec.europa.eu/en/policies/list-designated-vlops-and-vloses — fetched 2026-09-08, "Information updated on 31 August 2026". Newly designated platforms get a compliance deadline four months after designation, so the August 2026 cohort is not yet obliged. https://www.temu.com/ads-repository already returns HTTP 200; a headless browser was redirected to a bot-verification challenge, so its **contents** were not verified here. Guessed paths for Shein and Roblox returned 404, and no repository URL was established for them. Mozilla and Check First tested **twelve services from eleven companies** — "//AliExpress, Apple App Store, Bing, Booking.com, Alphabet (Google Search and YouTube), LinkedIn, Meta (Facebook and Instagram), Pinterest, Snapchat, TikTok, X, and Zalando//" — while their annex of repository links names **thirteen**, adding Amazon. Amazon was excluded from the testing for a specific reason the report states: "//We did not examine Amazon Store because Amazon had been granted an exemption from making its ad repository publicly available by the European Court of Justice//" — an interim order that was set aside six months later; see [[#DSA Article 39 is the floor]]. The table below follows the annex.))
  
 These are not comparable instruments, and the differences are exactly the ones that decide whether your study is possible. These are not comparable instruments, and the differences are exactly the ones that decide whether your study is possible.
Line 67: Line 70:
  
   * **Four archives out of thirteen have ever been used in these seven venues, and one paper accounts for three of the four.** Hand-assigning each of the 11 audited papers to the archive it actually queried gives **Meta 8, Google 3, X 1, Apple 1** (a paper can use several, and {[gkiouzepi2023_collaborative]} queries none — it critiques ad libraries and works from Facebook's Ads Manager instead). {[benzaamia2026_year]} alone supplies the only X use and one of the Google uses; {[breuer2026_ad]} supplies the only Apple use. **TikTok's Commercial Content Library, LinkedIn's, Bing's, Snapchat's, Amazon's, the four retail repositories and the EU political-ad repository have zero users.** That is not evidence they are unusable — Amazon's API is free and X's CSV needs no application at all. It means **there is no methods section to copy**, and if you use one you are close to first, in these venues.   * **Four archives out of thirteen have ever been used in these seven venues, and one paper accounts for three of the four.** Hand-assigning each of the 11 audited papers to the archive it actually queried gives **Meta 8, Google 3, X 1, Apple 1** (a paper can use several, and {[gkiouzepi2023_collaborative]} queries none — it critiques ad libraries and works from Facebook's Ads Manager instead). {[benzaamia2026_year]} alone supplies the only X use and one of the Google uses; {[breuer2026_ad]} supplies the only Apple use. **TikTok's Commercial Content Library, LinkedIn's, Bing's, Snapchat's, Amazon's, the four retail repositories and the EU political-ad repository have zero users.** That is not evidence they are unusable — Amazon's API is free and X's CSV needs no application at all. It means **there is no methods section to copy**, and if you use one you are close to first, in these venues.
-  * **The scopes do not line up, so a cross-platform rate is a cross-scope rate.** Meta's archive is worldwide for political ads and EU-only for everything else; TikTok's and X's are EU-only for everything; Google's political section is worldwide but restricted to verified advertisers. A "share of ads that are X" computed across platforms is comparing different populations unless you cut all of them down to the intersection, which in practice means the EU and one year. +  * **The scopes do not line up, so a cross-platform rate is a cross-scope rate.** Meta's archive is worldwide for political ads and EU-only for everything else; TikTok's and X's are EU-only for everything; Google's political section is limited both to **verified election advertisers** and to the regions where Google operates election-ads verification at all — "//Google has different requirements for political and election advertising based on region//", and its policy lists those regions one by one.((https://support.google.com/adspolicy/answer/6014595 — fetched 2026-09-08 with a headless browser: "//In some regions, election ads may run only if the advertiser is verified by Google//", followed by named per-region sections. Google Cloud's own introduction to the dataset says it "//supports election advertising verification in over 35 countries//" (https://cloud.google.com/blog/topics/developers-practitioners/how-get-started-political-ads-transparency-report-dataset, 25 January 2023). An earlier draft of this page called that section "worldwide"; it is not.)) A "share of ads that are X" computed across platforms is comparing different populations unless you cut all of them down to the intersection, which in practice means the EU and one year. 
-  * **Bulk access is much more widely available than the corpus's silence suggests.** Mozilla and Check First's annex records an ads-repository **API** for Alphabet, Amazon, Apple, Bing, LinkedIn, Meta, TikTok, X and Zalando as of March 2024, and none for AliExpress or Snapchat.((The annex table has three columns — //Ad Repository//, //Commercial Content// and //Ads Repository API// — and the API column is populated for nine of the thirteen rows; ''Booking.com'' and ''Pinterest'' rows carry only two links each in the extracted text, so their API status was read as unclear rather than absent. https://www.mozillafoundation.org/documents/363/Full_Disclosure_Stress_Testing_Tech_Platforms_Ad_Repositories_3FepU2u.pdf, page 53, extracted with ''pypdf'' on 2026-09-08.)) The four this page can confirm hands-on today are Meta's Graph API, Google's BigQuery datasets, TikTok's Commercial Content API and X's CSV, with Amazon's free Ad Library API verified from its own documentation. So the right reading is not "the data is locked up" — it is that **programmatic access exists across most of the thirteen and almost nobody in these venues has used it**.+  * **Bulk access is much more widely available than the corpus's silence suggests.** Mozilla and Check First's annex records an ads-repository **API** for Alphabet, Amazon, Apple, Bing, LinkedIn, Meta, TikTok, X and Zalando as of March 2024, and none for AliExpress or Snapchat.((The annex table has three columns — //Ad Repository//, //Commercial Content// and //Ads Repository API// — and the API column is populated for nine of the thirteen rows; ''Booking.com'' and ''Pinterest'' rows carry only two links each in the extracted text, so their API status was read as unclear rather than absent. https://www.mozillafoundation.org/documents/363/Full_Disclosure_Stress_Testing_Tech_Platforms_Ad_Repositories_3FepU2u.pdf, page 53, extracted with ''pypdf'' on 2026-09-08.)) The five confirmed from the provider's own current documentation are Meta's Graph API, Google's BigQuery datasets, TikTok's Commercial Content APIX's CSV and Amazon's free Ad Library API. **None of them was exercised for this page** — no query was issued and no file downloaded, and the provenance page says so. So the right reading is not "the data is locked up" — it is that **programmatic access exists across most of the thirteen and almost nobody in these venues has used it**.
  
 The one independent cross-platform comparison is not from these venues: Mozilla commissioned Check First to stress-test twelve of them against DSA-derived criteria, published **16 April 2024**, and concluded that "//even the best approaches don't meet our baseline//" and that "//our accuracy testing found many cases where ads in the user interface were not located in the ad repository//".((//Full Disclosure: Stress testing tech platforms' ad repositories//, Mozilla Foundation, dated 16 April 2024 on its research-library page, https://www.mozillafoundation.org/en/research/library/full-disclosure-stress-testing-tech-platforms-ad-repositories/ and the 55-page report at https://www.mozillafoundation.org/documents/363/Full_Disclosure_Stress_Testing_Tech_Platforms_Ad_Repositories_3FepU2u.pdf — both fetched 2026-09-08. Testing was "//as of March 18, 2024//". {[benzaamia2026_year]} cites this work as 2023; Mozilla's own page dates it 2024, and this page follows Mozilla.)) It is now two years old, which is why the statuses above were re-checked rather than taken from it — but it is the only source that has tried them side by side, and it should be your starting inventory rather than a vendor listicle. The one independent cross-platform comparison is not from these venues: Mozilla commissioned Check First to stress-test twelve of them against DSA-derived criteria, published **16 April 2024**, and concluded that "//even the best approaches don't meet our baseline//" and that "//our accuracy testing found many cases where ads in the user interface were not located in the ad repository//".((//Full Disclosure: Stress testing tech platforms' ad repositories//, Mozilla Foundation, dated 16 April 2024 on its research-library page, https://www.mozillafoundation.org/en/research/library/full-disclosure-stress-testing-tech-platforms-ad-repositories/ and the 55-page report at https://www.mozillafoundation.org/documents/363/Full_Disclosure_Stress_Testing_Tech_Platforms_Ad_Repositories_3FepU2u.pdf — both fetched 2026-09-08. Testing was "//as of March 18, 2024//". {[benzaamia2026_year]} cites this work as 2023; Mozilla's own page dates it 2024, and this page follows Mozilla.)) It is now two years old, which is why the statuses above were re-checked rather than taken from it — but it is the only source that has tried them side by side, and it should be your starting inventory rather than a vendor listicle.
Line 76: Line 79:
 Concrete, because the archive-by-archive table above does not say how long you will be waiting. Verified against each provider's own documentation on 2026-09-08. Concrete, because the archive-by-archive table above does not say how long you will be waiting. Verified against each provider's own documentation on 2026-09-08.
  
-  * **Meta Ad Library API.** You need a Facebook account, then identity **and location** verification at ''facebook.com/ID'' — the same process an advertiser must complete to run political ads, and Meta's own page says "//It can take a few days to confirm the information you submit//" — then a Meta for Developers account and an app. Queries then go through the Graph API, and Meta publishes a script repository with a command-line interface. So: **days, and a verified identity tied to your real name and location.** Plan it before the deadline, not after. +  * **Meta Ad Library API.** You need a Facebook account, then identity **and location** verification at ''facebook.com/ID'' — the same process an advertiser must complete to run political ads, and Meta's own page says "//It can take a few days to confirm the information you submit//" — then a Meta for Developers account and an app. So: **days, and a verified identity tied to your real name and location** — which is the measurement-relevant part, because it is not anonymous access and your institution may have a view. Plan it before the deadline, not after. 
-  * **Google's BigQuery public datasets.** No application, no account review: you need a Google Cloud project and you pay for bytes scanned, not for the data. ''bigquery-public-data.google_political_ads'' (table ''creative_stats'') is the political subset; ''bigquery-public-data.google-ads-transparency-center'' is the whole centre. **Minutes, and a credit card.** This is the cheapest bulk archive access described anywhere on this wiki.+  * **Google's BigQuery public datasets.** No application, no account review: you need a Google Cloud project and you pay for bytes scanned, not for the data. ''bigquery-public-data.google_political_ads'' (table ''creative_stats'') is the political subset; ''bigquery-public-data.google-ads-transparency-center'' is the whole centre. **Minutes, and a credit card.**
   * **TikTok Commercial Content API.** A TikTok for Developers account and an application TikTok reviews. [[Design:Platforms]] records a stated turnaround of "//within 4 weeks//" for the sibling Research API. **Weeks, and an institutional eligibility test.**   * **TikTok Commercial Content API.** A TikTok for Developers account and an application TikTok reviews. [[Design:Platforms]] records a stated turnaround of "//within 4 weeks//" for the sibling Research API. **Weeks, and an institutional eligibility test.**
   * **X Ads Repository.** No application at all: pick a country and a date on the page and download the commercial-communications CSV. **Minutes.** One paper in these venues has used it — {[benzaamia2026_year]} pulled the per-advertiser, per-campaign CSVs and matched **1,458 of 2,190** entries back to its own ads by Tweet ID — and nobody has characterised the whole file. Note the enforcement history in [[#The enforcement record is a published error budget]] before you treat it as complete.   * **X Ads Repository.** No application at all: pick a country and a date on the page and download the commercial-communications CSV. **Minutes.** One paper in these venues has used it — {[benzaamia2026_year]} pulled the per-advertiser, per-campaign CSVs and matched **1,458 of 2,190** entries back to its own ads by Tweet ID — and nobody has characterised the whole file. Note the enforcement history in [[#The enforcement record is a published error budget]] before you treat it as complete.
  
-Access friction does **not** explain who has published: Google'BigQuery route is the easiest of the four and has papers, Meta's is the hardest and has 8, and X's CSV — which needs nothing at all — has 1. Whatever is driving the distributionit is not the door.+Access friction alone does **not** explain who has published: Meta's route is the most gated and has papers, Google's is nearly free and has 3, and X's CSV needs nothing at all and has 1. Two obvious confounders are age and agenda — Meta's political archive has existed since 2018 while X's CSV and the all-ads BigQuery dataset date from 2023and almost all of this literature is about political advertising, which is where Meta's disclosures are richest. Do not read the distribution as a ranking of the archives.
  
 ===== What an Ad Archive Structurally Cannot Contain ===== ===== What an Ad Archive Structurally Cannot Contain =====
Line 105: Line 108:
 An archive built around a category — "political", "social issue", "election" — is the output of a classifier the platform runs and does not publish. That classifier has a false-positive rate and a false-negative rate, and **an archive study inherits both**. Papers that report only one of them are reporting half the error. An archive built around a category — "political", "social issue", "election" — is the output of a classifier the platform runs and does not publish. That classifier has a false-positive rate and a false-negative rate, and **an archive study inherits both**. Papers that report only one of them are reporting half the error.
  
-Five papers in this corpus have measured it. They used different platforms, years, regions and ground truths, and they agree that both rates are large.+Five papers in this corpus have measured it, on different platforms, years, regions and ground truths. They agree that the error is **material and strongly region-dependent** — but only **two** of the five measured both directions, and the rest could only reach one. That asymmetry is itself the finding: measuring over-inclusion needs an independent definition of the category, and measuring omission needs ads the archive never received.
  
 ^ Study ^ What was measured ^ Miss rate (false negatives) ^ Over-inclusion (false positives) ^ Denominator the paper used ^ ^ Study ^ What was measured ^ Miss rate (false negatives) ^ Over-inclusion (false positives) ^ Denominator the paper used ^
Line 111: Line 114:
 | {[silva2020_facebook]} TheWebConf 2020 | archive recall against an independent collector | **only 34 of 835** ads their classifier called political had a matching ad in the archive | — | 38,110 ads collected from volunteers in Brazil, 2018 | | {[silva2020_facebook]} TheWebConf 2020 | archive recall against an independent collector | **only 34 of 835** ads their classifier called political had a matching ad in the archive | — | 38,110 ads collected from volunteers in Brazil, 2018 |
 | {[lepochat2022_audit]} USENIX Sec 2022 | Facebook's own political-ad enforcement | **4.5%** worldwide; **0.85%** in the US against up to **45%** elsewhere | **55%** of ads Facebook flagged as political in the US were not political | 33.8 M ads observed; 4.2 M political and 29.6 M non-political from all 215,030 pages that ran political ads | | {[lepochat2022_audit]} USENIX Sec 2022 | Facebook's own political-ad enforcement | **4.5%** worldwide; **0.85%** in the US against up to **45%** elsewhere | **55%** of ads Facebook flagged as political in the US were not political | 33.8 M ads observed; 4.2 M political and 29.6 M non-political from all 215,030 pages that ran political ads |
-| {[bouchaud2024_beyond]} IMC 2024 | Meta's EU moderation under the DSA | **92.3%** of undeclared political ads were not moderated as political (recall **7.7%**| **60.4%** of the ads Meta did moderate did not match Meta's own criteria | all EU ads via the Ad Library API, Aug 2023 – Feb 2024; 29.5 M ads used for inference | +| {[bouchaud2024_beyond]} IMC 2024 | Meta's EU political //labelling//, not archive membership moderation recall **7.7%** — the other 92.3% of undeclared political ads were **in** the archive but not labelled political | **60.4%** of the ads Meta did moderate did not match Meta's own criteria | all EU ads via the Ad Library API, Aug 2023 – Feb 2024; 29.5 M ads used for inference | 
-| {[benzaamia2026_year]} PoPETs 2026 | repository completeness against donated impressions | **48%** of Facebook ads and **9%** of Instagram ads seen by real users were missing from the repository | — | 351 Facebook and 438 Instagram ads matched from 48,511 donated ad explanations |+| {[benzaamia2026_year]} PoPETs 2026 | repository completeness against donated impressions | **48%** of Facebook and **9%** of Instagram ads seen by real users could not be matched to a repository entry — an **upper bound** on missingness, because a failed match is not proof of absence | — | 351 Facebook and 438 Instagram ads matched from 48,511 donated ad explanations |
  
 <WRAP tip> <WRAP tip>
Line 136: Line 139:
 Two practical notes. **Matching is where these studies actually fail**, not collection: {[benzaamia2026_year]} links each donated ad to a repository record "//when a unique match can be determined using shared identifiers//", and its headline 48%-missing figure is really "48% we could not uniquely match", which is an upper bound on missingness. Say which it is. And **the donation route is the one that is growing** — 25 papers in this corpus mention data donation, 11 of them in 2025 — but only 4 mention //Who Targets Me// and 5 //Ad Observer// or an //Ad Observatory//, so the panels themselves are still nearly unused in these venues. Two practical notes. **Matching is where these studies actually fail**, not collection: {[benzaamia2026_year]} links each donated ad to a repository record "//when a unique match can be determined using shared identifiers//", and its headline 48%-missing figure is really "48% we could not uniquely match", which is an upper bound on missingness. Say which it is. And **the donation route is the one that is growing** — 25 papers in this corpus mention data donation, 11 of them in 2025 — but only 4 mention //Who Targets Me// and 5 //Ad Observer// or an //Ad Observatory//, so the panels themselves are still nearly unused in these venues.
  
-Both //Who Targets Me// and //Ad Observer// were still operating when checked on 2026-09-08; the Australian Ad Observatory's first phase closed in March 2024 and a second phase runs to 2027.((https://whotargets.me/en/ and https://adobserver.org/ both HTTP 200 on 2026-09-08. The Australian Ad Observatory's own page — https://www.admscentre.org.au/adobservatory/, fetched 2026-09-08 — states that "//The first phase of the Australian Ad Observatory project was completed in March 2024//", reports "//328,107 unique ads//" from "//1,909 citizens//" in phase one, and announces "//The Australian Ad Observatory P.2 (2024-2027)//". Note that {[chen2026_when]} describes the dataset it was given as "//over 700,000 ad observations ... from over 2,000 Australian Facebook users//": impressions are not unique ads and the two figures are not in conflict, but do not mix them. ''%%adobservatory.org%%'' does not resolve.))+Both //Who Targets Me// and //Ad Observer// still had live sites on 2026-09-08 — which is not the same as still collecting, and neither was tested; the Australian Ad Observatory's first phase closed in March 2024 and a second phase runs to 2027.((https://whotargets.me/en/ and https://adobserver.org/ both HTTP 200 on 2026-09-08. The Australian Ad Observatory's own page — https://www.admscentre.org.au/adobservatory/, fetched 2026-09-08 — states that "//The first phase of the Australian Ad Observatory project was completed in March 2024//", reports "//328,107 unique ads//" from "//1,909 citizens//" in phase one, and announces "//The Australian Ad Observatory P.2 (2024-2027)//". Note that {[chen2026_when]} describes the dataset it was given as "//over 700,000 ad observations ... from over 2,000 Australian Facebook users//": impressions are not unique ads and the two figures are not in conflict, but do not mix them. ''%%adobservatory.org%%'' does not resolve.))
  
 ===== The Regulatory Floor, and the October 2025 Cliff ===== ===== The Regulatory Floor, and the October 2025 Cliff =====
Line 150: Line 153:
   * **"Through application programming interfaces" is in the text.** A platform that offers only a search box is not obviously compliant. That is a lever, and Mozilla's stress test found several platforms with no API at all.   * **"Through application programming interfaces" is in the text.** A platform that offers only a search box is not obviously compliant. That is a lever, and Mozilla's stress test found several platforms with no API at all.
   * **One year is the floor, not the practice.** Political archives at Meta and Google run to seven years. The DSA all-ads repositories run to one.   * **One year is the floor, not the practice.** Political archives at Meta and Google run to seven years. The DSA all-ads repositories run to one.
-  * **Article 39 is not Article 40.** Art. 39 is the public ad repository, open to anyone; Art. 40 is the vetted-researcher regime for non-public data, which [[Design:Platforms]] covers and which nobody in this corpus has published through. Conflating them is the most common error in this area, and this corpus shows how thin the DSA literature really is: "//Digital Services Act//" appears in 29 of 5,855 papers, and Article 39 appears near a DSA mention in **4**.+  * **Article 39 is not Article 40.** Art. 39 is the public ad repository, open to anyone; Art. 40 is the vetted-researcher regime for non-public data, which [[Design:Platforms]] covers and which nobody in this corpus has published through. They are easy to conflate, and this corpus shows how thin the DSA literature really is: "//Digital Services Act//" appears in **25** of 5,855 papers, and Article 39 appears near a DSA mention in **4**.((25 is the case-sensitive count, which is the figure [[Design:Platforms]] publishes for the same probe. Case-insensitively it is 29; the four extras are lowercase renderings inside bibliography titles. The number is given case-sensitively here so the two pages agree.))
   * **It has been litigated, and the platform lost.** Amazon challenged its designation and obtained interim suspension of the ad-repository obligation from the President of the General Court on 27 September 2023 (Case T-367/23 R) — which is why Mozilla's stress test could not test it. On **27 March 2024** the Vice-President of the Court of Justice set that part of the order aside and dismissed the application: "//Amazon's request to suspend its obligation to make an advertisement repository publicly available is rejected//" (Case C-639/23 P(R), //Commission v Amazon Services Europe//).((https://curia.europa.eu/jcms/upload/docs/application/pdf/2024-03/cp240060en.pdf — Court of Justice press release No 60/24, 27 March 2024, fetched and extracted 2026-09-08. This is the interim-measures branch only; the outcome of Amazon's underlying action for annulment was not established for this page.)) Amazon now publishes a free Ad Library API. Article 39 is an obligation with a litigation history, not a voluntary transparency programme, and that is part of why these archives are worth building a method on.   * **It has been litigated, and the platform lost.** Amazon challenged its designation and obtained interim suspension of the ad-repository obligation from the President of the General Court on 27 September 2023 (Case T-367/23 R) — which is why Mozilla's stress test could not test it. On **27 March 2024** the Vice-President of the Court of Justice set that part of the order aside and dismissed the application: "//Amazon's request to suspend its obligation to make an advertisement repository publicly available is rejected//" (Case C-639/23 P(R), //Commission v Amazon Services Europe//).((https://curia.europa.eu/jcms/upload/docs/application/pdf/2024-03/cp240060en.pdf — Court of Justice press release No 60/24, 27 March 2024, fetched and extracted 2026-09-08. This is the interim-measures branch only; the outcome of Amazon's underlying action for annulment was not established for this page.)) Amazon now publishes a free Ad Library API. Article 39 is an obligation with a litigation history, not a voluntary transparency programme, and that is part of why these archives are worth building a method on.
  
Line 175: Line 178:
 For a measurement plan in 2026 the consequences are concrete, and they cut in opposite directions. For a measurement plan in 2026 the consequences are concrete, and they cut in opposite directions.
  
-  * **EU political-ad archives are now historical corpora.** The seven-year retention means {[bouchaud2024_beyond]}'and {[capozzi2023_thin]}'windows remain queryable, but nothing new arrivesA "current state of EU political advertising" study is not available through this instrument, and a longitudinal series must be labelled as ending in October 2025. +  * **EU political-ad archives are now closing corpora, not frozen ones.** Nothing new arrives, and **the retention window is rolling**: Meta's API covers political ads "//delivered anywhere in the world during the past 7 years//", and Google's election ads "//remain available for seven years after the ad's last impression//". So the archive is losing its oldest material at the same rate it always did, while gaining nothing in the EU. {[bouchaud2024_beyond]}'Aug 2023 – Feb 2024 window survives to about 2031; {[capozzi2023_thin]}'April–May 2019 ads and {[edelson2020_security]}'s May 2018 – June 2019 corpus have **already aged out****The foundational papers of this literature can no longer be reproduced against the live archive**, which is the reproducibility lesson: an ad archive is a moving window, so deposit your pull rather than the query that produced it (see [[:Artifacts]]), and a longitudinal EU series must be labelled as ending in October 2025. 
-  * **The EU commercial-ad repositories are unaffected and under-used.** DSA Art. 39 covers ads of every kind, and that obligation did not moveMeta's all-EU-ads scope, X's CSV, Amazon's API and TikTok's API are all still there; TikTok's and Amazon's have **zero** users in these venues and X's has exactly **one** {[benzaamia2026_year]}.+  * **The EU commercial-ad repositories are unaffected and under-used.** DSA Art. 39 covers ads of every kind, and that obligation did not moveMeta's all-EU-ads scope, X's CSV, Amazon's API and TikTok's API are all still there.
   * **Political-ad measurement outside the EU is unaffected**, and Meta's worldwide political archive still has its 7-year window. {[coelho2023_propaganda]} and {[lepochat2022_audit]} are US studies and their route is intact.   * **Political-ad measurement outside the EU is unaffected**, and Meta's worldwide political archive still has its 7-year window. {[coelho2023_propaganda]} and {[lepochat2022_audit]} are US studies and their route is intact.
   * **The corpus has not noticed any of this.** Regulation (EU) 2024/900 appears, under any phrasing, in **zero** of 5,855 papers — including the two 2026 papers whose subject is EU ad transparency. Whatever you write about the EU here in 2026, there is no prior methods section that accounts for the cliff.   * **The corpus has not noticed any of this.** Regulation (EU) 2024/900 appears, under any phrasing, in **zero** of 5,855 papers — including the two 2026 papers whose subject is EU ad transparency. Whatever you write about the EU here in 2026, there is no prior methods section that accounts for the cliff.
Line 186: Line 189:
 ==== The inclusion rule, and the homonym that decides the count ==== ==== The inclusion rule, and the homonym that decides the count ====
  
-A paper counts if an ad-transparency archive is a **data source** for it, or if the archive **itself** is what it measures. A bibliography entry does not count. A mention of an Android advertising SDK does not count, and that is the whole problem: **"ad library" is a homonym**, and the SDK sense is far more common in these seven venues than the archive sense.+A paper counts if an ad-transparency archive is a **data source** for it, if the archive **itself** is what it measures, or if the archive is the explicit baseline the paper critiques and sets out to replace. A bibliography entry does not count. A mention of an Android advertising SDK does not count, and that is the whole problem: **"ad library" is a homonym**, and the SDK sense is far more common in these seven venues than the archive sense
 + 
 +The third clause of that rule is load-bearing for exactly one paper and it is stated here rather than left implicit: {[gkiouzepi2023_collaborative]} queries **no** archive — it argues ad libraries cannot deliver targeting transparency and proposes a panel-based alternative. **Excluding it gives 10 rather than 11**, and a reader who prefers the narrower rule should read every "11" on this page as "10, plus one critique".
  
 ^ What %%/Ad Library|Ads? Archive API/i%% actually matched, over 5,855 full texts ^ Papers ^ ^ What %%/Ad Library|Ads? Archive API/i%% actually matched, over 5,855 full texts ^ Papers ^
Line 220: Line 225:
 | 2026 | PoPETs | {[breuer2026_ad]} | Google Ads Transparency Center via BigQuery, and Apple's Ad Repository | intended ground truth; rejected for Google, passed for Apple | | 2026 | PoPETs | {[breuer2026_ad]} | Google Ads Transparency Center via BigQuery, and Apple's Ad Repository | intended ground truth; rejected for Google, passed for Apple |
  
-Per venue: TheWebConf 3, IEEE S&P 2, IMC 2, PoPETs 2, USENIX Security 1, NDSS 1. Per year: 2020:2, 2021:1, 2022:1, 2023:3, 2024:1, 2025:1, 2026:2. **2025 and 2026 are provisional** venue-years — CCS 2026 and IMC 2026 have not been held, and IEEE S&P 2026 and TheWebConf 2026 abstracts are not in OpenAlexso those years are under-represented by construction rather than by relevance ([[literature:corpus]]).+Per venue: TheWebConf 3, IEEE S&P 2, IMC 2, PoPETs 2, USENIX Security 1, NDSS 1. Per year: 2020:2, 2021:1, 2022:1, 2023:3, 2024:1, 2025:1, 2026:2. **2025 and 2026 are provisional** venue-yearsCCS 2026 and IMC 2026 have not been held, and OpenAlex has DOIs before abstracts for much of 2026 — **563** papers have a DOI and no abstract, dominated by TheWebConf 2026 (324) and IEEE S&P 2026 (194) — so, since selection screens on abstracts, those venue-years are under-covered by construction rather than by relevance ([[literature:corpus]]).
  
 Hand-assigned per paper from its own methods section, the archives used are **Meta 8, Google 3, X 1, Apple 1** — a paper can use several, so these do not sum to 11, and {[gkiouzepi2023_collaborative]} queries no archive at all. Nine of the thirteen repositories have **no** user in these venues. The assignment is by hand and not by regex on purpose: {[zeng2021_polls]} writes "//Google's (or others') political ad transparency reports//", whose parenthetical defeats any pattern for //Google's political ad …//, and {[benzaamia2026_year]} studies four platforms' repositories at once. Set against a floor of **180** papers in this corpus whose measured phenomenon is advertising-shaped, the audited set is **10** of them — about **5.6%** of advertising measurement in these venues touches an archive.((The 180 is a candidate set over the free-text ''%%detection.phenomenon%%'' field and therefore a floor, not a population: a paper that measured advertising without using an ad-shaped word in that field is invisible to it. The field is only about 20% stable run-to-run, so treat 180 as an order of magnitude and 5.6% as a ratio between two candidate sets.)) Hand-assigned per paper from its own methods section, the archives used are **Meta 8, Google 3, X 1, Apple 1** — a paper can use several, so these do not sum to 11, and {[gkiouzepi2023_collaborative]} queries no archive at all. Nine of the thirteen repositories have **no** user in these venues. The assignment is by hand and not by regex on purpose: {[zeng2021_polls]} writes "//Google's (or others') political ad transparency reports//", whose parenthetical defeats any pattern for //Google's political ad …//, and {[benzaamia2026_year]} studies four platforms' repositories at once. Set against a floor of **180** papers in this corpus whose measured phenomenon is advertising-shaped, the audited set is **10** of them — about **5.6%** of advertising measurement in these venues touches an archive.((The 180 is a candidate set over the free-text ''%%detection.phenomenon%%'' field and therefore a floor, not a population: a paper that measured advertising without using an ad-shaped word in that field is invisible to it. The field is only about 20% stable run-to-run, so treat 180 as an order of magnitude and 5.6% as a ratio between two candidate sets.))
Line 232: Line 237:
 | **55%** false-positive rate among US ads Facebook flagged as political; **4.5%** false negatives worldwide (**116,963** ads), **0.85%** in the US against up to **45%** elsewhere; **40%** of ads detected in under a day {[lepochat2022_audit]} | 33.8 M unique ads observed; 4.2 M political and 29.6 M non-political from all 215,030 pages that ran a political ad, mid-2020 to early 2021 | | **55%** false-positive rate among US ads Facebook flagged as political; **4.5%** false negatives worldwide (**116,963** ads), **0.85%** in the US against up to **45%** elsewhere; **40%** of ads detected in under a day {[lepochat2022_audit]} | 33.8 M unique ads observed; 4.2 M political and 29.6 M non-political from all 215,030 pages that ran a political ad, mid-2020 to early 2021 |
 | **7.7%** moderation recall on undeclared political ads; **60.4%** of Meta's own moderations outside Meta's stated criteria; a conservative **1,217** undeclared political ads launched daily; **92** pages moderated 28+ times, **82** still active {[bouchaud2024_beyond]} | all EU ads via the Ad Library API, 16 countries, Aug 2023 – Feb 2024; 29.5 M ads for inference | | **7.7%** moderation recall on undeclared political ads; **60.4%** of Meta's own moderations outside Meta's stated criteria; a conservative **1,217** undeclared political ads launched daily; **92** pages moderated 28+ times, **82** still active {[bouchaud2024_beyond]} | all EU ads via the Ad Library API, 16 countries, Aug 2023 – Feb 2024; 29.5 M ads for inference |
-| **48%** of Facebook and **9%** of Instagram ads seen by real users were missing from the platform'DSA repository; **98.9%** of YouTube explanations "//cite only the main targeting form//" {[benzaamia2026_year]} | 351 Facebook and 438 Instagram ads matched, from 48,511 donated ad explanations across four platforms |+| **48%** of Facebook and **9%** of Instagram ads seen by real users could not be matched to a DSA repository entry (an upper bound on missingness); **98.9%** of YouTube explanations "//cite only the main targeting form//"; **1,458 of 2,190** X repository entries matched by Tweet ID {[benzaamia2026_year]} | 351 Facebook and 438 Instagram ads matched, from 48,511 donated ad explanations across four platforms |
 | Only **34 of 835** ads a classifier judged political had a matching archive entry {[silva2020_facebook]} | 38,110 ads collected from Brazilian volunteers, 2018 | | Only **34 of 835** ads a classifier judged political had a matching archive entry {[silva2020_facebook]} | 38,110 ads collected from Brazilian volunteers, 2018 |
 | Play Store ads were "//entirely missing from the Ad Transparency dataset//", and the web interface showed fewer ads than the BigQuery dataset for the same advertiser {[breuer2026_ad]} | 257,820 ads and 378,057 recommendations from 249 accounts on the EU Apple and Google app stores | | Play Store ads were "//entirely missing from the Ad Transparency dataset//", and the web interface showed fewer ads than the BigQuery dataset for the same advertiser {[breuer2026_ad]} | 257,820 ads and 378,057 recommendations from 249 accounts on the EU Apple and Google app stores |
Line 261: Line 266:
 | **Amazon** Store Ad Library API | **current, free, EU-only, zero published use** | own documentation 2026-09-08; obligation upheld by the Court of Justice on 27 March 2024 | | **Amazon** Store Ad Library API | **current, free, EU-only, zero published use** | own documentation 2026-09-08; obligation upheld by the Court of Justice on 27 March 2024 |
 | The **EU European repository** for political ads | **specified, not yet a usable instrument here** | Implementing Reg. (EU) 2026/818, 9 April 2026; 0 papers | | The **EU European repository** for political ads | **specified, not yet a usable instrument here** | Implementing Reg. (EU) 2026/818, 9 April 2026; 0 papers |
-| Treating an archive as **ground truth** | **superseded** | {[breuer2026_ad]} rejected it explicitly; {[silva2020_facebook]}, {[lepochat2022_audit]}, {[bouchaud2024_beyond]}{[benzaamia2026_year]} all measure the archive's error instead |+| Treating an archive as **ground truth** | **don't** | not a method that was standard and got replaced: {[breuer2026_ad]} tried it in 2026 and could notand four papers spanning 2020–2026 measure the archive's error instead of assuming it away. Using an archive as a **corpus** is a different and perfectly current thing — {[capozzi2023_thin]}, {[coelho2023_propaganda]} and {[bitaab2025_scammagnifier]} all do it |
 | **Two-sided error analysis** against the platform's published criteria | **current best practice** | {[lepochat2022_audit]}, {[bouchaud2024_beyond]} | | **Two-sided error analysis** against the platform's published criteria | **current best practice** | {[lepochat2022_audit]}, {[bouchaud2024_beyond]} |
 | **Donation panel** as the archive's control | **current and rising** | {[benzaamia2026_year]}, {[ali2023_problematic]}, {[chen2026_when]}; 25 papers mention data donation, 11 in 2025 | | **Donation panel** as the archive's control | **current and rising** | {[benzaamia2026_year]}, {[ali2023_problematic]}, {[chen2026_when]}; 25 papers mention data donation, 11 in 2025 |
Line 284: Line 289:
  
 <WRAP todo> <WRAP todo>
-  * **Nine of the thirteen repositories have never been used in these seven venues**, including TikTok'Commercial Content API, Amazon's free Ad Library API and every retail repository. X's has one user {[benzaamia2026_year]}which took the per-advertiser CSV and matched 1,458 of 2,190 entries by Tweet ID — so somebody has shown it can be done, and nobody has characterised the file as whole.+  * **Nine of the thirteen repositories have never been used in these seven venues** — TikTok's, LinkedIn's, Bing's, Snapchat's, Amazon's free API and every retail one. X's has a single user, who matched into it from a donated set rather than characterising the file itself. Any of the nine would be first methods section.
   * **What does the October 2025 cliff do to EU political-ad research?** The instrument that produced this literature's best work no longer receives new data in its most-studied jurisdiction, and zero papers in this corpus have noticed. Somebody has to write the post-TTPA design.   * **What does the October 2025 cliff do to EU political-ad research?** The instrument that produced this literature's best work no longer receives new data in its most-studied jurisdiction, and zero papers in this corpus have noticed. Somebody has to write the post-TTPA design.
   * **Is the European repository going to be usable?** Implementing Reg. (EU) 2026/818 specifies a data structure and an API. Whether it aggregates small publishers usefully, and whether it is queryable in bulk, is an empirical question nobody has answered.   * **Is the European repository going to be usable?** Implementing Reg. (EU) 2026/818 specifies a data structure and an API. Whether it aggregates small publishers usefully, and whether it is queryable in bulk, is an empirical question nobody has answered.
design/platforms/ad_archives.txt · Last modified: by karel.kubicek.claude