| Next revision | Previous revision |
| privacy:ads_txt [2026/09/02 17:05] – New page: measuring the ad supply chain (ads.txt, sellers.json, header bidding), corpus-backed. Authored by Claude karel.kubicek.claude | privacy:ads_txt [2026/09/02 17:33] (current) – Generic-review fixes: qualified universals, unit caveats on the HB split, bid-gap confounds, dating, non-crawl papers named. Authored by Claude karel.kubicek.claude |
|---|
| |
| <WRAP important> | <WRAP important> |
| **The load-bearing decision on this page is what you count as a "relationship".** The same pair (publisher, ad system) can be: **declared** — a line in the publisher's ''ads.txt''; **confirmed** — that line's seller ID also appears in the ad system's ''sellers.json'', and the ''sellers.json'' domain points back at the publisher; **observed** — a bid, or an ad inclusion chain, that actually passed through that ad system on a page your crawler loaded; or **paid** — money moved, which no external measurement sees. Papers in these venues report all of the first three and are not consistent about which. Bashir et al. {[bashir2019_adstxt]} needed 135M observed ad inclusion chains to test whether declarations were honoured; since 2022 the field has settled on the //confirmed// relationship — the ''ads.txt'' ∩ ''sellers.json'' intersection — because "each individual file of the above two files may provide fake information, their intersection provides the truth" {[papadogiannakis2025_before]}. State which of the four you measured, in the abstract. | **The load-bearing decision on this page is what you count as a "relationship".** The same pair (publisher, ad system) can be: **declared** — a line in the publisher's ''ads.txt''; **confirmed** — that line's seller ID also appears in the ad system's ''sellers.json'', and the ''sellers.json'' domain points back at the publisher; **observed** — a bid, or an ad inclusion chain, that actually passed through that ad system on a page your crawler loaded; or **paid** — money moved, which no external measurement sees. Papers in these venues report all of the first three and are not consistent about which. Bashir et al. {[bashir2019_adstxt]} needed 135M observed ad inclusion chains to test whether declarations were honoured; since 2022 ({[vekaria2024_darkpooling]}, then {[papadogiannakis2023_whofunds]} in these venues) every ''ads.txt'' paper — four in the corpus, three of them from one group — has reported the //confirmed// relationship, the ''ads.txt'' ∩ ''sellers.json'' intersection, because "each individual file of the above two files may provide fake information, their intersection provides the truth" {[papadogiannakis2025_before]}. State which of the four you measured, in the abstract. |
| </WRAP> | </WRAP> |
| |
| |
| * **A Longitudinal Analysis of the ads.txt Standard** {[bashir2019_adstxt]}, IMC 2019 — the reference measurement of the standard itself: 26 crawls of the Alexa Top-100K over 15 months, the record-error taxonomy, and the only compliance test in these venues that matched declarations against //observed// ad traffic (135M RTB inclusion chains). Read Section 3 for the crawl, Section 5 for why compliance cannot be tested without becoming a publisher, and Section 8 for the caveats you will be quoting. | * **A Longitudinal Analysis of the ads.txt Standard** {[bashir2019_adstxt]}, IMC 2019 — the reference measurement of the standard itself: 26 crawls of the Alexa Top-100K over 15 months, the record-error taxonomy, and the only compliance test in these venues that matched declarations against //observed// ad traffic (135M RTB inclusion chains). Read Section 3 for the crawl, Section 5 for why compliance cannot be tested without becoming a publisher, and Section 8 for the caveats you will be quoting. |
| * **The Inventory is Dark and Full of Misinformation** {[vekaria2024_darkpooling]}, IEEE S&P 2024 — the paper that turned ''ads.txt'' and ''sellers.json'' from an adoption question into a fraud-detection instrument. It names the misrepresentation classes, defines **dark pooling** (unrelated publishers sharing one seller ID), and is the only paper anywhere in this literature that parsed the OpenRTB SupplyChain object from captured bid requests. **It is not in the corpus behind this site** — the screening model labelled it neither a security nor a privacy measurement — so nothing in [[#Use in Publications]] counts it. Read it from the arXiv version. | * **The Inventory is Dark and Full of Misinformation** {[vekaria2024_darkpooling]}, IEEE S&P 2024 — the paper that turned ''ads.txt'' and ''sellers.json'' from an adoption question into a fraud-detection instrument. It names the misrepresentation classes, defines **dark pooling** (unrelated publishers sharing one seller ID), and is the only paper cited on this page that parsed the OpenRTB SupplyChain object from captured bid requests. **It is not in the corpus behind this site** — the screening model labelled it neither a security nor a privacy measurement — so nothing in [[#Use in Publications]] counts it. Read it from the arXiv version. |
| * **Welcome to the Dark Side** {[papadogiannakis2025_darkside]}, TheWebConf 2025 — the scale study: ''ads.txt'' from 456,971 of ~7M Tranco domains, ''sellers.json'' from 2,682 ad systems, 185,535 seller-ID pools, ~15,000 of them dark. Released crawlers and a public monitoring service. Read it for what a 2023-era pipeline looks like end to end, and for the ''sellers.json'' confidentiality figures. | * **Welcome to the Dark Side** {[papadogiannakis2025_darkside]}, TheWebConf 2025 — the scale study: ''ads.txt'' from 456,971 of ~7M Tranco domains, ''sellers.json'' from 2,682 ad systems, 185,535 seller-ID pools, ~15,000 of them dark. Released crawlers and a public monitoring service. Read it for what a 2023-era pipeline looks like end to end, and for the ''sellers.json'' confidentiality figures. |
| * **No More Chasing Waterfalls** {[pachilakis2019_waterfalls]}, IMC 2019 — the header-bidding reference: how to detect an auction from DOM events and requests, the client/server/hybrid split (server-side was already 48% of HB sites in 2019), and latency. Its tool link is dead (see [[#Tools, datasets and services that already exist]]). | * **No More Chasing Waterfalls** {[pachilakis2019_waterfalls]}, IMC 2019 — the header-bidding reference: how to detect an auction from DOM events and requests, the client/server/hybrid split (server-side was already 48% in 2019, in a breakdown whose unit the paper does not state), and latency. Its tool link is dead (see [[#Tools, datasets and services that already exist]]). |
| * **Inferring Tracker-Advertiser Relationships … using Header Bidding** {[cook2020_headerbidding]}, PoPETs 2020 — the paper that made bids a //dependent variable//: call ''getBidResponses()'', change one thing about the persona, watch who bids more. Every PETS paper since that uses bids as evidence of data use ({[musa2022_atom]}, {[iqbal2023_tracking]}, {[liu2024_opted]}, {[liu2025_fingerprinting]}) is a descendant. | * **Inferring Tracker-Advertiser Relationships … using Header Bidding** {[cook2020_headerbidding]}, PoPETs 2020 — the paper that made bids a //dependent variable//: call ''getBidResponses()'', change one thing about the persona, watch who bids more. Every paper in these venues since that uses bids as evidence of data use ({[musa2022_atom]}, {[iqbal2023_tracking]}, {[liu2024_opted]}, {[liu2025_fingerprinting]}) is a descendant. |
| |
| ===== The Three Observables, and What Each Cannot See ===== | ===== The Three Observables, and What Each Cannot See ===== |
| | ''ads.txt'' | ''https://<root domain>/ads.txt'', root domain = public suffix + 1; redirects are authoritative only if they stay inside that root domain((ads.txt 1.1 §3.1: "Crawlers should incorporate Public Suffix list [16] to derive the root domain"; "the advertising system should follow the redirect and consume the data as authoritative for the source of the redirect, if and only if the redirect is within scope of the original root domain". Default cache expiry 7 days if no cache-control headers.)) | The publisher | //ad system X may sell my inventory under account ID Y, as DIRECT or RESELLER// | Whether X ever did. Whether Y is really this publisher's account: 10% of publishers had ≥1 invalid record and ~10% of files were byte-identical copies distributed by more than one publisher in 2018–2019 {[bashir2019_adstxt]}; the DIRECT ID ''100141'' of one ad network appeared in 42,412 sites in 2023 {[papadogiannakis2025_darkside]} | | | ''ads.txt'' | ''https://<root domain>/ads.txt'', root domain = public suffix + 1; redirects are authoritative only if they stay inside that root domain((ads.txt 1.1 §3.1: "Crawlers should incorporate Public Suffix list [16] to derive the root domain"; "the advertising system should follow the redirect and consume the data as authoritative for the source of the redirect, if and only if the redirect is within scope of the original root domain". Default cache expiry 7 days if no cache-control headers.)) | The publisher | //ad system X may sell my inventory under account ID Y, as DIRECT or RESELLER// | Whether X ever did. Whether Y is really this publisher's account: 10% of publishers had ≥1 invalid record and ~10% of files were byte-identical copies distributed by more than one publisher in 2018–2019 {[bashir2019_adstxt]}; the DIRECT ID ''100141'' of one ad network appeared in 42,412 sites in 2023 {[papadogiannakis2025_darkside]} | |
| | ''app-ads.txt'' | Same file on the **developer website** named in the app-store listing, not on the app | The app developer | Same as above, for app inventory | Anything at all in these seven venues: no paper has crawled it. It appears only as background in two mobile ad-fraud papers | | | ''app-ads.txt'' | Same file on the **developer website** named in the app-store listing, not on the app | The app developer | Same as above, for app inventory | Anything at all in these seven venues: no paper has crawled it. It appears only as background in two mobile ad-fraud papers | |
| | ''sellers.json'' | ''https://<ad system domain>/sellers.json''; only ad systems publish one — 2,682 of 7,341,165 Tranco domains in March 2023 {[papadogiannakis2025_darkside]} | The ad system (SSP / exchange) | //seller ID Y is paid by us; it is a PUBLISHER, an INTERMEDIARY or BOTH; its name and domain are …// — unless ''is_confidential'' is set, in which case name and domain are withheld | Who the seller is, for the confidential majority: 75.53% of Google's 1,277,156 seller IDs were confidential in March 2023 {[papadogiannakis2025_darkside]}, and 70.96% of 972,929 on 2026-09-02.((Google's ''sellers.json'', fetched 2026-09-02, 108,724,970 bytes: 972,929 sellers, 690,433 with ''is_confidential = 1''; 972,347 typed PUBLISHER, 582 BOTH; the file's own ''ext.notice'' still reads "This file is a beta and is unverified." The canonical URL ''realtimebidding.google.com/sellers.json'' returned an empty 200 body on most fetches during the run; the byte-identical copy at ''storage.googleapis.com/adx-rtb-dictionaries/sellers.json'' did not. Fetch script and output on the provenance page.)) Whether the seller actually sold anything | | | ''sellers.json'' | ''https://<ad system domain>/sellers.json''; only ad systems publish one — 2,682 of 7,341,165 Tranco domains in March 2023 {[papadogiannakis2025_darkside]} | The ad system (SSP / exchange) | //seller ID Y is paid by us; it is a PUBLISHER, an INTERMEDIARY or BOTH; its name and domain are …// — unless ''is_confidential'' is set, in which case name and domain are withheld | Who the seller is, for the confidential majority: 75.53% of Google's 1,277,156 seller IDs were confidential in March 2023 {[papadogiannakis2025_darkside]}, and 70.96% of 972,929 on 2026-09-02.((Google's ''sellers.json'', fetched 2026-09-02, 108,724,970 bytes: 972,929 sellers, 690,433 with ''is_confidential = 1''; 972,347 typed PUBLISHER, 582 BOTH; the file's own ''ext.notice'' still reads "This file is a beta and is unverified." The canonical URL ''realtimebidding.google.com/sellers.json'' answers with a 302 to ''storage.googleapis.com/adx-rtb-dictionaries/sellers.json''; a fetcher that does not follow redirects records a 0-byte body. Fetch script and output on the provenance page.)) Whether the seller actually sold anything | |
| | **Header bidding (client-side)** | The page's Prebid.js global — ''pbjs.getBidResponses()'', ''getAllWinningBids()'' — or the auction's DOM events and bidder requests | Nobody declares it; you observe it | //bidder B offered CPM c for ad unit U on this page load, for this browser// | Server-side auctions: Prebid Server, Amazon TAM and Google's Exchange Bidding make "bid requests from the server rather than the browser"((''docs.prebid.org/prebid-server/overview/prebid-server-overview.html'', fetched 2026-09-02.)) and are invisible from the client; server-side and hybrid HB were already 48% and 34.7% of HB sites in 2019 {[pachilakis2019_waterfalls]}. What was //paid//: a winning bid is not a rendered ad — 25,764 winning bids but 7,117 rendered in {[zeng2022_factors]} | | | **Header bidding (client-side)** | The page's Prebid.js global — ''pbjs.getBidResponses()'', ''getAllWinningBids()'' — or the auction's DOM events and bidder requests | Nobody declares it; you observe it | //bidder B offered CPM c for ad unit U on this page load, for this browser// | Server-side auctions: Prebid Server — "Use Prebid Server to do the processing rather than the client browser"((''docs.prebid.org/overview/intro.html'', fetched 2026-09-02.)) — Amazon TAM and Google's Exchange Bidding are invisible from the client; server-side and hybrid HB were already 48% and 34.7% of the 2019 "facet breakdown" {[pachilakis2019_waterfalls]}, whose unit — sites or auctions — the paper does not state. What was //paid//: a winning bid is not a rendered ad — 25,764 winning bids but 7,117 rendered in {[zeng2022_factors]} | |
| | **SupplyChain object** (''schain'') | Inside OpenRTB bid requests, i.e. only if you are a bidder or you capture the requests a client-side wrapper sends | Each hop of the chain | //this impression passed through (asi, sid) nodes …// | Nearly everything, in practice: only 20.5% of captured bid requests carried one in 2022, all single-node {[vekaria2024_darkpooling]}. **Zero papers in the seven venues mention it** | | | **SupplyChain object** (''schain'') | Inside OpenRTB bid requests, i.e. only if you are a bidder or you capture the requests a client-side wrapper sends | Each hop of the chain | //this impression passed through (asi, sid) nodes …// | Nearly everything, in practice: only 20.5% of captured bid requests carried one in 2022, all single-node {[vekaria2024_darkpooling]}. **Zero papers in the seven venues mention it** | |
| |
| |
| ^ Unit ^ A published figure in that unit ^ Its denominator, stated ^ | ^ Unit ^ A published figure in that unit ^ Its denominator, stated ^ |
| | **Sites serving a valid ''ads.txt''** | 12.7% → 19.7% between January 2018 and April 2019 {[bashir2019_adstxt]} | Alexa Top-100K, growing to 240K sites as the list was re-fetched; 26 crawls | | | **Sites serving a valid ''ads.txt''** | 12.7% → 19.7% between January 2018 and April 2019 {[bashir2019_adstxt]} | Alexa Top-100K (the January-2018 set; a re-fetched 100K gives the same curve); 26 crawls. The paper's full crawl was the union of every list, 240K sites | |
| | **Sites serving one, conditioned on showing RTB ads** | 46.6% → 62.3% over the same 15 months {[bashir2019_adstxt]} | The subset of those sites on which the paper's own browser crawl saw RTB advertisements | | | **Sites serving one, conditioned on showing RTB ads** | 46.6% → 62.3% over the same 15 months {[bashir2019_adstxt]} | The subset of those sites on which the paper's own browser crawl saw RTB advertisements | |
| | **Sites serving one, at the scale of the whole list** | 456,971 domains {[papadogiannakis2025_darkside]} | ~7M Tranco domains, February–March 2023, fetched with the IAB reference crawler | | | **Sites serving one, at the scale of the whole list** | 456,971 domains {[papadogiannakis2025_darkside]} | ~7M Tranco domains, February–March 2023, fetched with the IAB reference crawler | |
| | **Bid ratio under a treatment** | up to 30× the control mean {[iqbal2023_tracking]}; 6.28× {[zhang2022_harpo]}; personas bid //higher// than control after opting out {[liu2024_opted]} | 200, 3 and 352 Prebid sites respectively; the treatment differs in each | | | **Bid ratio under a treatment** | up to 30× the control mean {[iqbal2023_tracking]}; 6.28× {[zhang2022_harpo]}; personas bid //higher// than control after opting out {[liu2024_opted]} | 200, 3 and 352 Prebid sites respectively; the treatment differs in each | |
| |
| **Three consequences.** The 2019 "20%" and the 2023 "456,971" are not two points on one adoption curve; one is a share of a top list, the other a count over a seven-million-domain tail. A bid from a stateless crawler is a price for //nobody//, and more than a hundred times lower than what real users' browsers received two years later — not because prices rose. And every ''sellers.json'' figure has a hidden denominator: whatever share of that exchange's sellers is confidential. | **Three consequences.** The 2019 "20%" and the 2023 "456,971" are not two points on one adoption curve; one is a share of a top list, the other a count over a seven-million-domain tail. A stateless crawler's bids are a price for //nobody//: 0.031 CPM median for a 300×250 slot in 2019, against $4.16 median //winning// bid from real users on ten US news sites in December 2021. The gap is two orders of magnitude, and slot mix, winning-versus-all bids, site selection and season are all inside it; no paper has run the two personas side by side. And every ''sellers.json'' figure has a hidden denominator: whatever share of that exchange's sellers is confidential. |
| |
| ===== Methods, and Which Ones Are Current ===== | ===== Methods, and Which Ones Are Current ===== |
| ^ Family ^ What it does ^ First / most recent in corpus ^ Papers ^ Status in 2026 ^ | ^ Family ^ What it does ^ First / most recent in corpus ^ Papers ^ Status in 2026 ^ |
| | **Fetch-and-parse ''ads.txt'' at list scale** | ''GET /ads.txt'' on every root domain of a top list; parse the four fields; count sites, records, systems, DIRECT vs RESELLER, errors | 2019 → 2025 | 5 — {[bashir2019_adstxt]}, {[papadogiannakis2023_whofunds]}, {[papadogiannakis2025_darkside]}, {[papadogiannakis2025_before]}, {[sheaib2025_unmasking]} (+ {[vekaria2024_darkpooling]} outside the corpus) | **Current and cheap.** Bashir's 100K sites took "2–3 hours" on a 16-node cluster; the 2025 study used the IAB reference crawler with only the User-Agent changed. Re-crawl: every 15–30 days {[bashir2019_adstxt]}, or one request per domain per month {[papadogiannakis2025_darkside]}. Alone it yields //declared// relationships only | | | **Fetch-and-parse ''ads.txt'' at list scale** | ''GET /ads.txt'' on every root domain of a top list; parse the four fields; count sites, records, systems, DIRECT vs RESELLER, errors | 2019 → 2025 | 5 — {[bashir2019_adstxt]}, {[papadogiannakis2023_whofunds]}, {[papadogiannakis2025_darkside]}, {[papadogiannakis2025_before]}, {[sheaib2025_unmasking]} (+ {[vekaria2024_darkpooling]} outside the corpus) | **Current and cheap.** Bashir's 100K sites took "2–3 hours" on a 16-node cluster; the 2025 study used the IAB reference crawler with only the User-Agent changed. Re-crawl: every 15–30 days {[bashir2019_adstxt]}, or one request per domain per month {[papadogiannakis2025_darkside]}. Alone it yields //declared// relationships only | |
| | **''ads.txt'' ∩ ''sellers.json'' intersection** | For each DIRECT record, look the seller ID up in the named ad system's ''sellers.json''; keep the relationship only if the exchange confirms it and its ''domain'' matches | 2023 → 2025 | 4 — {[papadogiannakis2023_whofunds]}, {[papadogiannakis2025_darkside]}, {[papadogiannakis2025_before]}, {[sheaib2025_unmasking]} (+ {[vekaria2024_darkpooling]}) | **Current default for a //confirmed// relationship.** Recursive: an INTERMEDIARY's domain should host its own ''sellers.json'', so crawl the graph, not a list {[vekaria2024_darkpooling]}. Bounded by confidentiality — of 188 DIRECT entries on piracy sites, 5 could be confirmed; 169 IDs were absent and 418 had a domain mismatch {[sheaib2025_unmasking]} | | | **''ads.txt'' ∩ ''sellers.json'' intersection** | For each DIRECT record, look the seller ID up in the named ad system's ''sellers.json''; keep the relationship only if the exchange confirms it and its ''domain'' matches | 2023 → 2025 | 4 — {[papadogiannakis2023_whofunds]}, {[papadogiannakis2025_darkside]}, {[papadogiannakis2025_before]}, {[sheaib2025_unmasking]} (+ {[vekaria2024_darkpooling]}) | **Current default for a //confirmed// relationship.** Recursive: an INTERMEDIARY's domain should host its own ''sellers.json'', so crawl the graph, not a list {[vekaria2024_darkpooling]}. Bounded by confidentiality — on piracy sites only 5 of 188 DIRECT seller entries could be confirmed; among the 587 further unique DIRECT entries the paper counts separately, 169 IDs were absent from ''sellers.json'' and 418 had a domain mismatch, and the paper does not reconcile the two totals {[sheaib2025_unmasking]} | |
| | **Pooling: shared seller IDs → ownership check** | Group sites by (ad system, seller ID); call a pool //dark// when its members have different owners, established from WHOIS, page-level Google publisher IDs, or by hand | 2025 → 2025 (2022 outside the corpus) | 1 — {[papadogiannakis2025_darkside]} (+ {[vekaria2024_darkpooling]}); ownership method from {[papadogiannakis2022_publisherids]} | **Current, and where the 2022–2026 findings come from.** The ownership step is the weak link: WHOIS was redacted for most of 185,535 pools' members (owner found for 3,981 sites), and page-level publisher IDs are Google-only | | | **Pooling: shared seller IDs → ownership check** | Group sites by (ad system, seller ID); call a pool //dark// when its members have different owners, established from WHOIS, page-level Google publisher IDs, or by hand | 2025 → 2025 (2022 outside the corpus) | 1 — {[papadogiannakis2025_darkside]} (+ {[vekaria2024_darkpooling]}); ownership method from {[papadogiannakis2022_publisherids]} | **Current, and where the 2022–2026 findings come from.** The ownership step is the weak link: WHOIS was redacted for most of 185,535 pools' members (owner found for 3,981 sites), and page-level publisher IDs are Google-only | |
| | **Declarations vs observed transactions** | Build ad inclusion chains from a browser crawl, attribute each RTB ad to a seller–buyer pair, check the seller is in the publisher's file; or parse the SupplyChain object out of captured bid requests | 2019 → 2019 (2022 outside the corpus) | 1 — {[bashir2019_adstxt]} (+ {[vekaria2024_darkpooling]}) | **The only test of whether the standard is honoured, and not repeated in these venues since 2019.** Expensive — 135M chains, EasyList to find the ads, WhoTracksMe to fold seller domains into 28 parents — and ethically awkward: the authors "attempted to become a publisher in order to conduct controlled experiments" and every exchange refused | | | **Declarations vs observed transactions** | Build ad inclusion chains from a browser crawl, attribute each RTB ad to a seller–buyer pair, check the seller is in the publisher's file; or parse the SupplyChain object out of captured bid requests | 2019 → 2019 (2022 outside the corpus) | 1 — {[bashir2019_adstxt]} (+ {[vekaria2024_darkpooling]}) | **The only test of whether the standard is honoured, and not repeated in these venues since 2019.** Expensive — 135M chains, EasyList to find the ads, WhoTracksMe to fold seller domains into 28 parents — and ethically awkward: the authors "attempted to become a publisher in order to conduct controlled experiments" and every exchange they contacted refused | |
| | **Client-side header-bidding instrumentation** | Detect Prebid.js (probe ''pbjs.version''), then read bids via the publisher API or by hooking auction events; record bidder, CPM, ad unit, latency | 2019 → 2025 | 8 — {[pachilakis2019_waterfalls]}, {[cook2020_headerbidding]}, {[zeng2022_factors]}, {[zhang2022_harpo]}, {[musa2022_atom]}, {[iqbal2023_tracking]}, {[liu2024_opted]}, {[liu2025_fingerprinting]} | **Current, with a visibility problem nobody has re-measured.** Client-side HB was 17.3% of HB sites in 2019 and ATOM's authors already wrote that bid-based inference "takes a hit as publishers migrate towards server side Header Bidding" {[musa2022_atom]}. No paper in these venues has measured the client/server split since 2019 | | | **Client-side header-bidding instrumentation** | Detect Prebid.js (probe ''pbjs.version''), then read bids via the publisher API or by hooking auction events; record bidder, CPM, ad unit, latency | 2019 → 2025 | 8 — {[pachilakis2019_waterfalls]}, {[cook2020_headerbidding]}, {[zeng2022_factors]}, {[zhang2022_harpo]}, {[musa2022_atom]}, {[iqbal2023_tracking]}, {[liu2024_opted]}, {[liu2025_fingerprinting]} | **Current, with a visibility problem nobody has re-measured.** Client-side HB was 17.3% of the 2019 facet breakdown (unit unstated) and ATOM's authors already wrote that bid-based inference "takes a hit as publishers migrate towards server side Header Bidding" {[musa2022_atom]}. No paper in these venues has measured the client/server split since 2019 | |
| | **Bids as a dependent variable** | Hold the site fixed, vary the persona (interests, opt-out, fingerprint, smart-speaker history), compare bid distributions against a control | 2020 → 2025 | 6 — {[cook2020_headerbidding]}, {[zeng2022_factors]}, {[zhang2022_harpo]}, {[iqbal2023_tracking]}, {[liu2024_opted]}, {[liu2025_fingerprinting]} | **Current, and PETS-centred.** It is the only way to see //server-side// data use from the client: advertisers who were never sent the interest still bid higher {[liu2024_opted]}. Confounds are the method's whole difficulty — day of week, slot, holidays {[iqbal2023_tracking]}; Christmas demand {[zeng2022_factors]} | | | **Bids as a dependent variable** | Hold the site fixed and vary the persona (interests, opt-out, fingerprint, smart-speaker history) against a control — or, in the one field study, compare bids across real users' attributes | 2020 → 2025 | 6 — {[cook2020_headerbidding]}, {[zeng2022_factors]}, {[zhang2022_harpo]}, {[iqbal2023_tracking]}, {[liu2024_opted]}, {[liu2025_fingerprinting]} | **Current, and PETS-centred.** It is the only way to see //server-side// data use from the client: advertisers who were never sent the interest still bid higher {[liu2024_opted]}. Confounds are the method's whole difficulty — day of week, slot, holidays {[iqbal2023_tracking]}; Christmas demand {[zeng2022_factors]} | |
| | **Prebid presence as a sampling frame** | Use a Prebid-detected site list (BuiltWith) to pick ad-supported sites, and a Prebid-aware extension to count ads | 2026 → 2026 | 1 — {[lukic2026_mv3]} | Legitimate, but it measures Prebid sites, not the web; say so in the population section | | | **Prebid presence as a sampling frame** | Use a Prebid-detected site list (BuiltWith) to pick ad-supported sites, and a Prebid-aware extension to count ads | 2026 → 2026 | 1 — {[lukic2026_mv3]} | Legitimate, but it measures Prebid sites, not the web; say so in the population section | |
| | **RTB win-notification prices** | Detect the exchange's price-notification URLs (''nURL'' macros) in users' traffic and read or model the cleartext/encrypted clearing price | 2017 → 2018 | 2 — {[papadopoulos2017_paying]}, {[papadopoulos2018_cost]} | **Historical.** It needs real users' traffic through a proxy and cleartext price macros; ~26% of mobile RTB prices were already encrypted in 2015 and encrypted prices ran ~1.7× higher {[papadopoulos2017_paying]}. No paper in these venues has used it since 2018; the Prebid API replaced it as the price observable | | | **RTB win-notification prices** | Detect the exchange's price-notification URLs (''nURL'' macros) in users' traffic and read or model the cleartext/encrypted clearing price | 2017 → 2018 | 2 — {[papadopoulos2017_paying]}, {[papadopoulos2018_cost]} | **Historical.** It needs real users' traffic through a proxy and cleartext price macros; ~26% of mobile RTB prices were already encrypted in 2015 and encrypted prices ran ~1.7× higher {[papadopoulos2017_paying]}. No paper in these venues has used it since 2018; the Prebid API replaced it as the price observable | |
| ==== What is superseded, and what has never been measured ==== | ==== What is superseded, and what has never been measured ==== |
| |
| * **''ads.txt'' alone as evidence of a relationship.** Bashir et al. could only test declarations against traffic. Since {[papadogiannakis2023_whofunds]} the intersection with ''sellers.json'' is the norm, and the 2025 papers treat a DIRECT line the exchange does not confirm as unverified, not as a relationship. If you can only fetch one file, you have measured //declarations//, and the title of your figure should say so. | * **''ads.txt'' alone as evidence of a relationship.** Bashir et al. could only test declarations against traffic. Since {[vekaria2024_darkpooling]} (crawl of February 2022) and {[papadogiannakis2023_whofunds]} the intersection with ''sellers.json'' is what every ''ads.txt'' paper reports — four in the corpus, three from one group — and the 2025 papers treat a DIRECT line the exchange does not confirm as unverified, not as a relationship. If you can only fetch one file, you have measured //declarations//, and the title of your figure should say so. |
| * **Alexa as the sampling frame.** 8 of the 16 measuring papers drew from Alexa; the list no longer exists, and the two 2019 reference studies are therefore unrepeatable as designed. See [[Design:Website selection]]. | * **Alexa as the sampling frame.** 8 of the 16 measuring papers drew from Alexa; the list no longer exists, and the two 2019 reference studies are therefore unrepeatable as designed. See [[Design:Website selection]]. |
| * **Assuming the Prebid global is called ''pbjs''.** It is only the default: the build option ''globalVarName'' renames it and ''defineGlobal: false'' removes it.((''github.com/prebid/Prebid.js'' README, build customisation table: ''globalVarName'' — "Prebid global variable name", default ''pbjs''; ''defineGlobal'' — "If false, do not set a global variable". Fetched 2026-09-02.)) {[liu2024_opted]} says explicitly that it did "not consider personalized custom API labels (i.e., other than pbjs)", so its 5,421 is a lower bound. Detect the auction from its events and bidder requests as well {[pachilakis2019_waterfalls]}, or accept and state the undercount. | * **Assuming the Prebid global is called ''pbjs''.** It is only the default: the build option ''globalVarName'' renames it and ''defineGlobal: false'' removes it.((''github.com/prebid/Prebid.js'' README, build customisation table: ''globalVarName'' — "Prebid global variable name", default ''pbjs''; ''defineGlobal'' — "If false, do not set a global variable". Fetched 2026-09-02.)) {[liu2024_opted]} says explicitly that it did "not consider personalized custom API labels (i.e., other than pbjs)", so its 5,421 is a lower bound. Detect the auction from its events and bidder requests as well {[pachilakis2019_waterfalls]}, or accept and state the undercount. |
| * **OWNERDOMAIN and MANAGERDOMAIN** (''ads.txt'' 1.1, August 2022) exist precisely to make pooling legible — the owner of a pooled site and its exclusive monetiser are meant to be declared. **No paper, in or out of the corpus, has measured their adoption**; both 2025 studies still infer ownership from WHOIS. This is the cheapest open measurement on the page. | * **OWNERDOMAIN and MANAGERDOMAIN** (''ads.txt'' 1.1, August 2022) exist precisely to make pooling legible — the owner of a pooled site and its exclusive monetiser are meant to be declared. **No paper in the seven venues, and none among the papers they cite, has measured their adoption**; both 2025 studies still infer ownership from WHOIS. This is the cheapest open measurement on the page. |
| * **The SupplyChain object** has one measurement, outside the seven venues, from 2022: present in 20.5% of bid requests, always one node {[vekaria2024_darkpooling]}. ''buyers.json'' and ''ads.cert'' have none. | * **The SupplyChain object** has one measurement that this page could find, outside the seven venues, from 2022: present in 20.5% of bid requests, always one node {[vekaria2024_darkpooling]}. ''buyers.json'' and ''ads.cert'' have none. |
| * **''app-ads.txt''** has none in these venues. The two CCS papers that name it are about app-side fraud detected from an ad network's own bidding logs and treat the file as background. | * **''app-ads.txt''** has none in these venues. The two CCS papers that name it are mobile ad-fraud studies — one from an ad network's invalid-traffic logs, one from static and dynamic analysis of apps — and both treat the file as background. |
| * **The client/server header-bidding split** has not been re-measured since February 2019. Every bids-as-instrument result since is conditioned on the sites that still expose the auction, and none of them reports what fraction of its candidate list that was. | * **The client/server header-bidding split** has not been re-measured since February 2019. Every bids-as-instrument result since is conditioned on the sites that still expose the auction. Two report what share of their candidate list that was — {[liu2024_opted]} (5,421 of the Alexa top-100K) and {[zeng2022_factors]} (703 of the Tranco top 10,000); {[iqbal2023_tracking]} stopped at 200 Prebid sites and does not. The client / hybrid / server split //among// HB sites is what nobody has re-measured. |
| * **LLM-based methods**: none. The classification steps in this literature — ad system name folding, brand extraction from landing pages, misinformation labels — are regex, curated lists and hand review, and the 2025–2026 slice does not change that. | * **LLM-based methods**: none. The classification steps in this literature — ad system name folding, brand extraction from landing pages, misinformation labels — are regex, curated lists and hand review, and the 2025–2026 slice does not change that. |
| |
| * **Root domain, not hostname.** The file is authoritative for the public-suffix-plus-one domain; a crawl keyed on ''www.'' hostnames or on subdomains double-counts and misses. Follow redirects only within that root domain — the specification makes an off-domain redirect non-authoritative, and the IAB reference crawler implements that rule.((''github.com/InteractiveAdvertisingBureau/adstxtcrawler'' — "A reference implementation in python of a simple crawler for Ads.txt"; created 2017-05-26, last pushed 2024-06-06, no licence file. Fetched via the GitHub API 2026-09-02.)) | * **Root domain, not hostname.** The file is authoritative for the public-suffix-plus-one domain; a crawl keyed on ''www.'' hostnames or on subdomains double-counts and misses. Follow redirects only within that root domain — the specification makes an off-domain redirect non-authoritative, and the IAB reference crawler implements that rule.((''github.com/InteractiveAdvertisingBureau/adstxtcrawler'' — "A reference implementation in python of a simple crawler for Ads.txt"; created 2017-05-26, last pushed 2024-06-06, no licence file. Fetched via the GitHub API 2026-09-02.)) |
| * **Validate before you count.** 56.4% of the 2,381 seller domains ever named in Alexa Top-100K files failed WHOIS/DNS validation {[bashir2019_adstxt]}; 10% of publishers had at least one syntactically invalid record. Report records //that follow the specification// and say how many you dropped. | * **Validate before you count.** 56.4% of the 2,381 seller domains ever named in Alexa Top-100K files failed WHOIS/DNS validation {[bashir2019_adstxt]}; 10% of publishers had at least one syntactically invalid record. Report records //that follow the specification// and say how many you dropped. |
| * **A User-Agent that looks like a browser.** The 2025 scale crawl kept the IAB crawler "as-is and only change the user-agent header so that we are not blocked by websites" {[papadogiannakis2025_darkside]}. The same is true of ''sellers.json'': Google's 109 MB file came back as an empty 200 on most fetches from its canonical URL during the check for this page, whatever the User-Agent; the Cloud Storage copy did not. | * **A User-Agent that looks like a browser.** The 2025 scale crawl kept the IAB crawler "as-is and only change the user-agent header so that we are not blocked by websites" {[papadogiannakis2025_darkside]}. And follow redirects: Google's canonical ''sellers.json'' URL is a 302 to a Cloud Storage object, and a fetcher that stops at the first response records a 0-byte file. |
| * **Fold the ad-system names before you rank them.** ''google.com'', ''doubleclick.net'' and friends are one seller; Bashir et al. mapped 101 domains to 28 parents with WhoTracksMe and call the clustering "not perfect". Print the residue. Use the entity maps on [[Design:Existing datasets]] rather than writing your own alias list. | * **Fold the ad-system names before you rank them.** ''google.com'', ''doubleclick.net'' and friends are one seller; Bashir et al. mapped 101 domains to 28 parents with WhoTracksMe and call the clustering "not perfect". Print the residue. Use the entity maps on [[Design:Existing datasets]] rather than writing your own alias list. |
| * **''sellers.json'' is a graph crawl.** Start from the ad-system domains your ''ads.txt'' corpus names, fetch each ''/sellers.json'', then recurse into every INTERMEDIARY / BOTH domain listed {[vekaria2024_darkpooling]}, {[papadogiannakis2025_darkside]}. Expect non-standard locations for the biggest systems — Google's is not at ''google.com/sellers.json'' — and expect copies: 28 domains served a copy of Google's file {[papadogiannakis2025_darkside]}. | * **''sellers.json'' is a graph crawl.** Start from the ad-system domains your ''ads.txt'' corpus names, fetch each ''/sellers.json'', then recurse into every INTERMEDIARY / BOTH domain listed {[vekaria2024_darkpooling]}, {[papadogiannakis2025_darkside]}. Expect non-standard locations for the biggest systems — Google's is not at ''google.com/sellers.json'' — and expect copies: 28 domains served a copy of Google's file {[papadogiannakis2025_darkside]}. |
| ==== Observing header bidding ==== | ==== Observing header bidding ==== |
| |
| * **Detect, then read, then wait.** The pattern used by every instrumenting paper since 2020: probe ''pbjs.version'' (or hook the auction events) to find Prebid sites, then call ''pbjs.getBidResponses()'' — and ''requestBids()'' yourself if no auction ran {[liu2024_opted]}. The documented response object carries ''bidder'', ''cpm'', ''currency'', ''adUnitCode'', ''timeToRespond'' and ''mediaType''.((''docs.prebid.org/dev-docs/publisher-api-reference/getBidResponses.html'', fetched 2026-09-02.)) Auctions are slow: median 600 ms, 4% of sites over 5 s {[pachilakis2019_waterfalls]}; a crawler that leaves at ''load'' has no bids. | * **Detect, then read, then wait.** The pattern in {[cook2020_headerbidding]}, {[iqbal2023_tracking]} and {[liu2024_opted]}: probe ''pbjs.version'' (or hook the auction events) to find Prebid sites, then call ''pbjs.getBidResponses()'' — and ''requestBids()'' yourself if no auction ran {[liu2024_opted]}. The documented response object carries ''bidder'', ''cpm'', ''currency'', ''adUnitCode'', ''timeToRespond'' and ''mediaType''.((''docs.prebid.org/dev-docs/publisher-api-reference/getBidResponses.html'', fetched 2026-09-02.)) Auctions are slow: median 600 ms, 4% of sites over 5 s {[pachilakis2019_waterfalls]}; a crawler that leaves at ''load'' has no bids. |
| * **A bid is not a rendered ad, and a rendered ad is not a paid one.** {[zeng2022_factors]} distinguishes ''getBidResponses()'' (all bids), ''getAllWinningBids()'' (rendered) and ''getAllPrebidWinningBids()'' (won the auction, then lost to the ad server's waterfall) and analyses only the 7,117 rendered winners of 25,764. Choose one and name it. | * **A bid is not a rendered ad, and a rendered ad is not a paid one.** {[zeng2022_factors]} distinguishes ''getBidResponses()'' (all bids), ''getAllWinningBids()'' (rendered) and ''getAllPrebidWinningBids()'' (won the auction, then lost to the ad server's waterfall) and analyses only the 7,117 rendered winners of 25,764. Choose one and name it. |
| * **Your persona is the price.** A stateless clean browser is a baseline nobody targets: 0.031 CPM median in 2019 {[pachilakis2019_waterfalls]} against $4.16 from real users in 2021 {[zeng2022_factors]}. Every treatment design in the table above runs a control persona alongside, from the same IP range, on the same day; {[liu2024_opted]} ran 17 personas per jurisdiction from Frankfurt and Northern California and visited each site nine times. Fewer repetitions than that and day-of-week and slot effects swamp the treatment {[iqbal2023_tracking]}. | * **Your persona is the price.** A stateless clean browser is a baseline nobody targets: 0.031 CPM median bid in 2019 {[pachilakis2019_waterfalls]} against $4.16 median winning bid from real users in 2021 {[zeng2022_factors]} — not the same unit, but the same direction as every treatment study below. Every crawler-based treatment design in the table above ({[cook2020_headerbidding]}, {[zhang2022_harpo]}, {[iqbal2023_tracking]}, {[liu2024_opted]}, {[liu2025_fingerprinting]}) runs a control persona alongside, from the same IP range, on the same day; {[liu2024_opted]} ran 17 personas per jurisdiction from Frankfurt and Northern California and visited each site nine times. Fewer repetitions than that and day-of-week and slot effects swamp the treatment {[iqbal2023_tracking]}. |
| * **Zero bids are data.** 22% of all bids in {[cook2020_headerbidding]} were zero, and the paper cannot say whether that is misconfiguration or intent. Do not drop them silently. | * **Zero bids are data.** 22% of all bids in {[cook2020_headerbidding]} were zero, and the paper cannot say whether that is misconfiguration or intent. Do not drop them silently. |
| * **Consent and vantage change who bids.** A European vantage brings a banner into the path, and 352 of the Alexa top-100K had both a CMP and Prebid — that intersection, not the 5,421 Prebid sites, is the auditable population for a consent study {[liu2024_opted]}. See [[Privacy:Consent]] and [[Design:Crawling location]]. Of the 16 measuring papers, 11 state a vantage location; 6 of those are the United States. | * **Consent and vantage change who bids.** A European vantage brings a banner into the path, and 352 of the Alexa top-100K had both a CMP and Prebid — that intersection, not the 5,421 Prebid sites, is the auditable population for a consent study {[liu2024_opted]}. See [[Privacy:Consent]] and [[Design:Crawling location]]. Of the 16 measuring papers, 11 state a vantage location; 6 of those are the United States. |
| * **Ad blockers and filter lists.** Prebid endpoints are on EasyList; a crawler with a blocking extension, or a Firefox profile with tracking protection on, sees no auction. {[cook2020_headerbidding]} and {[musa2022_atom]} disabled protections in OpenWPM and say so — see [[Privacy:Browser protection]]. | * **Ad blockers and filter lists.** Prebid endpoints are on EasyList, so a crawler with an EasyList-based blocker sees no auction; whether Firefox's tracking protection does the same has not been measured in these venues. {[musa2022_atom]} disabled tracking protection in OpenWPM and says so; {[cook2020_headerbidding]} only notes browsers' default protections as a caveat — see [[Privacy:Browser protection]]. |
| * **Real users beat crawlers here more than anywhere else on this site.** The only field study {[zeng2022_factors]} — a browser extension in 286 Prolific participants' own browsers — is the one bid dataset that is about people. Its cost is ten fixed sites and an eleven-day window before Christmas. | * **Real users are the only way to price a person rather than a persona.** The one field study {[zeng2022_factors]} — a browser extension in 286 Prolific participants' own browsers — is the only bid dataset in these venues that is about people. Its cost is ten fixed sites and an eleven-day window before Christmas. |
| |
| ===== Tools, datasets and services that already exist ===== | ===== Tools, datasets and services that already exist ===== |
| | AdSparency monitoring service | {[papadogiannakis2025_darkside]} | Live, ''adsparency.ics.forth.gr'' | Looking a domain up before you crawl it; the authors call its output "an indication of misuse", not evidence | | | AdSparency monitoring service | {[papadogiannakis2025_darkside]} | Live, ''adsparency.ics.forth.gr'' | Looking a domain up before you crawl it; the authors call its output "an indication of misuse", not evidence | |
| | Dark-pooling measurement code | {[vekaria2024_darkpooling]} | Live, MIT, ''github.com/Yash-Vekaria/ad-inventory-fraud-measurement'', last pushed May 2024 | The misrepresentation checks (Tables 2 and 3 of the paper) as code | | | Dark-pooling measurement code | {[vekaria2024_darkpooling]} | Live, MIT, ''github.com/Yash-Vekaria/ad-inventory-fraud-measurement'', last pushed May 2024 | The misrepresentation checks (Tables 2 and 3 of the paper) as code | |
| | ''sellers.json'' of the largest exchange | Google | Live; about 109 MB; 972,929 sellers, 70.96% confidential. The canonical URL ''realtimebidding.google.com/sellers.json'' intermittently returns an empty body; the copy at ''storage.googleapis.com/adx-rtb-dictionaries/sellers.json'' is byte-identical | Stream it; do not ''json.load'' it on a laptop without checking memory first | | | ''sellers.json'' of the largest exchange | Google | Live; about 109 MB; 972,929 sellers, 70.96% confidential. ''realtimebidding.google.com/sellers.json'' redirects (302) to ''storage.googleapis.com/adx-rtb-dictionaries/sellers.json'' — follow it | Stream it; do not ''json.load'' it on a laptop without checking memory first | |
| | Crawler that stores ''ads.txt'' alongside HTML, cookies and traffic ("Scrape Titan") | {[papadogiannakis2023_whofunds]} | Live, ''gitlab.com/papamano/scrape-titan'' | Collecting the file in the same visit as the page — the only way to join declarations to what the page actually loaded | | | Crawler that stores ''ads.txt'' alongside HTML, cookies and traffic ("Scrape Titan") | {[papadogiannakis2023_whofunds]} | Live, ''gitlab.com/papamano/scrape-titan'' | Collecting the file in the same visit as the page — the only way to join declarations to what the page actually loaded | |
| | HBDetector Chrome extension | {[pachilakis2019_waterfalls]} | **Dead** — ''github.com/mipach/HBDetector'' returns 404 | Read the paper's Section 3 for the event list and reimplement; the industry-maintained equivalent is Prebid's own ''header-bidder-expert'' / Professor Prebid extension | | | HBDetector Chrome extension | {[pachilakis2019_waterfalls]} | **Dead** — ''github.com/mipach/HBDetector'' returns 404 | Read the paper's Section 3 for the event list and reimplement; the maintained industry equivalent is Prebid's Professor Prebid extension (its older ''header-bidder-expert'' has not been touched since 2022) | |
| | HARPO source | {[zhang2022_harpo]} | **Dead** — ''github.com/bitzj2015/Harpo-NDSS22'' returns 404 | — | | | HARPO source | {[zhang2022_harpo]} | **Dead** — ''github.com/bitzj2015/Harpo-NDSS22'' returns 404 | — | |
| | Inclusion-chain crawler ("DeepCrawling") | {[bashir2019_adstxt]} | Moved to ''github.com/sajjadium/Crawlium'' | Declarations-vs-transactions designs | | | Inclusion-chain crawler ("DeepCrawling") | {[bashir2019_adstxt]} | Moved to ''github.com/sajjadium/Crawlium'' | Declarations-vs-transactions designs | |
| | **any of the seven core terms, at least once** | **41 (0.7%)** | | | | **any of the seven core terms, at least once** | **41 (0.7%)** | | |
| | **core terms summed, five or more times** | **18 (0.3%)** | | | | **core terms summed, five or more times** | **18 (0.3%)** | | |
| | //context:// real-time bidding / RTB | 64 | 19 | | | //context:// real-time bidding / RTB | 73 | 21 | |
| | //context:// dark pool(ing) | 8 | 4 | | | //context:// dark pool(ing) | 8 | 4 | |
| |
| This is a small literature in these venues: 41 papers ever name one of the observables and 18 do so substantively. The first mention of any of them is in 2017; there is nothing in 2010–2016 because ''ads.txt'' did not exist until June 2017. | This is a small literature in these venues: 41 papers ever name one of the observables and 18 name one five or more times. The first mention of any of them is in 2017; there is nothing in 2010–2016 because ''ads.txt'' did not exist until June 2017 and nobody in these venues wrote about header bidding or OpenRTB before it did. |
| |
| ==== By venue ==== | ==== By venue ==== |
| | Prebid presence as a sampling frame (''hb-sampling'') | 1 | 2026 | | | Prebid presence as a sampling frame (''hb-sampling'') | 1 | 2026 | |
| |
| Of the 16, **13 have a crawl-configuration record**; the other three are a browser extension in real users' browsers, a passive proxy, and industrial bidding logs. Among the 13: statefulness **6 stateless, 5 stateful, 1 both, 1 not stated**; consent action **5 no-interaction, 2 accept-all, 1 accept-and-reject, 1 CMP-specific, 1 not applicable, 3 not stated**; interaction depth **4 landing page only, 4 landing plus subpages, 3 single target page, 2 not stated**; headless mode **not stated by 10 of 13**. Population sources named, paper-counted after a short alias fold: **Alexa 8, Tranco 5, MediaBias/FactCheck 3, NextDNS piracy list 2, Prolific 1, BuiltWith 1**, with 30 unmapped source strings printed on the provenance page. **10 of 16 release a public artifact**, 4 mention none, 2 promised one. | Of the 16, **13 have a crawl-configuration record**; the other three are two passive-proxy studies of real users' traffic ({[papadopoulos2017_paying]}, {[papadopoulos2018_cost]}) and one browser extension in real users' browsers ({[zeng2022_factors]}). Among the 13: statefulness **6 stateless, 5 stateful, 1 both, 1 not stated**; consent action **5 no-interaction, 2 accept-all, 1 accept-and-reject, 1 CMP-specific, 1 not applicable, 3 not stated**; interaction depth **4 landing page only, 4 landing plus subpages, 3 single target page, 2 not stated**; headless mode **not stated by 10 of 13**. Population sources named, paper-counted after a short alias fold: **Alexa 8, Tranco 5, MediaBias/FactCheck 3, NextDNS piracy list 2, Prolific 1, BuiltWith 1**, with 30 unmapped source strings printed on the provenance page. **10 of 16 release a public artifact**, 4 mention none, 2 promised one. |
| |
| ==== What the corpus cannot tell you ==== | ==== What the corpus cannot tell you ==== |
| |
| The paper this whole literature cites for dark pooling is not in it, and neither is anything from EuroS&P, where its 2026 follow-up on notifying the ecosystem appeared,((Vekaria, Nithyanand, Shafiq, "Towards Multi-Stakeholder Vulnerability Notifications in the Ad-Tech Supply Chain", IEEE EuroS&P 2026 (Lisbon, 6–10 July 2026), listed on ''eurosp2026.ieee-security.org/program.html''; preprint ''arXiv:2406.06958'' (v2, 2 April 2026). Fetched 2026-09-02.)) nor the advertising-economics literature that measures Prebid configurations at panel scale.((Johnson & Neumann, "The advent of privacy-centric digital advertising: Tracing privacy-enhancing technology adoption", working paper dated 21 March 2024, measures Privacy Sandbox adoption from sites' Prebid settings on a panel of roughly sixty thousand sites drawn from a Tranco top-100K list. Working paper, not peer-reviewed; cited here only as an existence proof that the Prebid configuration is used as an observable outside computer science.)) Industry measurement of ''ads.txt'' adoption exists but publishes headline percentages without a stated denominator or method, and none of it is cited on this page. Any count above is a count over **CCS, IMC, NDSS, PoPETs, USENIX Security, TheWebConf and IEEE S&P**, 2010–2026. See [[Literature:Corpus]]. | The paper this whole literature cites for dark pooling is not in it, and neither is anything from EuroS&P, where its 2026 follow-up on notifying the ecosystem appeared,((Vekaria, Nithyanand, Shafiq, "Towards Multi-Stakeholder Vulnerability Notifications in the Ad-Tech Supply Chain", IEEE EuroS&P 2026 (Lisbon, 6–10 July 2026), listed on ''eurosp2026.ieee-security.org/program.html''; DOI ''10.1109/EuroSP68448.2026.00061'' (resolves to IEEE Xplore document 11624241); preprint ''arXiv:2406.06958'' (v2, 2 April 2026). Fetched 2026-09-02.)) nor the advertising-economics literature that measures Prebid configurations at panel scale.((Johnson & Neumann, "The advent of privacy-centric digital advertising: Tracing privacy-enhancing technology adoption", working paper dated 21 March 2024, measures Privacy Sandbox adoption from sites' Prebid settings on a panel of roughly sixty thousand sites drawn from a Tranco top-100K list. Working paper, not peer-reviewed; cited here only as an existence proof that the Prebid configuration is used as an observable outside computer science.)) Industry measurement of ''ads.txt'' adoption exists but publishes headline percentages without a stated denominator or method, and none of it is cited on this page. Any count above is a count over **CCS, IMC, NDSS, PoPETs, USENIX Security, TheWebConf and IEEE S&P**, 2010–2026. See [[Literature:Corpus]]. |
| |
| ===== What to Report ===== | ===== What to Report ===== |
| - **For pooling**: how ownership was established (WHOIS, publisher IDs, hand review) and for what share of pool members it could not be. | - **For pooling**: how ownership was established (WHOIS, publisher IDs, hand review) and for what share of pool members it could not be. |
| - **For header bidding**: how Prebid sites were detected (global name, events, or both), which API (''getBidResponses'' vs winning vs rendered), the wait after ''load'', the persona and control design, the vantage, the consent action, and the browser's protection state. | - **For header bidding**: how Prebid sites were detected (global name, events, or both), which API (''getBidResponses'' vs winning vs rendered), the wait after ''load'', the persona and control design, the vantage, the consent action, and the browser's protection state. |
| - **What fraction of your candidate sites exposed a client-side auction at all** — the number every bids paper since 2019 has omitted. | - **What fraction of your candidate sites exposed a client-side auction at all** — two of the seven bids papers report it; and, if you can, which of those auctions were client-only, hybrid or server-side, which nobody has reported since 2019. |
| |
| ===== Open Questions ===== | ===== Open Questions ===== |
| |
| * **Has the client-side share of header bidding fallen since 2019?** 17.3% client-only, 34.7% hybrid, 48% server-side is the only measurement. Every bids-as-instrument paper depends on the answer. | * **Has the client-side share of header bidding fallen since 2019?** 17.3% client-only, 34.7% hybrid, 48% server-side is the only measurement, and the paper never says whether its unit is sites or auctions. Every bids-as-instrument paper depends on the answer. |
| * **How widely are OWNERDOMAIN and MANAGERDOMAIN declared, and do they match ''sellers.json''?** Two fields, three years old, no measurement. A weekend with the IAB crawler would settle it. | * **How widely are OWNERDOMAIN and MANAGERDOMAIN declared, and do they match ''sellers.json''?** Two fields, four years old, no measurement. A weekend with the IAB crawler would settle it. |
| * **What does ''app-ads.txt'' look like?** Bashir et al. asked for "a separate study" in 2019; none has appeared in these venues. | * **What does ''app-ads.txt'' look like?** Bashir et al. asked for "a separate study" in 2019; none has appeared in these venues. |
| * **Do confirmed relationships predict served ads?** {[papadogiannakis2025_before]} ran both an intersection crawl and an ad-clicking crawl on the same sites; nobody has published the join. | * **Do confirmed relationships predict served ads?** {[papadogiannakis2025_before]} ran both an intersection crawl and an ad-clicking crawl on the same sites; no paper in these venues has published the join. |
| * **Is dark pooling declining after notification?** The 2026 EuroS&P notification study says notifications to ad networks work; a re-measurement of the 2023 pool census would show whether the ecosystem moved. | * **Is dark pooling declining after notification?** The 2026 EuroS&P notification study says notifications to ad networks work; a re-measurement of the 2023 pool census would show whether the ecosystem moved. |
| |