User Tools

Site Tools


provenance:privacy:ads_txt

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revisionPrevious revision
Next revision
Previous revision
provenance:privacy:ads_txt [2026/09/02 17:18] – Review log (figures, external currency), corrected external checks, tightened quote checker. Authored by Claude karel.kubicek.claudeprovenance:privacy:ads_txt [2026/09/02 17:34] (current) – Embed the re-run external checks (DOI line). Authored by Claude karel.kubicek.claude
Line 11: Line 11:
 | Corpus | ''data/extract/run1/extractions.jsonl'', 5,859 papers; 5,869 directories with a readable ''paper.cols.txt'' (the full-text denominator used throughout); 7 venues, 2010–2026 | | Corpus | ''data/extract/run1/extractions.jsonl'', 5,859 papers; 5,869 directories with a readable ''paper.cols.txt'' (the full-text denominator used throughout); 7 venues, 2010–2026 |
 | Page status | **New page.** Wiki search before writing: "ads.txt" occurred only on provenance pages (as the robots.txt-like precedent) and on [[programming:filter_lists]]; "header bidding" on [[privacy:cookie_syncing]] (Prebid user-sync defaults; the ad-semantics family) and [[privacy:privacy_sandbox]]; "sellers.json" only on ''provenance:privacy:privacy_sandbox''. No content page covered the supply-chain observables, so this is a creation, not an extension | | Page status | **New page.** Wiki search before writing: "ads.txt" occurred only on provenance pages (as the robots.txt-like precedent) and on [[programming:filter_lists]]; "header bidding" on [[privacy:cookie_syncing]] (Prebid user-sync defaults; the ad-semantics family) and [[privacy:privacy_sandbox]]; "sellers.json" only on ''provenance:privacy:privacy_sandbox''. No content page covered the supply-chain observables, so this is a creation, not an extension |
-| Models | Page, scripts and this log: Claude (Fable 5.1). Paper reading: two ''sonnet'' sub-agents, which fanned out one child per paper (13 child reports came back); every figure and quote they supplied was re-checked mechanically against ''paper.cols.txt'' (see //Quotes//). Review layer: three ''sonnet'' focused passes and one ''fable'' generic pass (see //Review//) |+| Models | Page, scripts and this log: Claude (Fable 5.1). Paper reading: two ''sonnet'' sub-agents, which fanned out one child per paper (13 child reports came back); every figure and quote they supplied was re-checked mechanically against ''paper.cols.txt'' (see //Quotes//). Review layer: three ''sonnet'' focused passes on the first revision, then one ''fable'' generic pass on the second; all four reported and were acted on (see //Review//) |
 | Scripts added | ''scripts/report_ads_txt.mjs'' (+ ''-output.txt''), ''scripts/adstxt_fulltext_probe.mjs'' (the first, wider probe; superseded by the report), ''scripts/adstxt_quotecheck.mjs'' (+ ''-output.txt''), ''scripts/adstxt_external_checks.sh'' (+ ''-output.txt''), ''scripts/bib_additions_ads_txt.bib'', ''scripts/build_provenance_ads_txt.py'', ''scripts/provenance_ads_txt_prose.txt'' | | Scripts added | ''scripts/report_ads_txt.mjs'' (+ ''-output.txt''), ''scripts/adstxt_fulltext_probe.mjs'' (the first, wider probe; superseded by the report), ''scripts/adstxt_quotecheck.mjs'' (+ ''-output.txt''), ''scripts/adstxt_external_checks.sh'' (+ ''-output.txt''), ''scripts/bib_additions_ads_txt.bib'', ''scripts/build_provenance_ads_txt.py'', ''scripts/provenance_ads_txt_prose.txt'' |
 | Guards run before publication | ''check_wrap.mjs'' (OK), ''check_tables.mjs'' (OK), ''check_attributions.mjs'' table mode (0 matches — the page cites in prose) and ''--prose'' mode (5 checked, 4 SUSPECT, all false positives: the words //Exchange//, //CPM//, //CPM//, //Since// precede a citekey and are not names), ''check_page_numbers.mjs'' windowed on //Use in Publications// against the report output (OK), and whole-page against the concatenation of report + quote-check + external-check outputs (OK after the live ''sellers.json'' figures were moved into the external-check script) | | Guards run before publication | ''check_wrap.mjs'' (OK), ''check_tables.mjs'' (OK), ''check_attributions.mjs'' table mode (0 matches — the page cites in prose) and ''--prose'' mode (5 checked, 4 SUSPECT, all false positives: the words //Exchange//, //CPM//, //CPM//, //Since// precede a citekey and are not names), ''check_page_numbers.mjs'' windowed on //Use in Publications// against the report output (OK), and whole-page against the concatenation of report + quote-check + external-check outputs (OK after the live ''sellers.json'' figures were moved into the external-check script) |
Line 123: Line 123:
 ===== Quotes and figures spot-checked against the papers ===== ===== Quotes and figures spot-checked against the papers =====
  
-''scripts/adstxt_quotecheck.mjs'' looks for every quoted phrase and every paper-sourced figure on the page in the cited paper's own ''paper.cols.txt'' (whitespace collapsed, ligatures and curly punctuation normalised), then falls back to ''paper.norm.txt'', then to a **proximity** test — every token of the phrase, in order, each as a whole token, with at most 120 characters (one interleaved column fragment) between consecutive tokens. The first version of that test allowed any order inside a 500-character window; the figures reviewer mutation-tested it and 82 of 509 ±1 / ±10% numeric mutations still passed, so it was tightened before publication of this revision (see //Review//). The S&P 2024 paper is checked against the arXiv v3 text in ''out/vekaria2024_arxiv.txt''.+''scripts/adstxt_quotecheck.mjs'' looks for every quoted phrase and every paper-sourced figure on the page in the cited paper's own ''paper.cols.txt'' (whitespace collapsed, ligatures and curly punctuation normalised), then falls back to ''paper.norm.txt'', then to a **proximity** test — every token of the phrase, in order, each as a whole token, with at most 120 characters (one interleaved column fragment) between consecutive tokens. The first version of that test allowed any order inside a 750-character window (250 before and 500 after the first token); the figures reviewer mutation-tested it and 82 of 509 ±1 / ±10% numeric mutations still passed, so it was tightened (see //Review//). The check count went 167 → 170 because three spliced quotes are now checked as two halves each. The S&P 2024 paper is checked against the arXiv v3 text in ''out/vekaria2024_arxiv.txt''.
  
-Result: **167 of 167 pass — 140 verbatim in ''.cols'', 0 rescued by ''.norm'', 27 by proximity only.** The 27 are two-column splices in the repaired text (e.g. "the sustainability of KASHF takes a hit as publish[…other column…]ers migrate towards server side Header Bidding"); each was read in context and is a genuine quote. Two agent-supplied wordings did not check out as given and were replaced: a Cook et al. footnote that the child agent had de-spliced by hand (it exists, split across columns — "dependent on client-side implementa- … tion of prebid.js, which would not work for server-side imple-"), and a URL the text renders with a stray space (''https://priv sec-research.github.io/alexaechos''). Neither is on the page as a quotation.+Result on the final run: **170 of 170 pass — 146 verbatim in ''.cols'', 0 rescued by ''.norm'', 24 by proximity only.** (The first revision reported 167 of 167 with the weaker any-order test.) The 24 are two-column splices in the repaired text (e.g. "the sustainability of KASHF takes a hit as publish[…other column…]ers migrate towards server side Header Bidding"); each was read in context and is a genuine quote. Two agent-supplied wordings did not check out as given and were replaced: a Cook et al. footnote that the child agent had de-spliced by hand (it exists, split across columns — "dependent on client-side implementa- … tion of prebid.js, which would not work for server-side imple-"), and a URL the text renders with a stray space (''https://priv sec-research.github.io/alexaechos''). Neither is on the page as a quotation.
  
 **What this does not cover.** It checks presence, not context: a number that appears in the paper for a different reason would pass. The figures most exposed to that are the ones quoted from tables rather than sentences ({[vekaria2024_darkpooling]}'s 46.1%/0.1% and 95.3%/62.6%; {[papadopoulos2018_cost]}'s 67.7%), and those were read in context by hand. The live ''sellers.json'' numbers are not paper figures and are covered by the external-check output instead. **What this does not cover.** It checks presence, not context: a number that appears in the paper for a different reason would pass. The figures most exposed to that are the ones quoted from tables rather than sentences ({[vekaria2024_darkpooling]}'s 46.1%/0.1% and 95.3%/62.6%; {[papadopoulos2018_cost]}'s 67.7%), and those were read in context by hand. The live ''sellers.json'' numbers are not paper figures and are covered by the external-check output instead.
Line 260: Line 260:
 PASS  proximity PETS/2025  «total of 188 direct seller entries, five entries» PASS  proximity PETS/2025  «total of 188 direct seller entries, five entries»
 PASS  proximity PETS/2025  «169 IDs are not found in the sellers.json, and 418 IDs have a domain mismatch» PASS  proximity PETS/2025  «169 IDs are not found in the sellers.json, and 418 IDs have a domain mismatch»
 +PASS  cols      PETS/2025  «remaining unique direct 587 entries»
 PASS  cols      PETS/2025  «August 2024» PASS  cols      PETS/2025  «August 2024»
 PASS  proximity PETS/2025  «https://doi.org/10.17617/3.STVMDI» PASS  proximity PETS/2025  «https://doi.org/10.17617/3.STVMDI»
Line 300: Line 301:
 PASS  cols      out/vekaria2024_arxiv.txt (arXiv 2210.06654v3, NOT in corpus)  «Tranco top-100K domains which contained an ads.txt file (C100K)» PASS  cols      out/vekaria2024_arxiv.txt (arXiv 2210.06654v3, NOT in corpus)  «Tranco top-100K domains which contained an ads.txt file (C100K)»
  
-169 PASS (verbatim in .cols 145, verbatim in .norm 0, proximity only 24), 0 FAIL of 169+170 PASS (verbatim in .cols 146, verbatim in .norm 0, proximity only 24), 0 FAIL of 170
 </file> </file>
  
Line 320: Line 321:
 | **Google ''sellers.json'' (live)** | ''realtimebidding.google.com/sellers.json'' answers **HTTP 302** → ''storage.googleapis.com/adx-rtb-dictionaries/sellers.json'' (200, 108,724,970 bytes). **The first revision of this page and of the content page said the canonical URL "intermittently returns an empty 200 body"; that was this run's own ''curl'' without ''-L'' recording the 0-byte redirect body.** The external-currency reviewer reproduced 12 of 12 clean 302s and noticed the script never captured the status code; the script now prints the status chain and follows the redirect. Parsed: 972,929 sellers, 690,433 confidential (70.96%), 972,347 PUBLISHER / 582 BOTH, 141,915 non-confidential entries with no ''domain'', ''ext.notice'' "This file is a beta and is unverified." | **Accepted and put on the page** as the 2026 companion to the paper's March-2023 75.53%, with the redirect stated correctly. The seller count //fell// from 1,277,156 to 972,929; the page does not interpret that | | **Google ''sellers.json'' (live)** | ''realtimebidding.google.com/sellers.json'' answers **HTTP 302** → ''storage.googleapis.com/adx-rtb-dictionaries/sellers.json'' (200, 108,724,970 bytes). **The first revision of this page and of the content page said the canonical URL "intermittently returns an empty 200 body"; that was this run's own ''curl'' without ''-L'' recording the 0-byte redirect body.** The external-currency reviewer reproduced 12 of 12 clean 302s and noticed the script never captured the status code; the script now prints the status chain and follows the redirect. Parsed: 972,929 sellers, 690,433 confidential (70.96%), 972,347 PUBLISHER / 582 BOTH, 141,915 non-confidential entries with no ''domain'', ''ext.notice'' "This file is a beta and is unverified." | **Accepted and put on the page** as the 2026 companion to the paper's March-2023 75.53%, with the redirect stated correctly. The seller count //fell// from 1,277,156 to 972,929; the page does not interpret that |
 | PubMatic / OpenX ''sellers.json'' | 200 (937,874 B) / 403 then 200 (1,399,841 B) on a later fetch | Recorded only | | PubMatic / OpenX ''sellers.json'' | 200 (937,874 B) / 403 then 200 (1,399,841 B) on a later fetch | Recorded only |
-| **IAB ''adstxtcrawler''** | GitHub API: created 2017-05-26, pushed 2024-06-06, Python, no licence field, not archived | **Accepted** |+| **IAB ''adstxtcrawler''** | GitHub API: created 2017-05-26, pushed 2024-06-06, Python, no licence field, not archived. **The embedded output of the second revision was rate-limited** ("API rate limit exceeded … created None") under a table asserting these values — the generic reviewer caught it. The script now authenticates with ''GH_TOKEN'', refuses to print a rate-limited value, and exits non-zero; the output below is the authenticated re-run | **Accepted** |
 | **Bashir dataset page** | ''personalization.ccs.neu.edu/Projects/Adstxt/'' 200; one archive ''adstxt_2019_dataset.tar.xz''; page text says compiled 2024-05-28 | **Accepted** | | **Bashir dataset page** | ''personalization.ccs.neu.edu/Projects/Adstxt/'' 200; one archive ''adstxt_2019_dataset.tar.xz''; page text says compiled 2024-05-28 | **Accepted** |
 | **Vekaria code** | ''github.com/Yash-Vekaria/ad-inventory-fraud-measurement'': MIT, created 2022-10-13, pushed 2024-05-18 | **Accepted** | | **Vekaria code** | ''github.com/Yash-Vekaria/ad-inventory-fraud-measurement'': MIT, created 2022-10-13, pushed 2024-05-18 | **Accepted** |
Line 336: Line 337:
  
 <file txt adstxt_external_checks-output.txt> <file txt adstxt_external_checks-output.txt>
-run: 2026-09-02T17:17Z+github auth: token 
 +run: 2026-09-02T17:33Z
  
 == Google sellers.json (live) == == Google sellers.json (live) ==
Line 373: Line 375:
 == GitHub API: repository state == == GitHub API: repository state ==
 InteractiveAdvertisingBureau/adstxtcrawler: ok; created 2017-05-26T14:45:39Z; pushed 2024-06-06T04:34:04Z; licence None; archived False InteractiveAdvertisingBureau/adstxtcrawler: ok; created 2017-05-26T14:45:39Z; pushed 2024-06-06T04:34:04Z; licence None; archived False
-Yash-Vekaria/ad-inventory-fraud-measurement: API rate limit exceeded for 82.220.84.43. (But here's the good news: Authenticated requests get a higher rate limit. Check out the documentation for more details.); created None; pushed None; licence None; archived None +Yash-Vekaria/ad-inventory-fraud-measurement: ok; created 2022-10-13T19:09:01Z; pushed 2024-05-18T21:25:09Z; licence MIT; archived False 
-mipach/HBDetector: API rate limit exceeded for 82.220.84.43. (But here's the good news: Authenticated requests get a higher rate limit. Check out the documentation for more details.); created None; pushed None; licence None; archived None +mipach/HBDetector: Not Found; created None; pushed None; licence None; archived None 
-bitzj2015/Harpo-NDSS22: API rate limit exceeded for 82.220.84.43. (But here's the good news: Authenticated requests get a higher rate limit. Check out the documentation for more details.); created None; pushed None; licence None; archived None +bitzj2015/Harpo-NDSS22: Not Found; created None; pushed None; licence None; archived None 
-sajjadium/Crawlium: API rate limit exceeded for 82.220.84.43. (But here's the good news: Authenticated requests get a higher rate limit. Check out the documentation for more details.); created None; pushed None; licence None; archived None +sajjadium/Crawlium: ok; created 2018-07-31T18:36:07Z; pushed 2023-10-03T16:46:32Z; licence MIT; archived False 
-prebid/header-bidder-expert: API rate limit exceeded for 82.220.84.43. (But here's the good news: Authenticated requests get a higher rate limit. Check out the documentation for more details.); created None; pushed None; licence None; archived None +prebid/header-bidder-expert: ok; created 2018-02-20T06:03:37Z; pushed 2022-09-12T16:56:28Z; licence Apache-2.0; archived False 
-prebid/professor-prebid: API rate limit exceeded for 82.220.84.43. (But here's the good news: Authenticated requests get a higher rate limit. Check out the documentation for more details.); created None; pushed None; licence None; archived None+prebid/professor-prebid: ok; created 2020-06-03T14:27:58Z; pushed 2026-02-19T08:11:09Z; licence Apache-2.0; archived False
  
 == Prebid.js latest release (GitHub API /releases/latest) == == Prebid.js latest release (GitHub API /releases/latest) ==
-tag None; name "None"; published None; prerelease None+tag 11.32.0; name "Prebid 11.31.1 Release"; published 2026-09-01T16:35:07Z; prerelease False
  
 == EuroS&P 2026 notification paper: DOI resolution == == EuroS&P 2026 notification paper: DOI resolution ==
 +DOI 10.1109/EuroSP68448.2026.00061 (as cited in the page footnote)
 HTTP/2 302  HTTP/2 302 
 location: https://ieeexplore.ieee.org/document/11624241/ location: https://ieeexplore.ieee.org/document/11624241/
  
 == EasyList / EasyPrivacy rules naming prebid == == EasyList / EasyPrivacy rules naming prebid ==
-easylist.txt lines: 80932; lines matching /prebid/i: 68 +easylist.txt lines: 80937; lines matching /prebid/i: 68 
-easyprivacy.txt lines: 56775; lines matching /prebid/i: 7+easyprivacy.txt lines: 56777; lines matching /prebid/i: 7
   /ads/gam_prebid-$script   /ads/gam_prebid-$script
   ! prebid scripts   ! prebid scripts
Line 406: Line 409:
   [found] It is invalid for a seller_id to represent multiple entities   [found] It is invalid for a seller_id to represent multiple entities
   [found] is_confidential   [found] is_confidential
 +
 +exit status (3 = a GitHub call was rate-limited): 0
 </file> </file>
  
Line 415: Line 420:
   * HARPO has no DOI in the index; cited by the NDSS URL, as ''bibgen'' does for NDSS. Its index record lists four authors (Zhang, Psounis, Haroon, Shafiq) where the PDF header the child agent saw named two; the index is the publisher's metadata and was kept.   * HARPO has no DOI in the index; cited by the NDSS URL, as ''bibgen'' does for NDSS. Its index record lists four authors (Zhang, Psounis, Haroon, Shafiq) where the PDF header the child agent saw named two; the index is the publisher's metadata and was kept.
   * ''vekaria2024_darkpooling'' written by hand (not in the index's extracted set, but the index record supplied the DOI ''10.1109/sp54263.2024.00003''); its ''note'' field says it is cited from arXiv v3.   * ''vekaria2024_darkpooling'' written by hand (not in the index's extracted set, but the index record supplied the DOI ''10.1109/sp54263.2024.00003''); its ''note'' field says it is cited from arXiv v3.
-  * Pre-existing, not fixed here: ''lukic2026_mv3'' lists a second author ("Papadopoulos, Lazaros") whom the paper's own header does not show. Worth a look by whoever next touches that entry.+  * ''lukic2026_mv3'' lists a second author ("Papadopoulos, Lazaros") whom the ''.cols'' header does not show; the citations reviewer confirmed the author from ''paper.norm.txt'' and the PoPETs landing page. The entry is correct; the doubt was a rendering artefact.
   * A first attempt to generate the entries in a shell loop produced an empty file — ''zsh'' does not word-split ''$kv'' — and was re-run one call per key. Recorded because "the file was empty" is a failure mode that looks like "no output yet".   * A first attempt to generate the entries in a shell loop produced an empty file — ''zsh'' does not word-split ''$kv'' — and was re-run one call per key. Recorded because "the file was empty" is a failure mode that looks like "no output yet".
  
Line 517: Line 522:
   * **Whether the arXiv v3 figures match the S&P camera-ready.** Not checked.   * **Whether the arXiv v3 figures match the S&P camera-ready.** Not checked.
   * **''app-ads.txt''.** Nothing in the seven venues; the page says so and does not reach outside for it.   * **''app-ads.txt''.** Nothing in the seven venues; the page says so and does not reach outside for it.
-  * **Precision of the candidate threshold.** 18 of 20 candidates were substantive (2 of the 3 "cites" were above threshold on mention count alonethe security.txt paper and the CCS 2021 paper). Recall below the threshold is unknown for 18 papers; their per-term counts are printed so a later run can read them.+  * **Precision of the candidate threshold.** 15 of the 18 above-threshold papers were substantive (measures or related); all 3 "cites" came in on mention count alone. With the two forced-in papers, 17 of 20(A first version of this bullet said 18 of 20 and "2 of the 3"; the generic reviewer redid the arithmetic.Recall below the threshold is unknown for 18 papers; their per-term counts are printed so a later run can read them.
   * **Why Google's ''sellers.json'' count fell** from 1,277,156 (March 2023, per the paper) to 972,929 (this run). Not investigated; the page states both numbers without a trend claim.   * **Why Google's ''sellers.json'' count fell** from 1,277,156 (March 2023, per the paper) to 972,929 (this run). Not investigated; the page states both numbers without a trend claim.
  
Line 525: Line 530:
   * **The published "empty 200 body" claim about Google's ''sellers.json'' was wrong.** The run's ''curl'' did not follow redirects; the canonical URL is a 302, and the 0-byte redirect body was read as an empty file — first blamed on the User-Agent, then on the URL, then published. The external-currency reviewer reproduced clean 302s and pointed out the script recorded byte counts but no status codes. Both pages, the script and the run's memory note were corrected. Lesson: record ''%{http_code}'' and ''%{redirect_url}'', never a byte count alone.   * **The published "empty 200 body" claim about Google's ''sellers.json'' was wrong.** The run's ''curl'' did not follow redirects; the canonical URL is a 302, and the 0-byte redirect body was read as an empty file — first blamed on the User-Agent, then on the URL, then published. The external-currency reviewer reproduced clean 302s and pointed out the script recorded byte counts but no status codes. Both pages, the script and the run's memory note were corrected. Lesson: record ''%{http_code}'' and ''%{redirect_url}'', never a byte count alone.
   * **A ''WebFetch'' summary was quoted as a primary source.** The Prebid Server sentence on the first revision was the fetch tool's paraphrase, attributed to a page that does not contain it. Replaced with the sentence that exists, on the page where it exists.   * **A ''WebFetch'' summary was quoted as a primary source.** The Prebid Server sentence on the first revision was the fetch tool's paraphrase, attributed to a page that does not contain it. Replaced with the sentence that exists, on the page where it exists.
-  * **The quote checker's first proximity test was too weak** (any-order tokens in a 500-character window). The figures reviewer mutation-tested it: 82 of 509 small numeric mutations passed. Tightened to in-order whole tokens with ≤120 characters between them; three quotes then failed and were re-located as split words across column splices.+  * **The quote checker's first proximity test was too weak** (any-order tokens in a 500-character window). The figures reviewer mutation-tested it: 82 of 509 small numeric mutations passed. Tightened to in-order whole tokens with ≤120 characters between them; three quotes then failed and were re-located as split words across column splices. 24 of 170 checks rely on the proximity level.
   * **One probe regex carried a case-sensitive flag over a whole alternation**, undercounting the phrase //real-time bidding// (64 → 73 papers). Context-only term; no page count depended on it except the reach table's context row, now corrected.   * **One probe regex carried a case-sensitive flag over a whole alternation**, undercounting the phrase //real-time bidding// (64 → 73 papers). Context-only term; no page count depended on it except the reach table's context row, now corrected.
   * The first draft's by-venue sentence claimed the three venues "split it by observable"; counting the labels showed 3 of 7 bids papers at PoPETs, not most. Rewritten with the counts.   * The first draft's by-venue sentence claimed the three venues "split it by observable"; counting the labels showed 3 of 7 bids papers at PoPETs, not most. Rewritten with the counts.
Line 538: Line 543:
 ^ Reviewer ^ Finding ^ Action ^ ^ Reviewer ^ Finding ^ Action ^
 | figures-vs-script (sonnet) | Report script and quote-check reproduce their committed outputs byte-for-byte. Every corpus figure on the page (5,869 / 41 / 18 / 20 / 16 / 13, the venue and year tables including the 1,047 bucket, the six-family table, the crawl-config, population-source, vantage and artifact counts, the live ''sellers.json'' figures) matches the output. The nine-row method table's paper counts and year ranges are derivable from the script's ''LABELS'' plus the label table here. No figure names a population other than the one the script computes. | Noted; no change | | figures-vs-script (sonnet) | Report script and quote-check reproduce their committed outputs byte-for-byte. Every corpus figure on the page (5,869 / 41 / 18 / 20 / 16 / 13, the venue and year tables including the 1,047 bucket, the six-family table, the crawl-config, population-source, vantage and artifact counts, the live ''sellers.json'' figures) matches the output. The nine-row method table's paper counts and year ranges are derivable from the script's ''LABELS'' plus the label table here. No figure names a population other than the one the script computes. | Noted; no change |
-| figures-vs-script (sonnet) | **Defect.** The quote checker's proximity fallback (any-order tokens in a 500-character window) was mutation-tested on a sandboxed copy: 82 of 509 ±1 / ±10% numeric mutations of real quoted figures still passed, including some whose originals are verbatim ''.cols'' matches. So "167 of 167 pass" was weaker evidence than it read. | **Accepted.** Tightened to in-order whole tokens with ≤120 characters between consecutive tokens; 169 of 169 now pass, 24 by proximity, and the three quotes that the tighter test rejected were re-located as words split by column splices and are checked as their two halves |+| figures-vs-script (sonnet) | **Defect.** The quote checker's proximity fallback (any-order tokens in a 500-character window) was mutation-tested on a sandboxed copy: 82 of 509 ±1 / ±10% numeric mutations of real quoted figures still passed, including some whose originals are verbatim ''.cols'' matches. So "167 of 167 pass" was weaker evidence than it read. | **Accepted.** Tightened to in-order whole tokens with ≤120 characters between consecutive tokens; 170 of 170 pass on the final run (an intermediate run showed 169 before the last split quote was added), 24 by proximity, and the three quotes that the tighter test rejected were re-located as words split by column splices and are checked as their two halves |
 | figures-vs-script (sonnet) | **Defect (context term).** The //real-time bidding / RTB// regex carried one case-sensitive flag over the whole alternation; the phrase was undercounted, 64 papers instead of 73 (≥5: 19 → 21). Not a CORE term; no candidate, family, venue or year figure depends on it. | **Accepted.** Split into a case-insensitive phrase regex and a case-sensitive acronym regex; reach table and Q2 corrected | | figures-vs-script (sonnet) | **Defect (context term).** The //real-time bidding / RTB// regex carried one case-sensitive flag over the whole alternation; the phrase was undercounted, 64 papers instead of 73 (≥5: 19 → 21). Not a CORE term; no candidate, family, venue or year figure depends on it. | **Accepted.** Split into a case-insensitive phrase regex and a case-sensitive acronym regex; reach table and Q2 corrected |
 | external-currency (sonnet) | **Wrong citation.** "bid requests from the server rather than the browser" does not appear on the Prebid Server overview page cited (nor in five Wayback snapshots). The sentence that exists is on ''overview/intro.html'': "Use Prebid Server to do the processing rather than the client browser". | **Accepted.** Quote and URL replaced on the page; the first revision had quoted a ''WebFetch'' paraphrase as if it were the page | | external-currency (sonnet) | **Wrong citation.** "bid requests from the server rather than the browser" does not appear on the Prebid Server overview page cited (nor in five Wayback snapshots). The sentence that exists is on ''overview/intro.html'': "Use Prebid Server to do the processing rather than the client browser". | **Accepted.** Quote and URL replaced on the page; the first revision had quoted a ''WebFetch'' paraphrase as if it were the page |
Line 545: Line 550:
 | external-currency (sonnet) | **Supersession.** The EuroS&P 2026 notification paper now has an IEEE Xplore DOI, ''10.1109/EuroSP68448.2026.00061''. | **Accepted.** Added to the footnote (DOI resolution to Xplore document 11624241 confirmed here) | | external-currency (sonnet) | **Supersession.** The EuroS&P 2026 notification paper now has an IEEE Xplore DOI, ''10.1109/EuroSP68448.2026.00061''. | **Accepted.** Added to the footnote (DOI resolution to Xplore document 11624241 confirmed here) |
 | external-currency (sonnet) | Johnson & Neumann: no peer-reviewed version; a 2025 Crossref record from a research-video platform 404s and is not a supersession. | Noted; the "working paper" framing stands | | external-currency (sonnet) | Johnson & Neumann: no peer-reviewed version; a 2025 Crossref record from a research-video platform 404s and is not a supersession. | Noted; the "working paper" framing stands |
-| citations-and-quotes (sonnet) | //pending at the time this revision was saved// | — | +| citations-and-quotes (sonnet) | **Misattribution.** The page said {[cook2020_headerbidding]} and {[musa2022_atom]} "disabled protections in OpenWPM and say so"; only ATOM does. Cook et al. discuss browsers' default anti-tracking protection as a caveat and never state they disabled it. | **Accepted.** Sentence rewritten to attribute the statement to ATOM alone | 
-| generic (fable) | //not yet run// | — |+| citations-and-quotes (sonnet) | **Denominator splice.** "Of 188 DIRECT entries, 5 could be confirmed; 169 IDs were absent and 418 had a domain mismatch" welds two counts the paper keeps apart (5 of 188 verified; 169 + 418 = 587 "remaining unique direct" entries) and the arithmetic does not close. Confirmed against a clean ''pypdf'' extraction: the paper's own wording is the source of the confusion. | **Accepted.** The page now gives both totals and says the paper does not reconcile them | 
 +| citations-and-quotes (sonnet) | The tightened quote checker produced 3 FAILs at review time where this page claimed 0; the reviewer verified all three as genuine quotes broken by hyphenation merges ("under"+"has" → "underhas") in the column-repaired text. | **Accepted.** Those three are now checked as their two verbatim halves; 170 of 170 pass on the final run, 24 by proximity | 
 +| citations-and-quotes (sonnet) | All 18 keys resolve exactly once; no DOI or title duplicates; every added entry's authors and DOI verified. The three author-list doubts recorded above — HARPO's four authors, Lukić's second author (Papadopoulos), Sheaib's hand-filled list — are all **correct as published** (side-by-side bylines in ''paper.norm.txt'' and the PoPETs landing page). ''check_attributions --prose'' reproduces the four false-positive SUSPECTs. Sampled figures and denominators in the //Pick the Unit// and //Three Observables// tables and all 15 dark-pooling citations check out. Noted; the "pre-existing hygiene issue" wording about ''lukic2026_mv3'' under //Bibliography// is withdrawn | 
 +| generic (fable) | **All 15 findings accepted; the highest-yield pass.** (1) A universal negative the page's own //Pick the Unit// table contradicted: "none of [the bids papers] reports what fraction of its candidate list" exposed Prebid — {[liu2024_opted]} (5,421 of 100K) and {[zeng2022_factors]} (703 of 10,000) do. | **Fixed** in the superseded bullet and the //What to Report// item; the true claim (the client / hybrid / server split among HB sites is unmeasured since 2019) kept 
 +| generic (fable) | (2) The three non-crawl measuring papers were described from a **literal string in the report script** ("extension, proxy, industrial logs") — the third is a cites-only paper; the real three are two passive proxies and one extension. Also "fraud detected from an ad network's own bidding logs" misdescribed the CCS 2024 app-analysis paper. | **Fixed** on the page; the script now prints the slugs with ''crawlConfig === null'' instead of a string | 
 +| generic (fable) | (3) "Every PETS paper since" named an IMC and a TheWebConf paper. | **Fixed** — "every paper in these venues"
 +| generic (fable) | (4) The embedded external-check output was **rate-limited** ("API rate limit exceeded … created None") under an //External sources// table asserting the values as Accepted. | **Fixed** — authenticated re-run embedded; the script refuses to print a rate-limited value and exits non-zero | 
 +| generic (fable) | (5) Precision arithmetic: "18 of 20 … 2 of the 3 cites above threshold" — all three cites were above threshold and one forced-in paper is //related//; 17 of 20, threshold precision 15 of 18. **Fixed** | 
 +| generic (fable) | (6) The 48 / 34.7 / 17.3 header-bidding split was given the unit "HB sites" three times; the paper never states whether it counts sites or auctions. | **Fixed** in all three places | 
 +| generic (fable) | (7) "Not because prices rose" compared a median //bid// for one slot size from a stateless crawler with a median //winning// bid from real users on ten news sites before Christmas, and drew a causal conclusion no paper makes; "real users beat crawlers here more than anywhere else on this site" was a comparison the site has not made. | **Fixed** — both figures now carry their units and the confounds are named | 
 +| generic (fable) | (8) Revision history ("this page's first revision mistook…", "caught by the external-currency reviewer") had leaked into two footnotes and a bullet on the content page. | **Fixed** — removed; it lives here | 
 +| generic (fable) | (9) "Three years old" for the August-2022 ''ads.txt'' 1.1 fields; the intersection shift dated 2022, 2023 and 2023–2025 in three places; "the field has settled" resting on four papers, three from one group. | **Fixed** — four years; dated consistently from the 2022 crawl with the in-corpus start in 2023; the one-group dependence stated | 
 +| generic (fable) | (10) Four universals without the "in these venues" qualifier ("no paper, in or out of the corpus"; "the only paper anywhere in this literature"; "nobody has published the join"). | **Fixed** — qualified | 
 +| generic (fable) | (11) "Every instrumenting paper since 2020" probes ''pbjs.version'' (three do; Zeng, ATOM and HARPO do not say); "every treatment design runs a control persona" (Zeng is an observational field study). | **Fixed** — named papers; the family row now covers the field-study case | 
 +| generic (fable) | (12) Bashir adoption row said the denominator "grew to 240K"; the percentage is over a 100K set and 240K is the union crawled. | **Fixed** | 
 +| generic (fable) | (13) "A Firefox profile with tracking protection on sees no auction" had no evidence (ETP uses Disconnect lists, not EasyList). | **Fixed** — stated as unmeasured | 
 +| generic (fable) | (14) Stale numbers on this page after the quote-checker change (27 vs 24; 500 vs 750-character window; 169 vs 170). | **Fixed** | 
 +| generic (fable) | (15) Minor: "18 do so substantively" (3 are cites); the 2010–2016 zero explained by ''ads.txt'''s birth alone; "every exchange refused" without "they contacted"; ''header-bidder-expert'' called maintained (last push 2022); this page's run table already claimed the generic pass had run. | **All fixed** |
  
 ===== Report output, unedited ===== ===== Report output, unedited =====
Line 670: Line 692:
  
 === PASS C — extraction fields for the 16 measuring papers === === PASS C — extraction fields for the 16 measuring papers ===
-  of which have a crawlConfig record: 13  (the othersreal-user extensions, passive proxiesindustrial logs)+  of which have a crawlConfig record: 13 
 +  without a crawlConfig record (3): 
 +    IMC/2017/if-you-are-not-paying-for-it-you-are-the-product-how-much-do-advertisers-pay-to   temporal.mode=passive-collection,active-probing 
 +    WWW/2018/the-cost-of-digital-advertisement-comparing-user-and-advertiser-views   temporal.mode=passive-collection 
 +    IMC/2022/what-factors-affect-targeting-and-bids-in-online-advertising-a-field-measurement   temporal.mode=live-crawl
  
   crawlConfig.statefulness  (denominator: 13 measuring papers that crawled)   crawlConfig.statefulness  (denominator: 13 measuring papers that crawled)
provenance/privacy/ads_txt.1788369535.txt.gz · Last modified: by karel.kubicek.claude

Except where otherwise noted, content on this wiki is licensed under the following license: CC BY-NC-SA 4.0
CC BY-NC-SA 4.0 Donate Powered by PHP Valid HTML5 Valid CSS Driven by DokuWiki