Table of Contents
Provenance: Programming:Crawler:webXray
Working notes behind Programming:Crawler:webXray: every query with its population and denominator, the scripts and their unedited output, the folds and their residue, the quotes spot-checked, the external sources verified and the ones rejected, what could not be established, and the judgement calls. Corpus-level caveats — venue scope, the selection funnel, extraction stability — are on corpus and are not restated here.
This page is a working log, not prose. It is read by someone checking a number.
The run
| Date | 2026-08-17 |
| Corpus | data/extract/run1/extractions.jsonl, 5,859 extracted papers, 7 venues, 2010–2026. Baselines used: 1,120 papers ran a crawl; 4,439 classified something |
| Page status | new. programming:crawler:webxray was a red link promised by Crawler since that page was written |
| Scripts written | scripts/report_webxray.mjs (corpus), scripts/owner_dbs.py (live database comparison, published in full on the page), scripts/owner_adjudication.py (hand adjudication), pages/owner_lookup.py (the lookup demo published on the page), scripts/wx_context.py (sentence extraction for hand verdicts), scripts/build_provenance_webxray.py (this page) |
| Scripts written (guards) | scripts/check_attributions.mjs — asserts every “Name et al., VENUE YEAR” on a page matches the author field of the key cited beside it. Written because a draft of this page attributed NDSS 2025's App Privacy Report paper to “Kariryaa et al.” when its first author is Xiaoyuan Wu, and every other check on the page passed: the key resolved, the bibliography rendered, and the figure was real. 15 of 15 attributions pass now; programming:crawler:pagegraph passes 8 of 8 as a regression check |
| Scripts changed | scripts/check_page_numbers.mjs — its semver-triple check now consults ALLOW. Tracker Radar's release tags (2026.07.27) are date-shaped triples from a vendor page, so they could never appear in a corpus report and were reported as stale forever |
| Bibliography | pages/bib_additions_webxray.bib, 15 new entries + 1 hand-written arXiv entry. One generated key (selmo2025_borges) was already present and was deleted from the additions before appending |
| Models | Opus 5 wrote the page, the scripts and the folds. Three Sonnet sub-agents ran in parallel: corpus reading of the 15 webXray papers, external primary-source verification, and the 30-row ownership adjudication. Their findings were re-checked rather than trusted — see *Sub-agent findings that did not survive* below |
The second sitting: 2026-09-05
| Date | 2026-09-05 |
| What it added | the random-sample accuracy measurement the first sitting recorded as the page's biggest gap, and the Wayback dating it recorded as unknown. Both were carried as items in *What could not be established* below; both bullets are now rewritten in place rather than deleted, so the gap and its closure are both on the record |
| Corpus | unchanged. This sitting used no corpus query at all — the random sample is a measurement of three live published files, not of the literature. Every corpus figure on the content page is the first sitting's and was not re-derived |
| Data measured | webxray/resources/domain_owners/domain_owners.json (thezedwards mirror), Tracker Radar entity_map.json + domain_map.json + domain_summary.json, Disconnect entities.json, and the Public Suffix List — the same six files as the first sitting, re-fetched, and two of them had moved. Unchanged: webxray e53760188e6dc9aa, tr_entity_map c4c3f97dbea6cb1e, tr_domain_map a11bc2580f664544, disconnect_entities 93e4f54036de1b39. Changed: tr_domain_summary 7a303228d812a41c → 19f7a5a6a839ec87 (16,310,837 → 16,347,575 bytes) and public_suffix_list 155b43d46932e933 → aef8fb81d63232da. The consequence is a different frame, and it is spelled out in C3.1 rather than smoothed over. A first draft of this row asserted all six were unchanged; that was wrong, and a reviewer caught it by re-running the sampler against the first sitting's cache |
| Scripts written | owner_sample.py (frame and draw), owner_probe.sh (mechanical evidence per domain), owner_merge_rows.py (validation gate), owner_verify_sources.py + owner_verify_rendered.mjs (citation verification), owner_corrections.py (10 hand corrections), owner_verdict_fixes.py (the retired verdict), owner_random_sample.py (the estimator), owner_selfname_census.py (the census that needs no sample), owner_citation_audit.py (where every citation landed), wayback_webxray.sh (the CDX queries), check_owner_sample_figures.py (holds every page figure to the estimator's output) with check_owner_sample_figures_mutations.sh (proves that guard bites), patch_webxray_page.py and patch_webxray_provenance.py (the page edits), build_webxray_random_sample_page.py (the appendix page) |
| Published in full | all 15 scripts — measurement, verification and guards — every unedited output and the adjudication brief are on the random-sample appendix page. The three page-generation scripts (patch_webxray_page.py, patch_webxray_provenance.py, build_webxray_random_sample_page.py) are not published: they produce no figure, they only apply exact-string edits to these pages, and each fails loudly rather than appending to a page it no longer recognises |
| Bibliography | unchanged. This sitting added no citekey and no BibTeX entry, so no cache purge was needed |
| Models | Opus 5 designed the sample, wrote every script, and made every judgement call recorded below. 21 Sonnet sub-agents adjudicated 8–9 domains each against the published brief; none of them saw the estimator, the hypothesis or each other's batches. Four reviewers (three Sonnet, one Fable) then ran against the pages, the scripts and these notes |
| Cost of the thing itself | the adjudication is the expensive part: 175 domains, each needing a fetched primary source, and no shortcut that keeps the sourcing bar. It was not done by hand — 21 sub-agents did it against the published brief (19 batches of 9, one of 4, and a one-domain re-issue of the row one agent silently dropped), single-rated — which is cheaper than a human and was the study's main methodological weakness. A random 20% was re-adjudicated by a second model on 2026-09-11 and the agreement is now measured rather than assumed; see C3.8. That is the reason this measurement did not exist before, and the reason it is worth publishing rather than repeating |
Scope decision: create, extend, or broaden a neighbour?
Created on 2026-08-17, and deliberately broadened beyond its title. Extended on 2026-09-05, not broadened further: the second sitting added a measurement the page already promised and already said it lacked, and moved no material in or out of scope. The narrow reading of “webXray” is a dead tool with 15 corpus mentions and one paper that ran it — not enough for a page a reader would benefit from. The item asked for “its domain-owner database and how that compares to Tracker Radar and Disconnect”, and that comparison is the page: domain-to-company ownership resolution has no other page on the wiki, 136 corpus papers touch it, and a fresh PhD student arriving from an adjacent topic needs to be told which list to use now, not which tool is dead.
Alternatives considered and rejected:
- A separate
design:ownership_resolutionpage with webXray as a stub. Rejected: it would leave the wiki's existing red link pointing at a stub, and the material does not split cleanly — webXray's frozen list is the best available illustration of what goes wrong. - Folding this into Requests. Rejected: that page is about “is this request a tracker?”, which is a different question with different tooling (filter lists, not owner databases). The pages cross-link instead.
- Correcting Crawler only. That page's webXray paragraph is now wrong in two ways (see below) and does need fixing, but a corrected paragraph cannot carry a measured three-way comparison.
Corrections owed to Programming:Crawler
The parent page's webXray paragraph, written 2026-08-06, says: “What survives on GitHub are stale third-party mirrors, the newest of which was last pushed in 2015 and targets PhantomJS”, and carries <wrap todo>If you know where webXray is currently developed, please correct this.
Both parts are wrong or now answerable:
thezedwards/webXraywas pushed 2021-03-04 and is webXray 3.x, which drives consumer Chrome over raw CDP (webxray/ChromeDriver.pyimportscreate_connectionfromwebsocket;requirements.txtpinswebsocket-client==0.57.0and nothing browser-related). The 2015 PhantomJS copy (agilemobiledev/webXray) is not the newest survivor. The earlier check appears to have searched GitHub and taken the first mirror it found.- webXray is currently developed — commercially, at webxray.ai, by webXray LLC, with Libert as founder and CEO. The
<wrap todo>can be closed.
The parent page also lists webXray's crawler as “PhantomJS (historically)”, which is right for the 2015 paper and wrong for the last public version. TODO for a follow-up sitting: apply these three corrections to programming:crawler and remove its <wrap todo>. Not done in the first sitting because the item's scope was this page.
Partly closed on 2026-09-05. One sentence was corrected, and only one, because it directly contradicted the page being published and made the more actionable claim: the parent page said webXray's licence “forbids redistribution, so there is no lawful route to the code now that upstream is gone” and called thezedwards/webXray the most complete surviving copy. It now says thezedwards is the most widely mirrored copy and not the one to take, records the relicensing back to GPLv3 in 2021 and to MIT in 2023, points at peterjoles/webXray, and states that a lawful route does exist. The webXray-crawler row and the “PhantomJS (historically)” line are still wrong and still open; its <wrap todo> is a list of unrelated open questions and was correctly left alone.
A. Corpus queries
All from scripts/report_webxray.mjs. Every count is of papers, never tuples.
| # | Question | Query | Population | Result |
|---|---|---|---|---|
| A1 | How many papers use webXray, per the extraction? | tools[].name matching /webx[\s-]?ray/i with usedOrMentioned ∈ {used, produced} | 5,859 extracted papers | 7 |
| A2 | …and per the full text? | same regex over paper.cols.txt, whitespace-normalised | 5,859 | 15 |
| A3 | Do the two signals nest? | set difference both ways | — | schema ⊂ sweep; sweep finds 8 the schema misses; 0 the other way |
| A4 | Which spellings? | distinct tools[].name values | 7 papers | webxray (3), webXray (2), WebXRay (1), WebXray (1). Fold residue 0 |
| A5 | What role does webXray play in each? | ROLE hand map, one deciding quote per paper | 15 | owner-list-only 7, citation 4, instrument 1, compared 1, subject 1, miscitation 1 |
| A6 | How many actually ran it or read its list? | role ∈ {instrument, owner-list-only} | 15 | 8; of those 7 list-only, 1 crawler |
| A7 | Which snapshot of the list did they use? | SNAPSHOT hand map over the 8 | 8 | 2 say anything; 1 names a commit |
| A8 | Per-year and per-venue shape | year/venue of the 15 | 15 | last use of crawler-or-list: 2022 |
| A9 | Which ownership resources does the corpus name at all? | six full-text sweeps | 5,859 | PSL 101, Disconnect-as-a-list 74, Tracker Radar 32, WhoTracks.me 23, Crunchbase 23, webXray 15 |
| A10 | Union of the five owner databases | union of A9 rows minus the PSL row | 5,859 / 1,120 crawled | 136 (2.3% / 12.1%) |
| A11 | Which Tracker Radar artefact did each of the 32 use? | TR_ROLE hand map | 32 | ownership 11, tracker/category 9, Collector crawler 9, citation 2, compared 1 |
| A12 | Ownership use of Tracker Radar over time | year of the 11 | — | 2021:1, 2023:2, 2024:1, 2025:4, 2026:3 |
Both hand maps are checked against their sweep at run time and the script prints FAILURE if a paper appears in one and not the other. Residue is 0 in both directions for this run.
A.13 Why the population is a full-text sweep, not the schema
The other crawler pages on this wiki take their population from tools[]. This one cannot: the schema finds 7 papers and the sweep finds 15, and the eight it misses are exactly the distinction the page is about — a paper that cites webXray in its reference list is not a paper that used it. Merging the two signals into one “webXray papers” count would have produced a number that means nothing. They are reported separately and each of the 15 carries a hand verdict.
A.14 Sweep patterns, and the one that had to be tightened
/disconnect/i matches 700 papers (a plain grep -ril gives 697; the sweep normalises whitespace first and catches three more across a column break), almost none of which mean the list — “disconnect” is a common English word and the CSP literature uses it as a technical term ([1Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] measures “the disconnect of code”). The sweep therefore requires a list sense: /disconnect'?s? (entit|list|block|black|tracking)|disconnect\.me|entities\.json/i → 74. That is still an upper bound and the page says so. A sub-agent independently hit the same problem and reported that a looser “disconnect near entit” co-occurrence gave “a useless 481”.
The webXray pattern allows a space or hyphen (web xray, web-xray) because the .cols repair can insert one at a column boundary. It found no additional papers, which is itself worth recording: there is no fork under another name in the corpus.
A.15 Figures verified against the source, not against the extraction
prevalence strings in the extraction are a model summary of a paper's result, so every literal per-paper figure on the page is checked against the paper's own paper.cols.txt by regex. 26 of 26 located. The list, its regexes and its results are in section G of the report output below.
A.16 Evidence quotes spot-checked
The 16 evidence.quote values behind the webXray tool and classification tuples, checked by 4-word-window coverage against paper.cols.txt: 4 exact, 11 partial (≥60% of windows), 1 below threshold. Full list in section H of the report output.
The one below threshold is [2Musa, Maaz Bin; Nithyanand, Rishab (2022): "ATOM: Ad-network Tomography", in: Proceedings on Privacy Enhancing Technologies. (DOI)] at 58%: “We used external data sources including WHOIS records, TLS certificates, and WebXray [61] to identify the parent organizations of each identified tracker.” Read by hand in the source: present and verbatim; the .cols rendering splices an adjacent column through it. Below-threshold is not “unsupported”.
Every Name et al., VENUE YEAR attribution in the 15-row table was also checked mechanically against the author field of the BibTeX key cited beside it, matching the first author's surname. 15 of 15 pass now; one failed on the first run (see *Sub-agent findings that did not survive*, last row).
Additionally hand-read in full, because the page quotes them at length or leans on them:
| Paper | What was checked | Verdict |
|---|---|---|
| [3Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] | the author's own coverage limitation, “primarily contains major ad networks rather than small clients” | present; the sentence is interleaved with Table 1 in the .cols rendering, so the page quotes it with the table text removed and no words changed |
| [4Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] | the priority order, the 3,913 figure, the “one single conflict”, the Disconnect bug report | all present verbatim |
| [1Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] | 2,175 → 1,146 pairs, eight person-hours, 133 additional relations, the twitch.tv example | all present verbatim |
| [5Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] | “Disconnect list commit eb817fb1 (2019-12-10) WebXRay commit 04c3c8e8 (2019-06-18)” | present verbatim, in the reproducibility table |
| [6Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)] | 411 of 762, and the five WHOIS-proxy strings in the top-10 table | present; the 88 total is arithmetic done on the page (34+25+14+8+7) and is labelled as such |
| [7Jakaria, Md; Huang, Danny Yuxing; Das, Anupam (2024): "Connecting the Dots: Tracing Data Endpoints in IoT Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)] | that reference [28] resolves to Held 1998 | present verbatim in the reference list |
| [8McQuistin, Stephen; Snyder, Peter; Perkins, Colin; Haddadi, Hamed; Tyson, Gareth (2023): "A First Look at the Privacy Harms of the Public Suffix List", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] | timlib/webXray 27 in the frozen-PSL table | present |
| [9Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)] | the PhantomJS criticism | present verbatim |
| [10Utz, Christine; Amft, Sabrina; Degeling, Martin; Holz, Thorsten; Fahl, Sascha; Schaub, Florian (2023): "Privacy Rarely Considered: Exploring Considerations in the Adoption of Third-Party Services by Websites", Proceedings on Privacy Enhancing Technologies 2023(1):5-28. (DOI)] | “categorizations differ in granularity and focus” | present verbatim |
| [11Jannett, Louis; Mayer, Andreas; Westers, Maximilian; Mladenov, Vladislav; Mainka, Christian; Schwenk, Jörg (2026): "The State of Passkeys: Studying the Adoption and Security of Passkeys on the Web", in: Proceedings of the USENIX Security Symposium. (Link)] | the Entity Map used so gmail.com and google.com are not counted as separate authentication systems | present verbatim |
| [12Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)] | “an Internet shaped by constant mergers, rebrandings, and regional variation” | present, but column-spliced in the .cols rendering: the words are interleaved with the adjacent column, so a naive substring search fails on part of it. The phrase is the paper's own and the page quotes only the contiguous part |
| [13Wu, Xiaoyuan; Hu, Lydia; Zeng, Eric; Habib, Hana; Bauer, Lujo (2025): "Transparency or Information Overload? Evaluating Users’ Comprehension and Perceptions of the iOS App Privacy Report", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] | “the webXray list includes information on domain ownership, allowing distinction between first- and third-party domains” | present verbatim |
A.17 A figure deliberately NOT published
[1Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] contains: “Contrarily, webXray's list does account for 1,096 of our 1,146 found connections meaning that it alone does not suffice for our purposes.”
1,096 of 1,146 is 95.6% coverage, which contradicts “does not suffice”, contradicts the same paragraph's complaint that the curated lists “frequently miss connections among two hostnames”, and contradicts webXray having contributed only 133 relations. The most likely reading is a dropped “not” in the published text. Checked in three renderings — paper.cols.txt (line 417), paper.txt (line 278, the pre-repair extraction) and the PDF's own text layer — and all three read the same, so this is the paper as published and not a column-repair artefact.
Decision: publish no coverage percentage from this sentence. The page quotes the unambiguous parts of the paragraph instead (2,175 candidates, 1,146 confirmed, ~8 person-hours, 133 additional relations, “it alone does not suffice”). A sub-agent reported the 95.6% reading as a headline finding; it was not used. Someone with access to the authors could settle it; that would be worth doing.
A.18 A miscitation, recorded as a fact and not as an accusation
[7Jakaria, Md; Huang, Danny Yuxing; Das, Anupam (2024): "Connecting the Dots: Tracing Data Endpoints in IoT Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)] writes “Some tools like WebXray [28] provide this information; however, WebXray targets web servers…” and its reference [28] is “Gilbert Held. 1998. Cinco Network's WebXRay. International Journal of Network Management 8, 4 (1998), 254-261.” The described tool is Libert's; the cited work is a 1998 network-management product review. Both halves were read in the source. This is on the page as a hand verdict of miscitation and as a warning that a full-text sweep for a tool name needs the citing sentence read — which is the methodological point, and the only reason it is mentioned.
B. The live database comparison
scripts/owner_dbs.py. Fetches each list, hashes the exact bytes, prints every figure with its denominator and every fold's residue. Snapshot hashes for this run are in the output below; quote them with any figure.
| # | Question | Method | Result |
|---|---|---|---|
| B1 | How big is each list? | count records and distinct domains | webXray 827/3,215; Tracker Radar 19,148/38,368; Disconnect 1,887/7,850 |
| B2 | How deep is webXray's tree? | walk parent_id with a cycle guard | 319 of 827 have a parent; depths 1:508 2:211 3:80 4:24 5:2 6:2; 2,175 of 3,215 domains (67.7%) have leaf ≠ root |
| B3 | How much of webXray's schema is filled? | per-field non-empty count | country 826 (99.9%), platforms 803 (97.1%), uses 761 (92.0%), site policy URLs 591 (71.5%), GDPR 132 (16.0%), trade_groups 104 (12.6%), crunchbase_id 34 (4.1%), opt_out_urls 13 (1.6%), ccpa_urls 4 (0.5%) |
| B4 | What is the universe of third-party domains? | Tracker Radar domain_summary.json, folded to eTLD+1 with the ICANN section of the live PSL | 47,836 rows → 32,369 registrable domains; the ICANN fold merges 15,651 rows, the ICANN+PRIVATE fold only 2,339; 19 rows dropped as not hostnames |
| B5 | Coverage, unweighted | domains in the universe the list can name, parent labels walked | webXray 669 (2.1%), Tracker Radar 5,581 (17.2%), Disconnect 2,268 (7.0%) |
| B6 | Coverage, prevalence-weighted | Σ prevalence of covered ÷ Σ prevalence of all | webXray 58.6%, Tracker Radar 84.3%, Disconnect 80.5% |
| B6b | How much does the fold/lookup rule move it? | all four combinations of ICANN-vs-full fold and walk-vs-exact lookup | weighted coverage spans 54.5–59.7% (webXray), 79.3–84.8% (TR), 75.8–80.6% (Disconnect). The rule moves the answer by more than the lists differ. |
| B7 | Head versus tail | coverage of the top 100 / 1,000 / 10,000 by prevalence | 71/26/5, 98/69/29, 94/68/17 |
| B8 | Pairwise agreement | owner strings after a legal-suffix fold | disagreement 32.4% / 46.8% / 38.0% (see the page table) |
| B9 | Does resolving webXray's tree help? | repeat B8 with root owners | disagreement rises to 39.1% and 49.6% |
| B10 | What does nobody cover? | universe minus webXray minus Disconnect | 29,896 (92.4%) of domains, 14.4% of prevalence weight; the top row is tiktokw.us at 0.041, which only Tracker Radar names |
B.11 Folding: what was folded, the rule, and the residue
One fold only: legal-form suffixes on company names. SUFFIXES is an explicit list — inc, llc, ltd, limited, plc, corporation, corp, company, co, gmbh, ag, kg, sa, sas, sarl, srl, spa, bv, nv, ab, as, oy, aps, sro, pty, pte, pvt, kk, kabushiki kaisha, holdings, holding, sl and a few spellings of those — applied repeatedly until stable, because “Foo Co., Ltd.” needs two passes.
Deliberately not folded: group, media, technologies, digital, networks, solutions. Each is part of a real company name often enough (“Almondnet Group”, “Zeta Global”, “Lotame Solutions”) that stripping it would manufacture agreement between lists that name different companies. Also no synonym merging at all: “Facebook” is not folded to “Meta”, “DoubleClick” is not folded to “Google”. That is the whole point — those are the disagreements being measured.
Residue, printed in full by the script: the fold leaves 782 of webXray's 827 owner names untouched (94.6%), i.e. it changes 45 — but only 16 of those 45 lose a legal-form suffix. The other 29 change because norm_name() strips punctuation before the suffix regex runs (AT&T, 56.com, JD.com, Dun & Bradstreet, “Here, There & Everywhere”). The first version of this page attributed all 45 to the suffix fold, overstating its reach by nearly 3×; the script now prints the two counts separately and lists the punctuation-only names. That is the intended behaviour and its consequence is stated on the page: every “agree” figure is a lower bound on real agreement and every “disagree” figure an upper bound on real disagreement. A synonym-merging fold would move the numbers in a direction the script cannot justify, so it was not written.
The second normalisation is the PSL fold on the universe (B4). Its residue is printed as the 19 dropped rows, listed individually: 18 bracketed IPv6 literals and the literal string “null” with prevalence 0.024 and a full behaviour profile attached. A row that is not a hostname is not a domain; it is printed rather than silently dropped.
B.12 Two denominator choices that change the answer
- The PSL's ICANN/PRIVATE split, and exact-key versus parent-label lookup. This is the choice that moves the answer most, and the first version of this script got it wrong in both halves at once — see the reviewer log below.
googleapis.comis a PRIVATE rule, so folding with the private section leavesfonts.googleapis.com(prevalence 0.369) standing as its own registrable domain, and an exact-key lookup then reports it unowned although all three lists namegoogleapis.com. The script now prints all four combinations; the recommended rule is an ICANN-section fold, and with that fold every key already is a registrable domain so the parent walk is a no-op — which is the cleanest possible confirmation that the fold, not the lookup, is the fix. - Do not use Tracker Radar's
entity_map.jsonordomain_map.jsonas the universe. That is Tracker Radar's answer key and scores it at 100% by construction.domain_summary.jsonis a crawl result and is a defensible universe — while still Tracker Radar's own view of the web, which the page states wherever a coverage figure appears. There is no neutral census of third-party domains; if one existed this comparison would use it. - A registrable domain's weight is the maximum prevalence among its hostnames, never the sum. One site can request
fonts.googleapis.comandajax.googleapis.com, so summing double-counts sites and can exceed 1. Using the max makes every weighted figure a conservative lower bound, and the script says so in a comment at the line that does it.
B.13 The lookup demo published on the page
pages/owner_lookup.py is published on the content page as downloadable code. It was run before publication against the same cached snapshots owner_dbs.py used, and the output on the page is that run's real output, unedited. Its four example domains were chosen to hit one lesson each: a pure granularity difference (doubleclick.net), a pure vintage difference (adnxs.com), a disagreement between the two live lists (facebook.net), and a domain no list covers (fonts.googleapis.com).
The paths in the published copy say cache/… where the local copy reads out/webxray/cache/…, which is the only difference between them.
B.14 Cross-check of one figure by hand
doubleclick.net has prevalence 0.4456 in domain_summary.json and is the highest-prevalence disagreement between webXray (“DoubleClick”) and Tracker Radar (“Google LLC”). Checked by hand: it is in webXray's domains array under owner doubleclick, whose parent_id chain is doubleclick → google → alphabet. So the page's claim that resolving to the root yields “Alphabet” against “Google LLC” is read off the file, not inferred.
C. The hand adjudication
scripts/owner_adjudication.py holds 30 rows: domain, the owner a primary source confirms for today, the date ownership changed, the source URL, and one verdict per list from {current, stale, granularity, error, absent}.
Sourcing bar: a company newsroom or press release, an SEC filing, or the domain's own legal document (privacy policy, imprint, data-processing agreement). Wikipedia, Crunchbase summaries, PitchBook, Tracxn and “list of ad-tech acquisitions” pages were used only to locate leads and never cited. Where only trade press could be found the row is marked UNRESOLVED and is excluded from every tally — 2 of 30, leaving 28. That exclusion was not in the first version, and it mattered: 2 of webXray's 3 “current” verdicts were exactly those two rows, so the figure most favourable to webXray rested on the evidence the script itself called insufficient. Corrected tallies: webXray current on 1 of 23 entries, Tracker Radar 8 of 28, Disconnect 21 of 21.
Corrected again on 2026-09-05, and the second correction is larger than the first. The generic reviewer checked the absent column against the list files and found 9 of the 90 (row, list) verdicts contradict them: Disconnect has an entry for all eight rows it was scored absent on — five of them under resources rather than properties, which the page's own definition of Disconnect's coverage includes — and webXray has “PulsePoint” for contextweb.com. absent is a fact about a file, not a judgement, and this script took the adjudicator's word for it. The tallies are now webXray 1 of 24, Tracker Radar 8 of 28, Disconnect 26 of 28, and scripts/owner_adjudication_absent_audit.py re-derives every one with owner_dbs.lookup (parent-label, because an exact-key lookup manufactures false absences) with the script refusing to print until it passes. The 2026-09-05 pipeline had this check from the start (owner_merge_rows.py), which is why the random sample is unaffected. Every guard on this page passed the wrong table, because the page matched the script and the script's data was wrong.
The selection bias, stated because it is the main limitation: these are the disagreements with the highest Tracker Radar prevalence, not a random sample. They are chosen precisely where the lists differ, so the tallies are not an error rate for any list — a random sample would be dominated by domains all three get right. What they do establish is the shape of the disagreement: concentrated in acquisitions and renames, pointing overwhelmingly one way, with the most conservatively regenerated list being the least current. The script prints this warning above its own tallies and the page repeats it in a <WRAP important>.
C.1 Rows where the reviewer's hypothesis was wrong
Two rows were sent to adjudication flagged as “looks like a Disconnect error”. Both came back the other way:
postrelease.com→ Disconnect says “Life360”, which looked absurd next to webXray's and Tracker Radar's “Nativo”. Life360's own newsroom records the acquisition of Nativo as completed 2026-01-05. Disconnect is current; the other two are seven months stale.crwdcntrl.net→ Disconnect says “PublicisGroupe” against “Lotame”. Publicis' own press release (2025-03-06) and Lotame's own “one year post-acquisition by Publicis” post confirm it. Disconnect is current.
Recorded because the initial reading was mine and it was wrong. The lesson is on the page: do not assume the outlier is the error.
C.2 Rows scored as an error rather than staleness
jsdelivr.net— Tracker Radar says “Prospect One”. jsDelivr's own data-processing agreement (jsdelivr.com/documents/data-processing-agreement.pdf) names Volentio JSD Limited. Prospect One's own portfolio page describes build and infrastructure work for jsDelivr, not ownership.stackadapt.com— Tracker Radar says “Collective Roll”, which is StackAdapt's own pre-2014 founding name, not a distinct owner.1rx.io— webXray's root is “Marimedia”, which no primary source corroborates. Weaker evidence than the other two and labelled as such in the script's comments.
C3. The random sample (2026-09-05)
Section C settled the 30 highest-prevalence disagreements. That set is drawn exactly where the lists differ, so it cannot be an accuracy rate for any list, and the first sitting said so in three places. This section is the measurement that can be: a stratified random sample of each list's own coverage.
Everything below is reproducible from the appendix page, which carries all 15 scripts, the adjudication brief the sub-agents worked from, and all 10 unedited outputs, in full.
C3.1 The frame, and what it excludes
| Frame | the 32,337 registrable domains in Tracker Radar's domain_summary.json as fetched on 2026-09-05, folded with the ICANN section only of the Public Suffix List and resolved by parent-label lookup. owner_sample.py imports owner_dbs.py and uses its universe function, so the frame is built by the same code as the first sitting's |
| …but it is not the same snapshot | the first sitting's frame, on 2026-08-17, was 32,369 domains. Tracker Radar's domain_summary.json and the Public Suffix List both changed between the two dates (see the run table above), and the frame moved by 32 domains. Per-list coverage moved with it: webXray 669→664, Tracker Radar 5,581→5,566, Disconnect 2,268→2,264, and webXray's prevalence mass 58.6%→57.9%. The coverage shares are unchanged to one decimal (2.1% / 17.2% / 7.0%), which is why the two sets of figures can sit on one page without looking inconsistent — and is exactly why the difference is stated here. Verify with python3 scripts/owner_sample.py –cache out/webxray/cache against –cache out/owner_cache |
| Which figures use which | every accuracy figure and the coverage counts in the random-sample section use the 2026-09-05 frame (32,337). Every figure in the earlier coverage and agreement sections, and all 30 rows of the hand adjudication, use the 2026-08-17 frame (32,369) and are not re-derived: they are pinned to the snapshot the adjudication was done against, and re-running them would orphan that work |
| What that is | Tracker Radar's view of the third-party surface, not a neutral census of the web. Every figure inherits it |
| Why not a neutral frame | there is no neutral frame to be had. A list of “all third-party domains” is itself a crawl result. Restricting to domains a real crawl saw is the deliberate choice: an entry for a domain nobody ever requests is not an error anybody meets |
| Who it favours | nobody straightforwardly. It is Tracker Radar's crawl, but Tracker Radar is scored on the same rows as the others and comes out worst on the encounter-weighted axis |
| What it excludes | any domain in webXray's or Disconnect's file that Tracker Radar's crawl never saw. Those entries are never scored — right or wrong. So the rates are conditional on the frame as well as on adjudicability |
Per-list coverage of the 2026-09-05 frame, a census and not an estimate
(owner_sample.py):
| List | Frame domains it names an owner for | Share of 32,337 | Share of prevalence mass |
|---|---|---|---|
| webXray | 664 | 2.1% | 57.9% |
| Tracker Radar | 5,566 | 17.2% | 83.9% |
| Disconnect | 2,264 | 7.0% | 79.9% |
The gap between the third and fourth columns is the whole reason the sample is stratified: webXray covers 2.1% of domains but 57.9% of encounters.
C3.2 Allocation: quartiles, not the deciles the brief asked for
The follow-up item this sitting implements specified “stratified by prevalence decile so the head is not swamped”. Quartiles were used instead, and that is a deviation from the brief, recorded rather than quietly made.
The reason is arithmetic: 60 draws over ten strata is 6 per stratum, and a stratum of 6 supports no statement about that stratum at all — it makes the head-versus-tail comparison, which is the thing the decile design was for, weaker rather than stronger. Four strata of 24/12/12/12 keep the head heavy (the top quartile carries 96.5–99.0% of each list's prevalence mass) while leaving each stratum large enough that its exclusion rate is readable.
| List | Q1 (head) | Q2 | Q3 | Q4 (tail) |
|---|---|---|---|---|
| webXray | N=166, prev 0.005377–0.578843, mass 96.5%, n=24 | N=166, 0.000390–0.005117, 3.3%, n=12 | N=166, 0.000034–0.000390, 0.2%, n=12 | N=166, 0.000007–0.000034, 0.0%, n=12 |
| Tracker Radar | N=1,391, 0.000185–0.578843, 99.0%, n=24 | N=1,392, 0.000027–0.000185, 0.7%, n=12 | N=1,391, 0.000007–0.000027, 0.2%, n=12 | N=1,392, 0.000007–0.000007, 0.1%, n=12 |
| Disconnect | N=566, 0.001767–0.450522, 96.5%, n=24 | N=566, 0.000192–0.001754, 3.1%, n=12 | N=566, 0.000021–0.000185, 0.3%, n=12 | N=566, 0.000007–0.000021, 0.1%, n=12 |
One sample per list, not one shared sample. A single draw from the union
would leave roughly ten webXray rows, which is an estimate with no usable
interval. Each list is drawn from its own coverage instead. A domain drawn for
one list is adjudicated once and its verdict recorded for all three, but
only a list's own draw enters that list's estimate; the other verdicts are
printed in section D of the estimator's output so the discarded observations
stay visible. 180 draws, 5 domains drawn for more than one list, 175 distinct
domains to adjudicate, seed 20260905, sample digest b9acc0cb20cf7ff3.
C3.3 The adjudication, and the sourcing bar
The 175 domains went out in 21 batches of 8–9 to 21 Sonnet sub-agents, each given the brief published in full on the appendix page and nothing else — no hypothesis, no estimator, no sight of another batch. The sourcing bar is section C's, unchanged: a company newsroom or press release, a regulator filing, a company register, the domain's own legal document, or the acquiring parent's own site naming the brand as theirs. Never Wikipedia, Crunchbase, ZoomInfo, PitchBook, Owler, Tracxn, LinkedIn, an acquisitions list or an SEO listicle — those may locate a lead and may not be cited.
One addition to the bar for this sitting: an OV/EV TLS certificate's O =
subject field is accepted as a register-grade source, because it is an
identity a CA validated against a company register. The row cites
tls://<domain>:443 and quotes the exact subject line, and
owner_verify_sources.py re-does the handshake to check it. 17 of the 135
sourced rows are of this kind, and all 17 re-verify. A DV certificate
carries no organisation and is never accepted; owner_probe.sh prints the
issuer next to the subject precisely so the difference is visible, and its raw
output for all 175 domains — 55 of which gave no TLS handshake at all — is
published on the appendix page as the last of its unedited outputs.
unresolved is a correct answer and it is common. 40 of the 175 domains
could not be settled to the bar. That is not a defect of the adjudication; it is
what the tail of the third-party surface is like — domains that serve nothing,
have no legal document, no newsroom, and only aggregators to go on. Every
unresolved row is excluded from every rate, which makes each rate
conditional on the entry being adjudicable, and the exclusion rate rises
toward the tail, so the tail estimates are the ones to distrust:
| List | Drawn | Unresolved | Eligible | Q1 | Q2 | Q3 | Q4 |
|---|---|---|---|---|---|---|---|
| webXray | 60 | 11 | 49 | 20/24 | 10/12 | 7/12 | 12/12 |
| Tracker Radar | 60 | 13 | 47 | 17/24 | 9/12 | 11/12 | 10/12 |
| Disconnect | 60 | 18 | 42 | 20/24 | 7/12 | 7/12 | 8/12 |
Every excluded row is printed in full, with the reason, in section E of the estimator's output — not summarised, not counted and dropped.
C3.4 The verdict retired mid-run
The brief offered six verdicts. One of them, self-named — “the list's owner
is just the domain itself or its bare label … even if a real company of that
name exists” — is wrong, and the sample is what showed it. It fired on
“TrustArc” for trustarc.com, “Klaviyo” for klaviyo.com and “Bilibili”
for bilibili.com, which are the companies, alongside “cdnbasket.net” for
cdnbasket.net, which is not. It fired on 26 of 525 verdicts and it fired
unevenly: 15 of Disconnect's 60 drawn rows against 4 of webXray's. Left in, it
would have cut Disconnect's accuracy denominator by a quarter for a reason that
is about naming style and not about ownership.
Two things replaced it, and both are published:
- Every affected row was re-scored on the ownership question alone into
current/stale/granularity/error/unknown.owner_verdict_fixes.pycarries all 26 re-scores, each with the entity string, the adjudicated owner and the reason. - “The entity name is a domain, not a company” is measured where it belongs — as a census over every domain→owner pair in every file (
owner_selfname_census.py), which needs no sample at all: 0.8% of the 3,215 domains webXray names an owner for, 0.7% of Tracker Radar's 38,368 and 3.2% of Disconnect's 7,850. (Those are distinct domains, not pairs: webXray's file has 3,224 domain→owner pairs, three of them carrying a trailing slash. The census script's column was mislabelled “pairs” until 2026-09-05.) Against the 25% of Disconnect's sample the verdict was firing on, that is the size of the error. The blame belongs to the brief and not to the adjudicators: 25% is close to the census's 27.5% label-only rate for Disconnect, so they were applying the definition as written — which is what a definition saying “even if a real company of that name exists” asks for.
The census also prints a second column — names identical to the domain's label
(klaviyo.com → “Klaviyo”), 29.1% / 4.5% / 27.5% — which is usually not a
defect, because the company really is called Klaviyo. It is printed so the first
column cannot be inflated by conflating the two.
C3.5 Every citation re-fetched, and the ten fixed by hand
The single failure mode that would invalidate the whole study is a plausible-looking URL that does not say what the row claims — the adjudication was done by sub-agents, and a paraphrase inside quotation marks is exactly what a language model produces when told to quote. So every cited source was re-fetched by machine and searched for the quoted sentence.
| Pass | What it does | Result |
|---|---|---|
owner_verify_sources.py | re-fetch the URL with curl-grade HTTP, or redo the TLS handshake for a tls:// source; match the quote after normalising whitespace, quote characters and hyphenation | 121 OK of the 136 rows that carried a source at that point; 13 NOTFOUND, 2 FETCHFAIL |
owner_verify_rendered.mjs | re-check the failures in a real browser (Playwright), because footers and Impressum text are frequently client-rendered | rescued 5: fontawesome.com, flashtalking.com, force.com, conviva.com, cartfulsolutions.com |
owner_corrections.py | by hand, for the 10 that survived both | 4 requote, 5 resource, 1 demote |
The 4 requotes are the finding worth keeping: in 4 of 175 rows the “quote” was a re-ordering or a summary of the page's own words, despite a brief that said verbatim. The claim held in all four and the source stayed; the quote was replaced with text copied out of the fetched page. That is the argument for verifying every citation by machine rather than trusting the adjudicator, and it is the same defect a WebFetch summary produces.
The 5 resource corrections replaced a wrong page for a right claim — a
geo-redirecting landing page, an SEC form URL instead of a document, a host
that serves nothing to non-browser clients. The 1 demote is
cnevids.com: Condé Nast's own privacy policy does not name the domain,
Tracker Radar's “Sabin, Bermant & Gould LLP” is the registrant law firm (a
textbook WHOIS trap), and calling that an error needs a source naming the
real owner. The row became unresolved and left every rate.
Where every row's citation ended up (owner_citation_audit.py):
| State | Rows |
|---|---|
| cited, confirmed in the fetched bytes | 132 — of which 119 are the whole quote, 8 matched as fragments either side of an ellipsis, and 5 matched only on a 60-character window of a long quote |
| cited, confirmed only in a rendered browser | 3 |
makes no ownership claim (unresolved) and needs no citation | 40 |
| uncited | 0 |
The three browser-only rows are fontawesome.com, flashtalking.com and
force.com. A bytes-only audit re-run today still reports them as failures,
which is why the audit script exists: the answer to “so are three figures
uncited?” otherwise lives in three separate JSON files.
Every citation was re-fetched from a cold cache after the review, because a
reviewer found that the fetch cache retained from the run was the first,
partial pass — it held a 526,545-byte body for legal.quantcast.com where
the recorded row has 3,054,755, so a reader reproducing per-row statuses from it
would not have got the published ones. Nothing published depended on the cache,
but the artefact was misleading, so the check was redone rather than explained
away: owner_verify_sources.py was re-run over all 175 rows against an empty
cache on 2026-09-05, hitting every source live again. The result is
identical — OK=119, OK-FRAGMENTS=8, OK-WINDOW=5, NOTFOUND=2, FETCHFAIL=1,
NOSOURCE=40, zero rows changed OK-ness, and zero uncited rows after the
browser pass. The stale cache is kept under out/verify_cache_partial_firstpass
and the reproducing one is out/verify_cache; the re-fetch's own output is
out/owner_verify_sources_refetch-output.txt. This is the strongest evidence
on this page that the citations hold: it is the whole check redone from nothing,
by a different invocation, on a different day's network.
Total changes after the adjudicators returned their rows: 29 of 525 verdicts
— 26 from retiring self-named, 2 from re-scoring webXray's ownership-tree
root rather than its entry (addthiscdn.com, tqlkg.com; both are in
owner_verdict_fixes.py with the reasoning), and 1 from the cnevids.com
demotion. Separately, 10 of 175 rows had their source or quote replaced, and
9 of those 10 left the verdicts untouched.
C3.5a A page-destroying DokuWiki defect, found before publication
Every list item on the drafts of these three pages was written wrapped — the bullet on one line, its continuation indented four spaces on the next — because that is how the source was readable. DokuWiki does not accept that, and what it does instead is much worse than dropping the indentation.
Tested on the live wiki on 2026-09-05 by writing to playground:playground,
rendering through core.getPageHTML, and reverting. Given this source:
* **A bullet whose continuation lines are indented four spaces.** This is the
second line of the same bullet and it is indented by four spaces.
A third line, same indent.
* A second bullet, one line only.
the renderer produces a one-line list item, then the second line as a separate paragraph outside the list, and then this:
<pre class="code"> the style used throughout the new webXray pages. A third line, same indent. * A second bullet, one line only.</pre>
— the third line onwards becomes a preformatted code block that swallows every following bullet as literal text. One wrapped item destroys the rest of its list, and the page source looks entirely reasonable.
The drafts carried 46 of these on the content page and 82 on this one.
Published as written, both pages would have rendered large sections as literal
source. Nothing in the existing tooling caught it: check_wrap.mjs passed, the
figure guard passed, every citekey resolved.
scripts/check_wrapped_lists.py now detects and joins them, and the three page
builders run it and fail if anything survives. Its controls are the thing
that makes it trustworthy: it flags 46 and 82 in the drafts and zero in the
two pages already live — which were written with single-line bullets and do
render correctly. It is also exercised against a nested list, an indented line
inside a published code block, and an indented paragraph after a blank line, and
flags none of them.
The one construct it deliberately does not flag is the 171 indented lines in section H below, which are pre-existing, are the first sitting's plain-text reviewer reports, and render as a code block on purpose.
C3.6 The estimator
| Point estimate | stratified with population weights. Domain-level: R = Σ_h (N_h/N) · x_h/m_h. Encounter-weighted: R = Σ_h (M_h/M) · Σ(prev·y_h)/Σ(prev_h), a ratio estimator within each stratum |
| Denominator | the eligible rows only: a list's own draws, minus unresolved, minus absent. absent is not an accuracy question and is counted over the whole frame instead |
| Intervals | stratified percentile bootstrap, 10,000 resamples, seed 20260905. A normal approximation is meaningless here: several strata have 0 or 1 observation in a cell |
| Zero cells | a bootstrap interval on a cell with no observations is [0,0] and asserts nothing. The script says so and prints the rule-of-three bound instead — below 6.1% for webXray's error rate (49 eligible, 0 seen) and below 7.1% for Disconnect's (42, 0) |
| Empty strata | dropped with the remaining weights renormalised, and printed when it happens. It did not happen in this run: every quartile of every list has at least 7 eligible rows |
Every figure the content page takes from this sample, with its population:
| Figure on the content page | Population | Value, after the 2026-09-11 second pass | Before it |
|---|
Refreshed 2026-09-11 after the second pass over the residue settled 28 of the 40 rows the 2026-09-05 pass could not adjudicate. The before column is kept, because a table of published figures that quietly replaces them is the thing this page exists to prevent. The second pass is documented on ownership_resolution §L; its code is on residue_pass.
| webXray names today's owner, domain-level | 56 eligible of webXray's 60 draws, weighted to its 664 covered domains | 69.2% (56.1–81.4) | was 69.9% (55.5–83.1) on 49 |
| webXray, encounter-weighted | same rows, weighted by Tracker Radar prevalence | 93.8% (80.9–98.6) | was 93.9% (80.9–98.8) |
| Tracker Radar, domain-level | 56 eligible of 60, weighted to 5,566 | 73.9% (61.6–85.3) | was 73.7% (60.8–85.8) on 47 |
| Tracker Radar, encounter-weighted | same | 74.3% (41.1–97.8) | was 68.6% (29.3–97.0) |
| Disconnect, domain-level | 54 eligible of 60, weighted to 2,264 | 89.2% (80.5–96.8) | was 94.4% (86.9–100.0) on 42 |
| Disconnect, encounter-weighted | same | 82.8% (53.8–100.0) | was 83.1% (53.7–100.0) |
current+granularity, domain-level | as above | webXray 78.0 (65.8–88.8), TR 76.4 (64.6–87.5), DC 96.8 (91.5–100.0) | was 80.2, 76.5, 98.8 |
webXray stale, domain-level | 56 eligible | 22.0% (11.4–33.8) | was 19.8% (9.2–32.3) |
error is uncommon, and no longer absent from two lists | 56 / 56 / 54 eligible | webXray 0 (rule-of-three bound 5.4%), Tracker Radar 3 (elfsight.com, cnevids.com, cedscdn.it), Disconnect 1 (km0trk.com) | was 0 / 1 / 0; all four of the new ones were unsettleable in the first pass |
| the three pairwise Fisher tests | eligible raw counts, current only, Bonferroni 0.0167 | none of the three survives: DC vs wx 0.031, DC vs TR 0.083, TR vs wx 0.83 | was DC vs wx 0.014, the one that did |
| coverage 2.1% / 17.2% / 7.0% | census over all 32,337 frame domains, no sampling error | 664 / 5,566 / 2,264 | unchanged — a census, not a sample |
| entity name identical to the domain | census over all pairs: 3,215 / 38,368 / 7,850 | 0.8% / 0.7% / 3.2% | unchanged — a census, not a sample |
| what could not be settled | 175 drawn domains, 180 draws | 12 domains, 14 draws; 4, 4 and 6 of each 60 | was 40 domains, 42 draws; 11, 13 and 18 |
| concentration of the encounter-weighted estimate | the eligible rows, prevalence-weighted | gstatic.com 38.8% of webXray's; omtrdc.net 20.0% of Tracker Radar's; id5-sync.com 16.6% of Disconnect's | was 39.4 / 24.9 / 16.7 |
The estimator was re-implemented from scratch and checked against itself.
Every guard on this wiki compares a page to a tool's output, so a bug inside
the tool passes all of them. To close that for the figures that matter, the
stratified estimator was written a second time — twenty lines, reading only
owner_sample.json and adj_rows_scored.json, importing nothing from
owner_random_sample.py — and run against both the domain-level rates and the
C2 concentration figures. It is committed as scripts/owner_reestimate_independent.py
and published in full on the appendix page with its unedited output, because a
cross-check nobody can run is an assertion rather than a check. All 24
figures reproduce exactly: webXray 69.9 / 19.8 / 80.2 / 0.0, Tracker Radar
73.7 / 22.1 / 76.5 / 1.5, Disconnect 94.4 / 1.2 / 98.8 / 0.0 for current /
stale / current+granularity / error; and gstatic.com 39.4%, omtrdc.net
24.9%, id5-sync.com 16.7% with flips to 54.4% / 43.7% / 66.4%. The script
carries the estimator's figures as a literal table and exits non-zero on any
disagreement, naming which figure — verified by mutating one of them. This does
not cover the bootstrap intervals, which were not re-implemented and are
therefore checked only by the estimator itself.
Those figures are held to the estimator by check_owner_sample_figures.py,
which asserts that each one is present on the page and equal to the script's
output, and fails if a figure has been deleted as well as if it disagrees. It
compares the page to the tool, so it cannot see a bug inside the tool or a
correct figure attached to the wrong sentence; its docstring says so, and the
prose was re-read by hand as well. The guard itself is mutation-tested by
check_owner_sample_figures_mutations.sh, which corrupts one figure at a time
and requires the guard to reject all ten corruptions — including the case where
the mutation changes nothing, which would mean the guard was never exercised.
Both, and the mutation run's unedited output, are on the appendix page.
The four Tracker Radar rows the content page names as stale renames —
zoominfo.com, mountain.com, 20min.ch, tvtime.com — are all in
Tracker Radar's own 60 draws, not borrowed from another list's sample, and
so is elfsight.com. That was checked by hand against owner_sample.json
before the sentence was written, because a verdict recorded on a domain drawn
for a different list is exactly the observation the design discards.
C3.7 Limits, in the order they matter
- Conditional on adjudicability — and this is the limit the 2026-09-11 second pass went after. It read: “11–18 of each 60 rows could not be settled. If unadjudicable entries are systematically worse than adjudicable ones … all three lists are flattered, by roughly the same amount. This is the limit that would most change the numbers and it cannot be bounded from inside the sample.” It was right that the sample could not bound it, and wrong that the flattering would be even: harder sourcing settled 28 of the 40, leaving 4, 4 and 6 of each 60, and Disconnect's domain-level rate fell 94.4% → 89.2% while webXray's and Tracker Radar's did not move. The recovered rows are not significantly worse than the rows already in the rates (p = 0.64 / 1.00 / 0.40 on 7, 9 and 12 rows — almost no power), but they were worse for the list this page recommends, and they cost it the one pairwise comparison that survived Bonferroni. The 12 domains still unresolved are unbounded for the same reason as before. See ownership_resolution §L.
- Conditional on the frame. Entries for domains Tracker Radar's crawl never saw are never scored.
- The encounter-weighted column is a ratio estimator dominated by a handful of head domains, and the estimator now measures by how much. Section C2 of its output reports, per list, the share of the encounter-weighted estimate carried by the single heaviest eligible row, and what the
currentrate would become if that one row's verdict were flipped tostale:gstatic.comcarries 38.8% of webXray's (93.8% would read 55.0%),omtrdc.net20.0% of Tracker Radar's (74.3% would read 54.3%),id5-sync.com16.6% of Disconnect's (82.8% would read 66.2%); the top-three shares are 52.2% / 52.4% / 45.6%. (All six moved on 2026-09-11: the second pass added eligible rows, which dilutes every head domain's share. Before it: 39.4 / 24.9 / 16.7 and 53.1 / 64.9 / 46.0.) The flip is a sensitivity check, not a result. Intervals reach 68 percentage points wide (Tracker Radar,stale+error: 2.9–71.2%). Prevalence in the frame spans five orders of magnitude, so this concentration is what encounter weighting is and not a defect in the estimator — but a headline resting on one observation should say so beside itself, and not only inside an interval. Read the point estimate as an ordering, not as a number. - n is 54–56 per list after the 2026-09-11 second pass over the residue (42–49 before it). The domain-level intervals are ±12 points at best. Any claim that two lists differ needs the intervals or the Fisher test, and webXray versus Tracker Radar domain-level (69.2 vs 73.9) is not a difference this sample can establish — p = 0.83. The figures in this bullet carry no
%sign, which is whysweep_moved_figures.pydid not flag them when the rates moved; the generic review pass did. - A verdict is one adjudicator's reading, and a second reading agrees with it about as often as a coin would on the hardest list. Measured on 2026-09-11 — see C3.8. A random 20% (35 of 175 domains, seed 20260911) was re-adjudicated by a different model against the identical brief and the identical per-domain evidence blocks. On the rows the brief actually asks about, Cohen kappa is +0.027 for webXray (n=12, 25.0% raw agreement, 95% CI -0.200 to +0.280), +0.478 for Tracker Radar (n=30, 70.0%, +0.244 to +0.706) and +0.460 for Disconnect (n=23, 60.9%, +0.193 to +0.707); pooled over the 65 cells, +0.400 at 58.5% raw agreement. 63.0% of the disagreement is one rater finding a source where the other did not, not two readings of the same evidence: on rows both raters settled, the pooled kappa is +0.680 and on
current-versus-rest +0.708. The citation check verifies that the evidence exists and contains the quote; it still does not verify that the verdict follows from it. What would close what remains: a subsample drawn from webXray's own coverage rather than from the 175, so its cell is not 12 rows. - Quartiles, not deciles — see C3.2.
- All three lists are moving targets except webXray's. These rates are of the six files hashed above, on 2026-09-05, and Tracker Radar and Disconnect regenerate. webXray's does not, so its rate degrades from 2021-03-04 forward and from no other cause.
C3.8 Inter-rater agreement: a second model on a random fifth (2026-09-11)
C3.7 recorded, as the limit that would be cheapest to close, that every verdict was one adjudicator's reading and that what would close it was to “re-adjudicate a random 20% with a different model and report agreement”. That was done on 2026-09-11 and this section is the result. It is not a flattering one, and the part that is not flattering is the part worth reading.
Design. scripts/owner_irr_sample.py draws a simple random sample of 35
of the 175 adjudicated domains under a separate seed, 20260911 — the
parent draw's seed is 20260905 and reusing it would have made the audit a
deterministic function of the thing it audits. The subsample is not
re-stratified: it is estimating agreement between two readings of one brief, not
anything about the web, so the population that matters is the sample itself.
Manifest digest 07208284bbd90fd6; the 35 domains are listed in the script's
output on the appendix
page, section AF.
The stimulus is identical, not merely equivalent. The second rater was given
out/adj/INSTRUCTIONS.md byte-for-byte, and the per-domain blocks — the three
lists' claims plus the TLS/RDAP/HTTP probe — were copied verbatim out of the
first rater's batch files rather than regenerated. Re-running owner_probe.sh
today would have handed rater 2 a different stimulus and any disagreement would
then be partly the web moving. The brief's own “Today is 2026-09-05” line was
left in place for the same reason.
A different model, and which way that cuts. Rater 1 was 21 Sonnet sub-agents; rater 2 was 5 Fable sub-agents, 7 domains each, no sight of rater 1's verdicts, of each other, or of the estimator. A different model family was chosen over a different instance of the same one deliberately: two Sonnet runs would share their failure modes, and an agreement figure whose two raters fail together overstates the instrument. The cost is that a capability difference and a reading difference are not separable, and section 2b below suggests most of what was found is the former.
| Population | What it is | webXray | Tracker Radar | Disconnect | Pooled cells |
|---|---|---|---|---|---|
| NAMED — headline | rows where the list file names an owner: the rows the brief asks a rater to judge | +0.027 (n=12, agree 25.0%) | +0.478 (n=30, agree 70.0%) | +0.460 (n=23, agree 60.9%) | +0.400 (n=65, agree 58.5%) |
| 95% CI of the above | stratified-free percentile bootstrap over domains, 10,000 resamples | -0.200 to +0.280 | +0.244 to +0.706 | +0.193 to +0.707 | +0.241 to +0.551 |
| ALL rows | all 35, including the cells owner_merge_rows.py forces to absent from the list file | +0.525 | +0.614 | +0.663 | — |
| ADJUDICABLE | NAMED rows where both raters reached an ownership verdict | +0.250 (n=6) | +0.718 (n=24) | +1.000 (n=9) | +0.680 (n=39, agree 84.6%) |
current vs rest | the same rows, collapsed to the one distinction every published accuracy figure rests on | +0.400 (n=6) | +0.710 (n=24) | +1.000 (n=9) | +0.708 (n=39, agree 87.2%) |
The ALL row is in the table to be argued with, not to be quoted. More than
half of webXray's cells on this subsample (23 of 35) are absent, and
absent is re-derived from the list file by the merge script rather than
judged, so the two raters cannot disagree on them. Counting those as agreement
lifts webXray's kappa from +0.027 to +0.525 without a single extra judgement
being made. Any inter-rater figure computed over a verdict set that includes a
mechanically-forced category is inflated by exactly this much, and that is the
form most such figures take.
The result, in one sentence
On the rows the brief actually asks about, two models given identical inputs agree on 58.5% of 65 verdicts, kappa +0.400 (95% CI +0.241 to +0.551) — and about two-thirds of the disagreement is one rater finding a source where the other did not, rather than the two reading the same evidence differently.
What the disagreements are made of
| List | NAMED cells | agreed | differ on sourcing (one rater unknown) | differ on reading (two different verdicts) |
|---|---|---|---|---|
| webXray | 12 | 3 | 4 | 5 |
| Tracker Radar | 30 | 21 | 6 | 3 |
| Disconnect | 23 | 14 | 7 | 2 |
| Pooled cells | 65 | 38 | 17 | 10 |
17 of the 27 disagreeing cells (63.0%) are one rater reaching unknown
where the other reached a verdict. The row-level picture is starker: rater 1
left 10 of 35 rows unresolved and rater 2 left 2, and they made the same
resolved/unresolved call on 27 of 35. Rater 2 settled every one of Tracker
Radar's 30 named rows; rater 1 settled 24.
That decomposition is the main finding of this section and it changes what the
figures mean. Restricted to rows both raters settled, the pooled kappa is
+0.680 with 84.6% raw agreement, and on current-versus-rest it is
+0.708 with 87.2%. So the verdict vocabulary is reasonably reliable and
the sourcing effort is not: given the same evidence, two models mostly read it
the same way; given the same brief, they do not find the same evidence. For a
reader planning their own adjudication, that says the thing to specify and
control is not the label set but how hard and how uniformly the adjudicator is
made to look.
Which way rater 2 moves the numbers
Unweighted current share among each rater's own adjudicable rows on this
subsample. This is not the estimator — 35 rows, no stratum weights, a
different denominator per rater — and it is here to answer one question only:
does the second rater make the published rates look better or worse?
| List | rater 1 | rater 2 |
|---|---|---|
| webXray | 50.0% (4/8) | 70.0% (7/10) |
| Tracker Radar | 66.7% (16/24) | 66.7% (20/30) |
| Disconnect | 80.0% (8/10) | 76.5% (13/17) |
Rater 2 is more favourable to webXray, identical on Tracker Radar and slightly less favourable to Disconnect. So the disagreement does not point in the direction of “the published accuracy figures are too kind”; if anything the opposite for the list the page is about. That is a statement about direction on 35 rows, not a re-estimate, and it must not be read as one.
The retired verdict came back
self-named was retired mid-run in the first sitting because it fired on
company names (C3.4). It is still in the published brief, and rater 2 —
which had never seen that decision — reached for it 6 times in 105 cells
against rater 1's 4, agreeing with rater 1 on 3 of them. Two models
independently finding the same wrong verdict attractive is evidence that the
defect was in the brief rather than in one adjudicator, which is what C3.4
claimed on weaker grounds. It also means the brief published on the appendix
page still contains a verdict the study does not use, and a reader reusing it
should delete that bullet.
Rater 2's own citation check
owner_verify_sources.py was re-run over rater 2's 35 rows from an empty
cache. It returned three failures — and all three turned out to be limits of
the instrument rather than bad citations. That is the uncomfortable part: on a
35-row sample the checker's false-failure rate was 3 in 33, and one of the three
was acted on before a reviewer caught it. After the fixes below the tally is
OK 30, OK-MARKUP 1, FETCHFAIL 2, NOSOURCE 2 — zero NOTFOUND, so every one of
rater 2's 33 sourced rows has a verified citation. The two NOSOURCE rows are the
two rater 2 declared unresolved, which is correct. What each failure actually
was (scripts/owner_irr_corrections.py):
| Row | Machine verdict | What a hand check found | Action |
|---|---|---|---|
adgrx.com | FETCHFAIL | samsungads.ca serves a Let's Encrypt certificate that expired 2026-03-12 and has not been renewed, so curl -sL and Chromium both refuse it; curl -sk returns 59,465 bytes carrying the quoted sentence verbatim | kept |
cedscdn.it | FETCHFAIL | the rater cited %whois://whois.nic.it/cedscdn.it%, a scheme neither instrument can open. A port-43 query returns Organization: CED DIGITAL & SERVIZI Srl verbatim | kept; the gap is the brief's, which offers %tls://% and no whois equivalent |
kameleoon.io | NOTFOUND in bytes and NOTFOUND in a rendered browser | the citation is sound and the check was wrong. The quote is a %<script src>% line, and both instruments are blind to markup by construction: norm() strips %<…>% before searching and owner_verify_rendered.mjs reads document.body.innerText, which excludes script elements. The string is verbatim in the 27,121 bytes curl returns | kept; this row had already been demoted on the machine verdict and the demotion is reversed |
This one was published as a demotion first and corrected after a reviewer
fetched the page. It is left in the record because it is the most instructive
failure of the sitting: two independent checks agreed, and they agreed because
they share an assumption — that evidence lives in a document's text. A second
instrument that fails the same way is not a second instrument. The verifier now
has an OK-MARKUP status for exactly this case.
Because nothing is demoted, the corrected and uncorrected inter-rater figures are
identical — owner_irr_corrections.py reports “verdict cells changed: 0”
and prints a warning if that ever stops being true. Both the headline table and
the correction pass therefore rest on the same 65 cells: Disconnect +0.460,
pooled +0.400.
Three defects this pass found in the instruments
owner_verify_sources.pywas deleting every non-Latin character. Itsnorm()ended in%re.sub(r“[^a-z0-9]+”, “ ”, s)%, so a Cyrillic, Greek, Hebrew or CJK quote was checked only on whatever Latin substring it happened to contain — and a wholly non-Latin quote normalised to the empty string and was reportedNOQUOTE, a status that reads like “the row has no quote” rather than “I could not check this quote”.owner_verify_rendered.mjshas always used%[^a-z0-9À-]+%, so the browser pass and the bytes pass disagreed about what a character is. Found because rater 2 cited a wholly Chinese sentence forqbox.me. The Python class is now the JavaScript one. Consequence for the first sitting: re-running the fixed script over all 175 rows moves three statuses, and only one of them is this fix biting —compass-fit.jp, whose Japanese quote was previously verified on the six lettersCOMPASS-FITalone and is not verbatim on the cited page. The mismatch is one character: the page's body reads台湾国内での、ネイティブ型広告配信サービス…and the row's quote drops the full-width comma after台湾国内での. Everything else, including the致しました。ending, is on the page. So this is a transcription slip rather than a paraphrase, and the row's claim is unaffected — MicroAd's own newsroom page carries the partnership, the service name and Taiwan — so no published rate moves. The row needs arequoteinowner_corrections.pyand that is recorded as open work rather than done here. A first draft of this bullet said the page's headline endsの提供を開始and implied the致しました。ending appears nowhere; a reviewer found the full sentence in the body and the real one-comma difference.tns-counter.ru, the other non-Latin row, verifies clean under the fix.- Neither citation check can see a quote that is markup.
owner_verify_sources.norm()strips%<…>%before searching, which is right for a prose quote and fatal for a citation to a%<script src>%, a%<link>%or an attribute;owner_verify_rendered.mjsreadsdocument.body.innerText, which excludes script elements for the same reason. So the two passes, which exist to catch each other's blind spots, share this one — and a false NOTFOUND from both was acted on. The bytes pass now also searches an un-stripped copy and reportsOK-MARKUPwhen that is where the quote lives. Re-running the fixed script over the first sitting's 175 rows movesfontawesome.comfrom NOTFOUND to OK-MARKUP; that row was already rescued by the browser pass inowner_citation_audit.py, so the published “0 uncited rows” is unchanged and now has two independent routes rather than one. - The citation check is time-sensitive, and by a measurable amount. The same 175 rows re-verified six days later move 3 statuses of 175, two of them for reasons that are purely the web moving:
cartfulsolutions.comwentOK→OK-WINDOW, andbrightcove.netwentOK-WINDOW→FETCHFAILon one run and back toOK-WINDOWon the next, which is a host being briefly unreachable rather than a citation changing. A “zero uncited rows” claim is therefore a claim about a date, and this page should be read as saying the citations verified on 2026-09-05 and 2026-09-11 rather than that they verify.
What this section does and does not license
- It does not rehabilitate the per-list accuracy estimates, and it does not overturn them either. Kappa is about two readings of the same rows; it says nothing about whether either reading is right. A sample on which both raters are confidently wrong scores kappa 1.
- webXray's +0.027 is the weakest cell in the table and n=12 is why. Its interval, -0.200 to +0.280, includes both “no agreement beyond chance” and “moderate agreement”. The honest reading is that this subsample cannot characterise agreement on webXray at all; it took 12 named rows to the 30 Tracker Radar got because webXray names an owner for 2.1% of the frame. Closing that needs a subsample drawn from webXray's coverage rather than from the 175.
- The pooled column counts cells, not domains. One domain contributes up to three verdicts and they are not independent — the same fetched document usually settles all three. The pooled interval is therefore narrower than it should be, and it is in the table for its point estimate, not its width.
- Kappa is not read here against the usual verbal bands. They are a convention, not a threshold, and at these marginals — one category carrying most of the mass — kappa is driven as much by the marginals as by the disagreements. Raw agreement is printed beside every kappa for that reason, and
ADJUDICABILITYfor Tracker Radar is the extreme case: rater 2 settled all 30 rows, so one rater used a single category, and kappa is 0.000 by construction at 80.0% raw agreement. A zero there means “undefined”, not “no agreement”. - Six days separate the two ratings. Corporate ownership rarely moves in six days, and the frozen probe removes the mechanical part of the drift, but a page that changed in that window is indistinguishable from a rater that read it differently.
D. External sources: verified, and rejected
Every external claim on the page, with how it was checked. All checks 2026-08-17.
| Claim on the page | How verified |
|---|---|
github.com/timlib/webXray → 404 | curl -o /dev/null -w '%{http_code}' https://api.github.com/repos/timlib/webXray → 404 |
github.com/timlib/webXray_Domain_Owner_List → 404 | same method → 404 |
the timlib account still exists with 0 public repos | api.github.com/users/timlib → 200, public_repos: 0, company: “webXray.ai”, blog: “https://timlibert.me” |
webxray.org is a placeholder | fetched; body is a heading and one line, “Public interest projects for the interested public.” No source link, no version, no download |
| no PyPI package | pypi.org/pypi/{webxray,web-xray,policyxray}/json → 404 each |
thezedwards/webXray last pushed 2021-03-04, 19 forks, newest fork activity 2023-03-12 | GitHub API repo + /forks?per_page=100 |
the surviving README still says git clone https://github.com/timlib/webXray.git | fetched raw.githubusercontent.com/thezedwards/webXray/master/README.md |
PolyForm Strict License 1.0.0 on the thezedwards snapshot | fetched LICENSE.md from that repo and read the licence text itself, not a summary. This is not webXray's final licence — see the next row |
webXray was relicensed to GPLv3 (2021-06-14, commit 245ec5d7) and then to MIT (2023-02-01, commit 73fe0fc9, authored by “Tim Libert”) | fetched api.github.com/repos/peterjoles/webXray (spdx_id: MIT), its LICENSE (“Copyright © 2023 Tim Libert”), the commits?path=LICENSE history, and compare/master…peterjoles:master (ahead_by: 36). The page's first version asserted webXray “is not open source and redistributing it is prohibited” from the thezedwards snapshot alone, which was the most restrictively licensed copy in existence — the exact mistake the page tells readers to avoid |
| PolyForm Strict 1.0.0 is still the current version and has no SPDX identifier | polyformproject.org/licenses lists strict/1.0.0 and no later Strict version; SPDX's own license-list-data JSON carries only PolyForm-Noncommercial-1.0.0 and PolyForm-Small-Business-1.0.0 |
| raw CDP, no Selenium | webxray/ChromeDriver.py imports create_connection from websocket; requirements.txt pins lxml==4.6.2, psycopg2-binary==2.8.6, textstat==0.7.0, websocket-client==0.57.0 |
| the bundled PSL patch header “current as of 20160428” | fetched webxray/resources/pubsuffix/ccSLD-patches.txt |
| the PSL's own warning against frozen copies | publicsuffix.org/learn/ |
RDBinns/webXray_Domain_Owner_List is GPL-3.0, all commits 2018-04-05, older schema | GitHub API (license.spdx_id: GPL-3.0) plus its README, which documents owner_name where the in-tool file has name and has no uses/platforms/trade_groups |
webxray.ai is a commercial litigation product; Libert is founder and CEO | fetched both webxray.ai and timlibert.me. The latter states verbatim: “(Dr.) Timothy Libert is founder and CEO of webXray LLC.” |
| Tracker Radar last commit 2026-08-12, monthly release tags, CC BY-NC-SA 4.0 | GitHub API pushed_at; /releases; LICENSE and README |
| Tracker Radar's method and “handled internally” | docs/DATA_MODEL.md and docs/FAQ.md in the repository |
| Disconnect last commit 2026-08-07, CC BY-NC-SA 4.0, GPL-3.0 until 2020-06-24 | GitHub API; then the commit history of LICENSE — GPLv3 in the initial commit 89d421e8 (2015-10-13), rewritten to CC BY-NC-SA by two commits on 2020-06-24 |
Firefox ships Disconnect's services.json via Mozilla's shavar-prod-lists | mozilla-services/shavar-prod-lists README: Firefox's ETP “rely on lists of trackers maintained by Disconnect… Mozilla does not maintain these lists”; disconnect-blacklist.json is “a version controlled copy of Disconnect's list of trackers” |
Ghostery trackerdb 2026-08-06, CC BY-NC-SA 4.0; WhoTracks.me data repo 2026-08-04 (“July update”); whotracks.me now redirects to ghostery.com/whotracksme | GitHub API on ghostery/trackerdb and whotracksme/whotracks.me (the ghostery/whotracks.me path 301s); followed the site redirect. The licence is from the repo's own LICENSE and package.json (“license”: “CC-BY-NC-SA-4.0”), not the GitHub API, which returns no SPDX id — a first pass read “MIT” off the API and was wrong |
| WhoTracks.me paper is arXiv 1804.08959, v2 revised 2019-04-25 | export.arxiv.org/api/query?id_list=1804.08959 — title, authors and updated field read from the Atom response |
| Crunchbase API is paid | api.crunchbase.com/api/v4/… → HTTP 401 “Unauthorized user_key” |
| PhantomJS is dead | api.github.com/repos/ariya/phantomjs → archived: true, last push 2022-11-26 |
| Disconnect assigns ownership “through DNS, WHOIS, and behavioral evidence”; corrections are not taken by pull request | fetched disconnect.me/trackerprotection (the six-step process page) and disconnect.me/domain_evaluations; the repository README says verbatim “Pull requests are not reviewed and will be closed” |
Disconnect's own headline count differs from entities.json | disconnect.me/trackerprotection claims “14,332 Verified domains and entity mappings”; entities.json holds 7,850. The page tells the reader to count the file they read rather than quote the vendor's number |
D.1 Sources rejected
- Every “top web privacy tools” listicle and blog roundup surfaced while looking for webXray's current home. None was used. The tool's status was established from GitHub API responses, PyPI 404s, the licence file and the author's own two websites.
- Wikipedia, Crunchbase profiles, PitchBook and Tracxn for the 30 acquisitions. Used to find which press release to look for; never cited. Two rows where only trade press existed are marked
UNRESOLVEDinstead of being filled from it. toolness/webxray(Mozilla's Web X-Ray Goggles, 2010–2013) — a genuinely different project that shares the string. Checked and excluded, because a name search finds it first: GitHub's repository search for “webxray” returns 18 repositories, of which the majority are unrelated (X-Ray Goggles, an X-ray physics database, an offensive web scanner, a Shopify evaluator). Recorded here so the next run does not include it.- Cinco Network's WebXRay (Held, IJNM 1998) — the homograph [7Jakaria, Md; Huang, Danny Yuxing; Das, Anupam (2024): "Connecting the Dots: Tracing Data Endpoints in IoT Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)] cites. Excluded from the tool discussion; mentioned only as the miscitation it is.
carlsaturnino/webxray-dockeron Docker Hub (a counter: 399 pulls on 2026-08-17, 405 on 2026-09-05;last_modified2024-10-16) was not presented as a maintained image: the timestamp is a metadata bump on an image registered 2016-04-26, andhub.docker.com/v2/repositories/carlsaturnino/webxrayand…/timlib/webxrayboth 404. The page says no maintained image exists.
D2. Wayback: dating webXray's disappearance
The first sitting recorded two questions as unknown because
web.archive.org returned 502/503 for the whole of 2026-08-17 — nine
attempts, two independent fetch paths, about thirty minutes apart. The Archive
was up on 2026-09-05 and answers both. Both answers are bounds, not dates,
because the Archive has no capture inside either window, and they are published
as bounds.
scripts/wayback_webxray.sh holds the seven CDX queries; its unedited output is
on the appendix page.
The method has two traps worth stating.
collapse=digestreturns one row per distinct response body, which is what makes a 200→404 transition visible in a handful of lines — and it also hides how many captures there were. So every transition found that way is re-queried uncollapsed over a narrow window before any date is published. That is query 3, and it is what turns “somewhere in 2023” into a two-row bracket with nothing between.- A CDX row is a capture, not a page. The first row of a domain's history is very often a registrar or host parking page, and reading a date off the index without fetching the capture is how a first draft of this section dated the pre-webXray product to 2010-10-08 when that capture is a DreamHost “Coming Soon” page. Every date below was read out of the fetched
id_body, not out of the index. Query 7 was added after review for the same class of reason: queries 5 and 6 stop in 2024, so the published output did not contain the 2026 capture thewebxray.aibound rests on — a bound whose evidence is not in the output is an uncited figure however true it is.
| Question | Answer | Evidence |
|---|---|---|
When did github.com/timlib/webXray first 404? | between 2023-03-31 and 2023-11-15. Not narrower | last archived 200 20230331144433; first archived 404 20231115075841; the uncollapsed query over from=20230101&to=20240601 returns exactly those two rows, so there is no capture between them |
| …anything that narrows it? | no | the github.com/timlib profile has captures at 2022-10-27 and then not again until 2025-08-16. The one corroborating fact is outside the Archive: peterjoles/webXray's last push is 2023-03-12, i.e. the fork was taken and then upstream went |
When did github.com/timlib/webXray_Domain_Owner_List first 404? | between 2022-10-06 and 2025-01-18. A much wider window; do not narrow it | last archived 200 20221006015001; first archived 404 20250118193856 |
What did webxray.org serve, and when did the demo die? | a live demo search engine from 2022-01-23 to at least 2024-03-28; 401 Authorization Required on 2024-05-24; a 301 to webxray.ai from 2024-07-24 to at least 2026-02-05 | the 20240328005456id_ capture carries “The search engine is drawn from scans of 100,000 sites” and the company drop-down (33Across, 360, 51La, AdGear, Adition Technologies, …). 20240524194734 is a 401 from nginx/1.18.0. Every capture from 20240724105426 on is a 301 whose Location is https://webxray.ai/ |
When did the domain come back off webxray.ai? | between 2026-02-05 and 2026-08-17. Cannot be narrowed | query 7 (uncollapsed, 2025–2026) returns 14 rows, every one of them the same 301 digest OM6ALWQTEGJ5NTY5RYSC2EXCKJOAHOVU, the last being 20260205035830. After that the Archive has no capture at all of webxray.org, so it has never seen the placeholder served today |
Two further findings from the same captures, both of which changed the content page:
webxray.orgwas a different product before it was webXray. From 2010-11-08 to at least 2011-09-27 it served a page whose<title>is “Web Xray - Free Website Data Mashup Tool” and which carries a “beta 0.5” badge beside the logo — a public directory of analysed sites with Google AdSense units (google_ad_client = “pub-2283090127531811”) and a Clicky analytics call. Nothing to do with third-party request measurement. By the next capture, 2014-12-18, the domain served webXray; there is no capture in between. Recorded because a reader dating the project from its domain registration would be off by four years.
A first draft of this bullet said 2010-10-08, taking the first row of the CDX query as the first Web Xray page. It is not:web/20101008070114id_is a DreamHost “Coming Soon” parking page (“The DreamHost customer who owns webxray.org has not yet uploaded their website”). Caught by a reviewer who fetched the capture instead of reading the timestamp, which is the whole lesson — a CDX row is a capture, not a page.- A fourth licence state, and it is a reversion. The 2015-05-30 capture describes webXray 1.0's cost as “Free!*” and “(*Subject to terms of the GNU Public License.)”, linking that phrase to
gnu.org/copyleft/gpl.html. That URL served GPLv3 in 2015 (archived atweb/20150531125231as “Version 3, 29 June 2007”) and still resolves tognu.org/licenses/gpl-3.0.htmltoday — a 302 on the path, after a 301 from the baregnu.orghost towww.gnu.org. So the sequence is GPLv3 (2015) → PolyForm Strict (by 2021-03) → GPLv3 again (2021-06-14) → MIT (2023-02-01), and the 2021 commit titled “Now open-source” was a return, not a first opening. The content page's licence paragraph was rewritten on that basis.
Two honesty notes on this evidence:
- The 2015 quote joins two adjacent
<p>elements — the page reads<p>Free!*</p>then<p>(*Subject to terms of the <a …>GNU Public License</a>.)</p>. It is the visible text as a reader saw it, not a contiguous string in the source. - The licence claim rests on the project's own website, not on a repository
LICENSEfile. No snapshot oftimlib/webXrayfrom 2015 survives — the earliest mirror is the 2021thezedwardscopy — so what the repository itself said in 2015 is not recoverable and is not asserted.
E. What could not be established
- ANSWERED — When
github.com/timlib/webXrayfirst returned 404, and whatwebxray.orgused to contain. Answered on 2026-09-05, when the Archive was up — see D2. The bullet is rewritten rather than deleted so the gap and its closure are both on the record. What the first sitting wrote: “web.archive.orgwas returning 502/503 for the whole of 2026-08-17 (nine attempts, two independent fetch paths, ~30 minutes apart). Both questions are marked unknown on the page rather than estimated.” What is still not established: both answers are bounds, not dates — the repository 404ed somewhere between 2023-03-31 and 2023-11-15, the owner-list repository between 2022-10-06 and 2025-01-18, and the domain came back offwebxray.aibetween 2026-02-05 and 2026-08-17. The Archive has no capture inside any of the three windows, so nothing short of a source outside the Archive will narrow them, and none is known. - Whether a paper documents policyXray specifically, as distinct from the module description in webXray's README. Not settled; scholarly indices were not searched exhaustively. [3Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)] is the closest thing and is what the page cites.
- Whether DuckDuckGo has published a paper on Tracker Radar's methodology. Nothing found; the documentation is repo-native plus a vendor blog post. The page says “no paper” on that basis, which is an absence-of-evidence claim and is phrased as one.
- What
propertiesversusresourcesmeans in Disconnect'sentities.json. Not documented anywhere in the repository — not in the README, the LICENSE orDomain evaluations.md. Empirically most entities have identical arrays, and some diverge (24TTL:properties: [“24ttl.net”]vsresources: [“24ttl.stream”]). The page reports both counts and tells the reader to state which key they read, rather than asserting a definition the vendor does not document. - The
1,096 of 1,146sentence in [1Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)] — see A.17. Unresolvable from the text. - The list-age and missing-hostname columns for
timlib/webXrayin [8McQuistin, Stephen; Snyder, Peter; Perkins, Colin; Haddadi, Hamed; Tyson, Gareth (2023): "A First Look at the Privacy Harms of the Public Suffix List", in: Proceedings of the ACM Internet Measurement Conference. (DOI)]'s table. The.colsrendering splits that table into separate column runs and the numeric columns cannot be aligned to the repository names reliably. Only the verifiable part is on the page — that webXray appears in the table, in the “Production” category, with 27 stars. What would close it: read the published PDF's table directly. - Whether any non-numeric claim on this page has drifted from its script.
check_page_numbers.mjscompares numerals only, so when a reviewer's fix changed “third of four” to “second of four” on the page, the hand map in the script and its committed output kept saying “third” and every guard still passed. The same happened with “MIT” for Ghostery's licence. Both were caught by a human reviewer reading the two side by side, not by a tool, and no tool here covers that class.check_attributions.mjscloses one narrow case of it (author names). What would close the rest: a guard that extracts the quoted strings a page attributes to a script's hand map and diffs them against the map. - ANSWERED — A random-sample accuracy rate for any of the three lists. Measured on 2026-09-05 — see C3. 175 domains, 525 verdicts, 138 settled, per-list rates with bootstrap intervals. What the first sitting wrote, and it was right about the cost: “tens of hours at the rate section C ran at”. What is still not established, and these are the parts a reader should hold against the figures:
- Whether unadjudicable entries are worse than adjudicable ones. 40 of 175 rows could not be settled to the sourcing bar, and the rate rises toward the tail. Every accuracy figure is conditional on adjudicability, and if the unsettled entries are worse — which is plausible — all three lists are flattered. This cannot be bounded from inside the sample. What would close it: a second, harder-sourcing pass over the 40, e.g. national register searches by company number rather than by name, and RDAP history where a registry publishes it.
- ANSWERED — Inter-rater agreement on the verdicts. Measured on 2026-09-11 — see C3.8. What the second sitting wrote: “Each row was adjudicated once. The citation check confirms the evidence exists and contains the quote; it does not confirm the verdict follows from it.” 35 domains re-adjudicated by a different model against the identical brief; pooled kappa +0.400 over 65 named cells. What is still not established: agreement on webXray specifically — its cell is 12 rows and its interval, -0.200 to +0.280, spans everything from chance to moderate. What would close it: a subsample drawn from webXray's own 60-domain draw rather than from the 175, which needs roughly 40 more adjudications.
- Whether the two raters are wrong together. Kappa measures whether two readings coincide, not whether either is right; a sample both raters get confidently wrong scores 1. Nothing here bounds that, and the only thing that would is an adjudication against a source neither rater could reach for — a company register queried by company number, not by name. What would close it: the same harder-sourcing pass named in the bullet above, used as a third rating rather than as a rescue.
- Whether the rates hold outside Tracker Radar's crawl. The frame is Tracker Radar's
domain_summary.json, so entries for domains that crawl never saw are never scored. What would close it: repeat the draw against a frame built from a different crawl, e.g. one of the corpus papers' own request logs. - A per-stratum statement. 12 eligible rows per tail stratum supports the pooled estimate and nothing finer. The head-versus-tail difference visible in the exclusion rates is real; a head-versus-tail difference in accuracy is not something this sample can assert.
- Whether anyone has built an LLM-based domain-owner resolver. Nothing in this corpus. [12Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)] and [14Gouda, Deepak; Dainotti, Alberto; Testart, Cecilia (2025): "Prefix2Org: Mapping BGP Prefixes to Organizations", in: Proceedings of the ACM Internet Measurement Conference. (DOI)] do the network-layer equivalent. Named on the page as a gap rather than as a method to use.
F. Judgement calls
- Broadening the page beyond its title. See *Scope decision* above. Anyone would notice the page spends more words on Tracker Radar and Disconnect than on webXray; that is deliberate, because a reader arriving at a dead tool needs to be sent somewhere.
- Calling webXray's owner list “historical” and Tracker Radar's / Disconnect's “current practice”. Evidence: last corpus use of webXray's list is 2022; ownership use of Tracker Radar runs 2021→2026 with 7 of its 11 papers in 2024–2026; the file itself has not changed since 2021-03-04; the tool is unobtainable; and 17 of the 25 adjudicated rows it covers are stale. The 2025–2026 corpus slice is provisional, so the year-by-year counts are given rather than a trend claim, and the page says the counts are too small to be a market share.
- Not calling any list “the best”. Restated 2026-09-05: the figures this bullet rested on were wrong — Disconnect is current on 26 of the 28 adjudicated rows and absent on none of them, not “22 of 22 … absent on 8 of 30”. The conclusion is unchanged and is now better supported. Disconnect is current on 26 of 28 rows and carries no hierarchy; Tracker Radar covers everything and is stale on renames. The page gives a per-question table instead of a ranking. A ranking would have been shorter and less useful.
- Publishing the
webXray-rootcomparison even though it makes agreement worse. The tempting move was to report only the leaf comparison, since “resolve the hierarchy” sounds like the obvious fix. Publishing both is what makes the granularity-versus-vintage decomposition visible, and that decomposition is the page's main analytical contribution. - Keeping the
nulldomain and the IPv6 literals in the write-up. They are a defect in a widely used published dataset, they are cheap to state, and a reader who does not know about them will silently include anullrow in a prevalence table. - Naming the miscitation. Recorded as a verifiable fact about a reference list, with the methodological point it illustrates, and with no characterisation of the authors. The alternative — a vague “watch out for homographs” — would have been unfalsifiable.
- Quoting webXray's licence at length. It is the single most actionable fact on the page: a student who plans a study around webXray has planned a study around software they cannot legally obtain. A one-line “non-commercial licence” would have understated it.
- Discoverability, and where this material really belongs. The wiki's only treatment of domain-to-company ownership resolution now sits under the name of a dead tool, where nobody asking “how do I attribute a domain?” will look. The generic reviewer was right that the earlier rejection of a separate
design:ownership_resolutionpage (“it would leave a red link pointing at a stub”) was a weak argument — a real webXray page and a topic page can coexist. Mitigated for now by linking this page from Requests, which is where a reader asking “whose request is this?” actually lands. Splitting out a topic page is worth doing and is recorded here rather than done. <WRAP important>boxes rather than<wrap todo>. No open TODOs were left on the content page: the unknowns are stated in its methodology section with what would close them, which is where a reader checking a number will look. The one real TODO — correcting Crawler — is recorded here, because it is work on a different page.- Drawing the sample at all, rather than extending the hand adjudication. The cheap move was another 30 high-prevalence disagreements, which would have made the existing table look better sourced and would have measured nothing new. The sample is the expensive move and it is the one that produced a result that contradicts the page's previous impression — webXray at 69.9% domain-level and 93.9% encounter-weighted as measured on 2026-09-05 (69.2% and 93.8% after the second pass over the residue), against “1 of 23” from the disagreement table. A page that only ever looks where it already knows the answer is where that impression came from.
- Publishing the contradiction prominently instead of quietly correcting the old wording. The content page now carries both numbers, in the same bullet, with the explanation that the first was a statement about where it had been checked. The alternative — deleting “badly stale wherever it has been checked” and moving on — would have hidden the most useful thing the sample found, which is how badly a selection-biased set misleads.
- Not letting the new figure rehabilitate webXray. The encounter-weighted rate — 93.9% on 2026-09-05, 93.8% after the second pass — is a flattering number and it would have been easy to lead with it. The content page instead puts three counter-facts in the same box: coverage is 2.1%, the file is frozen so the rate only degrades, and Disconnect is better on the measured axis. That last one is now metric-dependent: Disconnect beats webXray on current+granularity (p = 0.008) and no longer at this page's threshold on current only (p = 0.031). The recommendation did not change; only the evidence for it got more honest, and then more qualified.
- Quartiles instead of the deciles the follow-up item specified. See C3.2. A reasonable person could have kept deciles and reported the head stratum only.
- Retiring a verdict in the middle of the run.
self-namedwas in the brief 21 sub-agents worked from, so retiring it means re-scoring by hand after the fact — which is exactly the kind of post-hoc adjustment that should make a reader suspicious. It is published in full, with every re-score and its reason, precisely because it should. The alternative was worse: leaving a verdict that had fired on “TrustArc” and “Klaviyo” and that would have cut Disconnect's denominator by a quarter. - Accepting an OV/EV certificate subject as a primary source. It is an identity a CA validated against a company register, and the handshake is re-done by machine. 17 of 135 sourced rows rest on it. A stricter reading would have demoted all 17 to
unresolved; that would have cost about an eighth of the eligible rows for no gain in truth. - Correcting one sentence on Crawler, which is a different page. That page told a reader that webXray's licence “forbids redistribution, so there is no lawful route to the code now that upstream is gone”, and called
thezedwards/webXraythe most complete surviving copy. Both are false, and the page being published here says so — two live pages contradicting each other, with the wrong one making the more actionable claim. The scope rule for this sitting was one page and its provenance, and the broader rewrite that parent page needs is still the open TODO the first sitting recorded. But shipping a page while leaving a neighbour asserting its opposite is not a defensible reading of scope, so exactly that sentence was replaced and nothing else on that page was touched. Someone could reasonably have left it and filed an item. - Publishing the wrapped-list finding rather than just fixing it. It is embarrassing and it is the most useful thing on this page for anyone else writing a long DokuWiki page. See the entry below.
- Splitting the scripts and outputs onto an appendix page. 230 kB of code and raw output on top of a 140 kB page would make one 370 kB page. Raw size alone would not decide it —
literature:bibliographyis 388 kB and renders fine — but about 60 kB of the appendix is syntax-highlighted source, and a reader who came to check one denominator would have to scroll past all of it. The split is a readability call: no figure is derived only on the appendix, though several are only shown there — the 175-row adjudication table, the per-verdict tables, the concentration test and the raw Wayback output — and both this page and the content page send the reader there for them. Someone could reasonably have kept one page. - Rewriting two *What could not be established* bullets in place rather than deleting them. Both are now answered. Deleting them would have erased the fact that the page shipped with those gaps and named what would close them — which is the part that made this sitting possible.
- Not re-deriving the corpus figures. This sitting measured three live files and the Wayback index. Every corpus figure on the content page is the first sitting's, against the same 5,859-paper extraction, and was not re-run. Stated here so a reader does not read “2026-09-05” at the top of the page as a date on the corpus numbers.
- Using a different model family for the second rating, not a second run of the same model. Two Sonnet runs would share their failure modes and the agreement figure would flatter the instrument. The cost, stated in C3.8 rather than hidden: a capability difference and a reading difference are not separable afterwards, and the sourcing/reading decomposition suggests most of what was measured is the former. Someone could reasonably have run Sonnet twice and called the result a cleaner measurement of the brief.
- Publishing the uncorrected comparison as the headline. Rater 2's one disqualified citation (
kameleoon.io) is demoted in a sensitivity line rather than in the headline table, because rater 1's headline rows are pre-correction too and the symmetric comparison is the honest one. The corrected variant moves the pooled figure from +0.400 to +0.400 and is printed beside it. - Not re-estimating the accuracy rates from rater 2's verdicts. It would have been one line of code and it would have been meaningless: 35 rows is not the sample design, the stratum weights do not apply to it, and substituting 35 of 175 rows produces a hybrid that is neither rater's reading. What is published instead is the unweighted
currentshare per rater, labelled as a direction and not a rate. - Fixing
owner_verify_sources.pyrather than only recording its defect. The script is published, its hash is on the appendix manifest, and changing it means the appendix regenerates. The alternative — a note saying the check is blind to non-Latin quotes, beside a check that stays blind — would have left the next reader with an instrument the page knows is broken. The consequence for the first sitting's rows is stated in C3.8 and one row now needs a requote that this sitting did not apply. - No
~~DISCUSSION~~on this provenance page. Following the convention set by the earlier provenance pages: comments belong on the content page.
G. Sub-agent findings that did not survive
Three Sonnet sub-agents were run in parallel and every load-bearing claim was re-checked. Rejections are recorded because they are the only measure of whether a reviewer slot is worth having.
| Claim | Verdict |
|---|---|
| “IEEE-SP/2020 states a version: WebXRay commit 04c3c8e8, 2019-06-18” | accepted, and it corrected me. My draft said only 1 of 8 papers pinned a snapshot and that none named a commit. Verified in the source: the reproducibility table pins both webXray and Disconnect to commits. The report script's SNAPSHOT map and the page were fixed |
| “NDSS/2021 found webXray's list covers 1,096/1,146 (~95.6%) of same-party pairs” | rejected as a publishable figure. The sentence contradicts itself; see A.17. The unambiguous parts of the paragraph are used instead |
| “IMC/2020 used the owner list” (8 owner-list-only papers) | rejected. webXray appears in that paper only in its reference list; the method is a TLD + certificate-SAN + SOA heuristic. Verdict kept as citation, giving 7 |
| “CCS/2016 is a citation” | rejected as imprecise. The paper compares against webXray and criticises its browser. Verdict compared, which is what the page reports |
| “Tracker Radar: 31 papers” | rejected in favour of 32. The report's sweep normalises whitespace first and catches one match broken across a column boundary that a plain grep misses |
| “Disconnect-as-a-list: ~78 papers” | rejected in favour of 74. Different tightening regex; the page uses the one in the script, which is printed |
| “webXray is 'discontinued and superseded by webxray.ai'” | accepted after independent re-fetch of webxray.ai and timlibert.me. This is the finding that closes the parent page's open <wrap todo> |
| “Disconnect's list is GPL-3.0 with an attribution requirement” (my own briefing premise) | rejected — my premise was wrong. It was GPLv3 from 2015-10-13 and CC BY-NC-SA 4.0 since 2020-06-24, verified from the LICENSE commit history. The page carries the correction as a warning, since a paper quoting the GPL terms is quoting a dead licence |
“postrelease.com → Life360 and crwdcntrl.net → Publicis look like Disconnect errors” (my own hypothesis) | rejected — I was wrong on both. See C.1 |
| “Disconnect accepts community pull requests” (my own first draft) | rejected — wrong. Its README says “Pull requests are not reviewed and will be closed”. Caught by self-review, not by a reviewer. The corrected row also gave the page Disconnect's own documented method, which the first draft called undocumented |
| “webXray is the reason most pre-2022 papers could name a company” (my own first draft) | rejected as overstated. The corpus has 8 papers that used it against a 74-paper upper bound for Disconnect. Softened to “one of the three hand-curated lists the pre-2022 literature used”, which is what [4Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)] and [15Dambra, Savino; Sanchez-Rola, Iskander; Bilge, Leyla; Balzarotti, Davide (2022): "When Sally Met Trackers: Web Tracking From the Users' Perspective", in: Proceedings of the USENIX Security Symposium. (Link)] actually did |
| “the database … is wrong far more often than it is right” (my own first draft) | rejected as an unsupported generalisation. True of the selection-biased adjudicated sample, not of the file. Softened to “frozen at 2021 and badly stale wherever it has been checked”, with the selection bias stated in the same bullet |
| “over 1M sites” for Libert's 2018 crawl (my own first draft) | replaced with the paper's own figures — Alexa top one million, October 2017, 938,093 pages loaded, 248,029 policy links extracted — after reading the results section |
| “Kariryaa et al., NDSS 2025” (my own first draft) | a fabricated author name. The paper's first author is Xiaoyuan Wu. Caught by a self-review script that matches every Name et al., VENUE YEAR attribution on the page against the author field of the key cited next to it; 14 of 15 rows passed, this one did not. The citekey was renamed wu2025_appprivacyreport at the same time |
Reproducibility check run after publication
All three scripts were re-run and their output diffed byte-for-byte against the committed output: identical. owner_dbs.py was then re-run against a fresh cache, i.e. re-fetching all five inputs and the Public Suffix List live, and produced output identical to the published snapshot — so every figure in sections B and C of this page still reproduces from the live sources as of 2026-08-17, with the same five SHA-256 prefixes. That will stop being true: two of the three lists change weekly.
Publishing gotcha: the bibliography needs a cache purge
After appending to literature:bibliography and saving this page, every one of the 14 new {[key]} markers rendered as a broken citation and the reference list showed only the 6 keys that already existed. The bibliography source was well-formed; the bibtex plugin was serving a cached parse of the old bibliography.
Fix: curl -sL “https://measuretheweb.org/literature/bibliography?purge=true” and then the same on the citing page. After that all 22 references rendered. Do this after any bibliography append and re-read the rendered page, or you will publish a page full of broken citations and the raw source will look perfect.
H. Reviewers
Filled in after the review passes; see the page history for what changed.
Three focused reviewers (Sonnet) ran in parallel against the page text, the three scripts and their outputs, and the provenance notes. Every finding below was re-verified against the primary source before being accepted or rejected.
— Reviewer 2: citations and quotes ——————————————
Reported 4 blocking, 2 should-fix, 2 nits. Checked ~30 quotations, 22 citekeys and 16 BibTeX entries. Independently re-derived the numeric infrastructure from the cached JSON and matched every figure exactly.
ACCEPTED (blocking):
1. "Kariryaa et al., NDSS 2025" is a fabricated attribution; the first author is Xiaoyuan Wu. Already found and fixed by self-review before this report arrived, along with a new guard (scripts/check_attributions.mjs) and a key rename to wu2025_appprivacyreport. Independent confirmation of the same defect. 2. Yang & Yue try webXray SECOND, not third: the paper's order is CrunchBase, Tim Libert's library, TLS certificate, WHOIS. The page contradicted itself -- one sentence had the order right and the table row had it wrong. Fixed. 3. sanchezrola2021_journey is IEEE S&P **2021**, not 2022. Confirmed against Crossref for 10.1109/sp40001.2021.9796062: container "2021 IEEE Symposium on Security and Privacy (SP)", event 2021-05-24, issued 2021-05. The corpus files it under 2022 (a DBLP indexing quirk). BibTeX year and series fixed, citekey renamed from sanchezrola2022_journey, and the page now says 2021 with an inline note that the year table follows the corpus. 4. Ghostery trackerdb is **CC BY-NC-SA 4.0**, not MIT. Confirmed: its package.json says "CC-BY-NC-SA-4.0" and the GitHub API returns no SPDX id. Fixed, with a warning not to trust the API field. This strengthens the page's licence point -- all three live options are non-commercial.
ACCEPTED (should fix):
5. "all commits on 2018-04-05" for RDBinns/webXray_Domain_Owner_List. The reviewer found 12 commits spanning 2018-03-29 to 2018-04-05. GitHub's commits API was rate-limited when re-checking, so the claim was replaced with what the repo API itself returns and was fetched first-hand: created and last pushed on 2018-04-05. The unverifiable "all commits" claim is gone. 6. The PSL quotation was silently truncated mid-sentence. The full clause is "...without update mechanisms that are frequently checking for updates and incorporating them." Re-fetched and confirmed; the page now quotes it whole.
ALSO FIXED, found while acting on 3 and 8:
7. pages/bib_additions_webxray.bib carried bibgen's own diagnostic lines
("--- check these ---", "metadata source: openalex-doi", …) which are not
BibTeX. They would have been appended verbatim to literature:bibliography and
could have broken every page that renders it. Stripped; the file is now 16
entries and nothing else.
REJECTED:
8. Curly apostrophe in the steffens2021_blockparty title. bibgen takes titles from publisher metadata; normalising punctuation by hand is how titles drift from the record. Left as generated.
— Reviewer 3: external currency ——————————————-
Checked 22 claim groups against live primary sources. Independently confirmed both blocking findings reviewer 2 raised (trackerdb's licence, the RDBinns commit range), which had already been fixed, and confirmed 20 further claims exactly: the 404s, webxray.org's body text, the PyPI misses, the fork count and newest fork activity, the Docker Hub state, every licence including PolyForm Strict 1.0.0's own clauses, Disconnect's 2020-06-24 relicence, Tracker Radar's release tags, Disconnect's and Ghostery's last commits, the WhoTracks.me redirect chain, tracker-radar-collector still being Puppeteer-based, Crunchbase's 401, the PSL's warning and canonical URL, Mozilla's shavar pipeline being unreplaced, and five of the thirty ownership adjudications.
ACCEPTED (should fix):
1. lxml 4.6.2 and psycopg2-binary 2.8.6 have no PyPI wheel beyond cp39, and Python 3.9 has been end-of-life since October 2025 -- so even a reader who obtains the code cannot `pip install -r requirements.txt` on a current interpreter without system libxml2/libxslt and PostgreSQL headers. Added to the licence box, which is where a reader deciding whether to try is looking. websocket-client 0.57.0 and textstat 0.7.0 are fine.
ACCEPTED (currency confirmations worth recording):
2. PolyForm Strict 1.0.0 is still the only version: `strict/1.1.0` 404s and the project's Announcements page has a single entry, "Announcing Version 1.0.0 of the PolyForm Licenses" (2020-05-24). 3. CC BY-NC-SA 4.0 is still current; there is no 5.0. 4. No broken internal links anywhere. The only two unresolved targets are these two pages themselves, i.e. the red link being fulfilled.
REJECTED:
5. "PhantomJS's exact pushed_at could not be confirmed (GitHub rate limit)." Not a defect in the page: `api.github.com/repos/ariya/phantomjs` was fetched first-hand earlier in this run and returned `archived: true` with `pushed_at: 2022-11-26T19:43:12Z`. The reviewer's inability to re-fetch under a rate limit is not evidence against it.
— Reviewer 4 (generic, Fable): the three findings that mattered most ———-
Given no checklist and told to find what the focused three were not looking for. It found the most serious defect on the page, in the page's own method.
ACCEPTED (blocking – this was a real bug in owner_dbs.py):
1. **googleapis.com is a PRIVATE PSL rule, so the fold + exact-key lookup manufactured a coverage hole and depressed every weighted coverage figure.** Verified: the rule is at line 13935 of the snapshot, past BEGIN PRIVATE DOMAINS at 11276, and all three lists DO name googleapis.com (webXray "Google APIs", Tracker Radar "Google LLC", Disconnect "Google"). So fonts.googleapis.com -- which the page showcased three times as "the single most prevalent third-party name with no owner in any list" -- was an artefact of my own lookup rule, and it carries prevalence 0.369 on its own.
Fixed properly rather than patched: fetch_psl() now returns the ICANN and PRIVATE sections separately, a lookup() helper walks parent labels, and the script prints all FOUR combinations (ICANN/full fold x walk/exact lookup) so the reader can see that the choice moves the answer by more than the lists differ. Every figure in sections C and D of the page was re-derived: weighted coverage 54.5/79.3/75.8% -> 58.6/84.3/80.5%; domain share 1.4/12.2/5.0% -> 2.1/17.2/7.0%; top-100 70/95/91 -> 71/98/94; the no-owner-anywhere figure 94.6% of domains and 19.8% of weight -> 92.4% and 14.4%. The pairwise disagreement table moved by under a point.
The replacement example is better than the one it replaces: the most requested domain no list can name is now tiktokw.us at prevalence 0.041, which is a genuine, recent, mid-tail coverage hole. The googleapis rows are re-captioned as the trap they actually are, in the page's trap list and again in the lookup demo, which now shows the buggy version, says so, and gives the three-line fix.
ACCEPTED (blocking):
2. **The two rows that failed the sourcing bar were still counted in the adjudication tallies, and two of webXray's three "current" verdicts WERE those two rows.** Verified from the script: webXray's "current" rows were casalemedia.com, 360yield.com and fwmrm.net, and the latter two are exactly the UNRESOLVED pair. So the one figure most favourable to webXray rested on the evidence the script itself called insufficient -- while both pages claimed those rows were "recorded as unresolved rather than guessed". owner_adjudication.py now excludes them from every tally and prints what it excluded and why: 28 rows, webXray current on 1 of 23 (was 3 of 25), Tracker Radar 8 of 28, Disconnect 21 of 21.
Also accepted: the "Read as a share of each list's own entries ... 100%" sentence invited exactly the accuracy ranking the box below it disclaimed. Cut. The WRAP box now also says that Tracker Radar's zero in "No entry" and Disconnect's zero in "Stale" are artefacts of how the sample was drawn, and the script prints the same caveat above its own tallies.
ACCEPTED (blocking):
3. **The provenance page and the committed script output were not brought up to date after reviewer 2's fixes, while certifying that everything passed.** The ROLE hand map in report_webxray.mjs still said Yang & Yue tried webXray "third of four" where the content page had been corrected to "second", and the section D verification table still said Ghostery trackerdb was "(MIT)". Both fixed at the source and re-run, so the "unedited output" now agrees with the page it certifies. The reviewer is right that check_page_numbers.mjs cannot catch this class -- "third" and "MIT" are not numerals -- and that gap is now stated in *What could not be established*.
ACCEPTED (smaller):
4. "Treat webXray as **unobtainable**" was contradicted by the page's own practice: every figure on it was computed from the surviving copy. Reworded to "obtainable but not redistributable, and not installable", with the artefact- appendix consequence spelled out. 5. Ghostery trackerdb is named as a live option but is absent from the measured comparison. The omission and its reason (per-company .eno files plus a separate patterns layer, not a single domain->owner map) are now stated in the comparison section, labelled an omission rather than a judgement. 6. The reader could not assemble a pipeline. A new *Assembling the pipeline* section gives the six steps, including the one that was entirely missing: resolve the VISITED site to an owner too and drop same-owner requests. Without it, google.com embedding gstatic.com counts as third-party tracking. 7. The actionable answer sat ~120 lines after the lede. Jump links to *Choosing a resolution source now* and *Assembling the pipeline* added to the lede, which now says plainly that everything between is the evidence.
NOTED, not acted on in this sitting:
8. Discoverability: the wiki's only treatment of domain-to-company resolution lives under a dead tool's name. The reviewer is right that the earlier rejection of a separate topic page was weak. Mitigated by adding a link from [[Privacy:Requests]], which is where a reader asking "whose request is this?" actually lands. A ''design:ownership_resolution'' page that this one feeds is worth considering and is recorded as such rather than done. 9. The 16-row schema table "could halve". Kept: its fill-rate column is the evidence for "most of that ambition is unfilled", which is the section's claim, and a reader checking one field wants the row.
— Reviewer 1: figures versus script ——————————————
Landed after the first publication and after reviewer 4's fixes were applied, and it re-verified over 150 figures and table cells against freshly re-run script output, all byte-identical. It found two defects nobody else did, both in the script's own descriptions of its method rather than in any downstream figure.
ACCEPTED (blocking):
1. **"domain_summary.json is keyed by hostname for 16,396 of its 47,836 rows" was computed by a label-count heuristic (>=3 labels), which answers a different question and is wrong under either fold.** 14,033 of those 16,396 rows already ARE their own registrable domain, precisely because googleapis.com, s3.amazonaws.com and cloudfront.net are private-section PSL suffixes -- the same root cause as reviewer 4's finding, showing up in a second place. The honest figure is the number of rows the fold you used merges: 15,651 under the ICANN fold, 2,339 under ICANN+PRIVATE. The script now prints both and says explicitly that the label count answers neither question; the page carries the same correction. The reviewer also re-implemented the PSL algorithm from scratch and found zero mismatches against registrable(), so the fold function itself was never wrong -- only its description.
ACCEPTED (should fix):
2. **The suffix-fold residue conflated two operations.** "The fold ... fires on 45" was measured with the whole norm_name() pipeline, which strips punctuation before the suffix regex runs. Only 16 of the 45 lose a legal-form suffix; 29 change on punctuation alone (AT&T, JD.com, "Here, There & Everywhere"). The causal claim overstated the suffix fold's reach by nearly 3x. The script now prints both counts and names the punctuation-only cases; both pages say so.
ACCEPTED (nit, already fixed):
3. owner_adjudication.py's docstring said "29" where ROWS has 30. Never published; corrected.
Its process caveat is fair and worth recording: the repository was being edited throughout its run, so its findings are pinned to a snapshot. Both defects it found were still present in the last revision it sampled and are fixed now. Its brief was to re-run all three scripts and diff every figure against the real output. That check was also run directly, repeatedly, throughout the session: check_page_numbers.mjs passes whole-page (not windowed) on both pages against the concatenated output of all three scripts, check_tables.mjs passes on both, check_attributions.mjs passes 15 of 15, and the report's own figure verifier locates 26 of 26 literal per-paper figures in the cited sources. Anything reviewer 1 reports after publication goes into the page history, not into this log.
— Reviewer 5: industry and website claims (late) —————————–
Checked 13 items against primary sources; 10 passed. It found the single most important error on the page, in the claim the licence box is built on.
ACCEPTED (blocking – the page's central licence claim was wrong):
1. **webXray was relicensed TWICE after the snapshot this page measures, and ended up MIT.** Verified directly: commit 245ec5d7 (2021-06-14) "Update LICENSE.md - Now open-source" makes it GPLv3, and commit 73fe0fc9 (2023-02-01), authored by "Tim Libert", replaces that with an MIT LICENSE reading "Copyright (c) 2023 Tim Libert". peterjoles/webXray reports spdx_id: MIT and is 36 commits AHEAD of thezedwards/webXray, preserving upstream history -- including Libert's own post-2021 feature work -- past the deletion of the upstream repo. Several other forks report GPL-3.0.
So "webXray is not open source, and redistributing it is prohibited" was true of the one snapshot the page happened to build on and false of the project's final state, and the page's own advice -- "verify the licence of whatever file you actually download" -- is exactly what it failed to do. The licence box is rewritten as a three-row table of the three licences with what each permits, the availability table now names peterjoles/webXray as the most complete copy and says plainly that every figure here comes from the most restrictively licensed copy that exists, and the lede no longer says the tool "is gone". 2. **"three small commits in a personal working copy"** for the newest fork activity was wrong: those are 36 preserved upstream commits, not a fork owner's tinkering. Corrected. 3. **forks_count reports 19 while the forks endpoint returns 20 objects.** The page cited the endpoint for the number 19. Now states both.
ACCEPTED (should fix):
4. The Mozilla shavar footnote is attached to a row about entities.json, but disconnect-blacklist.json mirrors services.json; the file mirroring entities.json is disconnect-entitylist.json. The quote is verbatim, the attachment was to the sibling file. 5. "whotracks.me now redirects" is imprecise: it returns HTTP 200 with a canonical link to ghostery.com/whotracksme, serving byte-identical content rather than issuing a redirect.
Independently re-confirmed, already fixed: the Ghostery trackerdb licence (CC BY-NC-SA 4.0, and it is whotracks.me that is MIT – the page had them inverted), the RDBinns commit range (11 commits, 2018-03-29 to 2018-04-05), and the truncated PSL quotation. Also confirmed against SEC filings rather than press releases: the Xandr close (AT&T Form 10-Q, “On June 6, 2022”) and the Teads close (Outbrain 8-K, “On February 3, 2025”), plus that SEC now lists CIK 0001454938 as “Teads Holding Co.” with formerNames “Outbrain Inc.” – the acquirer took the target's name, as the page says.
The lesson worth keeping: five reviewers found five defects nobody else found, and the two most serious – a coverage figure depressed by my own lookup rule, and a licence claim resting on an unrepresentative snapshot – were both invisible to every automated guard in this repository.
Reviewers, second sitting (2026-09-05)
Four reviewers, all told explicitly that the briefing they were given might not be exhaustive and to verify from the files and the live web rather than from it. Three focused Sonnet passes ran in parallel against the three page drafts, the scripts and their outputs, and these notes; the generic Fable pass ran after their findings were applied. Findings are recorded with the decision, because a rejection is the only measure of whether a reviewer slot is worth having.
Reviewer: citations and quotes (Sonnet)
Checked ~25 quotes and all 22 citekeys, ran bib_dedup_scan.py and
check_attributions.mjs, re-fetched all six CDX queries and the id_
captures behind every archived quote, spot-checked all 15 named adjudication rows
against live sources, and diffed every published block against its file.
| Finding | Decision |
|---|---|
The 2010 bound is wrong and cites the wrong capture. web/20101008070114id_ is a DreamHost “Coming Soon” parking page, not the Web Xray mashup tool; that content first appears at 20101108204354 | ACCEPTED, blocking. Verified by fetching both captures. The bound is now 2010-11-08 to at least 2011-09-27, with the parking page named on both pages and the general lesson — a CDX row is a capture, not a page — written into the method notes as the second trap |
The “at least 2026-02-05” bound is true but uncited: queries 5 and 6 of wayback_webxray.sh stop at 2024-12-31, so no line of the published output contains a 2025 or 2026 capture | ACCEPTED, blocking. This is the sharpest finding of the run: the figure was right, independently re-verified by the reviewer, and not supported by the evidence the page publishes. Fixed at the source rather than in the prose: wayback_webxray.sh gained query 7 (uncollapsed, 2025-01-01 to 2026-12-31) and was re-run. Its output now carries 20260205035830 301 OM6ALWQTEGJ5NTY5RYSC2EXCKJOAHOVU |
gnu.org/copyleft/gpl.html → gpl-3.0.html is a 302, not a 301 | ACCEPTED. Verified: the bare-host hop is a 301, the path hop is a 302. Both pages now say 302 |
The content page's 2010 quote merges a <title> and a separate “beta 0.5” badge into one quoted string, while the same page discloses exactly that kind of join for the GPL quote — an inconsistent standard for one defect class | ACCEPTED. The footnote now names the two elements. The reviewer is right that the inconsistency was the real problem: the page had already decided this needed disclosing |
The published owner_random_sample.py output block is one trailing blank line short of the file whose SHA-256 the manifest quotes | ACCEPTED as a wording fix, not a behaviour change. The generator strips trailing newlines because DokuWiki discards them inside a block. The appendix now states that this is the only transformation, and that the hash and byte count are of the file on disk, so a reader reconstructing a file must add the newline back before hashing |
Confirmed clean by this reviewer, and recorded because these were the things most
likely to be wrong: all 22 citekeys resolve to distinct entries with no DOI or
title collisions and no citekey was added this sitting; 15/15 table
attributions pass; every other Wayback bound reconfirmed against fresh queries;
all 15 named adjudication rows verbatim against live sources, with
cnevids.com and cdnbasket.net correctly carrying no quote because they
are unresolved; no banned-tier source used as a citation anywhere in the 135
sourced rows; all 19 manifest hashes and byte counts correct. It also
independently re-derived 175 rows / 135 resolved / 40 unresolved / 17 tls://
sources / 29 changed verdicts, and confirmed that the second Tracker Radar
error verdict in the raw data (loopme.me) is correctly excluded from
Tracker Radar's rate because that domain was drawn for webXray's sample.
Reviewer: figures versus script (Sonnet)
Re-ran the whole pipeline, diffed every output against its committed artefact, and cross-checked every figure on all three pages. It also noted, correctly, that the pages were being edited while it read them, and said so rather than reporting a moving target as a fixed one.
| Finding | Decision |
|---|---|
“the same universe owner_dbs.py uses, so the two sittings' figures are on one denominator” is false. The first sitting's frame is 32,369 domains, the second's 32,337 | ACCEPTED, blocking, and the most valuable finding of the run. Reproduced by running owner_sample.py –cache out/webxray/cache (32,369; 669 / 5,581 / 2,268) against –cache out/owner_cache (32,337; 664 / 5,566 / 2,264). Both are real. The coverage shares round to the same 2.1% / 17.2% / 7.0%, which is precisely why two different frames could sit on one page unnoticed. C3.1 now carries three rows on it, and the content page carries a <WRAP tip> next to the figures themselves |
The run table's hash list is false. It claimed all six input files were unchanged between the sittings; tr_domain_summary went 7a303228d812a41c → 19f7a5a6a839ec87 (16,310,837 → 16,347,575 bytes) and the Public Suffix List went 155b43d46932e933 → aef8fb81d63232da | ACCEPTED, blocking. Verified against out/webxray/owner_dbs-output.txt. This was written from the second sitting's hashes without checking them against the first sitting's — an assumption stated as a measurement, which is the exact failure the page spends a section warning about. The row now lists the four that are unchanged and the two that moved, and says so |
The appendix's manifest hash for out/probe_all.txt matches no file. Published 7d2ca1e722836803; the file hashes 3def89e0d3fc8a14 | ACCEPTED, blocking. Root cause found on verification: probe_all.txt contains one lone CR, inside a copyright line scraped from a live site, and Python's text mode silently translates it to LF — so the builder hashed a string that was not the file. Fixed at the root: the builder now reads bytes, hashes the file on disk, discloses newline normalisation per file in a new manifest column, and self-checks every row and block before writing. The self-check was mutation-tested by reintroducing the original bug, which it rejects. All 23 manifest rows now verify independently |
owner_merge_rows.py does not fail on what its docstring promises: duplicates, out-of-vocabulary verdicts and resolved-rows-without-a-source were printed but returned exit 0 | ACCEPTED. Reproduced with a synthetic batch. Fixed, and mutation-tested with all three defects: each now exits 1. The reviewer is right that it changed no published number — the committed merge has zero problems — and right that the docstring invited a caller to trust the exit code |
The retained fetch cache does not reproduce the recorded statuses: legal.quantcast.com is 526,545 bytes in out/verify_cache against 3,054,755 in verify.json | ACCEPTED. Traced: the retained cache was the first, partial pass (verify_partial.json has that row as NOTFOUND at 526,545 bytes); the run that produced verify.json used a different cache. Nothing published depended on it, but a misleading artefact is still a defect. Rather than explain it, the whole check was redone from an empty cache — see C3.5. Result identical, zero rows changed, zero uncited |
| Live-network scripts are not byte-reproducible by design | Noted, not a defect. wayback_webxray.sh and the two verifiers hit third-party servers. The CDX queries were independently reconfirmed by two reviewers |
Confirmed clean by this reviewer: every figure from the new pipeline reproduces
byte-for-byte and matches the pages exactly (rates, intervals, n=49/47/42, the
concentration shares, the self-name census, 10 corrections, 29 changed verdicts,
the 132/3/40/0 citation tallies); every named example domain is a member of the
sample it is claimed to be in; all 13 published script blocks are byte-identical
to their files; the estimator's empty-stratum, zero-cell and bad-verdict paths
behave as documented; report_webxray.mjs and owner_dbs.py still reproduce
the first sitting's figures digit-for-digit; and check_owner_sample_figures.py
plus its mutation test both do what they claim.
Reviewer: external currency (Sonnet)
Fetched 189 URLs appearing literally on the three pages plus ~70 further live endpoints — roughly 260 requests. No blocking findings.
| Finding | Decision |
|---|---|
corporate.comcast.com/press/releases/freewheel-acquires-beeswax, cited by the bidr.io row of the 30-row adjudication, is a 404 | ACCEPTED. Confirmed. Replaced in owner_adjudication.py with the acquirer's own newsroom, freewheel.com/news/freewheel-to-acquire-ad-tech-leader-beeswax (HTTP 200, datePublished 2020-12-17, names Beeswax and “FreeWheel, A Comcast Company”), and the script re-run so section K's published output changes with it |
gnu.org/copyleft/gpl.html → gpl-3.0.html is a 302, not a 301 | ACCEPTED, independently and identically to the citations reviewer. Fixed on both pages |
| Four repository states in the “Choosing a resolution source now” table have moved since 2026-08-17 | ACCEPTED and updated, after re-fetching each myself: Tracker Radar 2026-08-12 → 2026-09-02 with newest release 2026.08.28 (checked against both /releases and /releases/latest, which agree — the newest tag is not a prerelease); Disconnect 2026-08-07 → 2026-08-28; Ghostery trackerdb 2026-08-06 → 2026-09-01; WhoTracks.me 2026-08-04 “July update” → 2026-09-02 “August update”. Not false before, but the page is being republished today and a stale date on a republished page is a new claim |
carlsaturnino/webxray-docker is recorded at “399 pulls”; it is 405 today | ACCEPTED as a wording problem, not a figure to chase. A monotonically increasing counter should not be published as a fixed number; the note now gives it as a counter with its check date. The conclusion (an image registered 2016-04-26, last_modified a 2024 metadata bump, not maintained) is unchanged |
| The PSL's ICANN/PRIVATE rule counts have drifted (6,941/3,290 → 6,949/3,372) | NOT a page defect, and it became evidence instead. The page already scopes those counts to “the snapshot used here”. But the reviewer's numbers are what let the figures reviewer's finding be confirmed: the PSL is one of the two files that moved between sittings |
mediaocean.com serves an incomplete TLS chain; nine investor-relations and newswire hosts 403 a bare curl but serve a browser | Recorded, no change. These are instrument artefacts, not link rot, and they are why the citation verification has a browser pass. Recorded so a future reviewer does not report them as dead links |
Reviewer: generic, no checklist (Fable)
Ran after the three focused passes and their fixes. It found the two most serious defects of the whole run, both of which the three focused passes had each passed over for a defensible reason: one is a claim no script computes, and the other is data inside a script, so “the page matches its script” was true and useless.
| Finding | Decision |
|---|---|
The “reversal” is confounded. The page attributed the whole gap between “webXray current on 1 of 24” and “69.9%” to selection. But seven domains were adjudicated in both sittings against a bit-identical file, and of the five settled in both, webXray's verdict differs in four, all towards current — because 2026-08-17 scored webXray's ownership-tree root and named the acquiring group as owner, where 2026-09-05 scored the entry and named the operating subsidiary | ACCEPTED, blocking, and it is the finding of the run. Reproduced: lijit.com is “Sovrn” at the entry and “Federated Media” at the root; simpli.fi is “simpli.fi” and “GTCR”; bidr.io and crwdcntrl.net differ on parent-versus-operator, not tree level; only smartadserver.com agrees. The lede, the section intro, the <WRAP important> box and the “60 percentage points” sentence were all rewritten. The page now says the two numbers are not comparable, gives both causes with the per-domain evidence, and publishes the uncomfortable number the reviewer surfaced: webXray verdicts agreed on 1 of the 5 domains adjudicated twice, which is the only inter-rater evidence this page has |
The 30-row table's “No entry” column is wrong. Disconnect has an entry for all eight rows scored absent, five under resources rather than properties; webXray has “PulsePoint” for contextweb.com | ACCEPTED, blocking. owner_adjudication_absent_audit.py was written to re-derive every absent from the files with the project's own parent-label lookup: 9 of 90 verdicts contradicted the files. All nine corrected, each with its reason in the script. Tallies move: webXray 1 of 24 (was 1 of 23), Disconnect 26 of 28 (was 21 of 21), Tracker Radar 8 of 28 unchanged. owner_adjudication.py now refuses to print until the audit passes — which is what the 2026-09-05 pipeline always did, and the reason this defect is a first-sitting one |
| The pages never say the adjudicators were language models, and “adjudicated by hand” implied otherwise | ACCEPTED, blocking. The content page now says “by Sonnet sub-agents working to a published brief, single-rated”. The provenance page said it already; the content page is where a reader deciding whether to cite 69.9% will look |
| “all three lists are flattered by the same amount” by the unresolved exclusion | ACCEPTED. The shares are 18% / 22% / 30%, so Disconnect — the list the page recommends — is the most flattered. Worst case 58.3 / 58.3 / 65.0%, which narrows the margin without reversing it. Both now on the page |
| The non-overlapping-intervals argument is method-fragile; a percentile bootstrap on 42 rows undercovers at the top | ACCEPTED. Replaced with Fisher's exact: p = 0.014 for Disconnect against webXray, p = 0.82 for webXray against Tracker Radar. Both recomputed here before publishing |
| “its own coverage” is a fifth of webXray's file: 664 of 3,215 domains, the rest never met by a 2026 crawl, and survivors are where “current” concentrates | ACCEPTED. Stated on the page. The reviewer's DNS probe (22 of 40 out-of-frame domains do not resolve, against 6 of 20 in-frame) is quoted as its finding, not re-run |
| The lede gave 93.9% bare while the section says to read it as an ordering | ACCEPTED. The lede now carries the gstatic.com concentration inline |
| Assorted stale cross-references: a pointer to the wrong section, “a single snapshot per list” when there are two, a table header dated 2026-08-17 above rows dated 2026-09-02 | ALL ACCEPTED and fixed |
| “cited, confirmed in the fetched bytes: 132” hides that 8 are fragment matches and 5 are 60-character window matches | ACCEPTED. The table row now breaks it down |
| “the sampled verdict was measuring the adjudicator's reading of a name” misplaces blame — 25% matches the census's own 27.5% label-only rate | ACCEPTED, and it is a fair correction. They applied the brief as written; the brief was wrong |
| “the part that cannot be automated” — it was done by 21 LLM agents | ACCEPTED. Reworded |
| “no figure lives only here, and no figure lives only there” is not true of the appendix | ACCEPTED. Now: no figure is derived only there, several are only shown there, and they are named |
| The census's “pairs” are distinct domains (3,224 pairs, 3,215 distinct, three with a trailing slash) | ACCEPTED. The script's column label and both pages corrected |
| “Prevalence” is never defined for a reader who does not know Tracker Radar | ACCEPTED |
| Cross-list rates are confounded by naming level — a brand name survives an acquisition, a legal-entity name does not | ACCEPTED as an opinion worth publishing. It explains why Tracker Radar's encounter-weighted rate is lowest, and the page now says the three rates are not strictly on one axis |
All 16 re-scores in owner_verdict_fixes.py moved towards the lists | ACCEPTED in substance, rejected as stated. Recounted: there are 28 verdict changes, not 16 — 14 to current, 2 to granularity, 1 to stale, 9 to unknown (which removes the row from every rate), plus both root-scoring fixes making the verdict less severe. The direction is as the reviewer says and the page now states it, with the real counts |
Reviewer: re-check of the applied fixes (Sonnet)
Ran against the fixes from the first three passes, and verified 8 of 10 outright.
| Finding | Decision |
|---|---|
Tracker Radar's “last commit 2026-09-02” is wrong: main is at 2026-08-28, the commit the 2026.08.28 tag points at. The 2026-09-02 commit is on an unmerged automation branch that the repository's pushed_at field picks up | ACCEPTED. I could not re-fetch the API to confirm directly (GitHub rate-limited this host by then), but the release tag I fetched earlier — 2026.08.28, published 2026-08-28T15:40:22Z — corroborates it, and a 2026-09-02 data commit on main would have produced a newer tag. The page now says “last commit on main 2026-08-28” and records the pushed_at trap, which is worth more than either date |
| The appendix builder's self-check never included the “block differs from the file” column, so a wrong normalisation flag could not be caught | ACCEPTED. Reproduced by hardcoding the flag; the check passed. Now included, and the mutation is rejected |
bidr.io carries “announced 2021-05-20” while the source it cites is dated 2020-12-17 | ACCEPTED. Corrected to 2020-12-17. Noted that the correct date was already written in the reviewer log while the code kept the wrong one |
| The two pages give different appendix sizes and neither addition is consistent | ACCEPTED. The appendix now states only its own measured size; the parent, which is built second and can measure both, does the arithmetic |
A stray owner_sample.json was left in the repository root by a reproduction run | ACCEPTED, removed |
check_owner_sample_figures.py can be defeated by swapping which list a row is attributed to, since it tests presence and not position | ACCEPTED as a disclosed limitation, not a bug. It is exactly what the script's docstring and this page already say it cannot see. Recorded rather than fixed, because the fix is a positional parser that would break on any table edit |
| The brief said 23 manifest rows; there are 24 | Correct — my count, not the page's |
Reviewers, third sitting (2026-09-11)
Three focused Sonnet passes ran in parallel against the page drafts, the scripts and their outputs; the generic pass was skipped for budget, which is recorded as a gap rather than dressed up as a full review layer. All three were told the briefing might not be exhaustive and to verify from the files and the live web rather than from it. The sitting's worst error was found here, not by any guard, and both reviewers who could have found it did.
| Finding | Decision |
|---|---|
The kameleoon.io demotion is wrong: the quoted <script src> line is verbatim in the bytes the checker itself cached. Raised independently by the external-currency reviewer (which fetched the live page twice with cache-busting headers) and by the evidence reviewer (which opened out/irr/verify_cache/eb6a82016d8e837fb6aee2c8.bin, the exact 94,416 bytes the checker read, and re-ran norm() on it to show the tag strip deletes the quote before the comparison) | ACCEPTED, blocking, and the most valuable finding of the run. The demotion is reversed, owner_verify_sources.py gains an OK-MARKUP status, and the failure is written up rather than smoothed away. The evidence reviewer's sharper point is also taken: the browser pass's NOTFOUND was structurally guaranteed for any markup quote and therefore carried no information at all, so citing it as corroboration was wrong twice over. The corrected-variant kappa figures now equal the uncorrected ones, and owner_irr_corrections.py asserts that rather than the page claiming it |
The compass-fit.jp explanation is misleading. The page implied the 致しました。 ending appears nowhere on the cited page; the reviewer found the full sentence in the body, differing from the row's quote by a single full-width comma after 台湾国内での | ACCEPTED. Verified directly: the body reads 台湾国内での、ネイティブ型広告配信サービス…. The bullet now says one character, names it, and carries a note that the first draft got the reason wrong. The finding is milder than first published — a transcription slip, not a paraphrase — and the filed follow-up item was corrected too |
The rater-2 note for cnn.com overstates the Warner Bros. Discovery position: WBD's own Q2-2026 10-Q calls the Discovery Global separation “previously proposed”, not pending, and conditions the PSKY merger on it not completing | ACCEPTED as a fact, no page change. The reviewer is right about the filing. That note is a field of one adjudication row and reaches no published block — cnn.com produced no inter-rater disagreement, so it is not printed in the estimator's section 7 either. Recorded here so the next reader of adj2_rows.json knows |
Every other external fact re-fetched and confirmed: the samsungads.ca certificate notAfter Mar 12 07:06:09 2026 GMT and a stable 59,465-byte body over four fetches; whois.nic.it answering on port 43 with the CED DIGITAL & SERVIZI Srl line at its exact five-space indent; microad.co.jp live; the PSKY merger still pending, paused by a federal judge and not closing before 2027-06-01. Nine rater-2 verdicts spot-checked against primary sources — cedscdn.it, tqlkg.com, awltovhc.com, sharethis.com, owneriq.net, mountain.com, pages02.net, eyeota.net, mediaset.es — all hold | No change needed. sa-as.com remains the weakest row, resting only on a MarkMonitor RDAP registrant field; the reviewer could not corroborate it further either, and rater 2 had flagged it itself |
Independence and stimulus-identity confirmed by hand: prompt00.txt is INSTRUCTIONS.md + batch00.txt + one output-path line, with no rater-1 verdict, no estimator, and no mention of the self-named retirement anywhere in it; per-domain blocks byte-identical between out/adj/batch*.txt and out/irr/adj2/batch*.txt for four spot-checked domains drawn from four different original batches | No change needed. This is the claim the whole section rests on and it is now checked rather than asserted |
All kappa figures, intervals, agreement percentages and the 63.0% sourcing/reading split reproduce exactly on an independent re-run; the self-named counts (6, 4, 3 shared) recomputed from the JSON; no citekeys added; the published <file> block matches the script byte-for-byte | No change needed |
Three findings from the figures-versus-script pass, none of them a wrong number on the page today. (a) owner_irr_kappa.py documented NAMED as “neither rater is forced to absent” while the code selected with or; the two are equivalent only because both raters' absent cells happen to coincide, which is empirical luck rather than an invariant. (b) The JSON artefact was not byte-stable: pe summed over a set(), whose iteration order is hash-seed randomised, so the 15th–16th decimal moved between runs. © The page pointed at appendix section AE for the subsample domain list; AE is a mutation harness and the list is in AF | ALL THREE ACCEPTED. (a) assert_absent_agrees() now raises before any population is built, and prints that absent agrees on all 105 cells; the docstring now says what the code does and why the equivalence is checked rather than assumed. (b) sorted(set(…)), verified byte-stable over three runs under a randomised hash seed. © Fixed to AF. The reviewer also noted that the pooled bootstrap rows resample cells rather than domain blocks — which the page already discloses, and the estimator's docstring now does too |
The guard was blind to every figure that lives in a script's text output rather than in the kappa JSON. The reviewer mutated the self-named counts, the sourcing/reading split table, the resolution counts, the unweighted current-share table and the whole citation tally, and check_irr_figures.py passed on every one | ACCEPTED, and the most useful finding after the kameleoon one. The guard now parses those figures out of owner_irr_kappa-output.txt and owner_verify_sources_irr-output.txt (a new –verify-output), and six mutations were added for them. Two of the six then survived: the self-named count and the split table were being checked with bare-digit needles like “17”, which match anywhere on a long page, and the self-named needle was additionally satisfied by the other page's untouched copy. Both fixed — the split table is matched as a whole row, and the self-named count is checked wherever its sentence shape occurs on every page. 18 mutations, all caught |
| Process finding, and the one this sitting should be judged on: the reviewer was mid-review when the kameleoon fix landed, so its report straddles two states of the artefact. It said so unprompted | ACCEPTED as a real methodological failure of this sitting. The rule is “do not edit a page under its reviewers” and it was broken — the fix was applied the moment two other reviewers reported, without waiting for the third. The reviewer re-checked everything against the post-fix state and its findings above are against that state, but the next sitting should hold fixes until every reviewer has reported, or re-run the ones whose ground moved |
What the guards did not catch, and why it matters. check_irr_figures.py
compares the page to the estimator and passed throughout; the estimator was never
wrong. The defect was one layer below, in a different script's verdict about
evidence, and no guard here looks at that. Its own mutation harness also reported
a survivor after the fix — a mutation aimed at a figure that had stopped being on
the page — which is the harness working as intended: a mutation that changes
nothing is reported as a survivor rather than silently passing.
I. Unedited output: scripts/report_webxray.mjs
Corpus: 5859 extracted papers from CCS, IMC, NDSS, PETS, USENIX Security, TheWebConf and IEEE S&P, 2010-2026.
Baseline populations: 1120 papers ran a crawl; 4439 classified or labelled something.
=== A. Population: the schema signal versus the full-text sweep ===
Schema signal: tools[].name matching /webx[\s-]?ray/i with usedOrMentioned in {used, produced} fires on 7 papers.
Not used/produced (compared/mentioned/unclear): 0 papers.
Full-text sweep of paper.cols.txt for the same pattern: 15 papers, 0.3% of the 5859-paper corpus and 1.3% of the 1120 that ran a crawl.
Schema papers not in the sweep: none (a sweep miss would mean the extractor read a rendering the sweep cannot see).
Sweep papers the schema misses: 8.
Spellings in tools[].name: "webxray" (3), "webXray" (2), "WebXRay" (1), "WebXray" (1).
Spelling-fold residue: 0 -- every spelling is a case variant of one token. There is no fork under another name in the corpus.
ROLE hand map: 15 entries; sweep papers with no verdict: 0; verdicts for papers no longer in the sweep: 0.
=== B. What role webXray plays, by hand verdict ===
Role Papers Share of 15
--------------- ------ -----------
owner-list-only 7 46.7%
citation 4 26.7%
compared 1 6.7%
subject 1 6.7%
miscitation 1 6.7%
instrument 1 6.7%
Papers that actually ran webXray or read its owner list: 8 of 15.
Of those 8, 7 used only the owner list and 1 ran the crawler.
The one paper that ran the crawler is WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w -- webXray's own author's paper.
Per paper:
Paper Role What was used
--------------------------------------------------------------------------------------------------------------------------- --------------- --------------------------------------------------------------------------------------------------------------------------------
CCS/2016/online-tracking-a-1-million-site-measurement-and-analysis compared crawler, as a baseline whose browser is criticised
WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w instrument crawler + owner list + policy extraction
WWW/2019/before-and-after-gdpr-the-changes-in-third-party-presence-at-public-and-private citation webxray.org cited as the crawler used by a related study; this paper uses its own
IEEE-SP/2020/do-cookie-banners-respect-my-choice-measuring-legal-compliance-of-banners-from-i owner-list-only owner list
IMC/2020/analyzing-third-party-service-dependencies-in-modern-web-services-have-we-learne citation owner list in the reference list only; the paper builds its own TLD + SAN + SOA heuristic
PETS/2020/a-comparative-measurement-study-of-web-tracking-on-mobile-and-desktop-environmen owner-list-only owner list (cited as “Tim Libert's library”), SECOND of four sources tried in order: CrunchBase, webXray, TLS certificate, WHOIS
NDSS/2021/whos-hosting-the-block-party-studying-third-party-blockage-of-csp-and-sri owner-list-only owner list, retrieved from the Internet Archive
IEEE-SP/2022/journey-to-the-center-of-the-cookie-ecosystem-unraveling-actors-roles-and-relati owner-list-only owner list, third by priority of three merged lists
PETS/2022/atom-ad-network-tomography owner-list-only owner list, alongside WHOIS and TLS certificates
PETS/2022/omnicrawl-comprehensive-measurement-of-web-tracking-with-real-desktop-and-mobile owner-list-only owner list
PETS/2022/who-knows-i-like-jelly-beans-an-investigation-into-search-privacy citation another paper's use of the crawler, described in related work
USENIX/2022/when-sally-met-trackers-web-tracking-from-the-users-perspective owner-list-only owner list, one of three merged lists
IMC/2023/a-first-look-at-the-privacy-harms-of-the-public-suffix-list subject webXray is one of the repositories measured as shipping a frozen Public Suffix List copy
PETS/2024/connecting-the-dots-tracing-data-endpoints-in-iot-devices miscitation describes Libert's webXray but cites Cinco Network's WebXRay (Held, IJNM 1998), an unrelated 1998 network-management product
NDSS/2025/transparency-or-information-overload-evaluating-users-comprehension-and-perceptions-of-the-ios-app-privacy-report citation owner list, described in related work
=== C. Which snapshot of the owner list did they use? ===
The owner list has no version field and no release tags: it is a single JSON file on a branch.
A paper can only pin it by naming a date, a commit or an archive snapshot.
IEEE-SP/2020/do-cookie-banners-respect-my-choice-measuring-legal-compliance-of-banners-from-i: COMMIT: "WebXRay commit 04c3c8e8 (2019-06-18)", in a reproducibility table that pins the Disconnect list the same way ("commit eb817fb1 (2019-12-10)")
IEEE-SP/2022/journey-to-the-center-of-the-cookie-ecosystem-unraveling-actors-roles-and-relati: nothing -- no date, commit or snapshot
NDSS/2021/whos-hosting-the-block-party-studying-third-party-blockage-of-csp-and-sri: partial: "as available in the Internet Archive", no date and no snapshot id
PETS/2020/a-comparative-measurement-study-of-web-tracking-on-mobile-and-desktop-environmen: nothing -- no date, commit or snapshot
PETS/2022/atom-ad-network-tomography: nothing -- no date, commit or snapshot
PETS/2022/omnicrawl-comprehensive-measurement-of-web-tracking-with-real-desktop-and-mobile: nothing -- no date, commit or snapshot
USENIX/2022/when-sally-met-trackers-web-tracking-from-the-users-perspective: nothing -- no date, commit or snapshot
WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w: nothing -- no date, commit or snapshot
(IMC/2020/analyzing-third-party-service-dependencies-in-modern-web-services-have-we-learne: reference dated "June 29, 2018" -- a citation, not a use, so it is not counted below)
Papers that used the list and say anything at all about which snapshot: 2 of 8 (25.0%).
Papers that name an exact commit: 1 of 8 (12.5%).
Matte et al. is the model to copy: it pins webXray AND Disconnect to commits in a reproducibility table.
=== D. Venue and year shape ===
Year Papers naming webXray Of which used the crawler or list
---- --------------------- ---------------------------------
2016 1 0
2018 1 1
2019 1 0
2020 3 2
2021 1 1
2022 5 4
2023 1 0
2024 1 0
2025 1 0
Venue Papers naming webXray
------- ---------------------
CCS 1
IEEE-SP 2
IMC 2
NDSS 2
PETS 5
USENIX 1
WWW 2
=== E. The ownership-resolution landscape: full-text sweeps ===
Each row is an UPPER BOUND: a full-text match, not a hand-verified use. The two
smallest are hand-verified above and below; the rest are not, and are labelled so.
Untightened /disconnect/i matches 700 papers -- almost none of which mean the list; "disconnect" is a common English word and the CSP literature uses it as a technical term. Hence the list-sense pattern, which gives 74.
Resource Papers matching Share of 5859 Hand-verified?
----------------------- --------------- ------------- -----------------
Public Suffix List 101 1.7% no -- upper bound
Disconnect (list sense) 74 1.3% no -- upper bound
Tracker Radar 32 0.5% yes, all 32
WhoTracks.me 23 0.4% no -- upper bound
Crunchbase 23 0.4% no -- upper bound
webXray 15 0.3% yes, all 15
Papers matching ANY of webXray / Tracker Radar / WhoTracks.me / Crunchbase / Disconnect-as-a-list:
136 papers (2.3% of the corpus; 12.1% of the 1120 that crawled).
Year Corpus papers Naming an ownership resource Share
----- ------------- ---------------------------- -----
2010 119 0 0.0%
2011 116 1 0.9%
2012 151 0 0.0%
2013 125 0 0.0%
2014 166 0 0.0%
2015 190 0 0.0%
2016 182 2 1.1%
2017 231 4 1.7%
2018 254 5 2.0%
2019 402 9 2.2%
2020 404 15 3.7%
2021 379 13 3.4%
2022 546 19 3.5%
2023 719 19 2.6%
2024 690 17 2.5%
2025* 770 21 2.7%
2026* 415 11 2.7%
* 2025-2026 are the provisional corpus edge: CCS 2026 and IMC 2026 have not been held,
and IEEE S&P / WWW 2026 abstracts are not in OpenAlex, so selection under-samples them.
=== F. Tracker Radar names three artefacts, not one ===
What "Tracker Radar" meant Papers Share of 32
--------------------------- ------ -----------
dataset-ownership 11 34.4%
dataset-tracker-or-category 9 28.1%
collector-crawler 9 28.1%
citation 2 6.3%
compared 1 3.1%
TR_ROLE hand map: 32 entries; unverdicted sweep hits: 0; stale verdicts: 0.
Papers using Tracker Radar for domain-to-company ownership: 11. Earliest: 2021.
Papers using the Collector, which is a Puppeteer crawler and carries no ownership data at all: 9.
The Collector has its own page (Programming:Crawler:Tracker Radar Collector); it is not an ownership resource.
Ownership use of Tracker Radar over time, against webXray:
Year Tracker Radar for ownership webXray crawler or list
----- --------------------------- -----------------------
2010 0 0
2011 0 0
2012 0 0
2013 0 0
2014 0 0
2015 0 0
2016 0 0
2017 0 0
2018 0 1
2019 0 0
2020 0 2
2021 1 1
2022 0 4
2023 2 0
2024 1 0
2025* 4 0
2026* 3 0
=== G. Per-paper figures quoted on the page, checked against paper.cols.txt ===
OK WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w | 91.27% of successfully loaded pages initiated a third-party request | /91\.27/
OK WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w | average of 10.89 unique third-party domains per page | /10\.89/
OK WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w | the crawl: Alexa top one million, October 2017 | /In October 2017, a computer based at a United States academic .{0,60}institution is used to scan one million popular websites/
OK WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w | 938,093 pages loaded and 248,029 policy links extracted | /938,093 are successfully loaded/
OK WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w | the policy-link denominator | /248,\s?029 pages are extracted/
OK IEEE-SP/2022/journey-to-the-center-of-the-cookie-ecosystem-unraveling-actors-roles-and-relati | 3,913 domains appear in at least one of the three lists | /3,913 domains in any of these lists/
OK IEEE-SP/2022/journey-to-the-center-of-the-cookie-ecosystem-unraveling-actors-roles-and-relati | one single conflict between the automated approach and the three lists | /just one single conflict/
OK IEEE-SP/2022/journey-to-the-center-of-the-cookie-ecosystem-unraveling-actors-roles-and-relati | priority order Disconnect, WhoTracks.me, webxray | /Disconnect first, WhoTracks\.me second, and webxray third/
OK PETS/2020/a-comparative-measurement-study-of-web-tracking-on-mobile-and-desktop-environmen | 411 of 762 mobile-specific trackers resolved to an organization | /411 trackers/
OK PETS/2020/a-comparative-measurement-study-of-web-tracking-on-mobile-and-desktop-environmen | 762 mobile-specific trackers is the denominator | /762 mobile-\s?specific trackers/
OK PETS/2020/a-comparative-measurement-study-of-web-tracking-on-mobile-and-desktop-environmen | WHOIS privacy proxies appear in the top-10 organization table | /Redacted For Privacy 34/
OK PETS/2020/a-comparative-measurement-study-of-web-tracking-on-mobile-and-desktop-environmen | the other four proxy strings in that table: Domains By Proxy 25, Whois Guard 14, Global Domain Privacy Services 8, Whois Privacy 7 (34+25+14+8+7 = 88 trackers, arithmetic done on the page) | /Domains By Proxy 25.{0,200}Whois Guard 14.{0,400}Global Domain Privacy Services 8 Whois Privacy 7/
OK PETS/2022/who-knows-i-like-jelly-beans-an-investigation-into-search-privacy | the related study webXray was used for: 22,484 adult websites | /22,484 adult websites/
OK CCS/2016/online-tracking-a-1-million-site-measurement-and-analysis | webXray described as PhantomJS-based, with the stripped-down-browser criticism | /WebXray is a PhantomJS based tool/
OK NDSS/2021/whos-hosting-the-block-party-studying-third-party-blockage-of-csp-and-sri | owner list taken from the Internet Archive | /as available in the Internet Archive/
OK NDSS/2021/whos-hosting-the-block-party-studying-third-party-blockage-of-csp-and-sri | 2,175 candidate site pairs, 1,146 confirmed same-entity, about eight person-hours | /2,175 site pairs for further checks, out of which 1,146 are operated by the same entity/
OK NDSS/2021/whos-hosting-the-block-party-studying-third-party-blockage-of-csp-and-sri | the owner list added 133 same-party relations their own method had not found | /133 additional same-party relations/
OK NDSS/2021/whos-hosting-the-block-party-studying-third-party-blockage-of-csp-and-sri | the misses that motivated the manual work: twitch.tv / twitchcdn.net | /frequently miss connections among two hostnames, e\.g\., twitch\.tv and twitchcdn\.net/
OK NDSS/2021/whos-hosting-the-block-party-studying-third-party-blockage-of-csp-and-sri | AMBIGUOUS -- NOT PUBLISHED AS A PERCENTAGE: "does account for 1,096 of our 1,146 found connections meaning that it alone does not suffice". The two halves of the sentence contradict each other and the published PDF reads the same way, so no coverage figure is taken from it. | /does account for 1,096 of our 1,146 found connections/
OK IEEE-SP/2020/do-cookie-banners-respect-my-choice-measuring-legal-compliance-of-banners-from-i | both lists pinned to commits in a reproducibility table | /Disconnect list commit eb817fb1 \(2019-12-10\) WebXRay commit 04c3c8e8 \(2019-06-18\)/
OK WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w | the author's own coverage limitation: the list holds major ad networks, not the long tail | /database of .{0,60}domain ownership primarily contains major ad networks rather .{0,80}than small clients/
OK IMC/2023/a-first-look-at-the-privacy-harms-of-the-public-suffix-list | timlib/webXray listed among repositories shipping a frozen Public Suffix List | /timlib\/webXray 27/
OK PETS/2024/connecting-the-dots-tracing-data-endpoints-in-iot-devices | the miscitation: Cinco Network's WebXRay, 1998 | /Cinco Network's WebXRay/
OK PETS/2023/privacy-rarely-considered-exploring-considerations-in-the-adoption-of-third-part | five categorisations compared; they "differ in granularity and focus" | /categorizations differ in granularity and focus/
OK USENIX/2026/the-state-of-passkeys-studying-the-adoption-and-security-of-passkeys-on-the-web | Tracker Radar Entity Map used to group same-organization domains | /Tracker Radar Entity Map/
OK IMC/2024/diffaudit-auditing-privacy-practices-of-online-services-for-children-and-adolesc | whois plus Tracker Radar to find the parent organization of an eSLD | /parent organization owner of this domain, using whois/
26 of 26 literal figures/quotes located in the cited paper's .cols rendering.
=== H. Evidence quotes behind the schema tuples ===
partial 69% WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w | webxray | methodology | "the webxray software platform is used to monitor third-party network traffic generated by loading a given web page and attribute such traffic to the entities which receive the data."
partial 70% WWW/2018/an-automated-approach-to-auditing-disclosure-of-third-party-data-collection-in-w | classification:webxray attribution database | methodology | "webxray searches for them in an internal database of domain ownership. The webxray attribution database is the product of years of detective work"
partial 79% NDSS/2021/whos-hosting-the-block-party-studying-third-party-blockage-of-csp-and-sri | webXray | methodology | "We augment our list with same-entity entries from the most up-to-date list used by webXray [14] as available in the Internet Archive."
partial 79% NDSS/2021/whos-hosting-the-block-party-studying-third-party-blockage-of-csp-and-sri | classification:webXray Domain Owner List | methodology | "We augment our list with same-entity entries from the most up-to-date list used by webXray [14] as available in the Internet Archive."
exact 100% IEEE-SP/2020/do-cookie-banners-respect-my-choice-measuring-legal-compliance-of-banners-from-i | WebXRay | results | "We matched tracking domains to company names using the Disconnect list [10]. We find whether they are part of the TCF by checking if any company name linked to a tracker domain in WebXRay's database [40] is present in the Global Vendor List"
exact 100% IEEE-SP/2020/do-cookie-banners-respect-my-choice-measuring-legal-compliance-of-banners-from-i | classification:IAB Global Vendor List | results | "We find whether they are part of the TCF by checking if any company name linked to a tracker domain in WebXRay's database [40] is present in the Global Vendor List (version 168)."
below-threshold 58% PETS/2022/atom-ad-network-tomography | WebXray | methodology | "We used external data sources including WHOIS records, TLS certificates, and WebXray [61] to identify the parent organizations of each identified tracker."
partial 92% PETS/2022/omnicrawl-comprehensive-measurement-of-web-tracking-with-real-desktop-and-mobile | webXray | methodology | "Finally, for every request, we used webXray [62] data to determine the provenance of the requests."
exact 100% PETS/2022/omnicrawl-comprehensive-measurement-of-web-tracking-with-real-desktop-and-mobile | classification:webXray | methodology | "we used webXray [62] data to determine the provenance of the requests"
partial 86% USENIX/2022/when-sally-met-trackers-web-tracking-from-the-users-perspective | webxray | methodology | "Once the tracking domains are identified, we map the domain names to organizations based on three manually-curated lists: Disconnect [13], WhoTracks.me [12] and webxray [31]."
partial 81% USENIX/2022/when-sally-met-trackers-web-tracking-from-the-users-perspective | classification:Disconnect, WhoTracks.me, and webxray | methodology | "We map the domain names to organizations based on three manually-curated lists: Disconnect [13], WhoTracks.me [12] and webxray [31]."
partial 67% IEEE-SP/2022/journey-to-the-center-of-the-cookie-ecosystem-unraveling-actors-roles-and-relati | webxray | methodology | "The first is based on three manually-curated lists (Disconnect [30], WhoTracks.me [31] and webxray [32])"
partial 67% IEEE-SP/2022/journey-to-the-center-of-the-cookie-ecosystem-unraveling-actors-roles-and-relati | classification:Disconnect | methodology | "The first is based on three manually-curated lists (Disconnect [30], WhoTracks.me [31] and webxray [32])"
partial 67% IEEE-SP/2022/journey-to-the-center-of-the-cookie-ecosystem-unraveling-actors-roles-and-relati | classification:WhoTracks.me | methodology | "The first is based on three manually-curated lists (Disconnect [30], WhoTracks.me [31] and webxray [32])"
partial 67% IEEE-SP/2022/journey-to-the-center-of-the-cookie-ecosystem-unraveling-actors-roles-and-relati | classification:webxray | methodology | "The first is based on three manually-curated lists (Disconnect [30], WhoTracks.me [31] and webxray [32])"
exact 100% IEEE-SP/2022/journey-to-the-center-of-the-cookie-ecosystem-unraveling-actors-roles-and-relati | classification:graph-based label propagation algorithm (custom) | methodology | "We validated our approach by using the manually compiled lists as a “ground truth” reference: we then looked for conflicts"
Summary: 16 quotes: 4 exact, 11 partial (>=60% of 4-word windows), 1 below threshold.
Below-threshold is not "unsupported": every one checked by hand was present, mangled by a column splice or a dropped citation marker.
=== Z. Every number on the page that does NOT come from this corpus ===
These are measured by scripts/owner_dbs.py against live snapshots of the three
ownership databases, or read out of a repository or licence. They are listed here
so scripts/check_page_numbers.mjs can tell them from corpus figures.
owner_dbs.py, snapshots of 2026-08-17: webXray 827 owners / 3,215 domains; Tracker Radar 19,148 / 38,368; Disconnect 1,887 / 7,850.
owner_dbs.py: webXray owners with a parent_id 319 of 827 (38.6%); tree depth up to 6; largest owner groupm with 620 domains.
owner_dbs.py: coverage of 45,525 registrable third-party domains -- webXray 657 (1.4%), Tracker Radar 5,539 (12.2%), Disconnect 2,277 (5.0%).
owner_dbs.py: prevalence-weighted coverage -- webXray 54.5%, Tracker Radar 79.3%, Disconnect 75.8%.
owner_dbs.py: top 100 by prevalence -- webXray 70%, Tracker Radar 95%, Disconnect 91%; top 1,000 -- 26 / 67 / 67%; top 10,000 -- 5 / 27 / 16%.
owner_dbs.py: pairwise "needs a human" disagreement -- webXray vs TR 197 of 601 (32.8%), webXray vs Disconnect 216 of 461 (46.9%), TR vs Disconnect 584 of 1,532 (38.1%).
owner_dbs.py: 16,396 of 47,836 domain_summary.json rows are hostnames not registrable domains; 19 rows are not hostnames at all, one of them the literal string "null".
webXray licence: PolyForm Strict License 1.0.0 (LICENSE.md in the surviving copy).
Tracker Radar and Disconnect licences: CC BY-NC-SA 4.0 (LICENSE and entities.json "license" field).
webXray requirements.txt pins lxml 4.6.2, psycopg2-binary 2.8.6, textstat 0.7.0, websocket-client 0.57.0.
webXray's bundled ccSLD-patches.txt header: "current as of 20160428".
github.com/timlib/webXray and .../webXray_Domain_Owner_List: HTTP 404 on 2026-08-17; the timlib user account itself returns 200.
J. scripts/owner_dbs.py, in full
Published here rather than on the content page: it is 396 lines and reading it is an audit task.
- owner_dbs.py
#!/usr/bin/env python3 """Compare the three domain-to-company ownership databases used in web measurement: webXray's ``domain_owners.json``, DuckDuckGo Tracker Radar's ``entity_map.json``, and Disconnect's ``entities.json``. The question the script answers is the one a measurement paper actually has to answer: **if I resolve a third-party domain to a company, does it matter which list I use?** It reports, for each list, how many owners and domains it carries, how much of the third-party surface it can name an owner for (unweighted and weighted by how often you actually meet the domain), and how often two lists that both know a domain disagree about who owns it. Inputs are fetched live and hashed, so a figure quoted from the output is pinned to a snapshot rather than to "the list". python3 owner_dbs.py --cache ./cache # fetch (or reuse cache), print report python3 owner_dbs.py --cache ./cache --disagreements 40 # + sample disagreements Prevalence weighting uses Tracker Radar's own ``domain_summary.json`` prevalence (share of crawled top sites that request the domain). That is a Tracker Radar measurement, so the weighting favours no list less than it favours Tracker Radar itself -- the coverage numbers for webXray and Disconnect are therefore measured on Tracker Radar's view of the third-party surface, and the script says so rather than pretending to a neutral denominator. """ import argparse import hashlib import json import os import re import sys import urllib.request SOURCES = { # webXray's owner list. The upstream repository github.com/timlib/webXray # is gone (HTTP 404); this is the newest surviving copy of webXray 3.x. "webxray": "https://raw.githubusercontent.com/thezedwards/webXray/master/webxray/resources/domain_owners/domain_owners.json", "tr_entity_map": "https://raw.githubusercontent.com/duckduckgo/tracker-radar/main/build-data/generated/entity_map.json", "tr_domain_map": "https://raw.githubusercontent.com/duckduckgo/tracker-radar/main/build-data/generated/domain_map.json", "tr_domain_summary": "https://raw.githubusercontent.com/duckduckgo/tracker-radar/main/build-data/generated/domain_summary.json", "disconnect_entities": "https://raw.githubusercontent.com/disconnectme/disconnect-tracking-protection/master/entities.json", } PSL_URL = "https://publicsuffix.org/list/public_suffix_list.dat" # Legal-form suffixes only. Deliberately not "group", "media", "technologies" # or "digital": those are part of a company's name often enough that stripping # them would manufacture agreement. SUFFIXES = [ "incorporated", "inc", "llc", "l l c", "ltd", "limited", "plc", "llp", "lp", "corporation", "corp", "company", "co", "gmbh", "ag", "kg", "kgaa", "mbh", "sa", "s a", "sas", "sarl", "srl", "spa", "bv", "b v", "nv", "n v", "ab", "as", "oy", "oyj", "aps", "sro", "s r o", "pty", "pte", "pvt", "kk", "kabushiki kaisha", "co ltd", "holdings", "holding", "sl", ] SUFFIX_RE = re.compile(r"\b(" + "|".join(sorted(SUFFIXES, key=len, reverse=True)) + r")\b") def fetch(name, cache): path = os.path.join(cache, name + ".json") if not os.path.exists(path): os.makedirs(cache, exist_ok=True) req = urllib.request.Request(SOURCES[name], headers={"User-Agent": "owner-dbs/1.0"}) with urllib.request.urlopen(req, timeout=120) as r, open(path, "wb") as f: f.write(r.read()) raw = open(path, "rb").read() return json.loads(raw), hashlib.sha256(raw).hexdigest()[:16], len(raw) def fetch_psl(cache): path = os.path.join(cache, "public_suffix_list.dat") if not os.path.exists(path): os.makedirs(cache, exist_ok=True) req = urllib.request.Request(PSL_URL, headers={"User-Agent": "owner-dbs/1.0"}) with urllib.request.urlopen(req, timeout=120) as r, open(path, "wb") as f: f.write(r.read()) raw = open(path, "rb").read() # The PSL has TWO sections and the split matters more than any other choice # in this script. ICANN rules are real TLDs and registry suffixes; PRIVATE # rules are suffixes companies asked to have treated as boundaries -- # `googleapis.com` is one of them. Folding with the private section leaves # `fonts.googleapis.com` standing as its own "registrable domain", and an # exact-key lookup then fails even though all three lists name # `googleapis.com` -> Google. So both sections are loaded separately. icann, private, exceptions = set(), set(), set() section = "icann" for line in raw.decode("utf8").splitlines(): stripped = line.strip() if "BEGIN PRIVATE DOMAINS" in stripped: section = "private" continue if "BEGIN ICANN DOMAINS" in stripped: section = "icann" continue if not stripped or stripped.startswith("//"): continue if stripped.startswith("!"): exceptions.add(stripped[1:]) else: (icann if section == "icann" else private).add(stripped) return icann, private, exceptions, hashlib.sha256(raw).hexdigest()[:16] def registrable(host, rules, exceptions): """eTLD+1 of a hostname per the PSL algorithm, or None if there is none. Needed because Tracker Radar's domain_summary.json is keyed by *hostname* for 16,396 of its 47,836 entries (fonts.googleapis.com and ajax.googleapis.com are separate rows), while all three ownership lists key on the registrable domain. Comparing coverage without this step charges webXray and Disconnect for subdomains they were never meant to hold. """ labels = host.split(".") for i in range(len(labels)): candidate = ".".join(labels[i:]) if candidate in exceptions: return ".".join(labels[i + 1:]) or None best = 0 for i in range(len(labels)): candidate = ".".join(labels[i:]) wildcard = ".".join(["*"] + labels[i + 1:]) if candidate in rules or wildcard in rules: best = max(best, len(labels) - i) if best == 0: best = 1 # unknown TLD: treat as one label if best >= len(labels): return None # the host *is* a public suffix return ".".join(labels[-(best + 1):]) def lookup(host, mapping): """Owner for a hostname, trying the host then each parent label sequence. A list keyed on `googleapis.com` must answer for `fonts.googleapis.com`. Exact-key lookup is the single easiest way to manufacture a coverage hole, and it is what the first version of this script did. Stops at two labels so it can never walk up to a bare TLD. """ labels = host.split(".") for i in range(len(labels) - 1): candidate = ".".join(labels[i:]) if candidate in mapping: return mapping[candidate], candidate return None, None def norm_name(s): s = str(s).lower() s = re.sub(r"[’']", "", s) s = re.sub(r"[^a-z0-9]+", " ", s).strip() prev = None while prev != s: # "Foo Co., Ltd." needs two passes prev = s s = SUFFIX_RE.sub(" ", s).strip() s = re.sub(r"\s+", " ", s) return s def main(): ap = argparse.ArgumentParser() ap.add_argument("--cache", default="./cache") ap.add_argument("--disagreements", type=int, default=0, help="print this many domain-level disagreements per pair") args = ap.parse_args() print("=" * 78) print("Domain-to-company ownership databases: shape, coverage, agreement") print("=" * 78) print() data = {} print("Inputs (sha256 prefix of the exact bytes these figures were computed from):") for name in SOURCES: obj, digest, nbytes = fetch(name, args.cache) data[name] = obj print(f" {name:20s} {digest} {nbytes:>10,} bytes {SOURCES[name]}") print() # ---- webXray ----------------------------------------------------------- wx = data["webxray"] wx_owner = {} # domain -> owner name wx_domains_per_owner = {} wx_by_id = {e["id"]: e for e in wx} for e in wx: wx_domains_per_owner[e["id"]] = len(e["domains"]) for d in e["domains"]: wx_owner[d.lower()] = e["name"] wx_children = sum(1 for e in wx if e["parent_id"] is not None) def depth(e, seen=()): if e["parent_id"] is None: return 1 if e["id"] in seen: return 99 # cycle guard, printed if it fires return 1 + depth(wx_by_id[e["parent_id"]], seen + (e["id"],)) wx_depths = {} for e in wx: wx_depths[depth(e)] = wx_depths.get(depth(e), 0) + 1 # ---- Tracker Radar ----------------------------------------------------- tr_dm = data["tr_domain_map"] tr_owner = {d.lower(): v["entityName"] for d, v in tr_dm.items()} tr_em = data["tr_entity_map"] # ---- Disconnect -------------------------------------------------------- dc = data["disconnect_entities"]["entities"] dc_owner_props, dc_owner_res = {}, {} for name, v in dc.items(): for d in v.get("properties", []): dc_owner_props[d.lower()] = name for d in v.get("resources", []): dc_owner_res[d.lower()] = name dc_owner = dict(dc_owner_res) dc_owner.update(dc_owner_props) # properties are the ownership claim print("--- A. Shape ---------------------------------------------------------") print() rows = [ ("webXray domain_owners.json", len(wx), len(wx_owner), "yes (parent_id)", "purpose, country, trade bodies, per-language policy URLs"), ("Tracker Radar entity_map.json", len(tr_em), len(tr_owner), "no (flat)", "displayName, aliases; prevalence in a sibling file"), ("Disconnect entities.json", len(dc), len(dc_owner), "no (flat)", "properties vs resources split; category in services.json"), ] w = max(len(r[0]) for r in rows) print(f"{'List':{w}} {'Owners':>7} {'Domains':>8} {'Hierarchy':16} What else per owner") for r in rows: print(f"{r[0]:{w}} {r[1]:>7,} {r[2]:>8,} {r[3]:16} {r[4]}") print() print(f"webXray owners with a parent_id: {wx_children} of {len(wx)} " f"({100*wx_children/len(wx):.1f}%). Tree depths: " + ", ".join(f"depth {k}: {v}" for k, v in sorted(wx_depths.items()))) print(f"webXray domains per owner: median " f"{sorted(wx_domains_per_owner.values())[len(wx)//2]}, " f"max {max(wx_domains_per_owner.values())} " f"({max(wx_domains_per_owner, key=wx_domains_per_owner.get)})") dc_prop_only = sum(1 for d in dc_owner_props if d not in dc_owner_res) dc_res_only = sum(1 for d in dc_owner_res if d not in dc_owner_props) print(f"Disconnect: {len(dc_owner_props):,} distinct 'properties' domains, " f"{len(dc_owner_res):,} 'resources' domains; " f"{dc_prop_only:,} properties-only, {dc_res_only:,} resources-only.") multi = sum(1 for name, v in dc.items() if len(set(v.get("properties", []))) > 1) print(f"Disconnect entities naming more than one property domain: {multi:,} of " f"{len(dc):,} ({100*multi/len(dc):.1f}%) -- the rest are one-domain entities, " f"i.e. an ownership claim that carries no grouping information.") print() print("--- B. webXray's purpose and jurisdiction vocabulary -----------------") print() for field in ("uses", "platforms", "trade_groups"): tally = {} owners_with = 0 for e in wx: if e[field]: owners_with += 1 for v in e[field]: tally[v] = tally.get(v, 0) + 1 print(f"{field}: {owners_with} of {len(wx)} owners ({100*owners_with/len(wx):.1f}%) carry at " f"least one value; {len(tally)} distinct values: " + ", ".join(f"{k} ({v})" for k, v in sorted(tally.items(), key=lambda kv: -kv[1]))) countries = {} for e in wx: if e["country"]: countries[e["country"]] = countries.get(e["country"], 0) + 1 with_country = sum(1 for e in wx if e["country"]) top = sorted(countries.items(), key=lambda kv: -kv[1])[:10] print(f"country: {with_country} of {len(wx)} owners ({100*with_country/len(wx):.1f}%) carry one; " f"{len(countries)} distinct codes; top 10: " + ", ".join(f"{k} ({v})" for k, v in top)) for field in ("aliases", "notes", "site_privacy_policy_urls", "gdpr_statement_urls", "ccpa_urls", "opt_out_urls", "health_segment_urls", "crunchbase_id"): n = sum(1 for e in wx if e[field]) print(f"owners with a non-empty {field}: {n} of {len(wx)} ({100*n/len(wx):.1f}%)") langs = {} for e in wx: for lang, _ in e["site_privacy_policy_urls"]: langs[lang] = langs.get(lang, 0) + 1 print(f"distinct languages in site_privacy_policy_urls: {len(langs)}; top: " + ", ".join(f"{k} ({v})" for k, v in sorted(langs.items(), key=lambda kv: -kv[1])[:8])) print() # ---- webXray, resolved up the ownership tree --------------------------- # webXray records DoubleClick as its own owner whose parent_id is google. # Tracker Radar records doubleclick.net as "Google LLC" directly. Comparing # the two without walking webXray's parent chain measures a difference in # granularity and calls it a difference in fact. def root_of(e): seen = set() while e["parent_id"] is not None and e["id"] not in seen: seen.add(e["id"]) e = wx_by_id[e["parent_id"]] return e wx_root_owner = {} for e in wx: for d in e["domains"]: wx_root_owner[d.lower()] = root_of(e)["name"] reparented = sum(1 for d in wx_owner if wx_owner[d] != wx_root_owner[d]) print(f"Domains where webXray's immediate owner differs from the root of its " f"ownership tree: {reparented:,} of {len(wx_owner):,} " f"({100*reparented/len(wx_owner):.1f}%).") print() print("--- C. Coverage of the third-party surface ---------------------------") print() summary = data["tr_domain_summary"] icann, private, exceptions, psl_hash = fetch_psl(args.cache) print(f"Public Suffix List: {psl_hash} {len(icann):,} ICANN rules, " f"{len(private):,} PRIVATE rules, {len(exceptions):,} exceptions {PSL_URL}") print("The ICANN/PRIVATE split is the most consequential choice in this script.") print("`googleapis.com` is a PRIVATE rule, so folding with the private section") print("leaves fonts.googleapis.com standing as its own 'registrable domain' --") print("and an exact-key lookup then finds no owner for it, even though all three") print("lists name googleapis.com -> Google. Both rules are therefore reported.") print() raw_keys = list(summary) bad = [k for k in raw_keys if "." not in k] def build(rules_set): prev, folded = {}, 0 for host, v in summary.items(): host = host.lower() if "." not in host: continue # see the 'null' note below reg = registrable(host, rules_set, exceptions) if reg is None: continue if reg != host: folded += 1 # A registrable domain's weight is the *largest* prevalence among its # hostnames, not the sum: one site can request fonts.googleapis.com # and ajax.googleapis.com, so summing would double-count sites. This # makes every weighted coverage figure a conservative lower bound. prev[reg] = max(prev.get(reg, 0.0), v["prevalence"]) return prev, folded prev_icann, folded_icann = build(icann) prev_full, folded_full = build(icann | private) # How many rows are "keyed by hostname rather than registrable domain" is not # a property of the data alone -- it depends on which PSL section you fold # with, because googleapis.com and s3.amazonaws.com are private-section # suffixes. A label-count heuristic (>=3 labels) says 16,396 and is wrong # under either fold; the honest figure is the number of rows the fold you # actually used merges. Both are printed. Caught by review 2026-08-17. print(f"domain_summary.json rows: {len(raw_keys):,}, of which " f"{sum(1 for k in raw_keys if k.count('.') >= 2):,} have three or more " f"labels -- a count that does NOT answer 'how many are keyed by hostname " f"rather than registrable domain', because that depends on the fold:") print(f" ICANN-section fold: merges {folded_icann:,} rows -> " f"{len(prev_icann):,} registrable domains") print(f" ICANN+PRIVATE fold: merges {folded_full:,} rows -> " f"{len(prev_full):,} registrable domains") print(f"Rows dropped as not hostnames at all: {len(bad)} {bad!r} " "-- a literal 'null' key with a prevalence attached is a defect in the " "published data, not a domain.") print() print("Denominator: those registrable third-party domains, i.e. domains Tracker") print("Radar's own crawl of regional top-site lists actually saw. This is Tracker") print("Radar's view of the surface, not a neutral one.") print() LISTS = (("webXray", wx_owner), ("Tracker Radar", tr_owner), ("Disconnect (properties+res)", dc_owner)) def coverage(prev, walk): rows = [] total = sum(prev.values()) for label, m in LISTS: if walk: hit = [d for d in prev if lookup(d, m)[0] is not None] else: hit = [d for d in prev if d in m] rows.append((label, len(hit), 100 * len(hit) / len(prev), 100 * sum(prev[d] for d in hit) / total)) return rows VARIANTS = [ ("ICANN fold + parent-label lookup <- USE THIS", prev_icann, True), ("ICANN fold + exact-key lookup", prev_icann, False), ("ICANN+PRIVATE fold + parent-label lookup", prev_full, True), ("ICANN+PRIVATE fold + exact-key lookup <- THE TRAP", prev_full, False), ] for name, prev, walk in VARIANTS: print(f"{name} (universe {len(prev):,})") print(f" {'List':30} {'Named':>8} {'Share':>7} {'Prevalence-weighted':>20}") for label, n, share, wshare in coverage(prev, walk): print(f" {label:30} {n:>8,} {share:>6.1f}% {wshare:>19.1f}%") print() fg = summary.get("fonts.googleapis.com", {}).get("prevalence") print(f"The single row that drives most of that gap: fonts.googleapis.com, prevalence " f"{fg:.3f}. googleapis.com is a PRIVATE PSL rule, and all three lists name it: " f"webXray={lookup('googleapis.com', wx_owner)[0]!r}, " f"TR={lookup('googleapis.com', tr_owner)[0]!r}, " f"Disconnect={lookup('googleapis.com', dc_owner)[0]!r}.") print("The bottom variant is what the first version of this script did, and it") print("understates every list. The gap between the top and bottom rows is a") print("measurement artefact, not a property of any list -- report which rule you") print("used, because it moves the answer by more than the lists differ.") print() # Everything below uses the recommended rule. prev = prev_icann total_prev = sum(prev.values()) universe = sorted(prev) owner_of = {label: {d: lookup(d, m)[0] for d in universe} for label, m in (("webXray", wx_owner), ("webXray-root", wx_root_owner), ("Tracker Radar", tr_owner), ("Disconnect", dc_owner))} for n in (100, 1000, 10000): topn = sorted(universe, key=lambda d: -prev[d])[:n] line = f"top {n:>5} by prevalence: " line += " ".join( f"{label} {100*sum(1 for d in topn if owner_of[label][d] is not None)/len(topn):.0f}%" for label in ("webXray", "Tracker Radar", "Disconnect")) print(line) print() named_by_none = [d for d in universe if owner_of["webXray"][d] is None and owner_of["Disconnect"][d] is None] print(f"Domains in the universe that neither webXray nor Disconnect names an owner " f"for: {len(named_by_none):,} ({100*len(named_by_none)/len(universe):.1f}%), " f"{100*sum(prev[d] for d in named_by_none)/total_prev:.1f}% of prevalence weight.") print("Top 15 of those by prevalence (Tracker Radar's owner in brackets):") for d in sorted(named_by_none, key=lambda x: -prev[x])[:15]: print(f" {d:34} prev={prev[d]:.4f} [{owner_of['Tracker Radar'][d] or 'TR: none'}]") top_unowned = max(named_by_none, key=lambda x: prev[x]) print(f"Most prevalent domain no list can name an owner for: {top_unowned}, " f"prevalence {prev[top_unowned]:.3f} (rounded to 3 dp for quoting).") dropped = {k: summary[k]["prevalence"] for k in bad} print("Dropped rows and their prevalence, rounded to 3 dp for quoting: " + ", ".join(f"{k}={v:.3f}" for k, v in sorted(dropped.items(), key=lambda kv: -kv[1])[:3])) print() print("--- D. Do two lists that both know a domain agree on the owner? ------") print() pairs = [("webXray", wx_owner, "Tracker Radar", tr_owner), ("webXray-root", wx_root_owner, "Tracker Radar", tr_owner), ("webXray", wx_owner, "Disconnect", dc_owner), ("webXray-root", wx_root_owner, "Disconnect", dc_owner), ("Tracker Radar", tr_owner, "Disconnect", dc_owner)] disagreement_samples = {} for a_label, _a_raw, b_label, _b_raw in pairs: # Parent-label lookup here too, for the same reason as section C. a, b = owner_of[a_label], owner_of[b_label] both = [d for d in universe if a[d] is not None and b[d] is not None] agree_raw = [d for d in both if a[d] == b[d]] agree_norm = [d for d in both if norm_name(a[d]) == norm_name(b[d])] sub = [d for d in both if d not in agree_norm and (norm_name(a[d]) in norm_name(b[d]) or norm_name(b[d]) in norm_name(a[d]))] dis = [d for d in both if d not in agree_norm and d not in sub] disagreement_samples[(a_label, b_label)] = sorted(dis, key=lambda x: -prev[x]) print(f"{a_label} vs {b_label}: {len(both):,} domains named by both") print(f" identical owner string {len(agree_raw):>7,} " f"({100*len(agree_raw)/len(both):.1f}%)") print(f" agree after legal-suffix fold {len(agree_norm):>7,} " f"({100*len(agree_norm)/len(both):.1f}%)") print(f" one name contains the other {len(sub):>7,} " f"({100*len(sub)/len(both):.1f}%) [e.g. Amazon / Amazon Technologies]") print(f" neither {len(dis):>7,} " f"({100*len(dis)/len(both):.1f}%) <- needs a human") wdis = sum(prev[d] for d in dis) / sum(prev[d] for d in both) print(f" ...weighted by prevalence, the 'needs a human' share is {100*wdis:.1f}%") print() if args.disagreements: print("--- E. Sampled disagreements (highest prevalence first) --------------") print() for (a_label, b_label), ds in disagreement_samples.items(): a, b = owner_of[a_label], owner_of[b_label] print(f"{a_label} vs {b_label}:") for d in ds[:args.disagreements]: print(f" {d:32} prev={prev[d]:.4f} {a_label}={a[d]!r} {b_label}={b[d]!r}") print() print("--- F. Named-entity residue of the suffix fold -----------------------") print() unchanged = [e["name"] for e in wx if norm_name(e["name"]) == e["name"].lower()] # norm_name() does two things -- strips punctuation, then removes legal forms -- # so "names the fold changes" is not "names a legal suffix was removed from". # AT&T, JD.com and "Here, There & Everywhere" change on punctuation alone. # Reported separately after a review pointed out the conflation, 2026-08-17. changed = [e["name"] for e in wx if norm_name(e["name"]) != e["name"].lower()] suffix_hit = [n for n in changed if SUFFIX_RE.search( re.sub(r"[^a-z0-9]+", " ", re.sub(r"[’']", "", n.lower())).strip())] punct_only = [n for n in changed if n not in suffix_hit] print(f"webXray owner names the fold leaves untouched: {len(unchanged)} of {len(wx)} " f"({100*len(unchanged)/len(wx):.1f}%).") print(f"Of the {len(changed)} it does change, only {len(suffix_hit)} lose a legal-form " f"suffix; the other {len(punct_only)} change on punctuation alone: " + ", ".join(sorted(punct_only)[:8]) + ", ...") print("The fold merges no synonyms, so every figure in section D is a lower bound on " "real agreement and an upper bound on real disagreement.") return 0 if __name__ == "__main__": sys.exit(main())
J2. Its unedited output
Run as python3 scripts/owner_dbs.py –cache out/webxray/cache –disagreements 25.
============================================================================== Domain-to-company ownership databases: shape, coverage, agreement ============================================================================== Inputs (sha256 prefix of the exact bytes these figures were computed from): webxray e53760188e6dc9aa 1,023,079 bytes https://raw.githubusercontent.com/thezedwards/webXray/master/webxray/resources/domain_owners/domain_owners.json tr_entity_map c4c3f97dbea6cb1e 4,741,528 bytes https://raw.githubusercontent.com/duckduckgo/tracker-radar/main/build-data/generated/entity_map.json tr_domain_map a11bc2580f664544 10,442,377 bytes https://raw.githubusercontent.com/duckduckgo/tracker-radar/main/build-data/generated/domain_map.json tr_domain_summary 7a303228d812a41c 16,310,837 bytes https://raw.githubusercontent.com/duckduckgo/tracker-radar/main/build-data/generated/domain_summary.json disconnect_entities 93e4f54036de1b39 412,191 bytes https://raw.githubusercontent.com/disconnectme/disconnect-tracking-protection/master/entities.json --- A. Shape --------------------------------------------------------- List Owners Domains Hierarchy What else per owner webXray domain_owners.json 827 3,215 yes (parent_id) purpose, country, trade bodies, per-language policy URLs Tracker Radar entity_map.json 19,148 38,368 no (flat) displayName, aliases; prevalence in a sibling file Disconnect entities.json 1,887 7,850 no (flat) properties vs resources split; category in services.json webXray owners with a parent_id: 319 of 827 (38.6%). Tree depths: depth 1: 508, depth 2: 211, depth 3: 80, depth 4: 24, depth 5: 2, depth 6: 2 webXray domains per owner: median 1, max 620 (groupm) Disconnect: 5,843 distinct 'properties' domains, 4,148 'resources' domains; 3,702 properties-only, 2,007 resources-only. Disconnect entities naming more than one property domain: 551 of 1,887 (29.2%) -- the rest are one-domain entities, i.e. an ownership claim that carries no grouping information. --- B. webXray's purpose and jurisdiction vocabulary ----------------- uses: 761 of 827 owners (92.0%) carry at least one value; 38 distinct values: marketing (484), hosting (102), audience_measurement (78), video (38), general (34), security (27), customer_relationship_management (27), social_media (27), design_optimization (23), code (16), ecommerce (8), compliance (7), location (6), search (5), information_services (5), font (5), holding_company (5), content_recommendation (5), government_licensing (5), weather (4), domain_registration (4), network_services (4), gaming (3), information_service (2), publishing (2), tag_manager (2), public_opinion_monitoring (2), search_engine (2), payment_platform (1), egovernment (1), health (1), financial_services (1), accesibility (1), uncategorized (1), content_reccomendation (1), trustmark (1), website_certification (1), malware (1) platforms: 803 of 827 owners (97.1%) carry at least one value; 5 distinct values: web (777), mobile (249), tv (176), iot (20), email (5) trade_groups: 104 of 827 owners (12.6%) carry at least one value; 9 distinct values: nai (85), daa (51), iab (36), amm (7), daac (2), edaa (1), mcma (1), ana (1), mma (1) country: 826 of 827 owners (99.9%) carry one; 36 distinct codes; top 10: US (487), CN (114), UK (46), DE (36), FR (15), CA (14), JP (13), RU (10), IL (9), SE (8) owners with a non-empty aliases: 275 of 827 (33.3%) owners with a non-empty notes: 263 of 827 (31.8%) owners with a non-empty site_privacy_policy_urls: 591 of 827 (71.5%) owners with a non-empty gdpr_statement_urls: 132 of 827 (16.0%) owners with a non-empty ccpa_urls: 4 of 827 (0.5%) owners with a non-empty opt_out_urls: 13 of 827 (1.6%) owners with a non-empty health_segment_urls: 52 of 827 (6.3%) owners with a non-empty crunchbase_id: 34 of 827 (4.1%) distinct languages in site_privacy_policy_urls: 69; top: eng (554), chi (90), ger (81), fre (71), spa (64), jpn (54), ita (46), por (44) Domains where webXray's immediate owner differs from the root of its ownership tree: 2,175 of 3,215 (67.7%). --- C. Coverage of the third-party surface --------------------------- Public Suffix List: 155b43d46932e933 6,941 ICANN rules, 3,290 PRIVATE rules, 8 exceptions https://publicsuffix.org/list/public_suffix_list.dat The ICANN/PRIVATE split is the most consequential choice in this script. `googleapis.com` is a PRIVATE rule, so folding with the private section leaves fonts.googleapis.com standing as its own 'registrable domain' -- and an exact-key lookup then finds no owner for it, even though all three lists name googleapis.com -> Google. Both rules are therefore reported. domain_summary.json rows: 47,836, of which 16,396 have three or more labels -- a count that does NOT answer 'how many are keyed by hostname rather than registrable domain', because that depends on the fold: ICANN-section fold: merges 15,651 rows -> 32,369 registrable domains ICANN+PRIVATE fold: merges 2,339 rows -> 45,525 registrable domains Rows dropped as not hostnames at all: 19 ['null', '[2a01:4f9:2a:26e0::2]', '[2604:2dc0:100:5ce5::]', '[2001:41d0:800:4623::]', '[2001:41d0:602:556f::]', '[2604:8380:2e00:5::2]', '[2604:4500:8:2ea::2]', '[2604:8380:3300:1::2]', '[2604:8380:2900:15::2]', '[2001:41d0:403:579b::]', '[2001:41d0:306:44e6::]', '[2604:4500:a:432::2]', '[2402:1f00:8201:4a2::]', '[2604:4500:6:5a0::2]', '[2604:4500:21:8::4]', '[2402:1f00:8001:2518::]', '[2402:1f00:8300:c97::]', '[2604:8380:2f00:16::2]', '[2001:41d0:700:782c::]'] -- a literal 'null' key with a prevalence attached is a defect in the published data, not a domain. Denominator: those registrable third-party domains, i.e. domains Tracker Radar's own crawl of regional top-site lists actually saw. This is Tracker Radar's view of the surface, not a neutral one. ICANN fold + parent-label lookup <- USE THIS (universe 32,369) List Named Share Prevalence-weighted webXray 669 2.1% 58.6% Tracker Radar 5,581 17.2% 84.3% Disconnect (properties+res) 2,268 7.0% 80.5% ICANN fold + exact-key lookup (universe 32,369) List Named Share Prevalence-weighted webXray 669 2.1% 58.6% Tracker Radar 5,581 17.2% 84.3% Disconnect (properties+res) 2,268 7.0% 80.5% ICANN+PRIVATE fold + parent-label lookup (universe 45,525) List Named Share Prevalence-weighted webXray 11,150 24.5% 59.7% Tracker Radar 18,099 39.8% 84.8% Disconnect (properties+res) 10,389 22.8% 80.6% ICANN+PRIVATE fold + exact-key lookup <- THE TRAP (universe 45,525) List Named Share Prevalence-weighted webXray 657 1.4% 54.5% Tracker Radar 5,539 12.2% 79.3% Disconnect (properties+res) 2,277 5.0% 75.8% The single row that drives most of that gap: fonts.googleapis.com, prevalence 0.369. googleapis.com is a PRIVATE PSL rule, and all three lists name it: webXray='Google APIs', TR='Google LLC', Disconnect='Google'. The bottom variant is what the first version of this script did, and it understates every list. The gap between the top and bottom rows is a measurement artefact, not a property of any list -- report which rule you used, because it moves the answer by more than the lists differ. top 100 by prevalence: webXray 71% Tracker Radar 98% Disconnect 94% top 1000 by prevalence: webXray 26% Tracker Radar 69% Disconnect 68% top 10000 by prevalence: webXray 5% Tracker Radar 29% Disconnect 17% Domains in the universe that neither webXray nor Disconnect names an owner for: 29,896 (92.4%), 14.4% of prevalence weight. Top 15 of those by prevalence (Tracker Radar's owner in brackets): tiktokw.us prev=0.0406 [ByteDance Ltd.] consentmanager.net prev=0.0359 [consentmanager AB] rapidedge.io prev=0.0247 [TR: none] raptivecdn.com prev=0.0231 [TR: none] tracookiepixel.xyz prev=0.0228 [TR: none] growplow.events prev=0.0228 [TR: none] shopifycdn.com prev=0.0181 [Shopify Inc.] digitalaudience.io prev=0.0152 [Social Audience B.V.] openwebmp.com prev=0.0130 [TR: none] userway.org prev=0.0122 [TR: none] copper6.com prev=0.0113 [TR: none] ahrefs.com prev=0.0112 [Ahrefs Pte Ltd] anyrtb.com prev=0.0111 [TR: none] sparteo.com prev=0.0108 [TR: none] axiom.co prev=0.0102 [TR: none] Most prevalent domain no list can name an owner for: tiktokw.us, prevalence 0.041 (rounded to 3 dp for quoting). Dropped rows and their prevalence, rounded to 3 dp for quoting: null=0.024, [2a01:4f9:2a:26e0::2]=0.000, [2604:2dc0:100:5ce5::]=0.000 --- D. Do two lists that both know a domain agree on the owner? ------ webXray vs Tracker Radar: 612 domains named by both identical owner string 29 (4.7%) agree after legal-suffix fold 305 (49.8%) one name contains the other 109 (17.8%) [e.g. Amazon / Amazon Technologies] neither 198 (32.4%) <- needs a human ...weighted by prevalence, the 'needs a human' share is 29.0% webXray-root vs Tracker Radar: 612 domains named by both identical owner string 34 (5.6%) agree after legal-suffix fold 276 (45.1%) one name contains the other 97 (15.8%) [e.g. Amazon / Amazon Technologies] neither 239 (39.1%) <- needs a human ...weighted by prevalence, the 'needs a human' share is 51.1% webXray vs Disconnect: 464 domains named by both identical owner string 171 (36.9%) agree after legal-suffix fold 213 (45.9%) one name contains the other 34 (7.3%) [e.g. Amazon / Amazon Technologies] neither 217 (46.8%) <- needs a human ...weighted by prevalence, the 'needs a human' share is 44.5% webXray-root vs Disconnect: 464 domains named by both identical owner string 164 (35.3%) agree after legal-suffix fold 195 (42.0%) one name contains the other 39 (8.4%) [e.g. Amazon / Amazon Technologies] neither 230 (49.6%) <- needs a human ...weighted by prevalence, the 'needs a human' share is 64.1% Tracker Radar vs Disconnect: 1,538 domains named by both identical owner string 82 (5.3%) agree after legal-suffix fold 634 (41.2%) one name contains the other 320 (20.8%) [e.g. Amazon / Amazon Technologies] neither 584 (38.0%) <- needs a human ...weighted by prevalence, the 'needs a human' share is 29.7% --- E. Sampled disagreements (highest prevalence first) -------------- webXray vs Tracker Radar: doubleclick.net prev=0.4456 webXray='DoubleClick' Tracker Radar='Google LLC' googlesyndication.com prev=0.2115 webXray='AdSense' Tracker Radar='Google LLC' adnxs.com prev=0.1463 webXray='Xandr' Tracker Radar='Microsoft Corporation' jsdelivr.net prev=0.1242 webXray='jsDelivr' Tracker Radar='Prospect One' rubiconproject.com prev=0.1223 webXray='Rubicon Project' Tracker Radar='Magnite, Inc.' bing.com prev=0.1160 webXray='Bing' Tracker Radar='Microsoft Corporation' amazon-adsystem.com prev=0.1067 webXray='Amazon Marketing Services' Tracker Radar='Amazon Technologies, Inc.' linkedin.com prev=0.1054 webXray='LinkedIn' Tracker Radar='Microsoft Corporation' smartadserver.com prev=0.0752 webXray='Smart AdServer' Tracker Radar='Smartadserver S.A.S' bidswitch.net prev=0.0746 webXray='Bidswitch' Tracker Radar='IPONWEB GmbH' turn.com prev=0.0661 webXray='Turn' Tracker Radar='Amobee, Inc' 2mdn.net prev=0.0616 webXray='DoubleClick' Tracker Radar='Google LLC' scorecardresearch.com prev=0.0584 webXray='ScorecardResearch' Tracker Radar='comScore, Inc' dotomi.com prev=0.0583 webXray='Dotomi' Tracker Radar='Conversant LLC' 1rx.io prev=0.0546 webXray='Blinkx' Tracker Radar='RhythmOne' sitescout.com prev=0.0540 webXray='SiteScout' Tracker Radar='Centro, Inc.' simpli.fi prev=0.0519 webXray='simpli.fi' Tracker Radar='Simplifi Holdings Inc.' googleadservices.com prev=0.0510 webXray='AdSense' Tracker Radar='Google LLC' youtube.com prev=0.0485 webXray='YouTube' Tracker Radar='Google LLC' rfihub.com prev=0.0429 webXray='Rocketfuel' Tracker Radar='Zeta Global' loopme.me prev=0.0423 webXray='Loopme' Tracker Radar='Online Media Solutions Ltd. dba Brightcom' licdn.com prev=0.0373 webXray='LinkedIn' Tracker Radar='Microsoft Corporation' agkn.com prev=0.0347 webXray='Neustar Marketing' Tracker Radar='TransUnion LLC' intentiq.com prev=0.0337 webXray='Intent IQ' Tracker Radar='Almondnet Group' ytimg.com prev=0.0327 webXray='YouTube' Tracker Radar='Google LLC' webXray-root vs Tracker Radar: googletagmanager.com prev=0.5754 webXray-root='Alphabet' Tracker Radar='Google LLC' google.com prev=0.4525 webXray-root='Alphabet' Tracker Radar='Google LLC' doubleclick.net prev=0.4456 webXray-root='Alphabet' Tracker Radar='Google LLC' gstatic.com prev=0.4013 webXray-root='Alphabet' Tracker Radar='Google LLC' googleapis.com prev=0.3692 webXray-root='Alphabet' Tracker Radar='Google LLC' google-analytics.com prev=0.3489 webXray-root='Alphabet' Tracker Radar='Google LLC' googlesyndication.com prev=0.2115 webXray-root='Alphabet' Tracker Radar='Google LLC' adnxs.com prev=0.1463 webXray-root='AT&T' Tracker Radar='Microsoft Corporation' rubiconproject.com prev=0.1223 webXray-root='Rubicon Project' Tracker Radar='Magnite, Inc.' rlcdn.com prev=0.1098 webXray-root='Acxiom' Tracker Radar='LiveRamp Holdings, Inc.' yahoo.com prev=0.0984 webXray-root='Verizon' Tracker Radar='Yahoo Inc.' tapad.com prev=0.0973 webXray-root='Telenor' Tracker Radar='Tapad, Inc.' smartadserver.com prev=0.0752 webXray-root='Smart AdServer' Tracker Radar='Smartadserver S.A.S' turn.com prev=0.0661 webXray-root='Singtel' Tracker Radar='Amobee, Inc' lijit.com prev=0.0656 webXray-root='Federated Media' Tracker Radar='Sovrn Holdings' 2mdn.net prev=0.0616 webXray-root='Alphabet' Tracker Radar='Google LLC' pippio.com prev=0.0594 webXray-root='Acxiom' Tracker Radar='LiveRamp Holdings, Inc.' dotomi.com prev=0.0583 webXray-root='Here, There & Everywhere' Tracker Radar='Conversant LLC' 1rx.io prev=0.0546 webXray-root='Marimedia' Tracker Radar='RhythmOne' simpli.fi prev=0.0519 webXray-root='GTCR' Tracker Radar='Simplifi Holdings Inc.' googleadservices.com prev=0.0510 webXray-root='Alphabet' Tracker Radar='Google LLC' teads.tv prev=0.0486 webXray-root='Altice SA' Tracker Radar='Teads ( Luxenbourg ) SA' youtube.com prev=0.0485 webXray-root='Alphabet' Tracker Radar='Google LLC' 360yield.com prev=0.0467 webXray-root='Azerion' Tracker Radar='Improve Digital BV' fwmrm.net prev=0.0439 webXray-root='Comcast' Tracker Radar='FreeWheel' webXray vs Disconnect: doubleclick.net prev=0.4456 webXray='DoubleClick' Disconnect='Google' googlesyndication.com prev=0.2115 webXray='AdSense' Disconnect='Google' facebook.net prev=0.1830 webXray='Facebook' Disconnect='Meta' facebook.com prev=0.1665 webXray='Facebook' Disconnect='Meta' adnxs.com prev=0.1463 webXray='Xandr' Disconnect='Microsoft' jsdelivr.net prev=0.1242 webXray='jsDelivr' Disconnect='Volentio JSD' rubiconproject.com prev=0.1223 webXray='Rubicon Project' Disconnect='Magnite' bing.com prev=0.1160 webXray='Bing' Disconnect='Microsoft' casalemedia.com prev=0.1114 webXray='Index Exchange' Disconnect='IndexExchange' linkedin.com prev=0.1054 webXray='LinkedIn' Disconnect='Microsoft' crwdcntrl.net prev=0.0972 webXray='Lotame' Disconnect='PublicisGroupe' liadm.com prev=0.0755 webXray='LiveIntent' Disconnect='ZetaGlobal' smartadserver.com prev=0.0752 webXray='Smart AdServer' Disconnect='Equativ' bidswitch.net prev=0.0746 webXray='Bidswitch' Disconnect='Criteo' turn.com prev=0.0661 webXray='Turn' Disconnect='Nexxen' sharethrough.com prev=0.0659 webXray='Sharethrough' Disconnect='Equativ' bidr.io prev=0.0619 webXray='Beeswax' Disconnect='Comcast' 2mdn.net prev=0.0616 webXray='DoubleClick' Disconnect='Google' scorecardresearch.com prev=0.0584 webXray='ScorecardResearch' Disconnect='comScore' dotomi.com prev=0.0583 webXray='Dotomi' Disconnect='PublicisGroupe' outbrain.com prev=0.0550 webXray='Outbrain' Disconnect='Teads' 1rx.io prev=0.0546 webXray='Blinkx' Disconnect='Nexxen' sitescout.com prev=0.0540 webXray='SiteScout' Disconnect='BasisTechnologies' postrelease.com prev=0.0513 webXray='Nativo' Disconnect='Life360' googleadservices.com prev=0.0510 webXray='AdSense' Disconnect='Google' webXray-root vs Disconnect: google.com prev=0.4525 webXray-root='Alphabet' Disconnect='Google' doubleclick.net prev=0.4456 webXray-root='Alphabet' Disconnect='Google' gstatic.com prev=0.4013 webXray-root='Alphabet' Disconnect='Google' googleapis.com prev=0.3692 webXray-root='Alphabet' Disconnect='Google' google-analytics.com prev=0.3489 webXray-root='Alphabet' Disconnect='Google' googlesyndication.com prev=0.2115 webXray-root='Alphabet' Disconnect='Google' facebook.net prev=0.1830 webXray-root='Facebook' Disconnect='Meta' facebook.com prev=0.1665 webXray-root='Facebook' Disconnect='Meta' adnxs.com prev=0.1463 webXray-root='AT&T' Disconnect='Microsoft' jsdelivr.net prev=0.1242 webXray-root='Prospect One' Disconnect='Volentio JSD' rubiconproject.com prev=0.1223 webXray-root='Rubicon Project' Disconnect='Magnite' casalemedia.com prev=0.1114 webXray-root='Index Exchange' Disconnect='IndexExchange' rlcdn.com prev=0.1098 webXray-root='Acxiom' Disconnect='LiveRamp' yahoo.com prev=0.0984 webXray-root='Verizon' Disconnect='Yahoo!' tapad.com prev=0.0973 webXray-root='Telenor' Disconnect='Tapad' crwdcntrl.net prev=0.0972 webXray-root='Lotame' Disconnect='PublicisGroupe' liadm.com prev=0.0755 webXray-root='LiveIntent' Disconnect='ZetaGlobal' smartadserver.com prev=0.0752 webXray-root='Smart AdServer' Disconnect='Equativ' bidswitch.net prev=0.0746 webXray-root='IPONWEB' Disconnect='Criteo' turn.com prev=0.0661 webXray-root='Singtel' Disconnect='Nexxen' sharethrough.com prev=0.0659 webXray-root='Sharethrough' Disconnect='Equativ' lijit.com prev=0.0656 webXray-root='Federated Media' Disconnect='Sovrn' bidr.io prev=0.0619 webXray-root='Beeswax' Disconnect='Comcast' 2mdn.net prev=0.0616 webXray-root='Alphabet' Disconnect='Google' pippio.com prev=0.0594 webXray-root='Acxiom' Disconnect='LiveRamp' Tracker Radar vs Disconnect: facebook.net prev=0.1830 Tracker Radar='Facebook, Inc.' Disconnect='Meta' facebook.com prev=0.1665 Tracker Radar='Facebook, Inc.' Disconnect='Meta' jsdelivr.net prev=0.1242 Tracker Radar='Prospect One' Disconnect='Volentio JSD' casalemedia.com prev=0.1114 Tracker Radar='Index Exchange, Inc.' Disconnect='IndexExchange' crwdcntrl.net prev=0.0972 Tracker Radar='Lotame Solutions, Inc.' Disconnect='PublicisGroupe' stackadapt.com prev=0.0781 Tracker Radar='Collective Roll' Disconnect='StackAdapt' liadm.com prev=0.0755 Tracker Radar='LiveIntent Inc.' Disconnect='ZetaGlobal' smartadserver.com prev=0.0752 Tracker Radar='Smartadserver S.A.S' Disconnect='Equativ' bidswitch.net prev=0.0746 Tracker Radar='IPONWEB GmbH' Disconnect='Criteo' creativecdn.com prev=0.0724 Tracker Radar='RTB House S.A.' Disconnect='RTBHouse' turn.com prev=0.0661 Tracker Radar='Amobee, Inc' Disconnect='Nexxen' sharethrough.com prev=0.0659 Tracker Radar='Sharethrough, Inc.' Disconnect='Equativ' bidr.io prev=0.0619 Tracker Radar='Beeswax' Disconnect='Comcast' dotomi.com prev=0.0583 Tracker Radar='Conversant LLC' Disconnect='PublicisGroupe' temu.com prev=0.0555 Tracker Radar='Pinduoduo Inc.' Disconnect='PDD Holdings' outbrain.com prev=0.0550 Tracker Radar='Outbrain' Disconnect='Teads' 1rx.io prev=0.0546 Tracker Radar='RhythmOne' Disconnect='Nexxen' sitescout.com prev=0.0540 Tracker Radar='Centro, Inc.' Disconnect='BasisTechnologies' blismedia.com prev=0.0532 Tracker Radar='Blis Global Ltd' Disconnect='DT' simpli.fi prev=0.0519 Tracker Radar='Simplifi Holdings Inc.' Disconnect='Simpli.fi' postrelease.com prev=0.0513 Tracker Radar='Nativo, Inc' Disconnect='Life360' ipredictive.com prev=0.0506 Tracker Radar='Adelphic, Inc.' Disconnect='Viant' contextweb.com prev=0.0481 Tracker Radar='Pulsepoint, Inc.' Disconnect='Internet Brands' a-mo.net prev=0.0474 Tracker Radar='Monet Engine Inc.' Disconnect='AdaptMX' 360yield.com prev=0.0467 Tracker Radar='Improve Digital BV' Disconnect='Azerion' --- F. Named-entity residue of the suffix fold ----------------------- webXray owner names the fold leaves untouched: 782 of 827 (94.6%). Of the 45 it does change, only 16 lose a legal-form suffix; the other 29 change on punctuation alone: 56.com, AT&T, Ask.com, Bootstrap_China, Clearstream.TV, Cm_browser, Dictionary.com, Dun & Bradstreet, ... The fold merges no synonyms, so every figure in section D is a lower bound on real agreement and an upper bound on real disagreement.
K. Unedited output: scripts/owner_adjudication.py
Run as python3 scripts/owner_adjudication.py –table.
30 disagreements adjudicated 2026-08-17 against primary sources; 2 failed the sourcing bar and are EXCLUDED from every tally below, leaving 28. Excluded: 360yield.com (UNRESOLVED -- no dated primary source), fwmrm.net (acquired 2014; post-2026 spinoff UNRESOLVED) Selected as the highest-prevalence disagreements, NOT sampled at random: these tallies describe this set and are not an error rate for any list. List current stale granularity error absent webXray 1 18 4 1 4 Tracker Radar 8 12 6 2 0 Disconnect 26 0 2 0 0 webXray: has an entry for 24 of 28; of those, 1 name today's owner (4% of its own entries). Tracker Radar: has an entry for 28 of 28; of those, 8 name today's owner (29% of its own entries). Disconnect: has an entry for 28 of 28; of those, 26 name today's owner (93% of its own entries). The 'absent' column cannot be read as a coverage result either: the rows were ranked by Tracker Radar's own prevalence field, so Tracker Radar covering all 28 of 28 is partly how the sample was drawn. Rows scored as an outright error rather than staleness: 3 -- 1rx.io (webXray), jsdelivr.net (Tracker Radar), stackadapt.com (Tracker Radar) adnxs.com Microsoft Corporation stale current current https://www.prnewswire.com/news-releases/att-agrees-to-microsoft-acquisition-of-xandr-301448996.html rubiconproject.com Magnite, Inc. stale current current https://investor.magnite.com/news-releases/news-release-details/rubicon-project-and-telaria-complete-merger-following yahoo.com Yahoo Inc. (Apollo-managed funds majority) stale current current https://www.apollo.com/insights-news/pressreleases/2021/09/apollo-funds-complete-acquisition-of-yahoo-161530593 facebook.com Meta Platforms, Inc. stale stale current https://about.fb.com/news/2021/10/facebook-company-is-now-meta/ bidr.io Comcast (FreeWheel/Beeswax) stale stale current https://www.freewheel.com/news/freewheel-to-acquire-ad-tech-leader-beeswax turn.com Nexxen International Ltd. stale stale current https://investors.nexxen.com/news-releases/news-release-details/tremor-international-tremor-announces-closing-amobee-acquisition 1rx.io Nexxen International Ltd. error stale current https://nexxen.com/tremor-international-group-rebrands-as-nexxen/ outbrain.com Teads Holding Co. stale stale current https://investors.teads.com/news-releases/news-release-details/outbrain-completes-change-corporate-name-teads smartadserver.com Equativ stale stale current https://www.equativ.com/press/smart-adserver-rebrands-as-equativ sharethrough.com Equativ stale stale current https://www.equativ.com/press/equativ-and-sharethrough-will-now-operate-under-equativ-brand-solidifying-global-position-as-leading-end-to-end-media-platform postrelease.com Life360, Inc. (Nativo) stale stale current https://ads.life360.com/newsroom/life360-completes-acquisition-of-nativo-and-surpasses-50-million-us-mau crwdcntrl.net Publicis Groupe S.A. (Epsilon/Lotame) stale stale current https://www.publicisgroupe.com/en/news/press-releases/publicis-to-acquire-lotame-the-world-s-leading-independent-end-to-end-data-solution dotomi.com Publicis Groupe S.A. (Epsilon/Conversant) stale granularity current https://www.globenewswire.com/news-release/2019/07/02/1877064/0/en/Publicis-Groupe-finalizes-the-acquisition-of-Epsilon.html agkn.com TransUnion LLC stale current current https://www.globenewswire.com/news-release/2021/12/01/2344146/0/en/TransUnion-and-Neustar-Announce-Transaction-Close.html rlcdn.com LiveRamp Holdings, Inc. stale current current https://www.sec.gov/Archives/edgar/data/0000733269/000119312518279280/d616065d8k12b.htm tapad.com Experian plc stale granularity granularity https://www.experianplc.com/newsroom/press-releases/2020/experian-acquires-tapad contextweb.com Internet Brands, Inc. (PulsePoint) stale stale current https://www.pulsepoint.com/press-releases/internet-brands-to-acquire-pulsepoint temu.com PDD Holdings Inc. absent stale current https://www.sec.gov/Archives/edgar/data/1737806/000110465924051610/pdd-20231231x20f.htm lijit.com Sovrn Holdings, Inc. stale current current https://lijit.com/ teads.tv Teads Holding Co. stale current current https://www.globenewswire.com/news-release/2025/02/03/3019410/0/en/Outbrain-Completes-the-Acquisition-of-Teads.html blismedia.com Deutsche Telekom AG (T-Mobile US) absent stale current https://report.telekom.com/interim-report-q2-2025/financial-statements/significant-events-and-transactions/changes-in-the-composition-of-the-group.html jsdelivr.net Volentio JSD Limited granularity error current https://www.jsdelivr.com/documents/data-processing-agreement.pdf stackadapt.com StackAdapt Inc. (independent) absent error current https://www.stackadapt.com/legal-document-centre bidswitch.net Criteo S.A. (IPONWEB) granularity granularity current https://criteo.investorroom.com/2022-08-03-CRITEO-REPORTS-STRONG-SECOND-QUARTER-2022-RESULTS intentiq.com AlmondNet Group granularity granularity granularity https://www.intentiq.com/who-we-are/ ipredictive.com Viant Technology LLC (Adelphic) absent granularity current https://www.viantinc.com/company/news/press-releases/viant-completes-integration-of-adelphic-into-viant-advertising-cloud/ simpli.fi Simplifi Holdings Inc. (GTCR + Blackstone) granularity granularity current https://www.blackstone.com/news/press/simpli-fi-a-leading-programmatic-advertising-platform-announces-completion-of-significant-investment-from-blackstone-at-1-5-billion-valuation/ casalemedia.com Index Exchange, Inc. (independent) current current current https://www.indexexchange.com/team/andrew-casale/ 360yield.com Azerion Group N.V. (Improve Digital) current granularity current https://improvedigital.com/about/ EXCLUDED fwmrm.net Comcast Corporation (FreeWheel) current current current https://www.sec.gov/cgi-bin/browse-edgar?action=getcompany&CIK=0001166691 EXCLUDED
Related
- webxray — the page these notes are behind.
- random_sample — the appendix: all eleven scripts of the 2026-09-05 random sample, all 10 unedited outputs, and the adjudication brief.
- corpus — corpus scope, the selection funnel and extraction stability.
References
- [1]
- Steffens, Marius; Musch, Marius; Johns, Martin; Stock, Ben (2021): "Who’s Hosting the Block Party? Studying Third-Party Blockage of CSP and SRI", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [2]
- Musa, Maaz Bin; Nithyanand, Rishab (2022): "ATOM: Ad-network Tomography", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [3]
- Libert, Timothy (2018): "An Automated Approach to Auditing Disclosure of Third-Party Data Collection in Website Privacy Policies", in: Proceedings of the ACM Web Conference. (DOI)
- [4]
- Sánchez-Rola, Iskander; Dell'Amico, Matteo; Balzarotti, Davide; Vervier, Pierre-Antoine; Bilge, Leyla (2021): "Journey to the Center of the Cookie Ecosystem: Unraveling Actors' Roles and Relationships", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [5]
- Matte, Célestin; Bielova, Nataliia; Santos, Cristiana Teixeira (2020): "Do Cookie Banners Respect my Choice? Measuring Legal Compliance of Banners from IAB Europe's Transparency and Consent Framework", in: Proceedings of the IEEE Symposium on Security and Privacy. (DOI)
- [6]
- Yang, Zhiju; Yue, Chuan (2020): "A Comparative Measurement Study of Web Tracking on Mobile and Desktop Environments", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [7]
- Jakaria, Md; Huang, Danny Yuxing; Das, Anupam (2024): "Connecting the Dots: Tracing Data Endpoints in IoT Devices", in: Proceedings on Privacy Enhancing Technologies. (DOI)
- [8]
- McQuistin, Stephen; Snyder, Peter; Perkins, Colin; Haddadi, Hamed; Tyson, Gareth (2023): "A First Look at the Privacy Harms of the Public Suffix List", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [9]
- Englehardt, Steven; Narayanan, Arvind (2016): "Online Tracking: A 1-million-site Measurement and Analysis", in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, pp. 1388–1401. Association for Computing Machinery, New York, NY, USA. (DOI) (Link)
- [10]
- Utz, Christine; Amft, Sabrina; Degeling, Martin; Holz, Thorsten; Fahl, Sascha; Schaub, Florian (2023): "Privacy Rarely Considered: Exploring Considerations in the Adoption of Third-Party Services by Websites", Proceedings on Privacy Enhancing Technologies 2023(1):5-28. (DOI)
- [11]
- Jannett, Louis; Mayer, Andreas; Westers, Maximilian; Mladenov, Vladislav; Mainka, Christian; Schwenk, Jörg (2026): "The State of Passkeys: Studying the Adoption and Security of Passkeys on the Web", in: Proceedings of the USENIX Security Symposium. (Link)
- [12]
- Selmo, Carlos; Carisimo, Esteban; Bustamante, Fabián E.; Alvarez-Hamelin, J. Ignacio (2025): "Learning AS-to-Organization Mappings with Borges", in: Proceedings of the 2025 ACM Internet Measurement Conference, pp. 120-133. (DOI)
- [13]
- Wu, Xiaoyuan; Hu, Lydia; Zeng, Eric; Habib, Hana; Bauer, Lujo (2025): "Transparency or Information Overload? Evaluating Users’ Comprehension and Perceptions of the iOS App Privacy Report", in: Proceedings of the Network and Distributed System Security Symposium. (Link)
- [14]
- Gouda, Deepak; Dainotti, Alberto; Testart, Cecilia (2025): "Prefix2Org: Mapping BGP Prefixes to Organizations", in: Proceedings of the ACM Internet Measurement Conference. (DOI)
- [15]
- Dambra, Savino; Sanchez-Rola, Iskander; Bilge, Leyla; Balzarotti, Davide (2022): "When Sally Met Trackers: Web Tracking From the Users' Perspective", in: Proceedings of the USENIX Security Symposium. (Link)
